1
Performance Characterization of CXL Memory and Its Use Cases.
Presenter: Xi (Sherry) Wang
Advisor: Dong Li
PASA Lab
Why Care About CXL?
2
PASA Lab
3
Memory Tiering Solutions in OS
Page Migration
Page Interleaving
DRAM
DRAM
CXL
High Capacity
Low $/GB
Low Performance
Low Capacity
High $/GB
High Performance
LLM & HPC
memory growth
Tiered Memory System
PASA Lab
4
Memory Tiering Solutions in OS
Page Migration
Page Interleaving
DRAM
DRAM
CXL
High Capacity
Low $/GB
Low Performance
Low Capacity
High $/GB
High Performance
LLM & HPC
memory growth
How does genuine CXL hardware behave? How to integrate with local DRAM node (LDRAM) & remote DRAM node (RDRAM)?
Key Research Questions
PASA Lab
5
Memory Tiering Solutions in OS
Page Migration
Page Interleaving
DRAM
DRAM
CXL
High Capacity
Low $/GB
Low Performance
Low Capacity
High $/GB
High Performance
LLM & HPC
memory growth
Which workloads benefit
—or not—from CXL?
Key Research Questions
How does genuine CXL hardware behave? How to integrate with local DRAM node (LDRAM) & remote DRAM node (RDRAM)?
PASA Lab
6
Memory Tiering Solutions in OS
Page Migration
Page Interleaving
DRAM
DRAM
CXL
High Capacity
Low $/GB
Low Performance
Low Capacity
High $/GB
High Performance
LLM & HPC
memory growth
Which workloads benefit
—or not—from CXL?
How should OS place data across LDRAM, RDRAM & CXL tiers?
Key Research Questions
How does genuine CXL hardware behave? How to integrate with local DRAM node (LDRAM) & remote DRAM node (RDRAM)?
PASA Lab
7
Memory Tiering Solutions in OS
Page Migration
Page Interleaving
DRAM
DRAM
CXL
High Capacity
Low $/GB
Low Performance
Low Capacity
High $/GB
High Performance
LLM & HPC
memory growth
Which workloads benefit
—or not—from CXL?
What page placement policies should be used?
How should OS place data across LDRAM, RDRAM & CXL tiers?
Key Research Questions
How does genuine CXL hardware behave? How to integrate with local DRAM node (LDRAM) & remote DRAM node (RDRAM)?
PASA Lab
8
| CPU | Memory | CXL | PCIe |
Sys A | 2× AMD EPYC 9354 @3.8 GHz, 32 cores per CPU | 12× DDR5-4800 channels, mem 768GB, max bandwidth 460.8 GB/s per socket | Single channel DDR5-4800, mem 128 GB, max bandwidth 38.4 GB/s per channel | PCIe 5.0, 32GT/s, 16 lanes |
Sys B | 2× Intel(R) Xeon(R) Platinum 8470 @2.0GHz, 52 cores per CPU | 8× DDR5-4800 channels, mem 1TB, max bandwidth 307.2 GB/s per socket | Single channel DDR5-8000, mem 64 GB, max bandwidth 64.0 GB/s per channel | PCIe 5.0, 32GT/s, 16 lanes |
Sys C | 2× Intel(R) Xeon(R) Gold 6438Y+ @2.0GHz, 32 cores per CPU | 8× DDR5-4800 channels, mem 512GB, max bandwidth 307.2 GB/s per socket | Dual channel DDR5-6200, mem 128 GB, max bandwidth 48.4 GB/s per channel | PCIe 5.0, 32GT/s, 16 lanes |
Experimental Platforms & CXL Devices
PASA Lab
9
Load latency with random and sequential accesses to a cache block.
Latency: CXL ≈ 2‑Hop NUMA
PASA Lab
10
Bandwidth scaling for data loading.
CXL saturates at 4-8 threads.
Bandwidth & Scalability
Intel MLC
PASA Lab
11
Bandwidth scaling for data loading.
CXL saturates at 4-8 threads.
CXL Peak = 10‑80 % of local DRAM bandwidth.
Bandwidth & Scalability
Intel MLC
PASA Lab
12
Bandwidth scaling for data loading.
CXL saturates at 4-8 threads.
CXL Peak = 10‑80 % of local DRAM bandwidth.
Optimal thread mapping leads to aggregate BW to 420 GB/s:
CXL : LDRAM : RDRAM = 6 : 23 : 23
Bandwidth & Scalability
Intel MLC
PASA Lab
13
Interleaving
Can LLM training benefit from the extra bandwidth provided by CXL?
LLM Training (ZeRO‑Offload) with CXL
PASA Lab
14
The optimizer is latency-sensitive!
Longer data movement latency!
LLM Training (ZeRO‑Offload) with CXL
PASA Lab
15
LLM Inference (FlexGen) with CXL
PASA Lab
16
Testbed: LLaMA, OPT
Comparison LLM inference throughput across memory systems, each with 324 GB capacity.
Data loading in prefill stage is latency‑sensitive
Decode stage is bandwidth‑sensitive
LLM Inference (FlexGen) with CXL
PASA Lab
17
Comparison of LLM inference throughput across various memory systems with different capacities.
Extra 128 GB CXL
Testbed: LLaMA, OPT
LLM Inference (FlexGen) with CXL
PASA Lab
18
The performance of various interleaving policies for HPC applications.
HPC Workload Spectrum
PASA Lab
19
The performance of various interleaving policies for HPC applications.
HPC Workload Spectrum
PASA Lab
20
……
DRAM
CXL
App Memory
Is blind uniform page-level interleaving good enough?
Physical memory placement
Object‑Level Interleaving (OLI)
PASA Lab
21
……
DRAM
CXL
App Memory
Physical memory placement
Selected Objects
Other Objects
Object‑Level Interleaving (OLI)
PASA Lab
22
Object‑Level Interleaving (OLI)
PASA Lab
23
Use sufficient LDRAM (128 GB)
Object‑Level Interleaving (OLI)
PASA Lab
24
Use sufficient LDRAM (128 GB)
Use insufficient LDRAM (64 GB)
Object‑Level Interleaving (OLI)
PASA Lab
25
Testbed: NoBalance, AutoNUMA, Tiering‑0.8, TPP
Memory Tiering via Page Migration
PASA Lab
26
Testbed: NoBalance, AutoNUMA, Tiering‑0.8, TPP
Memory Tiering via Page Migration
PASA Lab
27
Conclusion & Interesting Takeaway
PASA Lab