1 of 27

1

Performance Characterization of CXL Memory and Its Use Cases.

Presenter: Xi (Sherry) Wang

Advisor: Dong Li

PASA Lab

2 of 27

Why Care About CXL?

  • AI/ML & HPC memory footprints outpace CPU DRAM capacity

  • Compute Express Link (CXL) offers cache‑coherent memory expansion over PCIe

  • But real‑world performance trade‑offs and best‑practice usage remain unclear

2

PASA Lab

3 of 27

3

Memory Tiering Solutions in OS

Page Migration

Page Interleaving

DRAM

DRAM

CXL

High Capacity

Low $/GB

Low Performance

Low Capacity

High $/GB

High Performance

LLM & HPC

memory growth

Tiered Memory System

PASA Lab

4 of 27

4

Memory Tiering Solutions in OS

Page Migration

Page Interleaving

DRAM

DRAM

CXL

High Capacity

Low $/GB

Low Performance

Low Capacity

High $/GB

High Performance

LLM & HPC

memory growth

How does genuine CXL hardware behave? How to integrate with local DRAM node (LDRAM) & remote DRAM node (RDRAM)?

Key Research Questions

PASA Lab

5 of 27

5

Memory Tiering Solutions in OS

Page Migration

Page Interleaving

DRAM

DRAM

CXL

High Capacity

Low $/GB

Low Performance

Low Capacity

High $/GB

High Performance

LLM & HPC

memory growth

Which workloads benefit

—or not—from CXL?

Key Research Questions

How does genuine CXL hardware behave? How to integrate with local DRAM node (LDRAM) & remote DRAM node (RDRAM)?

PASA Lab

6 of 27

6

Memory Tiering Solutions in OS

Page Migration

Page Interleaving

DRAM

DRAM

CXL

High Capacity

Low $/GB

Low Performance

Low Capacity

High $/GB

High Performance

LLM & HPC

memory growth

Which workloads benefit

—or not—from CXL?

How should OS place data across LDRAM, RDRAM & CXL tiers?

Key Research Questions

How does genuine CXL hardware behave? How to integrate with local DRAM node (LDRAM) & remote DRAM node (RDRAM)?

PASA Lab

7 of 27

7

Memory Tiering Solutions in OS

Page Migration

Page Interleaving

DRAM

DRAM

CXL

High Capacity

Low $/GB

Low Performance

Low Capacity

High $/GB

High Performance

LLM & HPC

memory growth

Which workloads benefit

—or not—from CXL?

What page placement policies should be used?

How should OS place data across LDRAM, RDRAM & CXL tiers?

Key Research Questions

How does genuine CXL hardware behave? How to integrate with local DRAM node (LDRAM) & remote DRAM node (RDRAM)?

PASA Lab

8 of 27

8

CPU

Memory

CXL

PCIe

Sys

A

2× AMD EPYC 9354 @3.8 GHz,

32 cores per CPU

12× DDR5-4800 channels, mem 768GB,

max bandwidth 460.8 GB/s per socket

Single channel DDR5-4800, mem 128 GB,

max bandwidth 38.4 GB/s per channel

PCIe 5.0, 32GT/s, 16 lanes

Sys

B

2× Intel(R) Xeon(R) Platinum 8470 @2.0GHz,

52 cores per CPU

8× DDR5-4800 channels, mem 1TB,

max bandwidth 307.2 GB/s per socket

Single channel DDR5-8000, mem 64 GB,

max bandwidth 64.0 GB/s per channel

PCIe 5.0, 32GT/s, 16 lanes

Sys

C

2× Intel(R) Xeon(R) Gold 6438Y+ @2.0GHz,

32 cores per CPU

8× DDR5-4800 channels, mem 512GB,

max bandwidth 307.2 GB/s per socket

Dual channel DDR5-6200, mem 128 GB,

max bandwidth 48.4 GB/s per channel

PCIe 5.0, 32GT/s, 16 lanes

Experimental Platforms & CXL Devices

PASA Lab

9 of 27

  • +150‑210 ns over local DRAM; vendor spread >30 %
  • Under heavy load DRAM latencies converge toward CXL

9

Load latency with random and sequential accesses to a cache block.

Latency: CXL ≈ 2‑Hop NUMA

PASA Lab

10 of 27

10

Bandwidth scaling for data loading.

CXL saturates at 4-8 threads.

Bandwidth & Scalability

Intel MLC

PASA Lab

11 of 27

11

Bandwidth scaling for data loading.

CXL saturates at 4-8 threads.

CXL Peak = 10‑80 % of local DRAM bandwidth.

Bandwidth & Scalability

Intel MLC

PASA Lab

12 of 27

12

Bandwidth scaling for data loading.

CXL saturates at 4-8 threads.

CXL Peak = 10‑80 % of local DRAM bandwidth.

Optimal thread mapping leads to aggregate BW to 420 GB/s:

CXL : LDRAM : RDRAM = 6 : 23 : 23

Bandwidth & Scalability

Intel MLC

PASA Lab

13 of 27

  • GPU: performs ① forward and ②backward
  • ③ GPU offloads gradients: GPU → CPU
  • ④ CPU: performs optimization computation
  • ⑤ Move updated parameters before the next forward step: CPU → GPU.

13

Interleaving

Can LLM training benefit from the extra bandwidth provided by CXL?

LLM Training (ZeRO‑Offload) with CXL

PASA Lab

14 of 27

  • CXL gives ≤5 % speed‑up
    • Lose performance when the optimizer latency dominates performance
  • GPU⇄CXL path goes via CPU (CXL 1.1) → PCIe bottleneck
  • Needs peer‑to‑peer (CXL 3.x) & latency‑aware placement

14

The optimizer is latency-sensitive!

Longer data movement latency!

LLM Training (ZeRO‑Offload) with CXL

PASA Lab

15 of 27

  • Prefill Stage (latency-sensitive):
    • ① Transfer parameters: CPU → GPU
    • ② GPU: Executes attention and MLP computation layer-by-layer
    • ③ Offload KV cache: GPU → CPU

  • Decode Stage (bandwidth-sensitive):
    • ④ CPU: Executes attention computation
    • ⑤ Transfer parameters and activations for MLP comp: CPU → GPU
    • ⑥ Transfer MLP activations for subsequent comp : GPU → CPU

15

LLM Inference (FlexGen) with CXL

PASA Lab

16 of 27

16

Testbed: LLaMA, OPT

Comparison LLM inference throughput across memory systems, each with 324 GB capacity.

Data loading in prefill stage is latency‑sensitive

  • DRAM matters

Decode stage is bandwidth‑sensitive

  • CXL ≈ RDRAM, beats NVMe by 20‑27 %

LLM Inference (FlexGen) with CXL

PASA Lab

17 of 27

17

Comparison of LLM inference throughput across various memory systems with different capacities.

Extra 128 GB CXL

  • ↑ batch 1.1‑3× → 28‑86 % throughput gain

Testbed: LLaMA, OPT

LLM Inference (FlexGen) with CXL

PASA Lab

18 of 27

  • 6 NPB Class E kernels + XSBench (116‑210 GB memory)

18

The performance of various interleaving policies for HPC applications.

HPC Workload Spectrum

PASA Lab

19 of 27

  • 6 NPB Class E kernels + XSBench (116‑210 GB memory)
  • Interleave RDRAM+CXL saves DRAM with <9 % perf loss

19

The performance of various interleaving policies for HPC applications.

HPC Workload Spectrum

PASA Lab

20 of 27

  • Uniform Page-Level Interleaving:
    • Interleaves pages across the entire application

20

……

DRAM

CXL

App Memory

Is blind uniform page-level interleaving good enough?

Physical memory placement

Object‑Level Interleaving (OLI)

PASA Lab

21 of 27

  • Object-Level Interleaving :
    • Selectively interleaves pages for specific data objects
    • Other objects allocated to fast DRAM

21

……

DRAM

CXL

App Memory

Physical memory placement

Selected Objects

Other Objects

Object‑Level Interleaving (OLI)

PASA Lab

22 of 27

  • Object-Level Interleaving:
    • Selectively interleaves pages for specific data objects → requirements
      • Large memory footprint (≥10% of total memory consumption)
      • Intensive memory access
    • Other objects allocated to fast DRAM

22

Object‑Level Interleaving (OLI)

PASA Lab

23 of 27

  • When LDRAM is sufficient:
    • OLI ≈ LDRAM preferred
    • OLI cuts LDRAM footprint by 32 %
    • OLI is 65 % faster than uniform interleave

23

Use sufficient LDRAM (128 GB)

Object‑Level Interleaving (OLI)

PASA Lab

24 of 27

  • When LDRAM is sufficient:
    • OLI ≈ LDRAM preferred
    • OLI cuts LDRAM footprint by 32 %
    • OLI is 65 % faster than uniform interleave

24

Use sufficient LDRAM (128 GB)

  • When LDRAM is insufficient:
    • OLI performs 1.42× better than LDRAM preferred

Use insufficient LDRAM (64 GB)

Object‑Level Interleaving (OLI)

PASA Lab

25 of 27

  • Dynamic page migration and static page interleaving are not well-integrated.

25

Testbed: NoBalance, AutoNUMA, Tiering‑0.8, TPP

Memory Tiering via Page Migration

PASA Lab

26 of 27

  • Dynamic page migration and static page interleaving are not well-integrated.
  • Dynamic migration can degrade performance; OLI often wins without migration

26

Testbed: NoBalance, AutoNUMA, Tiering‑0.8, TPP

Memory Tiering via Page Migration

PASA Lab

27 of 27

  • Use CXL memory can save fast memory size (e.g., LDRAM) without largely losing performance

  • For AI training/inference workloads, wait for CXL 3.x peer-to-peer DMA or redesign offloading paths

  • Leverage semantic or object‑aware placement over naive page interleaving

27

Conclusion & Interesting Takeaway

PASA Lab