1 of 16

The 4th Data Prefetching Championship (DPC4)

Co-located with HPCA 2026

Organizers: Rahul Bera, Konstantinos Kanellopoulos, and Onur Mutlu

February 1, 2026

Sydney, Australia

2 of 16

Why Do We Care about Prefetching?

2

3 of 16

Why Do We Care about Prefetching?

3

Prefetching is important for providing performance

4 of 16

Why Another DPC?

  • It’s been a while ...

4

DPC1

HPCA 2009

DPC2

ISCA 2015

DPC4

HPCA 2026

ML Prefetching

ISCA 2021

DPC3

ISCA 2019

Once in ~4 years

~7 years

5 of 16

What Changed since Last DPC?

5

AI/ML inference,

datacenter-class workloads,

complex graph-based apps

Workloads

Ever increasing working dataset size

Data Footprint

32+ core datacenter-class processors with limited

per-core DRAM bandwidth

System Configuration

6 of 16

Goal of DPC4

6

Evaluate innovative ideas on prefetching

under a common evaluation framework

to tangibly advance the state-of-the-art

7 of 16

Format of DPC4

  • Simulator: ChampSim

  • Participants are free to implement prefetchers at L1/L2/LLC

  • Limited hardware budget
    • 32KB for L1 data cache prefetcher
    • 128KB for L2 prefetcher
    • 256KB for LLC prefetcher

  • No limitation in complexity

  • System-level feedback are available to enable dynamic adaptation
    • Prefetcher accuracy, memory bandwidth usage, rolling-window IPC, ...

7

8 of 16

New Things in DPC4

8

  • State-of-the-art AI/ML inference
    • LLama2 (LLM), ViT (image generation), Whisper (text-to-speech), ...
  • Real workloads from Google’s warehouse-scale computers
  • Graph mining workloads

New Workloads representative of our time

  • Every submission was evaluated on three system configurations:
    • Single-core with ample DRAM bandwidth (FullBW score)
    • Single-core with limited DRAM bandwidth (LimitBW score)
    • Four-core random mixes (MC score)
  • Dynamic system-level feedback, instead of static knobs (e.g., bandwidth usage)

Broader evaluation across diverse system configurations

9 of 16

DPC4 at a Glance

9

Overall

8

competing teams across the globe

Evaluated across

610

single-core workloads and

484

four-core workload mixes, totaling

3TB

of space

Over

50,000

CPU core-hour

spent on evaluation

9

Program committee members,

providing three reviews

to each submission

10 of 16

Did We Meet the Goal?

10

11 of 16

Sneak Peak: DPC4 Advances the State-of-the-Art

11

2.6% across 610 traces!

2.8% across 484 mixes!

Berti@L1 + Pythia@L2: Already +15.1% over no-prefetching

5.2% across 610 traces!

12 of 16

Sneak Peak: DPC4 Advances the State-of-the-Art

12

27.7%

4%

4.5%

1.1%

13 of 16

Sneak Peak: DPC4 Advances the State-of-the-Art

13

27.7%

4%

4.5%

1.1%

More opportunities for the future

14 of 16

Today’s Agenda (I)

  • 2 invited talks from industry veterans

  • 8 competitors’ talks

14

Dr. Leeor Peled, Huawei Research

Title: “Is Prefetcher Research

Still Alive?”

Dr. Akanksha Jain, Google

Title: “Data Prefetching:

A Data-Center Perspective”

15 of 16

Today’s Agenda (II)

  • Session 1 (13:45 – 15:35): 1 invited talk + 4 papers
    • Invited talk by Dr. Leeor Peled
    • Crossing the Boundary – Lee+, Yonsei University
    • SPPAM – Merrell+, Texas A&M
    • Emender -- Chen+, Tsinghua University
    • sBerti – Zhou+, HKUST, Guangzhou
  • Break
  • Session 2 (15:50 – 17:40): 1 invited talk + 4 papers
    • Invited talk by Dr. Akanksha Jain
    • Pushing the Limits of Berti – Singh+, University of Murcia and University of Zaragoza
    • Entangling Data Prefetcher – Torres+, University of Zaragoza , University of Murcia, and IIT Bombay
    • uMAMA – Block+, UIUC
    • Global Berti – Posluns+, University of Toronto

15

16 of 16

The 4th Data Prefetching Championship (DPC4)

Co-located with HPCA 2026

Organizers: Rahul Bera, Konstantinos Kanellopoulos, and Onur Mutlu

February 1, 2026

Sydney, Australia