1 of 18

Modular chiplets for HPC & AI in OCP OCE��State of the Art and Research Plan

Jan 7th 2025

Raghu Shankar, LinkedIn

1

2 of 18

Contents

  • Next steps and Speakers
  • Vision & Opportunity
  • Use Cases, Apps, Benchmarks
  • Profiling & Characterization
  • Chiplets, EDA, Software, and Sim tools
  • FPGAs – SOTA
  • GPUs & “Bespoke” – SOTA
  • Kernels – Deeper analysis
  • Parking Lot

2

Use cases

Apps

Cloud

Apps

Benchmarks

Portability Software

Profiling, EDA & Simulation Tools

Workflow scheduling

IO

Memory

Compute

“DEV OPS Stack”

3 of 18

Next Steps & Speakers – from July 23rd meeting, updated Aug 6, 2024 1 of 2

3

#

Next Step

Owner

Start Date

Status

1

Customer & market “pain point” input: Perlmutter: Applications (incl kernels) for jobs > 4 nodes & > 6 hours? Start with Slurm data

George

Aug

Jie Li and Popovici part 1

2

FPGAs for HPC kernels: Engage experts / authors for insights, Prof Ferraris

Raghu

Sep

Prof Ferraris invited

3

ASICs with HPC kernels – state of the art – development, software, performance, power, cost, etc.

George

Raghu

Sep

George – Lab work in progress

Raghu – industry sources

4

Vision: Invited speakers Reed, Gannon, or Dongarra "HPC Forecast: Cloudy and uncertain” & “Modular HPC using Chiplets”

Raghu

Oct

Bapi completed

Raghu will connect with listed speakers next

5

Modular Semi chiplet market opportunity, Semi market – processors, GPUs, memory, FPGAs

TBD

Q4

Based on key apps/kernels addressable oppty

4 of 18

Next Steps & Speakers – updated Dec 2024 2 of 2

4

#

Next Step

Owner

Start Date

Status

6

Boadi Shan (Stonybrook U) & Mauricio Araya-Polo, Total Energies EP Research & Tech, “Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures,” Aug 2024

​

​

​

7

Francesco Antici, Andrea Bartolini, Zeynep Kiziltan, etc. (Univ of Bologna, Italy), Yuetse Kodama (Riken Center of Comp Sci), “MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC Jobs,” SC24, Nov 2024

​

​

​

8

Filippo Barbari, Federico Ficarelli, Daniele Cesarini, “High-throughput drug discovery on Fujitsu A64FX architecture, Jan 2024”

​

​

​

9

LBNL – Nan Ding, Brian Austin, Yang Liu, Neil Mehta, Steven Farrell, Johannes, “Workflow Roofline Model for End-to-End Workflow Performance Analysis” SC24

​

​

​

10

Anzt Hartwig (U of Tenn at Knoxville), Axel Huebl & Xiaoye S. Li (LBNL), “Then and Now: Improving Software Portability, Productivity, and 100x Performance, Jan 2024

​

​

​

5 of 18

Profiling, Characterization, & Kernel

Raghu Shankar, LinkedIn

5

6 of 18

Analyzing Resource Utilization in an HPC System: A Case Study of NERSC’s Perlmutter�Sample from Nov 2022

6

​

Median

Mean

Max

Std Dev

Observations & Questions

CPU jobs

21% of CPU jobs > 1 hour

  • Identify the Applications of jobs with highest “pain point”?

Allocated Nodes

1

4.84

1,477

25.43

​

Job Duration (hours)

4.19

5.825

90.99

4.73

  • List of appl/jobs at > = 6 hours? What’re the jobs above standard policy of 12-hour limit?
  • What is the benefit of speeding up by multiple of X factor?

CPU Util %

51.0

56.68

100.0

35.89

  • What’re the appl/jobs at MAX CPU util? What is util% in #cores?

DRAM util %

18.61

33.69

98.62

30.88

  • What’re the appl/jobs skewing the DRAM util to the right? What jobs are utilizing the max DRAM? What is util% in DRAM size GB?

​

​

​

GPU jobs

23% of GPU jobs > 1 hour

​

Allocated Nodes

1

5.88

512

23.33

​

Job Duration (hours)

2.2

4.12

13.76

3.67

​

Host CPU Util %

4.0

18.00

100.0

24.81

  • Appl/jobs candidates for modular chiplet design? What is util% in #cores?
    • Low CPU util % & high GPU util%?
    • High CPU and high GPU util %?

Host DRAM util %

18.04

28.24

98.29

20.94

  • Appl/Jobs with high DRAM and high GPU util %? What is util% in DRAM size GB?

GPU Util %

100.0

83.73

100.0

30.45

​

GPU HBM2 util %

18.88

40.23

100.0

36.33

  • Appl/Jobs with high HBM util %? Util% in HBM size GB?

7 of 18

Apps (incl HPC kernels) with high Job nodes and high Duration on GPUs and CPUs?

Next Steps:

  • Separate – jobs, nodes, and hours. Rank order by #hours and # of nodes (“top 5” applications)
  • List of applications (kernels) matching the jobs > 4 nodes? Apprx 21% of CPU jobs & 11% of GPU jobs
    • CPU, Memory, IO utilization available stats
  • List of applications (kernels) matching the jobs > 6 hours? Apprx 40% of CPU jobs & 23% of GPU jobs

7

* excludes jobs < 1 hour

21K jobs

24K jobs

8 of 18

Node-hours by applications; Next-level detail by #Jobs, #Nodes, and #Hours�Node-hours is calculated by multiplying the total number of allocated nodes by the runtime (duration) of each job

  • Top four CPU-only applications account for 50% of node hours, with ATLAS alone accounting for over a quarter.
  • Over 600 CPU applications make up only 22% of the node hours, using less than 2% each (not labeled on the pie chart).
  • On GPU-accelerated nodes, the top 11 applications consume 75% of node hours, while the other 400+ applications make up the remaining 25%.
  • The top six GPU applications account for 58% of node hours, with usage roughly evenly divided
  • 37% of high memory intensity CPU job consume 54% of total node hours (almost same for GPU intense jobs)

8

~600 apps ~ 22%

Jobs x Nodes x Hours

~400 apps ~ 25%

9 of 18

Problem Identification, Data Request for NERSC’s Perlmutter (example shown)�Top ~5 applications matching CPU jobs and GPU jobs with highest #hours and #nodes

9

​

Appl /

Workflow

Kernel

# hours

# nodes

CPU Util%

DRAM util%

GPU util%

HBM util%

CPU jobs

​

​

​

​

​

​

​

​

Job #1

Appl #1

​

12

64

​

​

​

​

Job #

Atlas / HEP

FFT?

11

128

​

​

​

​

Job #

VASP

​

10

32

​

​

​

​

Job #

LAMMPS

​

10

64

​

​

​

​

Job #

Appl #5

​

8

16

​

​

​

​

GPU jobs

​

​

​

​

​

​

​

​

Job #1

Appl #1

​

12

16

​

​

​

​

Job #

Espresso

​

10

32

​

​

​

​

Job #

Gyro

​

10

8

​

​

​

​

Job #

Appl #4

​

8

16

​

​

​

​

Job #

Appl #5

​

6

8

​

​

​

​

Request: Roofline analysis for top ~2 jobs or app, data set sizes, cache/HBM sizes, and cache/HBM hit rates

10 of 18

Evaluation of Programming Models and Performance for Stencil Computation on Current�GPU Architectures, Aug 2024 – 1 of 2

  • Boadi Shan (Stonybrook U) & Mauricio Araya-Polo, Total Energies EP Research & Tech, Houston
  • .. present results and share insights about highly tuned stencil-based kernels for NVIDIA Ampere (A100) and Hopper (GH200) architectures
  • .. leveraged by many scientific applications which involve stencils computations
  • .. evaluation of 3 different programming models: CUDA, OpenACC, and OpenMP target offloading conducted on accelerators
  • .. study performance and portability of various kernels under each programming model and provide corresponding optimization recommendations.
  • .. compare performance of different programming models on the mentioned architectures.
  • Up to 58% performance improvement was achieved against the previous GPGPU’s architecture generation for highly optimized kernel of the same class, and up to 42% for all classes.
  • .. Portability constraint: optimized OpenACC implementation outperforms OpenMP implementation by 33%
  • If portability is not a factor, best tuned CUDA implementation outperforms the optimized OpenACC one by 2.1×

​

10

11 of 18

Evaluation of Programming Models and Performance for Stencil Computation on Current�GPU Architectures, Aug 2024 – 2 of 2

(MAX) Roofline plot of Ampere A100

  • L1: 42 TB/s
  • L2: 4 TB/s
  • HBM: 1.5 TB/s
  • Compute = 10.6 TF/s

11

(MAX) Roofline plot of Grace Hopper GH200

  • L1: 64 TB/s
  • L2: 8 TB/s
  • HBM: ~3 TB/s
  • Compute: 51.2 TF/s

Summary:

  • Stencil AI: ~1 Flop/Byte
  • Memory (DATA) Bound
  • 3 TB/s Memory BW -> 3 TFlops/s
  • Improving compute perf does not help

​

Questions

  • L1, L2, & HBM hit rates

​

​

1 Tflop @ 1 Tbyte

Per sec

Log Scale

3 Tflops @ 3 Tbytes

Per sec

A100

GH200

Ridge point = 18 flops/byte

Memory bound

Memory bound

AI = 18 Flops/Byte

51 Tflops/s of 2.8 Tbytes/s

12 of 18

MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC Jobs, SC24, Nov 2024, Fugaku characterization, 1 of 4

  • Francesco Antici, Andrea Bartolini, Zeynep Kiziltan, etc. (Univ of Bologna, Italy), Yuetse Kodama (Riken Center of Comp Sci)
  • HPC systems play fundamental role in driving scientific research, as they execute computationally intensive jobs originating from diverse domains
  • However, HPC jobs are characterized by conflicting computational requirements, which may cause inefficiencies in resource usage, system throughput and energy consumption. One approach to tackling this problem is to distinguish between memory-bound and compute-bound jobs at their submission time, with the goal of making informed decisions about their execution.
  • .. present MCBound, the first online data-driven framework to classify HPC jobs as memory/compute-bound before job execution, without user intervention.
  • propose systematic characterization technique to generate a reference dataset from historical data for initial classification model training.
  • Using characterization technique, analyze data of 2.2 million job runs on the Supercomputer Fugaku, production HPC system installed at RIKEN Center for Computational Science, in Japan
  • implement MCBound for Fugaku and classify the jobs executed during February 2024 – KNN and Random Forest
  • approach obtains F1-macro average score of at least 0.89 as prediction quality, while incurring a negligible overhead on the system’s operations
  • .. Python-based implementation of MCBound can be seamlessly configured and deployed in other HPC systems
  • Data publicly available at Zenodo: https://doi.org/10.5281/zenodo.11467483

12

13 of 18

MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC Jobs, Fugaku Characterization, SC24, Nov 2024 2 of 4

13

Log Scale

0.0625

1024

Memory bound: ~1.6M, Compute bound: ~0.5M

Ridge: 3.3 Flops/byte

14 of 18

MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC Jobs, Fugaku Characterization, SC24, Nov 2024 3 of 4

Highlights

  • Ridge point = ~3.3 Flops/Byte
  • Memory bound jobs ~3.5X of compute bound jobs
  • Proportion of memory bound jobs is constant over time
  • Many jobs are far from roofline -> some well-engineered jobs saturating resources fully, but is not the case for majority of the jobs
  • Compute bound jobs benefit from high frequencies but still below roofline
  • Memory bound jobs do not benefit from higher frequencies

14

Frequency

Memory Bound

Compute Bound

Total

2.0 GHz (Normal mode)

891,056

330,878

1,221,934

2.2 GHz (Boost Mode)

752,421

147,097

899,518

Total

1,643,477

477,975

2,121,452

Roofline model of jobs data, categorized by frequency

1 TF/s

15 of 18

Potential for speedups & modularity for HPC jobs 4 of 4�Identify jobs at the roofline of memory and compute boundary

15

Incr Memory (Data) bandwidth for higher perf

Incr compute for higher perf in conjunction with memory perf

Optimal

Oper Intensity = 2-4

.. 1 Flop compute for 16 bytes data

​

Oper Intensity = 2-10

.. 1 Flop compute / 1024 bytes data

​

​

Oper Intensity = 23

.. 8 Flops compute for 1 byte data

​

Oper Intensity = 210

.. 1024 Flops compute / 1 byte data

16 of 18

FER: Benchmark for Empirical Roofline Analysis of FPGA Based HPC Accelerators�Enrico Calore & Sebastiano Fabio Schifano, IEEE Applied Research, Sep 2022

  • Benchmarking tool to empirically measure computing performance of accelerators & bandwidth of on-chip and off-chip memories. Available: https://baltig.infn.it/EuroEXA/FER
  • .. draw Roofine plots for performance comparisons with other processors, CPUs and GPUs
  • .. estimate the performance upper-bounds that applications & kernels could achieve on target device
  • .. Dataflow mode and Data-local modes
  • .. describe theoretical model, its implementation details, results measured on Xilinx Alveo – 4 configs below

16

4x 16GB DRAM only

U250 card

8 GB HBM2 only

Multiple kernel instances on CUs

U50 card

32GB DRAM & 8 GB HBM

U280 Card

HBM & DRAM

Multiple CUs per HBM

U280 card

17 of 18

Empirical Roofline Analysis of kernels on 4 compute/memory profiles and 2 modes

  • URAM – Ultra RAM on-chip memory for Data-local mode
  • MMM = Matrix-Matrix Multiplication
  • HPCG = HPC Conjugate Gradients Benchmark, C/HLS version targeted for FPGAs – memory bound exploiting U280 HBM & URAM

​

  • Compute perf and Mem bandwidths

1. U250 Card: 444 GF/s

    • URAM: 4.22 TB/s
    • DRAM: 71 GBs

2. U50 Card: 191 GF/S

    • URAM: 2.27 TB/s
    • HBM: 257 GB/s

3. U280: 308 GF/s

    • URAM: 3.23 TB/s
    • HBM: 407 GB/s
    • DRAM: 36 GB/s

​

17

1. U250 card

U50

Double Precision

AI = 70

AI = 0.25

Source: Empirical Roofline Analysis of FPGA Based HPC Accelerators, Enrico Calore & Sebastiano Fabio Schifano, IEEE Applied Research, Sep 2022

18 of 18

Decision Points: (1) Competitive Comps & (2) FP precision for Apps & kernels

Empirical Roofline plots (DP-FP)

  1. Xilinx Alveo cards
  2. Intel Skylake CPU and
  3. Arm on Marvell ThunderX2 CPU

Different numerical precisions on Xilinx Alveo U250

Lower precision allows higher compute

18

HPCG

AI = 0.25

MMM

AI =70

Source: Empirical Roofline Analysis of FPGA Based HPC Accelerators, Encrico Calore

Alveo excels over Skylake and Thunder X2 for HPCG

No comp advantage for Matrix-Matrix Multiply

TF/s