Modular chiplets for HPC & AI in OCP OCE��State of the Art and Research Plan
Jan 7th 2025
Raghu Shankar, LinkedIn
1
Contents
2
Use cases | ||
Apps | Cloud Apps | Benchmarks |
Portability Software | ||
Profiling, EDA & Simulation Tools | ||
Workflow scheduling | ||
IO | Memory | Compute |
“DEV OPS Stack”
Next Steps & Speakers – from July 23rd meeting, updated Aug 6, 2024 1 of 2
3
# | Next Step | Owner | Start Date | Status |
1 | Customer & market “pain point” input: Perlmutter: Applications (incl kernels) for jobs > 4 nodes & > 6 hours? Start with Slurm data | George | Aug | Jie Li and Popovici part 1 |
2 | FPGAs for HPC kernels: Engage experts / authors for insights, Prof Ferraris | Raghu | Sep | Prof Ferraris invited |
3 | ASICs with HPC kernels – state of the art – development, software, performance, power, cost, etc. | George Raghu | Sep | George – Lab work in progress Raghu – industry sources |
4 | Vision: Invited speakers Reed, Gannon, or Dongarra "HPC Forecast: Cloudy and uncertain” & “Modular HPC using Chiplets” | Raghu | Oct | Bapi completed Raghu will connect with listed speakers next |
5 | Modular Semi chiplet market opportunity, Semi market – processors, GPUs, memory, FPGAs | TBD | Q4 | Based on key apps/kernels addressable oppty |
Next Steps & Speakers – updated Dec 2024 2 of 2
4
# | Next Step | Owner | Start Date | Status |
6 | Boadi Shan (Stonybrook U) & Mauricio Araya-Polo, Total Energies EP Research & Tech, “Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures,” Aug 2024 | | | |
7 | Francesco Antici, Andrea Bartolini, Zeynep Kiziltan, etc. (Univ of Bologna, Italy), Yuetse Kodama (Riken Center of Comp Sci), “MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC Jobs,” SC24, Nov 2024 | | | |
8 | Filippo Barbari, Federico Ficarelli, Daniele Cesarini, “High-throughput drug discovery on Fujitsu A64FX architecture, Jan 2024” | | | |
9 | LBNL – Nan Ding, Brian Austin, Yang Liu, Neil Mehta, Steven Farrell, Johannes, “Workflow Roofline Model for End-to-End Workflow Performance Analysis” SC24 | | | |
10 | Anzt Hartwig (U of Tenn at Knoxville), Axel Huebl & Xiaoye S. Li (LBNL), “Then and Now: Improving Software Portability, Productivity, and 100x Performance, Jan 2024 | | | |
Profiling, Characterization, & Kernel
Raghu Shankar, LinkedIn
5
Analyzing Resource Utilization in an HPC System: A Case Study of NERSC’s Perlmutter�Sample from Nov 2022
6
| Median | Mean | Max | Std Dev | Observations & Questions |
CPU jobs | 21% of CPU jobs > 1 hour |
| |||
Allocated Nodes | 1 | 4.84 | 1,477 | 25.43 | |
Job Duration (hours) | 4.19 | 5.825 | 90.99 | 4.73 |
|
CPU Util % | 51.0 | 56.68 | 100.0 | 35.89 |
|
DRAM util % | 18.61 | 33.69 | 98.62 | 30.88 |
|
| | | |||
GPU jobs | 23% of GPU jobs > 1 hour | | |||
Allocated Nodes | 1 | 5.88 | 512 | 23.33 | |
Job Duration (hours) | 2.2 | 4.12 | 13.76 | 3.67 | |
Host CPU Util % | 4.0 | 18.00 | 100.0 | 24.81 |
|
Host DRAM util % | 18.04 | 28.24 | 98.29 | 20.94 |
|
GPU Util % | 100.0 | 83.73 | 100.0 | 30.45 | |
GPU HBM2 util % | 18.88 | 40.23 | 100.0 | 36.33 |
|
Apps (incl HPC kernels) with high Job nodes and high Duration on GPUs and CPUs?
Next Steps:
7
* excludes jobs < 1 hour
21K jobs
24K jobs
Node-hours by applications; Next-level detail by #Jobs, #Nodes, and #Hours�Node-hours is calculated by multiplying the total number of allocated nodes by the runtime (duration) of each job
8
~600 apps ~ 22%
Jobs x Nodes x Hours
~400 apps ~ 25%
Problem Identification, Data Request for NERSC’s Perlmutter (example shown)�Top ~5 applications matching CPU jobs and GPU jobs with highest #hours and #nodes
9
| Appl / Workflow | Kernel | # hours | # nodes | CPU Util% | DRAM util% | GPU util% | HBM util% |
CPU jobs | | | | | | | | |
Job #1 | Appl #1 | | 12 | 64 | | | | |
Job # | Atlas / HEP | FFT? | 11 | 128 | | | | |
Job # | VASP | | 10 | 32 | | | | |
Job # | LAMMPS | | 10 | 64 | | | | |
Job # | Appl #5 | | 8 | 16 | | | | |
GPU jobs | | | | | | | | |
Job #1 | Appl #1 | | 12 | 16 | | | | |
Job # | Espresso | | 10 | 32 | | | | |
Job # | Gyro | | 10 | 8 | | | | |
Job # | Appl #4 | | 8 | 16 | | | | |
Job # | Appl #5 | | 6 | 8 | | | | |
Request: Roofline analysis for top ~2 jobs or app, data set sizes, cache/HBM sizes, and cache/HBM hit rates
Evaluation of Programming Models and Performance for Stencil Computation on Current�GPU Architectures, Aug 2024 – 1 of 2
10
Evaluation of Programming Models and Performance for Stencil Computation on Current�GPU Architectures, Aug 2024 – 2 of 2
(MAX) Roofline plot of Ampere A100
11
(MAX) Roofline plot of Grace Hopper GH200
Summary:
Questions
1 Tflop @ 1 Tbyte
Per sec
Log Scale
3 Tflops @ 3 Tbytes
Per sec
A100
GH200
Ridge point = 18 flops/byte
Memory bound
Memory bound
AI = 18 Flops/Byte
51 Tflops/s of 2.8 Tbytes/s
MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC Jobs, SC24, Nov 2024, Fugaku characterization, 1 of 4
12
MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC Jobs, Fugaku Characterization, SC24, Nov 2024 2 of 4
13
Log Scale
0.0625
1024
Memory bound: ~1.6M, Compute bound: ~0.5M
Ridge: 3.3 Flops/byte
MCBound: An Online Framework to Characterize and Classify Memory/Compute-bound HPC Jobs, Fugaku Characterization, SC24, Nov 2024 3 of 4
Highlights
14
Frequency | Memory Bound | Compute Bound | Total |
2.0 GHz (Normal mode) | 891,056 | 330,878 | 1,221,934 |
2.2 GHz (Boost Mode) | 752,421 | 147,097 | 899,518 |
Total | 1,643,477 | 477,975 | 2,121,452 |
Roofline model of jobs data, categorized by frequency
1 TF/s
Potential for speedups & modularity for HPC jobs 4 of 4�Identify jobs at the roofline of memory and compute boundary
15
Incr Memory (Data) bandwidth for higher perf
Incr compute for higher perf in conjunction with memory perf
Optimal
Oper Intensity = 2-4
.. 1 Flop compute for 16 bytes data
Oper Intensity = 2-10
.. 1 Flop compute / 1024 bytes data
Oper Intensity = 23
.. 8 Flops compute for 1 byte data
Oper Intensity = 210
.. 1024 Flops compute / 1 byte data
FER: Benchmark for Empirical Roofline Analysis of FPGA Based HPC Accelerators�Enrico Calore & Sebastiano Fabio Schifano, IEEE Applied Research, Sep 2022
16
4x 16GB DRAM only
U250 card
8 GB HBM2 only
Multiple kernel instances on CUs
U50 card
32GB DRAM & 8 GB HBM
U280 Card
HBM & DRAM
Multiple CUs per HBM
U280 card
Empirical Roofline Analysis of kernels on 4 compute/memory profiles and 2 modes
1. U250 Card: 444 GF/s
2. U50 Card: 191 GF/S
3. U280: 308 GF/s
17
1. U250 card
U50
Double Precision
AI = 70
AI = 0.25
Source: Empirical Roofline Analysis of FPGA Based HPC Accelerators, Enrico Calore & Sebastiano Fabio Schifano, IEEE Applied Research, Sep 2022
Decision Points: (1) Competitive Comps & (2) FP precision for Apps & kernels
Empirical Roofline plots (DP-FP)
Different numerical precisions on Xilinx Alveo U250
Lower precision allows higher compute
18
HPCG
AI = 0.25
MMM
AI =70
Source: Empirical Roofline Analysis of FPGA Based HPC Accelerators, Encrico Calore
Alveo excels over Skylake and Thunder X2 for HPCG
No comp advantage for Matrix-Matrix Multiply
TF/s