Vision and Opportunity Analysis:�Modular chiplets for HPC
July 2024
Raghu Shankar
1
OCP ODSA HPC & AI Modularity
(1) Market/customer, (2) Competitive Analysis & (3) Tech Advantage
..
2
Modular High-Performance computing using Chiplets – Bapi Vinnakota & John Shalf, Nov 2023
3
Case for reviving Bespoke HPC systems – Reed, Gannon, and Dongarra�Source: “HPC Forecast: Cloudy and Uncertain”, Feb 2023
4
Bespoke HPC
+ Bespoke HPC?
2020’s -2030’s
Commodity HPC & Clouds
TPU = ASIC
Vision for Bespoke HPC systems – Reed, Gannon, and Dongarra
5
Next Steps: Engage authors for insights and directions – either offline or this forum
Bespoke HPC systems – Opportunity for modular chiplets?
6
Whiteboard
1960’s – 1990’s
2000’s
2010’s
2020’s
2030’s
FPGAs | ASICs
X86 – general purpose scale-out
Bespoke HPC
Cray, Thinking machines, Beowulf, Blue Gene, KNL, ..
“Vertically Integrated” single vendor
GPU – SIMD/vector scale-out
Cloud
On-prem
Cloud
On-prem
Analyzing Resource Utilization in an HPC System: A Case Study of NERSC’s Perlmutter�Sample from Nov 2022
7
| Median | Mean | Max | Std Dev | Observations & Questions |
CPU jobs | 21% of CPU jobs > 1 hour |
| |||
Allocated Nodes | 1 | 4.84 | 1,477 | 25.43 | |
Job Duration (hours) | 4.19 | 5.825 | 90.99 | 4.73 |
|
CPU Util % | 51.0 | 56.68 | 100.0 | 35.89 |
|
DRAM util % | 18.61 | 33.69 | 98.62 | 30.88 |
|
| | | |||
GPU jobs | 23% of GPU jobs > 1 hour | | |||
Allocated Nodes | 1 | 5.88 | 512 | 23.33 | |
Job Duration (hours) | 2.2 | 4.12 | 13.76 | 3.67 | |
Host CPU Util % | 4.0 | 18.00 | 100.0 | 24.81 |
|
Host DRAM util % | 18.04 | 28.24 | 98.29 | 20.94 |
|
GPU Util % | 100.0 | 83.73 | 100.0 | 30.45 | |
GPU HBM2 util % | 18.88 | 40.23 | 100.0 | 36.33 |
|
Apps (incl HPC kernels) with high Job nodes and high Duration on GPUs and CPUs?
Next Steps:
8
* excludes jobs < 1 hour
Node-hours by applications; Next-level detail by #Jobs, #Nodes, and #Hours�Node-hours is calculated by multiplying the total number of allocated nodes by the runtime (duration) of each job
9
~600 apps ~ 22%
Jobs x Nodes x Hours
~400 apps ~ 25%
HPC systems: Total Available Market $37B in 2023 – source Hyperion�Companies Need On-Premise HPC – And For More Than AI, Too (nextplatform.com)
Next steps:
$14.9B
HPC WW systems: Total On-prem by Vendor and vertical 2023
Next steps: Modular HPC opportunity and value proposition in key verticals:
Hyperion HPC Segmentation
11
Competitive Positioning & Total Value Prop – Qualitative
12
| General purpose | FPGAs | ASICs | Modular chiplets TARGET / GOALS | GPUs |
| X86, ARM, RISC-V | AMD Xilinx & Intel Altera | | |
| Vector support | Custom & reconfigurable vector | Modular mix of CPU & vector | Scaleout vector only |
Development cycle per gen | ~2 years | ~2 years | Derivatives in months | ~2 years |
Software support | Libraries | RTL, HDL, and HLS | Match GPU / FPGA SW | Cuda |
Implementation of specific kernel | SW changes in weeks? | Reconfigurable in hours | ~Weeks | ~ Software changes in Days/weeks TBD |
Performance | Medium – SVE instr | High for vector data | High for vector data | Very High for vector data |
Memory Bandwidth | .. | 800 Mhz – 1 GHz | > 2 GHz | 1 GHz – 1.7 GHz |
High Dataflow / low Memory latency | No (Instr fetch arch) | High | High | Medium |
Cost | | | | |
Power | High | Low | Low | High |
Scalability | | ESSPER – FPGA cluster | | NVlink |
DRAFT – WIP
Scientific Kernels in Computing & HPC �FPGAs
July 2024
Raghu Shankar
13
OCP ODSA HPC & AI Modularity
Role of FPGAs in Acceleration of HPC workloads (Objective: Lessons for reconfigurable chiplets)�Manuel de Castro, .. IEEE Computer Jul 2024
14
Next Steps: Engage authors for insights and directions – either offline or this forum
Toward FPGA-Based HPC: Advancing Interconnect Technologies, 1 of 3�Joshua Lant, .. IEEE Micro Hot Interconnects, Jan 2020
15
Toward FPGA-Based HPC: Advancing Interconnect Technologies, 2 of 3�Joshua Lant, .. IEEE Micro Hot Interconnects, Jan 2020
16
Toward FPGA-Based HPC: Advancing Interconnect Technologies, 3 of 3 �Joshua Lant, .. IEEE Micro Hot Interconnects, Jan 2020
17
Vector units
In modular chiplets
Control
Data
FPGA 2
FPGA 1
FPGA-based HPC accelerators: An evaluation on performance and energy efficiency�https://doi.org/10.1002/cpe.6570 - WILEY – NEED ACCESS
18
Scientific Kernels in Computing & HPC �Existing Solutions
July 2024
Raghu Shankar
19
OCP ODSA HPC & AI Modularity
A100 block diagrams
20
144 Streaming Multiprocessors (SMs)
Evaluating HPC Kernels for Processing in Memory; https://dl.acm.org/doi/abs/10.1145/3565053.3565054
21
Acceleration with long vector architectures: Implementation and evaluation of the FFT kernel on NEC SX-Aurora and RISC-V vector extension 1 of 2
22
Hardware Accelerator Integration Tradeoffs for HPC: Case Study of GEMM Acceleration in N-Body Methods – Need access
23
Acceleration with long vector architectures: Implementation and evaluation of the FFT kernel on NEC SX-Aurora and RISC-V vector extension 2 of 2
BENCHMARKS
24
Scientific Kernels in Computing & HPC �Lit Survey
June 2024
Raghu Shankar
25
OCP ODSA HPC & AI Modularity
References – Scientific kernels Lit Survey
Objective: Build and share understanding of kernels to initiate discussions in OCP forum
NOTE: This is not a comprehensive lit search/survey
26
Arithmetic Intensity Spectrum of select 8 scientific kernels
27
Source: Exploring and Analyzing the Real Impact of Modern On-Package Memory on HPC Scientific Kernels, Ang Li, et al., SC17, Nov 2017
Scientific kernels characteristics
28
Source: Exploring and Analyzing the Real Impact of Modern On-Package Memory on HPC Scientific Kernels, Ang Li, et al., SC17, Nov 2017
All the kernels use double-precision (DP). Threads are the optimal thread number running on Broadwell and KNL
Scientific Kernels – Details – Ordered by Arithmetic Intensity
29
General Matrix-Matrix Multiplication (GEMM) |
|
Cholesky Decomposition |
|
Stencil |
|
Fast Fourier Transform (FFT) |
|
Sparse Matrix-Vector Multiplication (SpMV) |
|
Sparse Matrix Transposition (SpTRANS) |
|
Sparse Triangular Solve (SpTRSV) |
|
Package Repositories built on kernels
30
Kernel | Package | Dataset | Output | Repository |
GEMM | Parallel Linear Algebra Software for Multicore Arch | Dense matrices | Dataset stats, elapsed exec time, GFLOPS throughput | |
Cholesky | PLASMA | Dense matrices | Same | Same |
SpMV | Benchmark SpMV using CSR5 | 968 matrices | Same | Github TBD |
SpTRANS | ScanTrans, MergeTrans | 968 matrices | Same | |
SpTRSV | SpMP: Sparse Matrix pre-processing library | 968 matrices | Same | |
Stencil | YASK (Yet Another Stencil Kernel) | 3D grid of given size | Same | |
FFT | FFTW v3.3.5 | 3D grid of given size | Same | |
Stream | Stream | Array of given size | |
Platform Configuration – On-package memory at 236 Gflops/sec
31
Note performance and bandwidth listed are the theoretic values calculated from the spec sheet; the actual number can be worse
Source: Exploring and Analyzing the Real Impact of Modern On-Package Memory on HPC Scientific Kernels, Ang Li, et al., SC17, Nov 2017
236 Gflops/s
Intel Haswell and Broadwell processor with eDRAM
eDRAM Analysis for reference: Performance Gaps and Speedup ~2X
32
Throughput Perf gap (Gflops/sec) limits Speedup
References – Scientific kernels Lit Survey
Objective: Build and share understanding of kernels to initiate discussions in OCP forum
NOTE: This is not a comprehensive lit search/survey
33
GenArchBench: genomics benchmark suite – Summary
34
Workflow diagram of common genome analysis pipelines: �(1) Sequencing, (2) Basecalling, (3a) Genome resequencing or (3b) and genome assembly.�Figure shows different computational kernels used within each stage or tool
35
A
G
GenArchBench: 13 multi-threaded CPU kernels genomic tools: 1 of 2
36
2. Basescaling | Basescaling | Adaptive Banded Signal to Event Alignment (ABEA) | Redesigned version of the Suzuki-Kasahara dynamic programming algorithm used to compare raw nanopore signals, produced by ONT sequencing machines to a reference genome sequence. |
Neural Network-based Base Calling (NN-BASE) | ONT sequencing machines monitor changes in an electrical current as single strands of DNA or RNA pass through a protein nanopore. These changes in the electrical current are then converted to a sequence of nucleotide bases in the basecalling process | ||
3a. Genome Resequencing | Seed | FM-Index Search (FMI) | compressed sub-string index based on the Burrows-Wheler transform. Given sub-string “s”, FM-index finds location of “s” in reference genome in O(|s|) time, length of sub-string |
Chain | Seed Chaining (CHAIN) | Given set of subsequences (seeds) from a DNA sequence (read), chaining consists of linking overlapping seeds to form larger ones (uses heuristics) | |
SIMD Seed Chaining (FAST-CHAIN) | An x86-vectorized version of CHAIN that removes the heuristics to exploit SIMD computation | ||
Extend | Bit-Parallel Myers (BPM) | dynamic programming algorithm that finds all locations a query string of size m matches a reference string of size n with k or fewer differences | |
Banded Smith-Waterman (BSW) | dynamic programming algorithm that computes local sequence alignment of two sequences of length m and n, respectively, in O(mn) time and space | ||
Wavefront Alignment (WFA) | pairwise alignment algorithm that takes advantage of homologous regions between the sequences to accelerate the alignment process. .. traditional dynamic programming algorithms run in quadratic time, WFA time complexity is O(ns), proportional to read length n and the alignment score s, using O(s2) memory. |
GenArchBench: 13 multi-threaded CPU kernels genomic tools: 2 of 2
37
Reseq uencing | Variant Calling | Neural Network-based Variant Calling (NN-VARIANT) | process of detecting the differences (variants or mutations) between the aligned reads and the reference genome. This is a costly process .. |
Pileup Counting (PILEUP) | Starting with alignment data of a set of aligned reads to a region of a reference genome, usually a SAM or BAM file, pileup counting is the process of summarizing the 13 base-pair information at each chromosomal position | ||
3b. Genome Assembly | | K-mer Counting (KMER-CNT) | aims to count the number of occurrences of each k-mer in an input Sequence |
| De-Bruijn Graph Construction (DBG) | .. of an input set of reads is used to represent the overlaps between the sub-strings of length k (k-mers) found in the input | |
Multiple Seq Alignment (MSA) | Partial-Order Alignment (POA) | construction of an overlap graph from a set of reads leads to an approximate representation of the original sample’s genome To determine consensus genome of the sample, alignment of all the reads against each other is performed in process called multiple sequence alignment |
Speedups: 1 of 2
38
Speedups: 2 of 2
39
References – Scientific kernels Lit Survey
Objective: Build and share understanding of kernels to initiate discussions in OCP forum
NOTE: This is not a comprehensive lit search/survey
40
Performance characterization of 64-core SG2042 RISC-V CPU for HPC
41
NAS Parallel Benchmark (NPB): ~5 Kernels & ~3 Pseudo Applications
42
Kernel | Brief Description |
Integer Sort | tests indirect, random, memory accesses which it can be seen stalls a significant fraction of the CPU due to cache accesses |
Multi Grid | heavily memory bound both in terms of time stalled on cache and main memory accesses, and also the percentage of execution time where DDR is under high utilization |
Embarassingly Parallel | designed to test compute performance and there are far fewer cycles stalled on memory access, and no time spent with high DDR bandwidth utilization |
Conjugate Gradient | irregular memory access and nearest neighbour communication, which results in around 37% of clock ticks stalled on cache or DDR accesses |
Fast Fourier Transform | requires all-to-all communications between ranks to undertake a parallel transposition of data |
Pseudo Applications: combine multiple kernels to provide more complicated workloads – compute finite difference solution to the 3D compressible Navier Stokes equations | |
Block Tridiagonal | Based on a Beam-Warming approximation, Gaussian elim., resulting equations are block-tridiagonal, stalls the least on memory access |
LU Gauss Seidel | LU benchmark solves via a block-lower block-upper triangular approximation based upon Gauss Seidel iterative method |
Scalar Pentadiagonal | Beam-Warming approximation, Gaussian elim., but resulting equations are fully diagonalized |
Performance Characterization – 2 kernels
43
Integer Sort
Multi Grid
AMD EPYC has 8 memory controllers and 8 memory channels, connected to DDR4-3200 memory
Skylake performs the best has the largest L2 cache, 1MB per core
Scientific apps: Molecular dynamics
44
| |
AMBER |
|
GROMACS |
|
LAMMPS |
|
NAMD | Nanoscale Molecular Dynamics program, is a parallel molecular dynamics code high-performance simulation of large biomolecular systems |
VASP | Vienna Ab initio Simulation Package (VASP) – atomic scale materials modeling, electronic structure calculations and quantum-mechanical molecular dynamics from first principles |
CHARMM | |
X-PLOR | |
Scientific applications: Engineering & Physics
45
| |
Abaqus |
|
Alphafold |
|
ANSYS |
|
Gaussian |
|
Chroma | Physics |
BerkeleyGW | Physics |
FUN3D | Engineering |
RTM | Geo |
SpecFEM3D | Geo |
Parking Section
46
The landscape of parallel computing research: A view from Berkeley,”
47
A Survey of Big Data, High Performance Computing, and Machine Learning Benchmarks�https://link.springer.com/chapter/10.1007/978-3-030-94437-7_7
48