1 of 48

Vision and Opportunity Analysis:�Modular chiplets for HPC

July 2024

Raghu Shankar

1

OCP ODSA HPC & AI Modularity

2 of 48

(1) Market/customer, (2) Competitive Analysis & (3) Tech Advantage

  • Modular High-Performance computing using Chiplets – Bapi Vinnakota & John Shalf, Nov 2023
  • Case for reviving Bespoke HPC systems – Reed, Gannon, and Dongarra�Source: “HPC Forecast: Cloudy and Uncertain”, Feb 2023
  • Analyzing Resource Utilization in an HPC System: A Case Study of NERSC’s Perlmutter
  • HPC systems: Total Available Market $37B per year global in 2023 – source Hyperion
  • CPU, GPU, FPGA and Modular Arch positioning – Template Draft

..

  • Learning from FPGAs HPC kernels:
  • “Role of FPGAs in Acceleration of HPC workloads,” Manuel de Castro, .. IEEE Computer Jul 2024
  • “Toward FPGA-Based HPC: Advancing Interconnect Technologies,” Joshua Lant, .. IEEE HotI Jan 2020

2

3 of 48

Modular High-Performance computing using Chiplets – Bapi Vinnakota & John Shalf, Nov 2023

  • Problem statement: Perf improvement of Top500 HPC systems dropped precipitously in last decade
  • Hypothesis: use chiplets to extend the functional and physical modularity of modern HPC systems to within the semiconductor package
  • Modular chiplet based designs promise economic development of many variants
  • Derivatives can be developed at lower cost vs monolithic designs
    1. Creating new component chiplet + new package design + reusing rest of chiplets
    2. Changing the composition of different type of chiplets – cost of new package design
  • Challenges:
    • Chiplets require more complex packaging vs monolithic designs
    • Multiple dies in one package make power delivery, cooling, mechanical, and signal analysis more complex
    • Manufacturing alignment
    • Packaging tech alignment
    • Inventory management
    • Software co-design to take advantage of modularity

3

4 of 48

Case for reviving Bespoke HPC systems – Reed, Gannon, and Dongarra�Source: “HPC Forecast: Cloudy and Uncertain”, Feb 2023

4

Bespoke HPC

+ Bespoke HPC?

2020’s -2030’s

Commodity HPC & Clouds

TPU = ASIC

5 of 48

Vision for Bespoke HPC systems – Reed, Gannon, and Dongarra

  • Bespoke systems designed and built collaboratively to support key scientific and engineering workload needs
  • The space of leading-edge HPC applications is far broader now than in the past (new domains and new approaches) – Maxim #4
  • Future advances require embracing end-to-end design, testing, evaluating advanced prototypes (Maxim #2)
  • Partnering with traditional chip and HPC vendors, cloud ecosystem vendors, academia, and government laboratories
  • Combination of commercially designed and constructed technology, but each also contains large numbers of custom elements for which there is no sustainable business model

5

Next Steps: Engage authors for insights and directions – either offline or this forum

6 of 48

Bespoke HPC systems – Opportunity for modular chiplets?

6

Whiteboard

1960’s – 1990’s

2000’s

2010’s

2020’s

2030’s

FPGAs | ASICs

X86 – general purpose scale-out

Bespoke HPC

Cray, Thinking machines, Beowulf, Blue Gene, KNL, ..

“Vertically Integrated” single vendor

  • Resurgence of bespoke HPC with modular chiplets
  • Multiple vendor eco-system

GPU – SIMD/vector scale-out

Cloud

On-prem

Cloud

On-prem

7 of 48

Analyzing Resource Utilization in an HPC System: A Case Study of NERSC’s Perlmutter�Sample from Nov 2022

7

Median

Mean

Max

Std Dev

Observations & Questions

CPU jobs

21% of CPU jobs > 1 hour

  • Identify the Applications of jobs with highest “pain point”?

Allocated Nodes

1

4.84

1,477

25.43

Job Duration (hours)

4.19

5.825

90.99

4.73

  • List of appl/jobs at > = 6 hours? What’re the jobs above standard policy of 12-hour limit?
  • What is the benefit of speeding up by multiple of X factor?

CPU Util %

51.0

56.68

100.0

35.89

  • What’re the appl/jobs at MAX CPU util?

DRAM util %

18.61

33.69

98.62

30.88

  • What’re the appl/jobs skewing the DRAM util to the right? What jobs are utilizing the max DRAM?

GPU jobs

23% of GPU jobs > 1 hour

Allocated Nodes

1

5.88

512

23.33

Job Duration (hours)

2.2

4.12

13.76

3.67

Host CPU Util %

4.0

18.00

100.0

24.81

  • Appl/jobs candidates for modular chiplet design?
    • Low CPU util % & high GPU util%
    • High CPU and high GPU util %?

Host DRAM util %

18.04

28.24

98.29

20.94

  • Appl/Jobs with high DRAM and high GPU util %?

GPU Util %

100.0

83.73

100.0

30.45

GPU HBM2 util %

18.88

40.23

100.0

36.33

  • Appl/Jobs with high HBM util %?

8 of 48

Apps (incl HPC kernels) with high Job nodes and high Duration on GPUs and CPUs?

Next Steps:

  • List of applications (kernels) matching the jobs > 4 nodes? Apprx 21% of CPU jobs & 11% of GPU jobs
  • List of applications (kernels) matching the jobs > 6 hours? Apprx 40% of CPU jobs & 23% of GPU jobs

8

* excludes jobs < 1 hour

9 of 48

Node-hours by applications; Next-level detail by #Jobs, #Nodes, and #Hours�Node-hours is calculated by multiplying the total number of allocated nodes by the runtime (duration) of each job

  • Top four CPU-only applications account for 50% of node hours, with ATLAS alone accounting for over a quarter.
  • Over 600 CPU applications make up only 22% of the node hours, using less than 2% each (not labeled on the pie chart).
  • On GPU-accelerated nodes, the top 11 applications consume 75% of node hours, while the other 400+ applications make up the remaining 25%.
  • The top six GPU applications account for 58% of node hours, with usage roughly evenly divided
  • 37% of high memory intensity CPU job consume 54% of total node hours (almost same for GPU intense jobs)

9

~600 apps ~ 22%

Jobs x Nodes x Hours

~400 apps ~ 25%

10 of 48

HPC systems: Total Available Market $37B in 2023 – source Hyperion�Companies Need On-Premise HPC – And For More Than AI, Too (nextplatform.com)

  • Changing HPC space in 2023 with prominent on-premises business but that also fast growth in cloud and in AI
  • Global HPC market was $37.2 billion last year, with on-prem servers accounting for 40%
  • Storage: 17%
  • Cloud: 20%
  • Areas driving growth, particularly in predictive and generative AI and large language models (LLMs).
  • Cloud computing an option for growing number of HPC workloads
  • New supercomputers being designed with AI model training as a priority workload among traditional HPC codes

Next steps:

  • Semi market size – processors, GPUs, memory, FPGAs – XX% of systems
  • Modular semi & chiplet opportunity YY% of Semi

$14.9B

11 of 48

HPC WW systems: Total On-prem by Vendor and vertical 2023

Next steps: Modular HPC opportunity and value proposition in key verticals:

  • Govt Labs: $3.2B
  • Univ/Academia: $2.5B
  • Defense: $1.5B
  • Bio Sciences: $1.4B

Hyperion HPC Segmentation

  • Supercomputers: > $500k
  • Divisional servers: $250k - $500k
  • Departmental servers: $100k - $250k
  • Workgroup servers: <$100K

11

12 of 48

Competitive Positioning & Total Value Prop – Qualitative

12

General purpose

FPGAs | ASICs

Modular chiplets

TARGET / GOALS

GPUs

X86, ARM, RISC-V

AMD Xilinx & Intel Altera

Vector support

Custom & reconfigurable vector

Modular mix of CPU & vector

Scaleout vector only

Development cycle per gen

~2 years

~2 years

Derivatives in months

~2 years

Software support

Libraries

RTL, HDL, and HLS

Match GPU / FPGA SW

Cuda

Implementation of specific kernel

SW changes in weeks?

Reconfigurable in hours

~Weeks

~ Software changes in Days/weeks TBD

Performance

Medium – SVE instr

High for vector data

High for vector data

Very High for vector data

Memory Bandwidth

..

800 Mhz – 1 GHz

> 2 GHz

1 GHz – 1.7 GHz

High Dataflow / low Memory latency

No (Instr fetch arch)

High

High

Medium

Cost

Power

High

Low

Low

High

Scalability

ESSPER – FPGA cluster

NVlink

DRAFT – WIP

13 of 48

Scientific Kernels in Computing & HPC �FPGAs

July 2024

Raghu Shankar

13

OCP ODSA HPC & AI Modularity

14 of 48

Role of FPGAs in Acceleration of HPC workloads (Objective: Lessons for reconfigurable chiplets)�Manuel de Castro, .. IEEE Computer Jul 2024

  • FPGAs adoption low in HPC; contribution to accelerating HPC workloads is unclear in both potential and extent; only TOP500 is Noctua2
  • Xilinx (AMD) Alveo & Versal; Altera (Intel) Stratix10 and Agilex
  • Advantages:
    • Reconfigurable (but cumbersome): highly customized for multiple apps that need to be accelerated
    • Highly parallelizable & power efficient (skipping general purpose processor overheads)
    • Beneficial in scientific computing where latency and predictability of exec times are crucial
  • Disadvantages:
    • Lower clock speeds (drops to ~30% for HPC kernels) & low off-chip memory bandwidth
    • Lower FP perf
    • Multi-FPGA weak due to lower scaling
    • Lack of software support like Cuda, OpenCL, SYCL, Data Parallel C++;
    • Requires working at RTL and HDL – less user friendly, error prone, non-familiarity for HPC community
    • Programming HPC kernels with HDL leads to very high development costs and time, several hours to compile
    • HLS tools bridging above gap

14

Next Steps: Engage authors for insights and directions – either offline or this forum

15 of 48

Toward FPGA-Based HPC: Advancing Interconnect Technologies, 1 of 3�Joshua Lant, .. IEEE Micro Hot Interconnects, Jan 2020

  • HPC architects are currently facing myriad challenges from ever tighter power constraints and changing workload characteristics
  • Discuss current state of FPGAs within HPC systems. Recent technological advances show that they are well placed for penetration into the HPC market.
  • However, there are still a number of research problems to overcome; address the requirements for system architectures and interconnects to enable their proper exploitation, highlighting the necessity of allowing FPGAs to act as full-fledged peers within a distributed system rather than attached to the CPU
  • Half HPC apps limited by memory latency instead of memory throughput – offers opportunity for custom memory layouts (.. Validate this further.. )
  • This model requires a reliable, connectionless, hardware-offloaded transport supporting a global memory space; results show how fully fledged hardware implementation gives latency improvements of up to 25% versus a software-based transport,
  • Drastic reduction of latency for small memory operations, Data transfer:
    • (1) RDMA transfers and (2) shared-memory operations
  • .. solution can outperform the state of the art in HPC workloads such as matrix–matrix multiplication achieving a 10% higher computing throughput

15

16 of 48

Toward FPGA-Based HPC: Advancing Interconnect Technologies, 2 of 3�Joshua Lant, .. IEEE Micro Hot Interconnects, Jan 2020

  • GPUs have catered to HPC market by providing increasingly powerful architectures, with higher FP perf and memory bandwidth. Unfortunately, these massive developments in performance have come at the cost of much higher power consumption
  • While the memory bandwidth of GPUs is incredibly high and is important for many HPC applications, there exist numerous applications that are more sensitive to memory latency, messaging rate, and overlap – these have been overlooked in GPUs ..

  • Main limitation with FPGAs is their floating-point capabilities. Typical number crunching algorithms:
    • Dense matrix–matrix/matrix–vector multiplication,
    • FFTs
    • N-body simulations
  • FP ops obtain best performance on GPUs
  • However, FPGAs offer clear advantage with much higher FLOPs/watt

  • Workloads exhibiting irregular memory access patterns combined with high levels of arithmetic computation already provide for more efficient FPGA-based implementations over CPU or GPU solutions
  • Workloads such as stencil codes are suited to the FPGA due to the high volume of on-chip memory, reducing DRAM accesses

16

17 of 48

Toward FPGA-Based HPC: Advancing Interconnect Technologies, 3 of 3 �Joshua Lant, .. IEEE Micro Hot Interconnects, Jan 2020

  • FPGAs operate independently within global shared-memory space
  • Network interface permits scaling-out FPGA resources, by allowing FPGA to read and write directly into the remote memory of other FPGAs or CPU resources
  • encapsulate the system-bus protocol for use over the network, so that any traffic arriving from the NI is dealt with as a local memory access, no matter the origin of the transaction

  • Flow of data and control when using distributed FPGA resources using different transport layers
  • Using our hardware offloaded transport layer
  • F1 completes work and issues RDMA directly to NI & writes shared memory ops to F2’s work buffer – inform new work to be performed
  • F2 completes work & notifies local CPU
  • Built for dataflow type processing

17

Vector units

In modular chiplets

Control

Data

FPGA 2

FPGA 1

18 of 48

FPGA-based HPC accelerators: An evaluation on performance and energy efficiency�https://doi.org/10.1002/cpe.6570 - WILEY – NEED ACCESS

  • Aug 2021, Tan Nguyen, Colin MacLean, Marco Siracusa, Douglas Doerfler, Nicholas J. Wright, ..
  • Hardware specialization is a promising direction for the future of digital computing. Reconfigurable technologies enable hardware specialization with modest non-recurring engineering cost, but their performance and energy efficiency compared to state-of-the-art processor architectures remain an open question.
  • we use FPGAs to evaluate the benefits of building specialized hardware for numerical kernels found in scientific applications
  • .. evaluate performance, we compare Intel Arria 10 and Xilinx U280 performance against Intel Xeon, Intel Xeon Phi, and NVIDIA V100 GPUs, and extend the Empirical Roofline Toolkit (ERT) to FPGAs in order to assess results in terms of Roofline model
  • show design optimization and tuning techniques for peak FPGA performance at reasonable hardware usage and power consumption.
  • As FPGA peak performance is known to be far less than that of a GPU, we also benchmark the energy efficiency of each platform for the scientific kernels comparing against microbenchmark and technological limits.
  • Results show that while FPGAs struggle to compete in absolute terms with GPUs on memory- and compute-intensive kernels, they require far less power and can deliver nearly the same energy efficiency.

18

19 of 48

Scientific Kernels in Computing & HPC �Existing Solutions

July 2024

Raghu Shankar

19

OCP ODSA HPC & AI Modularity

20 of 48

A100 block diagrams

20

144 Streaming Multiprocessors (SMs)

21 of 48

Evaluating HPC Kernels for Processing in Memory; https://dl.acm.org/doi/abs/10.1145/3565053.3565054

  • Kazi Asifuzzaman, Mohammad Alaul Haque Monil, Frank Liu, and Jeffrey S Vetter
  • MEMSYS '22: Proceedings of the 2022 International Symposium on Memory Systems, October 2022
  • https://doi.org/10.1145/3565053.3565054
  • Memory subsystems contribute significantly to the performance and energy efficiency of high-performance computing (HPC) applications
  • Traditional memory technologies with conventional organization (e.g., DRAM) are struggling to keep up with the increasing memory requirements of modern applications
  • Techniques such as multilayer cache hierarchy and out-of-order execution are still falling short of mitigating the penalty incurred by memory accesses
  • Processing-in-memory (PIM) involves moving memory-intensive kernels to memory for execution instead of bringing the data to the processing unit is emerging as a promising technique
  • PIM has recently received traction among computer architecture researchers and the increasing research activity surrounding this technique indicates its potential to alleviate main memory performance bottlenecks
  • we characterize and identify memory-intensive HPC kernels, perform a first-order evaluation of the PIM technique for selected HPC kernels, quantify performance deviation, and analyze the key factors that affect PIM efficiency

21

22 of 48

Acceleration with long vector architectures: Implementation and evaluation of the FFT kernel on NEC SX-Aurora and RISC-V vector extension 1 of 2

  • Aug 2022, Pablo Vizcaino, Barcelona Supercomputing, et. al.
  • Novel architectures leveraging long and variable vector lengths like the NEC SX-Aurora or the vector extension of RISCV are appearing as promising solutions on the supercomputing market.
  • These architectures often require re-coding of scientific kernels; traditional implementations of algorithms for computing fast Fourier transform (FFT) cannot take full advantage of vector architectures.
  • present the implementation of FFT algorithms able to leverage these novel architectures:
    • Evaluate these codes on NEC SX-Aurora , comparing them with the optimized NEC libraries; and
    • In a prototype of a RISC-V core with a vector processing unit
  • present the benefits and limitations of 2 approaches of RADIX-2 FFT vector implementations
    • Pease FFT
    • Stockham FFT
  • our approach makes better use of the vector unit of the NECSX-Aurora , reaching higher or equal performance than the optimized NEC library
  • prove the importance of maximizing the vector length usage of the algorithm, taking advantage of the FFT properties to reduce long-latency vector operations, and reordering the instructions according to the specific hardware features to boost the performance of FFT-like computational kernels

22

23 of 48

Hardware Accelerator Integration Tradeoffs for HPC: Case Study of GEMM Acceleration in N-Body Methods – Need access

  • IEEE trans on Parallel and Distributed systems, Aug 2021
  • Mochamad Asri; Dhairya Malhotra; Jiajun Wang; George Biros; Lizy K. John; Andreas Gerstlauer,
  • performance and energy saving benefits of hardware acceleration under different hardware configurations and usage scenarios for a state-of-the-art Fast Multipole Method (FMM), which is a popular N-body method.
  • Use dedicated ASIC to accelerate General Matrix-Matrix Multiply (GEMM) operations.
  • FMM is widely used in applications and is representative example of the workload for many HPC applications. We compare architectures that integrate the GEMM ASIC next to, in or near main memory with an on-chip coupling aimed at minimizing or avoiding repeated round-trip transfers through DRAM for communication between accelerator and CPU.
  • study tradeoffs using detailed and accurately calibrated x86 CPU, accelerator and DRAM simulations.
  • results show that simply moving accelerators closer to the chip does not necessarily lead to performance/energy gains.
  • demonstrate that, while careful software blocking and on-chip placement optimizations can reduce DRAM accesses by 2X over a naive on-chip integration, these dramatic savings in DRAM traffic do not automatically translate into significant total energy or runtime savings.
  • .. chiefly due to app characteristics, the high idle power and effective hiding of memory latencies in modern systems.
  • more aggressive co-optimizations such as software pipelining and overlapping are applied, additional performance and energy savings can be unlocked by 37% and 35% over baseline acceleration.
  • When similar optimizations (pipelining and overlapping) are applied with an off-chip integration, on-chip integration delivers up to 20% better performance and 17% less total energy consumption than off-chip integration

23

24 of 48

Acceleration with long vector architectures: Implementation and evaluation of the FFT kernel on NEC SX-Aurora and RISC-V vector extension 2 of 2

BENCHMARKS

24

25 of 48

Scientific Kernels in Computing & HPC �Lit Survey

June 2024

Raghu Shankar

25

OCP ODSA HPC & AI Modularity

26 of 48

References – Scientific kernels Lit Survey

Objective: Build and share understanding of kernels to initiate discussions in OCP forum

  1. “Exploring and Analyzing the Real Impact of Modern On-Package Memory on HPC Scientific Kernels,” Ang Li, et al., SC17, Nov 2017

  • “GenArchBench: A genomics benchmark suite for arm HPC processors, Future Generation Computer Systems,” Lorien Lopez-Villellas, et. al., Mar 2024

  • “Performance characterization of the 64-core SG2042 RISC-V CPU for HPC,” Nick Brown, et.al. Jun 2024

NOTE: This is not a comprehensive lit search/survey

26

27 of 48

Arithmetic Intensity Spectrum of select 8 scientific kernels

27

Source: Exploring and Analyzing the Real Impact of Modern On-Package Memory on HPC Scientific Kernels, Ang Li, et al., SC17, Nov 2017

28 of 48

Scientific kernels characteristics

28

Source: Exploring and Analyzing the Real Impact of Modern On-Package Memory on HPC Scientific Kernels, Ang Li, et al., SC17, Nov 2017

All the kernels use double-precision (DP). Threads are the optimal thread number running on Broadwell and KNL

29 of 48

Scientific Kernels – Details – Ordered by Arithmetic Intensity

29

General Matrix-Matrix Multiplication (GEMM)

  • One of the most widely used computation kernels in scientific computing. It calculates the product of two dense matrices: C = αA∗B+βC
  • Generally compute-bound, algorithm complexity is O(n3)
  • One implementation in PLASMA

Cholesky Decomposition

  • Decomposes a positive-definite Hermitian matrix A into a lower triangular matrix L, and its conjugate transpose L∗: A = L ∗ L∗
  • More efficient way to solve systems of linear equations than normal LU decomposition
  • Compute-bound and high arithmetic intensity

Stencil

  • Class of iterative kernels, which constitute the core of many scientific applications.
  • kernels sweep the same stencil computation on cells in a given multi-dimensional grid iteratively
  • Stencils could exhibit diverse arithmetic intensities depending on the specific stencil functions adopted
  • However, optimized stencil algorithms often see high arithmetic intensity under orchestrated spatial and temporal blocking techniques

Fast Fourier Transform (FFT)

  • Algorithm to compute the discrete Fourier transform (DFT) of a sequence, or the inverse
  • Rapidly converts signal from its original time or space domain into frequency domain or vice versa, by factorizing the DFT matrix into a product of sparse factors

Sparse Matrix-Vector Multiplication (SpMV)

  • Probably the most used and studied sparse linear algebra algorithm.
  • Multiplies a sparse matrix with a dense vector and returns a dense vector, low arithmetic intensity, normally bounded by memory bandwidth
  • Compressed Sparse Row (CSR) data structure-based implementation uses load balanced data partitioning and SIMD-friendly operations to achieve performance

Sparse Matrix Transposition (SpTRANS)

  • Transposes a sparse matrix A of order m×n to AT of order n×m, in which the CSR format is converted to the compressed sparse column (CSC) format, or vise versa.
  • Mainly rearranges nonzero entries of the input matrix, thus requiring little computation.

Sparse Triangular Solve (SpTRSV)

  • Computes dense solution vector x from a sparse linear system Lx = b, where L is square lower triangular sparse matrix and b is dense vector
  • Unlike other sparse BLAS routines, SpTRSV generally more difficult to parallelize as it is inherently sequential
  • Complexity of O(nnz) and very low arithmetic intensity as SpMV, but is often much slower than SpMV due to the dependencies among components of x

30 of 48

Package Repositories built on kernels

30

Kernel

Package

Dataset

Output

Repository

GEMM

Parallel Linear Algebra Software for Multicore Arch

Dense matrices

Dataset stats, elapsed exec time, GFLOPS throughput

Cholesky

PLASMA

Dense matrices

Same

Same

SpMV

Benchmark SpMV using CSR5

968 matrices

Same

Github TBD

SpTRANS

ScanTrans, MergeTrans

968 matrices

Same

SpTRSV

SpMP: Sparse Matrix pre-processing library

968 matrices

Same

Stencil

YASK (Yet Another Stencil Kernel)

3D grid of given size

Same

FFT

FFTW v3.3.5

3D grid of given size

Same

Stream

Stream

Array of given size

31 of 48

Platform Configuration – On-package memory at 236 Gflops/sec

31

  • “SP Perf.” and “DP Pref.” are the theoretical maximum single- and double-precision floating-point operation throughput, respectively
  • “C/” = memory capacity.
  • “B/” = memory bandwidth
  • “OPM” is on-package memory
  • “Cache” refers to the upper-level cache respect to OPM in the memory hierarchy.

Note performance and bandwidth listed are the theoretic values calculated from the spec sheet; the actual number can be worse

Source: Exploring and Analyzing the Real Impact of Modern On-Package Memory on HPC Scientific Kernels, Ang Li, et al., SC17, Nov 2017

236 Gflops/s

Intel Haswell and Broadwell processor with eDRAM

32 of 48

eDRAM Analysis for reference: Performance Gaps and Speedup ~2X

32

Throughput Perf gap (Gflops/sec) limits Speedup

33 of 48

References – Scientific kernels Lit Survey

Objective: Build and share understanding of kernels to initiate discussions in OCP forum

  1. “Exploring and Analyzing the Real Impact of Modern On-Package Memory on HPC Scientific Kernels,” Ang Li, et al., SC17, Nov 2017

  • “GenArchBench: A genomics benchmark suite for arm HPC processors, Future Generation Computer Systems,” Lorien Lopez-Villellas, et. al., Mar 2024

  • “Performance characterization of the 64-core SG2042 RISC-V CPU for HPC,” Nick Brown, et.al. Jun 2024

NOTE: This is not a comprehensive lit search/survey

33

34 of 48

GenArchBench: genomics benchmark suite – Summary

  • Fugaku A64FX processors, Graviton3, Intel Xeon Skylake Platinum, and AMD EPYC Rome
  • Most genomics applications tested and optimized for x86 systems; few are prepared to perform efficiently on Arm machines. Moreover, these applications do not exploit the newly introduced Scalable Vector Extensions (SVE)
  • computationally demanding kernels from the most widely used tools in genome data analysis and ported them to ARM
  • GenArch benchmark suite comprises 13 multi-core kernels from critical stages of widely-used genome analysis pipelines, incl
    • Base-calling, Read mapping, Variant calling, and Genome assembly.
  • Suite includes different input data sets per kernel (small and large), each with a corresponding regression test to verify the correctness of each execution automatically.
  • .. optimizations implemented in each kernel, performance & energy evaluations comparisons of 4 different HPC machines
  • evaluation shows that Graviton3 outperforms other machines on average,
  • performance of A64FX is significantly constrained by its small memory hierarchy and latencies
  • https://github.com/LorienLV/genarchbench

34

35 of 48

Workflow diagram of common genome analysis pipelines: �(1) Sequencing, (2) Basecalling, (3a) Genome resequencing or (3b) and genome assembly.�Figure shows different computational kernels used within each stage or tool

35

A

G

36 of 48

GenArchBench: 13 multi-threaded CPU kernels genomic tools: 1 of 2

36

2. Basescaling

Basescaling

Adaptive Banded Signal to Event Alignment (ABEA)

Redesigned version of the Suzuki-Kasahara dynamic programming algorithm used to compare raw nanopore signals, produced by ONT sequencing machines to a reference genome sequence.

Neural Network-based Base Calling (NN-BASE)

ONT sequencing machines monitor changes in an electrical current as single strands of DNA or RNA pass through a protein nanopore. These changes in the electrical current are then converted to a sequence of nucleotide bases in the basecalling process

3a. Genome Resequencing

Seed

FM-Index Search (FMI)

compressed sub-string index based on the Burrows-Wheler transform. Given sub-string “s”, FM-index finds location of “s” in reference genome in O(|s|) time, length of sub-string

Chain

Seed Chaining (CHAIN)

Given set of subsequences (seeds) from a DNA sequence (read), chaining consists of linking overlapping seeds to form larger ones (uses heuristics)

SIMD Seed Chaining (FAST-CHAIN)

An x86-vectorized version of CHAIN that removes the heuristics to exploit SIMD computation

Extend

Bit-Parallel Myers (BPM)

dynamic programming algorithm that finds all locations a query string of size m matches a reference string of size n with k or fewer differences

Banded Smith-Waterman (BSW)

dynamic programming algorithm that computes local sequence alignment of two sequences of length m and n, respectively, in O(mn) time and space

Wavefront Alignment (WFA)

pairwise alignment algorithm that takes advantage of homologous regions between the sequences to accelerate the alignment process.

.. traditional dynamic programming algorithms run in quadratic time, WFA time complexity is O(ns), proportional to read length n and the alignment score s, using O(s2) memory.

37 of 48

GenArchBench: 13 multi-threaded CPU kernels genomic tools: 2 of 2

37

Reseq uencing

Variant Calling

Neural Network-based Variant Calling (NN-VARIANT)

process of detecting the differences (variants or mutations) between the aligned reads and the reference genome. This is a costly process ..

Pileup Counting (PILEUP)

Starting with alignment data of a set of aligned reads to a region of a reference genome, usually a SAM or BAM file, pileup counting is the process of summarizing the 13 base-pair information at each chromosomal position

3b. Genome Assembly

K-mer Counting (KMER-CNT)

aims to count the number of occurrences of each k-mer in an input Sequence

De-Bruijn Graph Construction (DBG)

.. of an input set of reads is used to represent the overlaps between the sub-strings of length k (k-mers) found in the input

Multiple Seq Alignment (MSA)

Partial-Order Alignment (POA)

construction of an overlap graph from a set of reads leads to an approximate representation of the original sample’s genome

To determine consensus genome of the sample, alignment of all the reads against each other is performed in process called multiple sequence alignment

38 of 48

Speedups: 1 of 2

38

39 of 48

Speedups: 2 of 2

39

40 of 48

References – Scientific kernels Lit Survey

Objective: Build and share understanding of kernels to initiate discussions in OCP forum

  1. “Exploring and Analyzing the Real Impact of Modern On-Package Memory on HPC Scientific Kernels,” Ang Li, et al., SC17, Nov 2017

  • “GenArchBench: A genomics benchmark suite for arm HPC processors, Future Generation Computer Systems,” Lorien Lopez-Villellas, et. al., Mar 2024

  • “Performance characterization of the 64-core SG2042 RISC-V CPU for HPC,” Nick Brown, et.al. Jun 2024

NOTE: This is not a comprehensive lit search/survey

40

41 of 48

Performance characterization of 64-core SG2042 RISC-V CPU for HPC

  • .. RISC-V to gain significant traction HPC ...
  • Sophon’s SG2042 is the first mass produced, commodity available, high-core count RISC-V CPU designed for high performance workloads
  • NASA’s NAS Parallel Benchmark (NPB) suite to characterise performance of the SG2042 against other CPUs implementing RISC-V, x86-64, and AArch64 ISAs
  • NPB suite is a collection of benchmarks developed by NASA’s Advanced Supercomputing (NAS) division to characterize Computational Fluid Dynamics (CFD) applications
  • Two models of exec: (1) OpenMP uses shared memory and (2) MPI uses message passing

  • SG2042 consistently outperforms all other RISC-V solutions, delivering between 2.6 and 16.7 performance improvement at the single core level.
  • When compared against the x86-64 and AArch64 CPUs, SG2042 performs comparatively well with computationally bound algorithms but decreases in relative performance on memory bandwidth or latency bound algorithms
  • .. identify that performance of the SG2042’s memory subsystem is the greatest bottleneck.

41

42 of 48

NAS Parallel Benchmark (NPB): ~5 Kernels & ~3 Pseudo Applications

42

Kernel

Brief Description

Integer Sort

tests indirect, random, memory accesses which it can be seen stalls a significant fraction of the CPU due to cache accesses

Multi Grid

heavily memory bound both in terms of time stalled on cache and main memory accesses, and also the percentage of execution time where DDR is under high utilization

Embarassingly Parallel

designed to test compute performance and there are far fewer cycles stalled on memory access, and no time spent with high DDR bandwidth utilization

Conjugate Gradient

irregular memory access and nearest neighbour communication, which results in around 37% of clock ticks stalled on cache or DDR accesses

Fast Fourier Transform

requires all-to-all communications between ranks to undertake a parallel transposition of data

Pseudo Applications: combine multiple kernels to provide more complicated workloads – compute finite difference solution to the 3D compressible Navier Stokes equations

Block Tridiagonal

Based on a Beam-Warming approximation, Gaussian elim., resulting equations are block-tridiagonal, stalls the least on memory access

LU Gauss Seidel

LU benchmark solves via a block-lower block-upper triangular approximation based upon Gauss Seidel iterative method

Scalar Pentadiagonal

Beam-Warming approximation, Gaussian elim., but resulting equations are fully diagonalized

43 of 48

Performance Characterization – 2 kernels

43

Integer Sort

Multi Grid

AMD EPYC has 8 memory controllers and 8 memory channels, connected to DDR4-3200 memory

Skylake performs the best has the largest L2 cache, 1MB per core

44 of 48

Scientific apps: Molecular dynamics

44

AMBER

  • Molecular Dynamics

GROMACS

  • GROningen MAchine for Chemical Simulations molecular dynamics package simulates Newtonian equations of motion for systems with hundreds to millions of particles
  • primarily designed for biochemical molecules like proteins, lipids and nucleic acids that have a lot of complicated bonded interactions,
  • non-bonded interactions like polymers

LAMMPS

  • classical molecular dynamics code developed at Sandia National Labs.
  • LAMMPS (Large-scale Atomic/Molecular Massively Parallel Simulator) makes use of spatial-decomposition techniques to partition the simulation domain
  • parallel using MPI
  • capable of modeling systems with millions or even billions of particles on a large High Performance Computing machine.
  • variety of force fields and boundary conditions are provided in LAMMPS which can be used to model atomic, polymeric, biological, metallic, granular, and coarse-grained systems

NAMD

Nanoscale Molecular Dynamics program, is a parallel molecular dynamics code high-performance simulation of large biomolecular systems

VASP

Vienna Ab initio Simulation Package (VASP) – atomic scale materials modeling, electronic structure calculations and quantum-mechanical molecular dynamics from first principles

CHARMM

X-PLOR

45 of 48

Scientific applications: Engineering & Physics

45

Abaqus

  • from Dassault Systems used for finite element analysis (FEM) and computer-aided engineering (CAE) – domains petroleum engineering, biomedical engineering, and aerospace engineering

Alphafold

  • protein structure prediction tool developed by DeepMind;
  • uses a novel machine learning approach to predict 3D protein structures from primary sequences alone
  • AlphaFold depends on ~2.9 TB of databases and model parameters

ANSYS

  • comprehensive software suite that spans the entire range of physics, providing access to virtually any field of engineering simulation that a design process requires
  • used to simulate computer models of structures, electronics, or machine components for analyzing strength, toughness, elasticity, temperature distribution, electromagnetism, fluid flow, …

Gaussian

  • quantum mechanics package for calculating molecular properties from first principles
  • predicts the energies, molecular structures, vibrational frequencies and molecular properties of compounds and reactions in a wide variety of chemical environments

Chroma

Physics

BerkeleyGW

Physics

FUN3D

Engineering

RTM

Geo

SpecFEM3D

Geo

46 of 48

Parking Section

46

47 of 48

The landscape of parallel computing research: A view from Berkeley,”

  • K. Asanovic et al., “Dept. Elect. Eng. Comput. Sci., Univ. California, Berkeley, Berkeley, CA, USA, Tech. Rep. UCB/EECS-2006-183, 2006
  • Next Step Action: Is there more recent taxonomy?

47

48 of 48

A Survey of Big Data, High Performance Computing, and Machine Learning Benchmarks�https://link.springer.com/chapter/10.1007/978-3-030-94437-7_7

  • Springer conf paper, Jan 2022
  • Nina Ihde, Paula Marten, Ahmed Eleliemy, et. al.
  • convergence of Big Data (BD), High Performance Computing (HPC), and Machine Learning (ML) systems.
  • due to the increasing complexity of long data analysis pipelines on separated software stacks.
  • With the increasing complexity of data analytics pipelines comes a need to evaluate their systems, in order to make informed decisions about technology selection, sizing and scoping of hardware.
  • While there are many benchmarks for each of these domains, there is no convergence of these efforts.
  • First step, understand how the individual benchmark domains relate analyze some of the most expressive and recent benchmarks of BD, HPC, and ML systems
  • propose taxonomy of those systems based on individual dimensions such as accuracy metrics and common dimensions such as workload type.
  • aim at enabling the usage of our taxonomy in identifying adapted benchmarks for their BD, HPC, and ML system
  • identify challenges and research directions related to the future of converged BD, HPC, and ML system benchmarking

48