1 of 19

Programming paradigms, parallel�computing concepts for GCMs, and parallel performance aspects�

Mark Taylor

Sandia National Labs

mataylo@sandia.gov

DCMIP 2025, Boulder, June 5, 2025

2 of 19

Atmosphere Model Resolution

2

100 km: typical global climate model resolution

E3SM v2 Atmosphere: 64 SYPD on 85 CPU nodes

25 km: high resolution climate models

Only a few full CMIP-style climate simulation campaigns completed at this resolution

3 km: cloud resolving

E3SM/SCREAM: 1 SYPD on 32,000 GPUs

3 of 19

Why Cloud Resolving

  •  

How do we parameterize this sub-grid variability?

Movie: Precipitation (colors) and integrated water vapor (gray) for an atmospheric river from E3SM’s DYAMOND2 SCREAM simulation. By Paul Ullrich/UC Davis

4 of 19

  • Ability to capture cloud structures is impressive
  • Example: cold air outbreaks, extra-tropical cyclones are well-represented
  • Fig: 2d into a SCREAM DYAMOND simulation (January 22, 2020 at 2:00:00 UTC). Himawari visible satellite image and shortwave cloud radiative effect from SCREAMv0.

5 of 19

Exascale Computing

6 of 19

Exascale = GPUs

9/10 in the Top500 list of the world’s fasters computers are GPU based

DOE Office of Science machines are nearly all GPU based:

  • NERSC Perlmutter (8 MW, #19)
    • 6000 NVIDIA GPUs
    • 3000 CPU-only nodes
  • OLCF Frontier (30 MW, #2)
    • 38,000 AMD MI250X GPUs
    • 1 CPU + 4 GPUs per node
  • ALCF Aurora (40 MW, #3)
    • 64,000 Intel Data Center GPU Max
    • 2 CPU + 6 GPUs per node

https://www.top500.org/

7 of 19

  • Expect GPU trend to continue, but now driven by AI/ML, not HPC
  • Latest GPU systems from AMD (MI300) and NVIDIA (Grace Blackwell) continue to be hybrid, combining improved GPUs and CPUs
  • Some risk as HPC market continues to shrink relative to AI/ML market – i.e. support for double precision

8 of 19

GPU Programming Models

  • C++ with a parallel array class (Kokkos)
    • Robust and well supported solution across most hardware
    • Requires minimal vendor support
  • Fortran with OpenACC or OpenMP offload
    • Relies heavily on (lagging) vendor compiler support
    • Remains immature w.r.t. advanced Fortran features
    • Good performance requires major code refactoring
  • Domain Specific Languages
    • Promising approach ( e.g. GT4Py, GridTools, Psyclone, Julia)
    • Today’s DSLs need additional investments to support algorithms & meshes in E3SM components and to run on AMD and Intel GPUs

E3SM & SCREAM approach

9 of 19

Other global atmosphere model GPU porting approaches

  • JAMSTEC: NICAM
    • Fortran code, running on Fugaku/a64fx CPU
  • MPI: ICON
    • Fortran/openACC (versions with GT4Py?)
  • NOAA/GFDL/Allen Institute – PACE
    • Port of GFDL FV3 model
    • GT4Py: Python based DSL
  • UKMO: Upcoming LFRic model
    • PSYCLONE (DSL)
  • CLIMA: Climate Modeling Alliance
    • Julia + GPU backend
  • NCAR:
    • Internal efforts (physics) + collaborations: Earthworks, STORMspeed

10 of 19

Successes of the C++/Kokkos approach

  • Simple Cloud Resolving E3SM Atmosphere Model (SCREAM)
    • New code written in C++/Kokkos obtained excellent performance portability,
    • First major atmosphere model to run well on CPUs, AMD, NVIDIA and Intel GPUs
    • Rewriting code from scratch lead to other benefits: modern design using C++ abstractions and improved testing including property tests
    • Awarded 2023 Gordon Bell Prize for Climate Modelling for obtaining 1.4 SYPD at cloud resolving resolution on an Exascale system
    • Access to large amounts of GPU cycles: several multi-year and decadal length atmosphere (3km, “ne1024”) simulations

11 of 19

The Downside of C++:

11

Original F90

Ported to C++/Kokkos

  • C++ and Kokkos make the code more complicated (see example)

  • Code complexity may drive E3SM towards a model where only computer scientists can change code – but we hope not!
  • Opportunity to remove decades of outdated code is a major benefit of switching languages
  • Most of the complexity is due to “packs” needed to get C++ code to vectorize when running on CPUs

12 of 19

SCREAM Timeline

  • V0: Fortran implementation
    • NH spectral element dycore: Taylor et al., JAMES 2020
    • Physics: SHOC, P3, RRTMGP, prescribed aerosols
    • DYAMOND results: Caldwell et al., JAMES 2021
  • V1: Rewrite in C++/Kokkos:
    • 2017-2019: Hydrostatic dycore & prototyping Kokkos approach (Bertagna et al., GMD 2019)
    • 2020: Nonhydrostatic dycore (Bertagna et al., SC 2020
    • 2021: Physics (cloud resolving only)
    • 2022: Atmospheric driver
    • 2023: Cloud resolving DYAMOND simulations running on Frontier (Taylor et al., 2023 Climate Gordon Bell submission)

13 of 19

Performance: C++ vs Fortran

14 of 19

C++/Kokkos: Performance Portability� �

SCREAM GRCM, but running at 1 deg resolution: 128L, NH dycore, 10 tracers, P3/SHOC physics with prescribed aerosols, no convective parameterization

  • C++ code is as fast or faster than Fortran code on CPUs
  • Performance portability
    • IBM P9, AMD EYPC
    • NVIDIA V100, A100
    • AMD MI250
  • GPU performance:
    • Large scaling range where GPU nodes are 4-10x faster than CPU nodes

Bertagna, et al., Performance-Portable Nonhydrostatic Atmospheric Dycore for the Energy Exascale Earth System Model Running at Cloud-Resolving Resolutions, in 2020 SC20: International Conference for High Performance Computing, Networking, Storage and Analysis (SC), 2020

15 of 19

Performance: GPUs vs CPUs

16 of 19

GPUs vs CPUs

  • Our data always plotted per node basis.
  • GPU nodes cost more and consume more power than CPU nodes. How should we normalize this comparison?
    • Performance per total cost of ownership?
    • Performance per purchase cost
    • Performance per watt
  • Actual power consumption measured on Perlmutter running SCREAM at 3km resolution:
    • PM-GPU (1xAMD EYPC + 4xA100): 1150 W/node
    • PM-CPU (2xAMD EPYC) 690 W/node
    • GPU nodes consume ~ 1.7x more power

17 of 19

C++/Kokkos: Performance Portability� �

SCREAM GRCM, but running at 1 deg resolution: 128L, NH dycore, 10 tracers, P3/SHOC physics with prescribed aerosols, no convective parameterization

  • C++ code is as fast or faster than Fortran code on CPUs
  • Performance portability
    • IBM P9, AMD EYPC
    • NVIDIA V100, A100
    • AMD MI250
  • GPU performance:
    • Large scaling range where GPU nodes are 4-10x faster than CPU nodes

18 of 19

Strong scaling: Frontier, Summit, PM

  • DYAMOND simulations with C++ code
  • Resolution: 3.25km/128L
  • AMIP configuration: Active atmosphere, land, ice thermodynamics, prescribed SST and ice extent
  • Frontier (4 MI250s per node)
  • Summit ( 6 V100s per node)
  • Perlmutter ( 4 A100s per node)
  • Perlmutter CPU (2 EPYCs per node)

Perlmutter CPU nodes vs GPU nodes speedup:

  • 5.8x faster on 1536 nodes (full model)
  • 3.5x faster per Watt

19 of 19

References

  • Donahue et al., To exascale and beyond—The Simple Cloud‐ Resolving E3SM Atmosphere Model (SCREAM), a performance portable global atmosphere model for cloud‐resolving scales. Journal of Advances in Modeling Earth Systems, 16 (2024) e2024MS004314
  • Taylor et al., The Simple Cloud-Resolving E3SM Atmosphere Model Running on the Frontier Exascale System, SC23: International Conference for High Performance Computing, Networking, Storage and Analysis (2023) doi:10.1145/3581784
  • Caldwell et al., Convection-Permitting Simulations with the E3SM Global Atmosphere Model, J. Adv. Model. Earth Syst., 13 (2021) e2021MS002544
  • Bertagna, et al., Performance-Portable Nonhydrostatic Atmospheric Dycore for the Energy Exascale Earth System Model Running at Cloud-Resolving Resolutions, in 2020 SC20: International Conference for High Performance Computing, Networking, Storage and Analysis (SC), 2020
  • Bertagna et al., HOMMEXX 1.0: A Performance Portable Atmospheric Dynamical Core for the Energy Exascale Earth System Model, Geosci. Model Dev., 12 (2019)
  • Taylor, Guba, Steyer, Ullrich, Hall, Eldred, An energy consistent discretization of the nonhydrostatic equations in primitive variables, JAMES, 2020

Thanks!