1 of 18

Scaling Stormer Training: An Empirical Study of GPU Communication

Overhead Optimization

1

Nazmun Nahar Tui

Ph.D Student

University of Texas at El Paso

Email: ntui@miners.utep.edu

03 March, 2026

2 of 18

Stormer’s Core Idea

  • Problem: Medium-Range Weather Forecasting
    • Forecast future weather XT from initial condition X0​
  • Vision Transformer backbone
  • Weather-specific embedding
    • Tokenize each variable separately
    • Cross-attention across variables
    • Capture physical relationships

(wind ↔ pressure ↔ humidity)

  • Randomized iterative dynamic forecasting(6h, 12h, 24h intervals)
  • Multi-step fine-tuning
  • Pressure-weighted loss
    • prioritizes near-surface variables�

2

Fig: Stormer Architecture

3 of 18

Training Strategy

  • Phase 1: Train single-step (short δt)
  • Phase 2: Finetune multi-step (K=4 rollouts)
  • Phase 3: Finetune longer rollouts (K=8)
  • Builds stability for long forecasts

3

4 of 18

Inference Strategy

  • To forecast T days ahead:
    • Homogeneous: use one δt repeatedly (e.g., [24h × 7])
    • Best m in n: try many δt combinations, pick best ones
  • Final forecast = average forecasts (ensemble effect)

4

5 of 18

Replicating Stormer

Our Baseline

  • ERA5 / WeatherBench-2 dataset
  • Image Resolution: 240x120
  • Data Split:
    • Train: 2000-2018
    • Validation: 2019
    • Test: 2020
  • Patch size: 8
  • GPUs: 64 A100 GPUs

Stormer

  • ERA5 / WeatherBench-2 dataset
  • Image Resolution: 128x256
  • Data Split:
    • Train: 1979-2018
    • Validation: 2019
    • Test: 2020
  • Patch size: 2
  • GPUs: 128 A100 GPUs

5

6 of 18

Experimental Setup & System Configuration

Hardware Platform

  • System: NERSC Perlmutter
  • GPU Architecture: NVIDIA A100
  • High-speed interconnect for multi-node distributed training

Software Stack

  • PyTorch: 2.9.0
  • CUDA: 12.8
  • NCCL: 2.27.5
  • Distributed Framework: PyTorch Distributed Data Parallel (DDP)

6

7 of 18

Baseline Scaling Behaviour(Phase 2)

7

Metrics

1 Node

2 Nodes

4 Nodes

16 Nodes

NCCL Allreduce GPU Time (%)

11.0

73.8

76.7

86.6

Total NCCL Allreduce Time (s)

8.98

21.51

252.89

475.78

AllReduce Calls

18,870

18,870

18,870

18,870

Avg. cudaStreamSynchronize Time

53.5 𝜇s

98.0 𝜇s

1.32 ms

2.69 ms

8 of 18

Communication Optimization Knobs

  • Mixed Precision
    • Same communication volume as FP16, but improved numerical stability for large-scale training.
  • Gradient Bucket Size
    • Controls how gradients are grouped before AllReduce.
    • Larger buckets → fewer collective calls (better bandwidth efficiency),
    • Smaller buckets → earlier overlap but more synchronization overhead.
  • Gradient Accumulation
    • Reduces synchronization frequency by delaying optimizer steps.
    • Higher GA → fewer AllReduce operations per effective batch → lower communication overhead.�

8

9 of 18

Impact of Communication Optimization Knobs

9

Precision

NCCL AllReduce GPU Time(%)

Total NCCL AllReduce Time(s)

Default (Precision 16)

77.8

267.282

BF16-mixed

77.0

255.383

Bucket Size

NCCL AllReduce GPU Time(%)

Total NCCL AllReduce Time(s)

Default (25 MB)

77.8

267.282

50 MB

75.6

231.265

100 MB

75.5

234.066

Impact of precision on Communication (Baseline, 4 nodes):

Impact of bucket size on Communication (Baseline, 4 nodes):

10 of 18

Impact of Communication Optimization Knobs

10

Grad. Accu.

NCCL AllReduce GPU Time (%)

Total NCCL AllReduce Time(s)

Default (1)

86.0

468.37

2

65.0

256.59

4

49.4

254.60

8

33.8

257.97

16

28.7

84.93

Impact of grad. accu. on Communication (Baseline, 4 nodes):

11 of 18

Best Configuration at Large Scale (16 Nodes)

11

  • Precision: BF16-mixed
  • Bucket Size: 50 MB
  • Gradient Accumulation: 16

Metric

Baseline

Optimized

NCCL GPU Time (%)

86.8

28.1

Total AllReduce Time (s)

33.77

26.88

AllReduce Calls

1326

676

cudaStreamSynchronize Time (s)

31.55

25.7

12 of 18

Ablation Studies(16 Nodes)

12

Phase 2

NCCL AllReduce GPU Time

Total NCCL AllReduce Time

Baseline

86.8 %

33.77 s

Optimized

28.1 %

26.88 s

Study 1 (Removing BF16-Mixed)

25.4 %

26.61 s

Study 2 (Removing bucket size 50 MB)

29.0 %

32.32 s

Study 3 (Removing GA 16)

81.6 %

26.50 s

Study 4 (Removing GA 16, bucket size 50 MB)

84.2 %

31.61 s

Study 5 (Removing BF16-mixed, GA 16 )

81.7%

26.63 s

Study 6 (Removing BF16-mixed, bucket size 50 MB)

32.7 %

38.06 s

13 of 18

Accuracy Comparison

13

14 of 18

Accuracy Comparison

14

15 of 18

Future Research Direction

  • Theoretical Communication Modelling for Transformers
    • Can we generalize our communication model to any large Transformer architecture?
    • Goal: Predict communication cost before running large-scale experiments
  • Predictive Memory Modeling to Avoid OOM
    • Estimate GPU memory requirements beforehand
    • Enable safe resource allocation and scheduling on HPC systems
  • Beyond Data Parallelism: Pipeline Parallelism Analysis
    • Does pipeline parallelism reduce AllReduce cost or introduce new stage-to-stage communication bottlenecks?

15

16 of 18

References

[1] Nguyen, T., Shah, R., Bansal, H., Arcomano, T., Maulik, R., Kotamarthi, R., ... & Grover, A. (2024). Scaling transformer neural networks for skillful and reliable medium-range weather forecasting. Advances in Neural Information Processing Systems, 37, 68740-68771.

[2] Stormer GitHub Repository: https://github.com/tung-nd/stormer

16

17 of 18

Acknowledgements

  • This research was partially supported by the U.S. Department of Energy, Office of Science under � Grant #: DE-SC0024352
  • We acknowledge the use of NERSC Perlmutter Supercomputer under a NERSC ERCAP award

17

18 of 18

Thanks!

18