Scaling Stormer Training: An Empirical Study of GPU Communication
Overhead Optimization
1
Stormer’s Core Idea
(wind ↔ pressure ↔ humidity)
2
Fig: Stormer Architecture
Training Strategy
3
Inference Strategy
4
Replicating Stormer
Our Baseline
Stormer
5
Experimental Setup & System Configuration
Hardware Platform
Software Stack
6
Baseline Scaling Behaviour(Phase 2)
7
Metrics | 1 Node | 2 Nodes | 4 Nodes | 16 Nodes |
NCCL Allreduce GPU Time (%) | 11.0 | 73.8 | 76.7 | 86.6 |
Total NCCL Allreduce Time (s) | 8.98 | 21.51 | 252.89 | 475.78 |
AllReduce Calls | 18,870 | 18,870 | 18,870 | 18,870 |
Avg. cudaStreamSynchronize Time | 53.5 𝜇s | 98.0 𝜇s | 1.32 ms | 2.69 ms |
Communication Optimization Knobs
8
Impact of Communication Optimization Knobs
9
Precision | NCCL AllReduce GPU Time(%) | Total NCCL AllReduce Time(s) |
Default (Precision 16) | 77.8 | 267.282 |
BF16-mixed | 77.0 | 255.383 |
Bucket Size | NCCL AllReduce GPU Time(%) | Total NCCL AllReduce Time(s) |
Default (25 MB) | 77.8 | 267.282 |
50 MB | 75.6 | 231.265 |
100 MB | 75.5 | 234.066 |
Impact of precision on Communication (Baseline, 4 nodes):
Impact of bucket size on Communication (Baseline, 4 nodes):
Impact of Communication Optimization Knobs
10
Grad. Accu. | NCCL AllReduce GPU Time (%) | Total NCCL AllReduce Time(s) |
Default (1) | 86.0 | 468.37 |
2 | 65.0 | 256.59 |
4 | 49.4 | 254.60 |
8 | 33.8 | 257.97 |
16 | 28.7 | 84.93 |
Impact of grad. accu. on Communication (Baseline, 4 nodes):
Best Configuration at Large Scale (16 Nodes)
11
Metric | Baseline | Optimized |
NCCL GPU Time (%) | 86.8 | 28.1 |
Total AllReduce Time (s) | 33.77 | 26.88 |
AllReduce Calls | 1326 | 676 |
cudaStreamSynchronize Time (s) | 31.55 | 25.7 |
Ablation Studies(16 Nodes)
12
Phase 2 | NCCL AllReduce GPU Time | Total NCCL AllReduce Time |
Baseline | 86.8 % | 33.77 s |
Optimized | 28.1 % | 26.88 s |
Study 1 (Removing BF16-Mixed) | 25.4 % | 26.61 s |
Study 2 (Removing bucket size 50 MB) | 29.0 % | 32.32 s |
Study 3 (Removing GA 16) | 81.6 % | 26.50 s |
Study 4 (Removing GA 16, bucket size 50 MB) | 84.2 % | 31.61 s |
Study 5 (Removing BF16-mixed, GA 16 ) | 81.7% | 26.63 s |
Study 6 (Removing BF16-mixed, bucket size 50 MB) | 32.7 % | 38.06 s |
Accuracy Comparison
13
Accuracy Comparison
14
Future Research Direction
15
References
[1] Nguyen, T., Shah, R., Bansal, H., Arcomano, T., Maulik, R., Kotamarthi, R., ... & Grover, A. (2024). Scaling transformer neural networks for skillful and reliable medium-range weather forecasting. Advances in Neural Information Processing Systems, 37, 68740-68771.
[2] Stormer GitHub Repository: https://github.com/tung-nd/stormer
16
Acknowledgements
17
Thanks!
18