Rapid growth of GPUs in DCs: �With great power comes great responsibility
Esha Choukse
Azure Research- Systems
Microsoft
How rapid is this growth?
350,000 NVIDIA H100 GPUs
Across Azure, AWS, GCP, and other DCs very conservatively, assuming 8 GW per year
>0.6% of total US grid
Projected annual growth in US electricity demand until 2050 is 1% per year
Virginia’s data center alley
Source: PJM grid operator data
Prime power
Until then?
Datacenters today
Power efficiency in CPUs
Power-performance
1 core active
n core active
Power
Frequency
Feature | Outcome | CPUs | GPUs |
Process priority | Fewer cores | Yes | No |
Per core voltage domain | Lower idle power | Yes | No |
Per core frequency control | Better power-perf | Yes | No |
Performance vs energy efficient modes | Easy power-perf | Yes | No |
| | 700 W node 4 W per core | 10,000 W node 1200 W per GPU |
We need to act fast
CPUs took 2 decades of slowly evolving their energy and power management features to get here.
With GPUs, we do not have the luxury of time, and need solutions now.
We need to ensure that the 0.5% of US grid per year number does not exponentially increase
Bridge the gap using the existing features.
How rapid is this growth?
This just in, yesterday:
350,000 NVIDIA H100 GPUs
Across Azure, AWS, GCP, and other DCs very conservatively, assuming 6 GW per year
~0.5% of total US grid
Projected annual growth in US electricity demand until 2050 is 1% per year
Do we need all that power?
Power oversubscription despite the limitations of control/telemetry
Characterizing Power Management Opportunities for LLMs in the Cloud, ASPLOS 2024 – with our intern Pratyush Patel, UW
Do we need all that hardware?
1 core active
n core active
Power
Frequency
Break open the box to find efficiency
All the transformer-based generative LLMs in the past 4-5 years have the same behaviors from power/energy perspective
Motivation I
Metric | Prompt phase | Token generation phase |
Memory capacity needs | Low | High |
Memory Bandwidth needs | Low | High |
Compute needs | High | Low |
Power intensity | High | Low |
Time spent | Generally less | Generally more |
Motivation II
| H100/A100 ratio |
TFLOPs | 3.43x |
Memory capacity | 1x |
Memory Bandwidth | 1.64x |
Power | 1.75x |
Cost* | 2.16x |
| H100/A100 ratio |
Prompt throughput | 2.2x |
Cost per prompt | 1x |
Energy per prompt | 0.77x |
| H100/A100 ratio |
Token throughput | 1.42x |
Cost per token | 1.5x |
Energy per token | 1.2x |
BLOOM-176B with 4k input & 512 output – * From per-VM pricing in CoreWeave
Prompt
Token
H100 better
A100 better
Break open the box to find efficiency
We split the inference into these two phases and run them on separate machine pools that are
Sized independently
Can be different hardware
Can run at different power/perf point
Unlocks upto 80% more throughput at same power!
Important to optimize for power, and not just cost
Splitwise: Efficient generative LLM inference using phase splitting, ISCA, 2024 – with our intern Pratyush Patel, UW
Splitwise: Layerwise KV-cache transfer
17
CPU
GPU
GPU
GPU
GPU
GPU
GPU
GPU
GPU
CPU
GPU
GPU
GPU
GPU
GPU
GPU
GPU
GPU
: 1/8th of the KV-cache for an 8-GPU model
Layer 1
Layer 2
Splitwise has less than 0.8% impact on the end-to-end inference time
Prompt Machine
Token Generation Machine
Splitwise: Evaluated designs
| Prompt machine | Token machine | Prompt-token interconnect bandwidth | ||||
| Type | Cost | Power | Type | Cost | Power | |
Splitwise-AA | DGX-A100 | 1.00x | 1.00x | DGX-A100 | 1.00x | 1.00x | 1.00x |
Splitwise-HH | DGX-H100 | 2.35x | 1.75x | DGX-H100 | 2.50x | 1.75x | 2.00x |
Splitwise-HHcap | DGX-H100 | 2.35x | 1.75x | DGX-H100 capped | 2.50x | 1.23x | 2.00x |
Splitwise-HA | DGX-H100 | 2.35x | 1.75x | DGX-A100 | 1.00x | 1.00x | 1.00x |
Evaluated Splitwise designs all normalized to DGX-A100
SW-only solutions
Need infrastructure changes
Splitwise: System diagram
Splitwise: High-level results
Same power, but maximize throughput (requests per second are shown)
Trace | Baseline | Splitwise-AA | Splitwise-HH | Splitwise-HHcap | Splitwise-HA |
Github | Baseline-A100 | 1.75x | 2.25x | 2.25x | 2x |
Github | Baseline-H100 | 1.4x | 1.8x | 1.8x | 1.6x |
Bing | Baseline-A100 | 2.3x | 2.3x | 2.8x | 2.6x |
Bing | Baseline-H100 | 1.05x | 1.05x | 1.27x | 1.18x |
Other directions we are looking at
Layers |
Model – use the right model and sharding |
Platform – scheduling at the cluster level and instance level |
Driver – more efficient kernels for long context attention |
Hardware – trying out alternate hardware for lower energy |
Datacenter management – Power oversubscription, Idle power reduction, Gridflex, Better cooling (25% power) |
Azure Research – Systems
AI will make the world more efficient, but we need to keep making AI efficient
Thanks!
AI will make the world more efficient, but we need to keep making AI efficient
Splitwise Simulator
Splitwise Simulator
Step 1: Cluster provisioning under SLO, optimizing for the chosen function
Optimized design space search
Step 2: Simulate full trace of incoming requests on chosen cluster size
Request scheduling
Prompt/token pool management
Splitwise: High-level results
Same cost, but maximize throughput
Same throughput, but
Power vs Energy