1 of 29

Rapid growth of GPUs in DCs: �With great power comes great responsibility

Esha Choukse

Azure Research- Systems

Microsoft

2 of 29

How rapid is this growth?

350,000 NVIDIA H100 GPUs

    • 437 MW of additional power in 1 year at Meta
    • Enough for a town of 150,000 humans

Across Azure, AWS, GCP, and other DCs very conservatively, assuming 8 GW per year

>0.6% of total US grid

Projected annual growth in US electricity demand until 2050 is 1% per year

3 of 29

Virginia’s data center alley

Source: PJM grid operator data

4 of 29

Prime power

  • We are running out, and utility companies are unable to meet demands
  • In 5-10 years, we will see a lot of
    • Requirements for flexibility and load shaping from the grid
    • DCs with their own energy sources

Until then?

5 of 29

Datacenters today

6 of 29

Power efficiency in CPUs

  • Has been around for decades
  • Several works from CSPs have used various power features for efficiency

7 of 29

Power-performance

1 core active

n core active

Power

Frequency

Feature

Outcome

CPUs

GPUs

Process priority

Fewer cores

Yes

No

Per core voltage domain

Lower idle power

Yes

No

Per core frequency control

Better power-perf

Yes

No

Performance vs energy efficient modes

Easy power-perf

Yes

No

700 W node

4 W per core

10,000 W node

1200 W per GPU

8 of 29

We need to act fast

CPUs took 2 decades of slowly evolving their energy and power management features to get here.

With GPUs, we do not have the luxury of time, and need solutions now.

We need to ensure that the 0.5% of US grid per year number does not exponentially increase

Bridge the gap using the existing features.

9 of 29

How rapid is this growth?

This just in, yesterday:

350,000 NVIDIA H100 GPUs

    • 437 MW of additional DC power in 1 year for Meta
    • Enough for a town of 150,000 humans
    • Do we really need all that hardware?
      • And does that hardware really need all that power?

Across Azure, AWS, GCP, and other DCs very conservatively, assuming 6 GW per year

~0.5% of total US grid

Projected annual growth in US electricity demand until 2050 is 1% per year

10 of 29

Do we need all that power?

  • For training … yes
  • We saw opportunity in LLM inference
    • Max power utilization in a DC row was < 80%
    • Inherent nature of the workload: statistical multiplexing of low and high power phases
    • We should be able to oversubscribe power!
  • And then we see the challenges
    • GPUs are direct-device-access mode into VMs – no in-band access to cloud provider
    • Out-of-band mechanism takes 40 seconds!

11 of 29

Power oversubscription despite the limitations of control/telemetry

  • Despite the 40 seconds, we guarantee no power overdraw
    • Choose the thresholds based on known maximum power spikes in 40 seconds
    • Use heavy hammer of power brake, but sparingly
  • While being able to oversubscribe power by 30%
    • With just 5% latency impact to P99 latency of high priority workloads
  • Can be better with better telemetry and control

Characterizing Power Management Opportunities for LLMs in the Cloud, ASPLOS 2024 – with our intern Pratyush Patel, UW

12 of 29

Do we need all that hardware?

  • Better to use 70 servers at 70% utilization than 100 servers at 50% utilization

  • LLMs can run memory capacity bound
    • Memory is just 20-25% of GPU power
      • Forces the GPU compute to run underutilized

1 core active

n core active

Power

Frequency

13 of 29

Break open the box to find efficiency

All the transformer-based generative LLMs in the past 4-5 years have the same behaviors from power/energy perspective

14 of 29

Motivation I

Metric

Prompt phase

Token generation phase

Memory capacity needs

Low

High

Memory Bandwidth needs

Low

High

Compute needs

High

Low

Power intensity

High

Low

Time spent

Generally less

Generally more

15 of 29

Motivation II

H100/A100 ratio

TFLOPs

3.43x

Memory capacity

1x

Memory Bandwidth

1.64x

Power

1.75x

Cost*

2.16x

H100/A100 ratio

Prompt throughput

2.2x

Cost per prompt

1x

Energy per prompt

0.77x

H100/A100 ratio

Token throughput

1.42x

Cost per token

1.5x

Energy per token

1.2x

BLOOM-176B with 4k input & 512 output – * From per-VM pricing in CoreWeave

Prompt

Token

H100 better

A100 better

16 of 29

Break open the box to find efficiency

We split the inference into these two phases and run them on separate machine pools that are

Sized independently

Can be different hardware

Can run at different power/perf point

Unlocks upto 80% more throughput at same power!

Important to optimize for power, and not just cost

Splitwise: Efficient generative LLM inference using phase splitting, ISCA, 2024 – with our intern Pratyush Patel, UW

17 of 29

Splitwise: Layerwise KV-cache transfer

17

CPU

GPU

GPU

GPU

GPU

GPU

GPU

GPU

GPU

CPU

GPU

GPU

GPU

GPU

GPU

GPU

GPU

GPU

: 1/8th of the KV-cache for an 8-GPU model

Layer 1

Layer 2

Splitwise has less than 0.8% impact on the end-to-end inference time

Prompt Machine

Token Generation Machine

18 of 29

Splitwise: Evaluated designs

Prompt machine

Token machine

Prompt-token interconnect bandwidth

Type

Cost

Power

Type

Cost

Power

Splitwise-AA

DGX-A100

1.00x

1.00x

DGX-A100

1.00x

1.00x

1.00x

Splitwise-HH

DGX-H100

2.35x

1.75x

DGX-H100

2.50x

1.75x

2.00x

Splitwise-HHcap

DGX-H100

2.35x

1.75x

DGX-H100 capped

2.50x

1.23x

2.00x

Splitwise-HA

DGX-H100

2.35x

1.75x

DGX-A100

1.00x

1.00x

1.00x

Evaluated Splitwise designs all normalized to DGX-A100

SW-only solutions

Need infrastructure changes

19 of 29

Splitwise: System diagram

  1. Route incoming requests
  2. Maintain pending queue and batching
  3. Mixed pool to avoid perf impact at high load
  4. Pending queue for local batching
  5. KV-cache transfer using MSCCL++

20 of 29

Splitwise: High-level results

Same power, but maximize throughput (requests per second are shown)

Trace

Baseline

Splitwise-AA

Splitwise-HH

Splitwise-HHcap

Splitwise-HA

Github

Baseline-A100

1.75x

2.25x

2.25x

2x

Github

Baseline-H100

1.4x

1.8x

1.8x

1.6x

Bing

Baseline-A100

2.3x

2.3x

2.8x

2.6x

Bing

Baseline-H100

1.05x

1.05x

1.27x

1.18x

21 of 29

Other directions we are looking at

Layers

Model – use the right model and sharding

Platform – scheduling at the cluster level and instance level

Driver – more efficient kernels for long context attention

Hardware – trying out alternate hardware for lower energy

Datacenter management – Power oversubscription, Idle power reduction, Gridflex, Better cooling (25% power)

22 of 29

Azure Research – Systems

  • Project: Efficient AI
  • Extremely fast research 🡪 Product impact
    • Sometimes faster than paper acceptances ☺
  • We are hiring!

AI will make the world more efficient, but we need to keep making AI efficient

23 of 29

24 of 29

Thanks!

AI will make the world more efficient, but we need to keep making AI efficient

25 of 29

Splitwise Simulator

  • Optimization functions
    • Throughput, cost, and provisioned power
  • Traces from production clusters running GPT-4
    • Different distributions of prompt and generated sizes
      • Bing (conversation)
      • Github (coding)
  • Service level objectives
    • P50, P90, P99 for TTFT, TBT, E2E

26 of 29

Splitwise Simulator

Step 1: Cluster provisioning under SLO, optimizing for the chosen function

Optimized design space search

Step 2: Simulate full trace of incoming requests on chosen cluster size

Request scheduling

Prompt/token pool management

27 of 29

Splitwise: High-level results

Same cost, but maximize throughput

    • Splitwise-AA gets 40% more throughput than Baseline-H100 at 25% more power

Same throughput, but

    • Lower power – win for Azure
      • Splitwise-AA needs 50% less power (fewer machines) than Baseline-A100
      • Splitwise-HH needs 9% less power (fewer machines) than Baseline-H100
      • Splitwise-HHcap needs 20% less power than Baseline-H100
    • Lower cost – win for the customer
      • Splitwise-AA costs 25% less than Baseline-H100

28 of 29

29 of 29

Power vs Energy

  • Peak power determines provisioning
    • We reserve power for DCs
  • Reducing peak power has a direct impact on the efficient use of the reserved power sources
  • In a DC, there are always some nodes active, some not
    • Energy efficiency still brings down peak power
    • Secondary effect