AI Data Centres:�Networking Problems and Solutions
Shraddha Hegde
Connections 2026
Agenda
Problem Space
AI Training workloads powered by collective communications
GPU-GPU training traffic
Frontier Models
High bandwidth
Low entropy
High cost GPUs
Training step dependency
Tail latency governs completion
Size of the models
Scale out limits
Scale across DCs
Multi-site training
Traffic Characteristics
Failure Sensitivity
Scale out/Scale across
Synchronous bursts
Traffic management solutions
Deterministic forwarding: SRv6 micro-sid
Micro-Sid stack on the NIC
Alternate paths on failure/congestion
draft-filsfils-srv6ops-srv6-ai-backend
BGP based routing planes
Micro-SID based SR policies within the planes
Per routing plane locators
draft-hss-bgp-srv6-routing-planes
Multi-path DAG, unequal cost load-bala
Reduced signalling and fwd state
Strategically bandwidth managed
Multicast for CCL
draft-kompella-teas-mpte/draft-kompella-teas-mcte
Congestion Mitigation
“pause” frame to be sent to upstream neighbors
C-SIG
FANN WG
Multi-path Reliable Connection
Random Network Graph
References
https://arxiv.org/pdf/2605.04333v1
THANKS
shraddha.hegde@hpe.com