What is inside a ML Platform?
Scheduler & Federated Learning Intro
Background
About
About
Cost
Content
Non Scheduler
Non Scheduler
I am Reco ->
I am Ads ->
I am video ->
I am NLU ->
CUDA_VISIBLE_DEVICES
But…
BERT-large -> 340M
GPT-2 -> 1.5B
Megatron -> 8B
DALL-E -> 12B
GPT-3 -> 175B
???
Problems
Solution
Scheduler for
Scheduler itself should not too slow
Cluster Scheduler
Cluster Scheduler
Is Cluster Scheduler Good Enough?
Bottlenecks?
Bottlenecks?
Bottlenecks?
Bottlenecks
Operations
Bottleneck is relative
Bottlenecks
Bottlenecks
IO improved a lot!
IO is still a bottleneck.
Solution
Preload
Pipeline
Stream
Low-cost nodes
Pruning
Precision
Distill
…
Solution
Solution
Industry
Industry
Content
Security Cost
Hard restrictions
Soft restrictions
Security is hard in ML
Security is hard in ML
Industry Practise
Hack case study
DB.getPassword.equals(inputPwd)
Differential privacy
Federated Learning
Federated Learning
Why Federated Learning?
Why Federated Learning?
Why Federated Learning?
IS federated Learning silver bullet?
IS federated Learning silver bullet?
IS federated Learning silver bullet?
Thanks!