An Overview of Distributed Training
2022.07.15
�Sponsored by 系统设计开荒小分队/东哥IT笔记
Why do we need distributed training?
Type of Parallelism
Data Parallelism
Data Parallelism
Synchronized training
All-Reduce Algorithm
Ring All-Reduce
Tree All-Reduce
Asynchronized training
In synchronized training, worker has to wait for other workers in order to move ahead, thus whole process is only as fast as the slowest worker in the cluster.
Parameter Server
Parameter Server
Steps:
Parameter Server
Parameter Server
Parameter Server
Parameter Server
1. Efficient communication
2. Flexible consistency models
3. Elastic Scalability
4. Fault Tolerance and Durability
5. Ease of Use
Model Parallelism
Model Parallelism
Model Parallelism
Centralized and Decentralized Training
*Distributed Training of Deep Learning Models: A Taxonomic Perspective
Benefits of Distributed Training
Distributed Training Frameworks
Distributed Training Frameworks
References