1 of 24

An Overview of Distributed Training

2022.07.15

�Sponsored by 系统设计开荒小分队/东哥IT笔记

2 of 24

Why do we need distributed training?

  • Deep learning is training data heavy
  • Complex model needs to store intermediate results

3 of 24

Type of Parallelism

  • Data Parallelism
  • Model Parallelism
  • Graph Parallelism
  • Task Parallelism
  • Hybrid Parallelism/ Mixed Parallelism

4 of 24

Data Parallelism

5 of 24

Data Parallelism

  • Synchronized training
  • Asynchronous training

6 of 24

Synchronized training

  1. Forward pass begins at the same
  2. After computing gradients, start communicating, aggregating (all-reduce)
  3. After all the gradients are combined, the copy of these updated gradients is sent to all the workers. 
  4. Now after getting the updated gradients using the all-reduce algorithm, each worker continues with the backward pass and updates the local copy of the weights normally.

7 of 24

All-Reduce Algorithm

  • Instead of having a single machine to perform the aggregation task, we can distribute the aggregation task on all machines.

  • All machines share the load of storing and maintaining global parameters

8 of 24

Ring All-Reduce

9 of 24

Tree All-Reduce

10 of 24

Asynchronized training

In synchronized training, worker has to wait for other workers in order to move ahead, thus whole process is only as fast as the slowest worker in the cluster.

  • Asynchronized training:
  • Workers train independently
  • Tolerance towards machine computation power difference

11 of 24

Parameter Server

  • Workers that act as parameter servers�- maintain the globally shared parameters

  • Workers that train the model�- Perform training tasks

12 of 24

Parameter Server

Steps:

  1. Replicate the model in all of our workers and each worker uses a subset of data for training.
  2. Each training worker fetches the parameters from the parameter servers.
  3. Each training worker performs a training loop and sends the gradients back to all the parameter servers which then update the model parameters.

13 of 24

Parameter Server

14 of 24

Parameter Server

  • Disadvantages:
  • Using stale version of the model during training
  • Parameter server as a bottleneck �- Single point of failure�- Bottleneck for communication�Solution: introducing multiple parallel servers

15 of 24

Parameter Server

  • Disadvantages:
  • Using stale version of the model during training
  • Parameter server as a bottleneck �- Single point of failure�- Bottleneck for communication�Solution: introducing multiple parallel servers

16 of 24

Parameter Server

  • Advantages:

1. Efficient communication

2. Flexible consistency models

3. Elastic Scalability

4. Fault Tolerance and Durability

5. Ease of Use

17 of 24

Model Parallelism

  • When model is too large to fit on a single worker node�
  • Depends on specific model architecture

18 of 24

Model Parallelism

19 of 24

Model Parallelism

20 of 24

Centralized and Decentralized Training

*Distributed Training of Deep Learning Models: A Taxonomic Perspective

21 of 24

Benefits of Distributed Training

  • Efficiency
  • Scalability 
  • Cost effectiveness
  • Fault tolerance and reliability

22 of 24

Distributed Training Frameworks

  • Pytorch
  • Tensorflow
  • Amazon Sagemaker
  • Keras
  • Horovod

23 of 24

Distributed Training Frameworks

  • Pytorch
    • Provided torch.distributed package API
  • Tensorflow
    • Built-in support for distributed training, provided  tf.distribute.Strategy API
  • Amazon Sagemaker
    • Both data parallelism and model parallelism are supported
  • Keras
    • Simple and easy to use
  • Horovod
    • Distributed deep learning training framework for TensorFlow, Keras, and PyTorch.

24 of 24

References