1 of 162

Large Language Models

Lecture 12

Mixture of Experts & Scaling Laws

Krishnendu Ghosh

2 of 162

Large Language Models

3 of 162

Large Language Models

4 of 162

Large Language Models

5 of 162

Large Language Models

6 of 162

Large Language Models

7 of 162

Why care about model size?

8 of 162

Neural Scaling Laws

9 of 162

Neural Scaling Laws

10 of 162

Neural Scaling Laws

11 of 162

How to efficiently

increase model size?

12 of 162

How to efficiently

increase model size?

13 of 162

Mixture of Experts

14 of 162

Mixture of Experts

15 of 162

Mixture of Experts

16 of 162

Mixture of Experts

17 of 162

Mixture of Experts

18 of 162

Mixture of Experts

19 of 162

Mixture of Experts

20 of 162

Mixture of Experts

21 of 162

Mixture of Experts

22 of 162

Mixture of Experts

23 of 162

Mixture of Experts

24 of 162

Mixture of Experts

25 of 162

Mixture of Experts - Chronology

26 of 162

Mixture of Experts - Chronology

27 of 162

Mixture of Experts - Chronology

28 of 162

Mixture of Experts - Chronology

29 of 162

Mixture of Experts as a Layer

30 of 162

Mixture of Experts as a Layer

31 of 162

Mixture of Experts as a Layer

32 of 162

Mixture of Experts as a Layer

33 of 162

Mixture of Experts as a Layer

34 of 162

Mixture of Experts as a Layer

35 of 162

Mixture of Experts as a Layer

36 of 162

Mixture of Experts as a Layer

37 of 162

MoE Layer

Backpropagation

38 of 162

MoE Layer

Backpropagation

39 of 162

MoE Layer

Backpropagation

40 of 162

MoE Layer

Backpropagation

41 of 162

MoE Layer

Backpropagation

42 of 162

MoE Layer

Backpropagation

43 of 162

MoE Layer

Backpropagation

44 of 162

MoE Layer

Backpropagation

45 of 162

MoE Layer

Backpropagation

46 of 162

Sparse MoE Layer

Greedy expert selection

47 of 162

Sparse MoE Layer

Greedy expert selection

48 of 162

Sparse MoE Layer

Greedy expert selection

49 of 162

Sparse MoE Layer

Greedy expert selection

50 of 162

Sparse MoE Layer

Noisy Top-K Gating

51 of 162

Sparse MoE Layer

Noisy Top-K Gating

52 of 162

Sparse MoE Layer

Noisy Top-K Gating

53 of 162

Sparse MoE Layer

Noisy Top-K Gating

54 of 162

Sparse MoE Layer

Noisy Top-K Gating

55 of 162

Sparse MoE Layer

Noisy Top-K Gating

56 of 162

Sparse MoE Layer

Noisy Top-K Gating

57 of 162

Sparse MoE Layer

Noisy Top-K Gating

58 of 162

Sparse MoE Layer

Noisy Top-K Gating

59 of 162

Sparse MoE Layer

Noisy Top-K Gating

60 of 162

Sparse MoE Layer

Noisy Top-K Gating

61 of 162

Sparse MoE Layer

Noisy Top-K Gating

62 of 162

Why do experts not collapse

in practice?

63 of 162

Why do experts not collapse

in practice?

64 of 162

Why do experts not collapse

in practice?

65 of 162

Why do experts not collapse

in practice?

66 of 162

Why do experts not collapse

in practice?

67 of 162

Why do experts not collapse

in practice?

68 of 162

Why do experts not collapse

in practice?

69 of 162

Why do experts not collapse

in practice?

70 of 162

Learning Dynamics of Experts

and Router

71 of 162

Learning Dynamics of Experts

and Router

72 of 162

Learning Dynamics of Experts

and Router

73 of 162

Learning Dynamics of Experts

and Router

74 of 162

Pros and Cons of Sparse MoE Layer

75 of 162

Example MoE Layer in LLMs

76 of 162

Example MoE Layer in LLMs

77 of 162

Example MoE Layer in LLMs

78 of 162

Switch Transformer Layer

79 of 162

Switch Transformer Layer

Greedy routing to only 1 expert

80 of 162

Expert Parallel for Sparse MoEs

81 of 162

Expert Parallel for Sparse MoEs

82 of 162

Expert Parallel for Sparse MoEs

83 of 162

Switch Transformer Layer

84 of 162

Switch Transformer Layer

85 of 162

Switch Transformer Layer

86 of 162

Switch Transformer Layer

87 of 162

Switch Transformer Layer

88 of 162

Switch Transformer Layer

89 of 162

Expert Parallel for Sparse MoEs

90 of 162

Model Parallelism for

Larger Dense Model

91 of 162

Model Parallelism for

Larger Dense Model

92 of 162

Model Parallelism for

Larger Dense Model

93 of 162

Switch Transformer Layer

94 of 162

Switch Transformer Layer

95 of 162

Switch Transformer Layer

96 of 162

Switch Transformer Layer

97 of 162

Load Balancing Loss

98 of 162

Load Balancing Loss

99 of 162

Switch Transformer Layer

100 of 162

Selective Precision

101 of 162

Switch Transformer Layer

102 of 162

Smaller parameter initialization

for stability

103 of 162

Switch Transformer Layer

104 of 162

Higher regularization for Experts

during fine-tuning

105 of 162

Switch Transformer Layer

106 of 162

Distributed Switch Implementation

107 of 162

Modulating Expert Capacity

via Capacity Factor

108 of 162

Modulating Expert Capacity

via Capacity Factor

109 of 162

Modulating Expert Capacity

via Capacity Factor

110 of 162

No token left behind!

111 of 162

No token left behind!

112 of 162

Benchmarking Switch (Top-1) vs

MoE (Noisy Top-2)

113 of 162

Mixture of Experts: 8x7B

114 of 162

Mixture of Experts: 8x7B

115 of 162

Mixture of Experts: 8x7B

116 of 162

Reasoning vs Knowledge Intensive

Tasks

117 of 162

Reasoning vs

Knowledge Intensive Tasks

118 of 162

Interpreting Routing Decisions

119 of 162

Interpreting Routing Decisions

120 of 162

Routing of Consecutive Tokens

121 of 162

Which experts are active for

different tokens?

122 of 162

Interpreting Experts

123 of 162

LLMs: few-shot Learners

124 of 162

LLMs: few-shot Learners

125 of 162

LLMs: few-shot Learners

126 of 162

LLMs: few-shot Learners

127 of 162

Google BIG-BENCH benchmark

128 of 162

LLMs become intelligent with

129 of 162

Parameter Size & FLOPs

130 of 162

If that’s true, we never get

to know the following:

131 of 162

Is the value of scaling laws only in predicting?

132 of 162

Can there be a curve

that fits “emergence” Intuition

133 of 162

What about fitting “emergence” in non-parametric setting

134 of 162

Notations

135 of 162

Loss-drop (Emergence) & Power-Law

136 of 162

A more reliable version:

137 of 162

We are in luck!

Turns out that scale is predictable

138 of 162

Observation 1a:

Universality of Overfitting

139 of 162

Observation 1b: Sample Efficiency

140 of 162

Key takeaway 1:

Both Parameter and Dataset to be scaled

141 of 162

Observation 2:

What about training time (Steps & FLOPs)

142 of 162

Key takeaway 2:

Universality of training

143 of 162

Key Takeaway 3:

Model shape does not matter!

144 of 162

Key Takeaway 4:

Embedding matrix does not matter!

145 of 162

Key Takeaway 5:

Dataset composition does not matter!

146 of 162

Kaplan Scaling Laws

147 of 162

Chinchilla (Hoffman) Scaling Law

148 of 162

Chinchilla (Hoffman) Scaling Law

149 of 162

Chinchilla Scaling Law vs

Kaplan Scaling Law

150 of 162

(Revised) Chinchilla Scaling Law

151 of 162

Motivation:

Not all metrics score same (Emergence Score)

152 of 162

Is your accuracy metric non-linear

or discontinuous?

153 of 162

Power Law in play!

154 of 162

Problem with Non-linear Measure:

Exact string match

155 of 162

Change of Perspective:

Measure: Edit distance

156 of 162

Problem with Discontinuous

Measure: MCG

157 of 162

Change of Perspective:

Measure: Brier Score

158 of 162

Prediction:

Power Law vs Near-Linear Counterpart

159 of 162

Results on GPT3.5/3

Task: 2-digit integer multiplication

160 of 162

Does the claim work for

Google BIG-BENCH benchmark?

161 of 162

Key Takeaways

162 of 162

Source: https://lcs2-iitd.github.io/ELL881-AIL821-2401/lectures/