Large Language Models
Lecture 12
Mixture of Experts & Scaling Laws
Krishnendu Ghosh
Large Language Models
Large Language Models
Large Language Models
Large Language Models
Large Language Models
Why care about model size?
Neural Scaling Laws
Neural Scaling Laws
Neural Scaling Laws
How to efficiently
increase model size?
How to efficiently
increase model size?
Mixture of Experts
Mixture of Experts
Mixture of Experts
Mixture of Experts
Mixture of Experts
Mixture of Experts
Mixture of Experts
Mixture of Experts
Mixture of Experts
Mixture of Experts
Mixture of Experts
Mixture of Experts
Mixture of Experts - Chronology
Mixture of Experts - Chronology
Mixture of Experts - Chronology
Mixture of Experts - Chronology
Mixture of Experts as a Layer
Mixture of Experts as a Layer
Mixture of Experts as a Layer
Mixture of Experts as a Layer
Mixture of Experts as a Layer
Mixture of Experts as a Layer
Mixture of Experts as a Layer
Mixture of Experts as a Layer
MoE Layer
Backpropagation
MoE Layer
Backpropagation
MoE Layer
Backpropagation
MoE Layer
Backpropagation
MoE Layer
Backpropagation
MoE Layer
Backpropagation
MoE Layer
Backpropagation
MoE Layer
Backpropagation
MoE Layer
Backpropagation
Sparse MoE Layer
Greedy expert selection
Sparse MoE Layer
Greedy expert selection
Sparse MoE Layer
Greedy expert selection
Sparse MoE Layer
Greedy expert selection
Sparse MoE Layer
Noisy Top-K Gating
Sparse MoE Layer
Noisy Top-K Gating
Sparse MoE Layer
Noisy Top-K Gating
Sparse MoE Layer
Noisy Top-K Gating
Sparse MoE Layer
Noisy Top-K Gating
Sparse MoE Layer
Noisy Top-K Gating
Sparse MoE Layer
Noisy Top-K Gating
Sparse MoE Layer
Noisy Top-K Gating
Sparse MoE Layer
Noisy Top-K Gating
Sparse MoE Layer
Noisy Top-K Gating
Sparse MoE Layer
Noisy Top-K Gating
Sparse MoE Layer
Noisy Top-K Gating
Why do experts not collapse
in practice?
Why do experts not collapse
in practice?
Why do experts not collapse
in practice?
Why do experts not collapse
in practice?
Why do experts not collapse
in practice?
Why do experts not collapse
in practice?
Why do experts not collapse
in practice?
Why do experts not collapse
in practice?
Learning Dynamics of Experts
and Router
Learning Dynamics of Experts
and Router
Learning Dynamics of Experts
and Router
Learning Dynamics of Experts
and Router
Pros and Cons of Sparse MoE Layer
Example MoE Layer in LLMs
Example MoE Layer in LLMs
Example MoE Layer in LLMs
Switch Transformer Layer
Switch Transformer Layer
Greedy routing to only 1 expert
Expert Parallel for Sparse MoEs
Expert Parallel for Sparse MoEs
Expert Parallel for Sparse MoEs
Switch Transformer Layer
Switch Transformer Layer
Switch Transformer Layer
Switch Transformer Layer
Switch Transformer Layer
Switch Transformer Layer
Expert Parallel for Sparse MoEs
Model Parallelism for
Larger Dense Model
Model Parallelism for
Larger Dense Model
Model Parallelism for
Larger Dense Model
Switch Transformer Layer
Switch Transformer Layer
Switch Transformer Layer
Switch Transformer Layer
Load Balancing Loss
Load Balancing Loss
Switch Transformer Layer
Selective Precision
Switch Transformer Layer
Smaller parameter initialization
for stability
Switch Transformer Layer
Higher regularization for Experts
during fine-tuning
Switch Transformer Layer
Distributed Switch Implementation
Modulating Expert Capacity
via Capacity Factor
Modulating Expert Capacity
via Capacity Factor
Modulating Expert Capacity
via Capacity Factor
No token left behind!
No token left behind!
Benchmarking Switch (Top-1) vs
MoE (Noisy Top-2)
Mixture of Experts: 8x7B
Mixture of Experts: 8x7B
Mixture of Experts: 8x7B
Reasoning vs Knowledge Intensive
Tasks
Reasoning vs
Knowledge Intensive Tasks
Interpreting Routing Decisions
Interpreting Routing Decisions
Routing of Consecutive Tokens
Which experts are active for
different tokens?
Interpreting Experts
LLMs: few-shot Learners
LLMs: few-shot Learners
LLMs: few-shot Learners
LLMs: few-shot Learners
Google BIG-BENCH benchmark
LLMs become intelligent with
Parameter Size & FLOPs
If that’s true, we never get
to know the following:
Is the value of scaling laws only in predicting?
Can there be a curve
that fits “emergence” Intuition
What about fitting “emergence” in non-parametric setting
Notations
Loss-drop (Emergence) & Power-Law
A more reliable version:
We are in luck!
Turns out that scale is predictable
Observation 1a:
Universality of Overfitting
Observation 1b: Sample Efficiency
Key takeaway 1:
Both Parameter and Dataset to be scaled
Observation 2:
What about training time (Steps & FLOPs)
Key takeaway 2:
Universality of training
Key Takeaway 3:
Model shape does not matter!
Key Takeaway 4:
Embedding matrix does not matter!
Key Takeaway 5:
Dataset composition does not matter!
Kaplan Scaling Laws
Chinchilla (Hoffman) Scaling Law
Chinchilla (Hoffman) Scaling Law
Chinchilla Scaling Law vs
Kaplan Scaling Law
(Revised) Chinchilla Scaling Law
Motivation:
Not all metrics score same (Emergence Score)
Is your accuracy metric non-linear
or discontinuous?
Power Law in play!
Problem with Non-linear Measure:
Exact string match
Change of Perspective:
Measure: Edit distance
Problem with Discontinuous
Measure: MCG
Change of Perspective:
Measure: Brier Score
Prediction:
Power Law vs Near-Linear Counterpart
Results on GPT3.5/3
Task: 2-digit integer multiplication
Does the claim work for
Google BIG-BENCH benchmark?
Key Takeaways
Source: https://lcs2-iitd.github.io/ELL881-AIL821-2401/lectures/