Large Language Models
Lecture 10
Advanced Attention Mechanisms
Krishnendu Ghosh
Self Attention
Causal (Forward Masked) Attention
Why do we need to do better?
Why do we need to do better?
Why do we need to do better?
Row 1: "I" predicts: love
Row 2: "I love" predicts: machine
Row 3: "I love machine" predicts: learning
Row 4: "I love machine learning" predicts: next token
Q. Is it only for Causal models?
Why do we need to do better?
KV Cache based (Forward Masked)
Attention
Sliding Window Attention
Sliding Window Attention
What happens to the KV Cache?
Multi-Head Self Attention
Multi-Query Attention (MQA)
Do we lose out on something?
Uptraining: Converting MHA to MQA
What can still go wrong?
Grouped Query Attention
What did we gain?
So are we all set?
Key Takeaways
A bit more about the GPU
What was happening so far
The magic: Fused Kernel
The magic does not end here!
More optimization
The magic does not end here!
Tiling
Does the story end here?
What’s the problem?
The softmax denominator problem
Summary Statistics: the final touch!
Summary Statistics: the final touch!
How well did they do?
So, finally solved attention hurdle?
Key Takeaways
Want more? Follow:
Source: https://lcs2-iitd.github.io/ELL881-AIL821-2401/lectures/