1 of 33

Large Language Models

Lecture 10

Advanced Attention Mechanisms

Krishnendu Ghosh

2 of 33

Self Attention

3 of 33

Causal (Forward Masked) Attention

4 of 33

Why do we need to do better?

5 of 33

Why do we need to do better?

6 of 33

Why do we need to do better?

Row 1: "I" predicts: love

Row 2: "I love" predicts: machine

Row 3: "I love machine" predicts: learning

Row 4: "I love machine learning" predicts: next token

Q. Is it only for Causal models?

7 of 33

Why do we need to do better?

8 of 33

KV Cache based (Forward Masked)

Attention

9 of 33

Sliding Window Attention

10 of 33

Sliding Window Attention

11 of 33

What happens to the KV Cache?

12 of 33

Multi-Head Self Attention

13 of 33

Multi-Query Attention (MQA)

14 of 33

Do we lose out on something?

15 of 33

Uptraining: Converting MHA to MQA

16 of 33

What can still go wrong?

17 of 33

Grouped Query Attention

18 of 33

What did we gain?

19 of 33

So are we all set?

Key Takeaways

20 of 33

A bit more about the GPU

21 of 33

What was happening so far

22 of 33

The magic: Fused Kernel

23 of 33

The magic does not end here!

More optimization

24 of 33

The magic does not end here!

Tiling

25 of 33

Does the story end here?

What’s the problem?

26 of 33

The softmax denominator problem

27 of 33

Summary Statistics: the final touch!

28 of 33

Summary Statistics: the final touch!

29 of 33

How well did they do?

30 of 33

So, finally solved attention hurdle?

31 of 33

Key Takeaways

32 of 33

Want more? Follow:

33 of 33

Source: https://lcs2-iitd.github.io/ELL881-AIL821-2401/lectures/