1 of 69

Large Language Models

Lecture 7

Introduction to Transformer: Self-Attention and Multi-Head Attention

Krishnendu Ghosh

2 of 69

Recap: Attention

3 of 69

Recap: Attention

4 of 69

Recap: Attention

5 of 69

Recap: Attention

6 of 69

Self-Attention

7 of 69

Self-Attention

8 of 69

Self-Attention

9 of 69

Self-Attention

10 of 69

Self-Attention

11 of 69

Self-Attention

12 of 69

Self-Attention

13 of 69

14 of 69

Self-Attention

15 of 69

Self-Attention

16 of 69

Self-Attention

17 of 69

Self-Attention

18 of 69

Self-Attention

19 of 69

Self-Attention

20 of 69

Self-Attention

21 of 69

Self-Attention

22 of 69

Self-Attention

23 of 69

Self-Attention

24 of 69

Self-Attention

25 of 69

Self-Attention

26 of 69

Self-Attention

27 of 69

Self-Attention

28 of 69

Self-Attention

29 of 69

30 of 69

From Self-Attention to Transformers

• We will talk about a class of models for processing sequences that does not use recurrent connections but instead relies entirely on attention and will build up towards a class of models called Transformers.

• To address a few key limitations, we need to add certain elements:

1. Positional encoding addresses lack of sequence information

2. Multi-headed attention allows querying multiple positions at each layer

3. Adding nonlinearities so far, each successive layer is linear in the previous one

4. Masked decoding how to prevent attention lookups into the future?

31 of 69

Positional Encoding: Motivation

32 of 69

33 of 69

Sinusoidal Positional Encoding

34 of 69

Multi-Head Attention

35 of 69

Multi-Head Attention

36 of 69

Multi-Head Attention

37 of 69

Multi-Head Attention

38 of 69

Multi-Head Attention

39 of 69

Multi-Head Attention

40 of 69

Multi-Head Attention

41 of 69

42 of 69

Self-Attention (In Encoder)

43 of 69

Self-Attention (In Encoder)

44 of 69

Self-Attention (In Encoder)

45 of 69

Self-Attention (In Encoder)

46 of 69

Self-Attention (In Encoder)

47 of 69

Self-Attention (In Encoder)

48 of 69

Self-Attention (In Encoder)

49 of 69

Self-Attention (In Encoder)

50 of 69

Multi-Head Self-Attention

51 of 69

Multi-Head Self-Attention

52 of 69

53 of 69

Self-Attention Is “Linear”

54 of 69

55 of 69

Position-wise

Feed-Forward Networks

56 of 69

Position-wise

Feed-Forward Networks

57 of 69

58 of 69

Self-attention can see the future!

59 of 69

Self-attention can see the future!

60 of 69

Masked Attention

61 of 69

Example Tokens

62 of 69

Self-Attention

63 of 69

Self-Attention

64 of 69

Multi-Head Attention

65 of 69

Cross-Attention

66 of 69

Cross-Attention

67 of 69

Masked Attention

68 of 69

Masked Attention

69 of 69