1 of 46

Attention�Self-Attention�Transformer

Lecture : 06

Google Colab Notebook

2 of 46

Few Reminders

  • Midterm: October 25
  • Project proposal: October 18 (group)

Project proposal template [You can add other points if you have]

      • Problem statement
      • Significance
      • Solution approach
        1. Preprocessing
        2. Model architecture
        3. Hardware requirement to train
      • Dataset
      • Evaluation criteria

3 of 46

RNN/LSTM

4 of 46

RNN based seq2seq translation

Sutskever et al, “Sequence to sequence learning with neural networks”, NeurIPS 2014

We are learning Language

Context vector

Initial decoder state

5 of 46

RNN based seq2seq translation

We are learning Language [START]

estamos

6 of 46

RNN based seq2seq translation

estamos Aprendiendo

We are learning Language [START] estamos

7 of 46

RNN based seq2seq translation

We are learning Language [START]

estamos Aprendiendo español [END]

estamos Aprendiendo español

8 of 46

RNN based seq2seq translation

We are learning Language [START]

estamos Aprendiendo español [END]

estamos Aprendiendo español

-How about long-context and dependency on a fixed-size vector?

-Dependency of previous state and Un-Parallel .

9 of 46

Attention

  • Attention is a mechanism that helps models focus on specific parts of the input sequence while generating output.

The animal didn't cross the street because it was too tired

10 of 46

Decoder to Encoder Attention

We are learning Language

Context vector

Initial decoder state

 

Bahdanau et al, “Neural machine translation by jointly learning to align and translate”, ICLR 2015

11 of 46

Decoder to Encoder Attention

 

We are learning Language

Context vector

Initial decoder state

Normalize attention (sum attention = 0)

Bahdanau et al, “Neural machine translation by jointly learning to align and translate”, ICLR 2015

12 of 46

Decoder to Encoder Attention

 

We are learning Language

Context vector

Initial decoder state

Normalize attention (sum attention = 0)

Bahdanau et al, “Neural machine translation by jointly learning to align and translate”, ICLR 2015

13 of 46

Decoder to Encoder Attention

 

We are learning Language

Context vector

Initial decoder state

Normalize attention (sum attention = 0)

Bahdanau et al, “Neural machine translation by jointly learning to align and translate”, ICLR 2015

 

14 of 46

Decoder to Encoder Attention

 

We are learning Language

Context vector

Initial decoder state

Normalize attention (sum attention = 0)

Bahdanau et al, “Neural machine translation by jointly learning to align and translate”, ICLR 2015

 

estamos Aprendiendo español [END]

15 of 46

Attention weight

Example: English to French translation

Input: “The agreement on the European Economic Area was signed in August 1992.”

Output: “L’accord sur la zone économique européenne a été signé en août 1992.”

Bahdanau et al, “Neural machine translation by jointly learning to align and translate”, ICLR 2015

16 of 46

Attention weight

Example: English to French translation

Input: “The agreement on the European Economic Area was signed in August 1992.”

Output: “L’accord sur la zone économique européenne a été signé en août 1992.”

Bahdanau et al, “Neural machine translation by jointly learning to align and translate”, ICLR 2015

17 of 46

Decoder to Encoder Attention

 

We are learning Language

Context vector

Initial decoder state

Normalize attention (sum attention = 0)

Bahdanau et al, “Neural machine translation by jointly learning to align and translate”, ICLR 2015

 

estamos Aprendiendo español [END]

-How about long-context and dependency on a fixed-size vector?

-Dependency of previous state and Un-Parallel .

18 of 46

Decoder to Encoder Attention

 

We are learning Language

Context vector

Initial decoder state

Normalize attention (sum attention = 0)

Bahdanau et al, “Neural machine translation by jointly learning to align and translate”, ICLR 2015

 

estamos Aprendiendo español [END]

-How about long-context and dependency on a fixed-size vector? Fixed

-Dependency of previous state and Un-Parallel .

19 of 46

Self-Attention

  • Given a token from an input sequence, self-attention tries to measure attention wight from the input sequence.
  • Its all about Keys, Queries , and Values.

Embedding

a3

a2

a1

a4

20 of 46

Self-Attention

  • Given a token from an input sequence, self-attention tries to measure attention wight from the input sequence.
  • Its all about Keys, Queries , and Values.

Embedding

a3

a2

a1

a4

X4

Keys

Values

X1

v1

x2

v2

x3

v3

x4

v4

v4

v4

Query

21 of 46

Self-Attention as a look-up analogy

X4

Keys

Values

X1

v1

x2

v2

x3

v3

x4

v4

v4

X4

Keys

Values

X1

v1

x2

v2

x3

v3

x4

v4

Query

Query

Weighted sum of values

Value = Attention

22 of 46

Hypothetical example

Attention for word “learning”

Embedding

k

K

K

k

V

V

V

V

Q

Self-Attention

Q, K, and V are three learnable weight matrices

23 of 46

Hypothetical example

  •  

Embedding

k

K

K

k

V

V

V

V

Q

Self-Attention

Q, K, and V are three learnable weight matrices

Attention for word “learning”

24 of 46

What is happening in matrix space?

https://jalammar.github.io/illustrated-transformer/

25 of 46

What is happening in matrix space?

https://jalammar.github.io/illustrated-transformer/

26 of 46

Can we use self-attention only for downstream task? ( prediction ..)

  • Positional information not available
  • Lack of nonlinearity. Just weighted average.
  • Future sequence get exposed in machine translation.

27 of 46

Can we use self-attention only for downstream task? ( prediction ..)

  • Positional information not available 🡪 Positional embedding
  • Lack of nonlinearity. Just weighted average. 🡪 Adding FFN layer
  • Future sequence get exposed in machine translation. 🡪 Masking

28 of 46

Positional Encoding

  •  

Position: Binary

29 of 46

Positional Encoding

  •  

Position: Binary

30 of 46

Positional Encoding

  •  

31 of 46

Hypothetical example with positional encoding

  •  

Embedding

k

K

K

k

V

V

V

V

Q

Self-Attention

Q, K, and V are three learnable weight matrices

Attention for word “learning”

+ + + +

Positional vector

32 of 46

Adding FF layer with self-attention

https://medium.com/@umbertofontana/nlp-part-7-self-attention-and-transformers-c770dea1283a

33 of 46

Multi-Head Attention

🡪

34 of 46

Multi-Head Attention

https://jalammar.github.io/illustrated-transformer/

35 of 46

Masking in self-attention ( Machine translation )

  • Why?

36 of 46

Masking in self-attention ( Machine translation )

  • Why?

Embedding

k

K

K

k

V

V

V

V

Q

Self-Attention

+ + + +

Positional vector

estamos Aprendiendo español [END]

Decoder

37 of 46

Transformer

Attention is all you need

Image source: Walmart.com

38 of 46

Transformer

Neuron

39 of 46

Transformer

40 of 46

Transformer

Thinking

Machine

0.31

0.14

0.93

0.14

0.88

0.98

Learning

Language

0.85

0.20

0.14

0.46

0.61

0.49

Batch Normalization

41 of 46

Transformer

Thinking

Machine

0.31

0.14

0.93

0.14

0.88

0.98

Learning

Language

0.85

0.20

0.14

0.46

0.61

0.49

Layer Normalization

42 of 46

Transformer Feed-Forward Layers Are Key-Value Memories

https://arxiv.org/pdf/2012.14913

43 of 46

Transformer Feed-Forward Layers Are Key-Value Memories

https://arxiv.org/pdf/2012.14913

44 of 46

Transformer Decoder

https://towardsdatascience.com/transformers-explained-visually-part-3-multi-head-attention-deep-dive-1c1ff1024853

45 of 46

Transformer training

https://www.youtube.com/watch?v=XowwKOAWYoQ&t=1589s

46 of 46

Further reading