1 of 67

Large Language Models

Lecture 6

Neural LMs: Seq-to-Seq and Attention

Krishnendu Ghosh

2 of 67

LSTM

3 of 67

LSTM

4 of 67

LSTM

5 of 67

LSTM

6 of 67

LSTM

7 of 67

LSTM

8 of 67

LSTM

9 of 67

LSTM

10 of 67

LSTM

11 of 67

LSTM

12 of 67

Attention based Models

13 of 67

Attention based Models

14 of 67

Attention based Models

15 of 67

Attention based Models

16 of 67

Attention based Models

17 of 67

Neural Machine Translation?

• Neural Machine Translation (NMT) is a way to do Machine Translation with a single neural network.

• The neural network architecture is called sequence-to-sequence (aka seq2seq) and it involves two RNNs.

18 of 67

Neural Machine Translation (NMT)

19 of 67

Neural Machine Translation (NMT)

20 of 67

Sequence-to-Sequence is Versatile!

• The general notion here is an encoder-decoder model

• One neural network takes input and produces a neural representation

• Another network produces output based on that neural representation

• If the input and output are sequences, we call it a seq2seq model

• Sequence-to-sequence is useful for more than just MT

• Many NLP tasks can be phrased as sequence-to-sequence:

• Summarization (long text → short text)

• Dialogue (previous utterances → next utterance)

• Parsing (input text → output parse as sequence)

• Code generation (natural language → Python code)

21 of 67

Neural Machine Translation (NMT)

22 of 67

Training an NMT System

23 of 67

Greedy Decoding

24 of 67

Problems: Greedy Decoding

• Greedy decoding has no way to undo decisions!

• Input: il a m’entarté

• → he

• → he hit

• → he hit a

(he hit me with a pie)

(whoops! no going back now...)

How to fix this?

25 of 67

Exhaustive Search Decoding

26 of 67

Beam Search Decoding

27 of 67

Beam Search Decoding:

Stopping Criterion

28 of 67

Beam Search Decoding: Example

29 of 67

Beam Search Decoding: Example

30 of 67

Beam Search Decoding: Example

31 of 67

Beam Search Decoding: Example

32 of 67

Beam Search Decoding: Example

33 of 67

Beam Search Decoding: Example

34 of 67

Beam Search Decoding: Example

35 of 67

Beam Search Decoding: Example

36 of 67

Beam Search Decoding: Example

37 of 67

Beam Search Decoding: Example

38 of 67

Beam Search Decoding: Example

39 of 67

Beam Search Decoding: Example

40 of 67

Beam Search Decoding: Example

41 of 67

Beam Search Decoding:

Stopping Criterion

42 of 67

Beam Search Decoding:

Finishing Up

43 of 67

Sequence-to-Sequence:

The Bottleneck Problem

• Linear interaction distance

• Bottleneck problem

• Lack of parallelizability

44 of 67

Attention

• Attention provides a solution to the bottleneck problem.

• Core idea: on each step of the decoder, use direct connection to the encoder to focus on a particular part of the source sequence

• Let’s start with the visualization of the attention mechanism.

45 of 67

Sequence-to-Sequence

with Attention

46 of 67

Sequence-to-Sequence

with Attention

47 of 67

Sequence-to-Sequence

with Attention

48 of 67

Sequence-to-Sequence

with Attention

49 of 67

Sequence-to-Sequence

with Attention

50 of 67

Sequence-to-Sequence

with Attention

51 of 67

Sequence-to-Sequence

with Attention

52 of 67

Sequence-to-Sequence

with Attention

53 of 67

Sequence-to-Sequence

with Attention

54 of 67

Attention: In Equations

55 of 67

Attention is Great

Attention significantly improves NMT performance

• It’s very useful to allow decoder to focus on certain parts of the source

• Attention solves the bottleneck problem

• Attention allows decoder to look directly at source; bypass bottleneck

• Attention helps with vanishing gradient problem

• Provides shortcut to faraway states

• Attention provides some interpretability

• By inspecting attention distribution, we can see what the decoder was focusing on

• We get (soft) alignment for free!

• This is cool because we never explicitly trained an alignment system

• The network just learned alignment by itself

56 of 67

Seq2Seq+Attention for LM

57 of 67

Attention is General Deep Learning

• We’ve seen that attention is a great way to improve the sequence-to-sequence model for Machine Translation.

• However: You can use attention in many architectures (not just seq2seq) and many tasks (not just MT)

More general definition of attention:

• Given a set of vector values, and a vector query, attention is a technique to compute a weighted sum of the values, dependent on the query.

• We sometimes say that the query attends to the values.

• For example, in the seq2seq + attention model, each decoder hidden state (query) attends to all the encoder hidden states (values).

Intuition:

• The weighted sum is a selective summary of the information contained in the values, where the query determines which values to focus on.

• Attention is a way to obtain a fixed-size representation of an arbitrary set of representations (the values), dependent on some other representation (the query).

58 of 67

Attention

59 of 67

Attention

60 of 67

Attention

61 of 67

Attention

62 of 67

Attention

63 of 67

Attention: Examples

64 of 67

Attention: Examples

65 of 67

Attention: Examples

66 of 67

Variants of Attention

67 of 67