1 of 44

Deep Learning (DEEP-0001)�

13 – Transformers

2 of 44

Transformers

  • Motivation
  • Dot-product self-attention
  • Matrix form
  • The transformer
  • NLP pipeline
  • Decoders
  • Large Language models

3 of 44

Motivation

Design neural network to encode and process text:

The word their must “attend to” the word restaurant.

4 of 44

Motivation

The word their must “attend to” the word restaurant.

Conclusions:�

  • There must be connections between the words.
  • The strength of these connections will depend on the words themselves.

Design neural network to encode and process text:

5 of 44

Motivation

  • RNN models become critical at longer sequence lengths, as memory constraints limit batching across examples.

6 of 44

Architecture

7 of 44

Input Embedding

8 of 44

Positional encoding

Order is important in language:

The man ate the fish

vs.

The fish ate the man

9 of 44

Positional Encoding

10 of 44

Positional Encoding

Problem: the scale of these numbers may limit optimization!

Problem: for length 5, 0.8=4/5, meaning it would be the 4th element. For sequence length 20, 0.8=16/20 means 0.8 represents the 16th element!

11 of 44

Positional Encoding

12 of 44

Positional Encoding

13 of 44

Positional Encoding

14 of 44

Transformers

  • Motivation
  • Dot-product self-attention
  • Matrix form
  • The transformer
  • NLP pipeline
  • Decoders
  • Large Language models

15 of 44

Dot product = measure of similarity

16 of 44

Self Attention

17 of 44

Query, Key, and Value

18 of 44

Query, Key, and Value

Search request

Possible similar videos (e.g., same title)

Content of a video

Similarity: proxy to attention!

19 of 44

Derive key, query, and value

20 of 44

Attention scores (filter)

21 of 44

All the inputs

22 of 44

Multi-head

Instead of one single attention head, Q, K, and V are split into multiple heads. It allows the model to jointly attend to information from different representation subspaces at different positions.

As we encode the word "it", one attention head is focusing most on "the animal", while another is focusing on "tired" - in a sense, the model's representation of the word "it" bakes in some of the representation of both "animal" and "tired".

23 of 44

General overview

24 of 44

Residual Connections

25 of 44

Add and Norm

26 of 44

Transformers

  • Motivation
  • Dot-product self-attention
  • Matrix form
  • The transformer
  • NLP pipeline
  • Decoders
  • Large Language models

27 of 44

Matrix form

  • Store N input vectors in matrix X

  • Compute values, queries and keys:�

  • Combine self-attentions

28 of 44

Matrix form

29 of 44

Transformers

  • Motivation
  • Dot-product self-attention
  • Matrix form
  • The transformer
  • NLP pipeline
  • Decoders
  • Large Language models

30 of 44

Encoder/Decoder

31 of 44

Encoder/Decoder

  • Decoder has two inputs:
    • Output from encoder;
    • Input text.

32 of 44

Decoder

Output from encoder:

split into two copies

33 of 44

Decoder

Logits

34 of 44

Training

35 of 44

Training

36 of 44

Training

37 of 44

Training

38 of 44

softmax

39 of 44

Masked Attention

40 of 44

Masked Attention

41 of 44

Masked Attention

42 of 44

Masked Attention

43 of 44

Transformers

  • Motivation
  • Dot-product self-attention
  • Matrix form
  • The transformer
  • NLP pipeline
  • Decoders
  • Large Language models

44 of 44

GPT3 (Brown et al. 2020)

  • Sequence lengths are 2048 tokens long
  • Batch size is 3.2 million tokens.
  • 96 transformer layers (some of which implement a sparseversion of attention), each of which processes a word embedding of size 12288.
  • 96 heads in the self-attention layers and the value, query, and key dimension is 128.
  • 300 billion tokens
  • 175 billion parameters