1 of 57

STATS / DATA SCI 315

Transformers

slides credit: Dr. Simon J. D. Prince

https://udlbook.github.io/udlbook/

2 of 57

Natural language processing (NLP)

  • Translation
  • Question answering
  • Summarizing
  • Generating new text
  • Correcting spelling and grammar
  • Finding entities
  • Classifying bodies of text
  • Changing style etc.

3 of 57

Transformers

  • Motivation
  • Dot-product self-attention
  • Matrix form
  • The transformer
  • NLP pipeline
  • Decoders
  • Large Language models

4 of 57

Motivation

Design neural network to encode and process text:

5 of 57

Motivation

Design neural network to encode and process text:

Word embeddings convert words into fixed dimensional vectors

word2vec is a popular family of word embeddings

6 of 57

Motivation

Design neural network to encode and process text:

x N

7 of 57

Standard fully-connected layer

8 of 57

Standard fully-connected layer

9 of 57

Standard fully-connected layer

Problem:�

  • A very large number of parameters
  • Can’t cope with text of different lengths

Conclusion:

  • We need a model where parameters don’t increase with input length

10 of 57

Motivation

Design neural network to encode and process text:

The word their must “attend to” the word restaurant.

11 of 57

Motivation

The word their must “attend to” the word restaurant.

Conclusions:�

  • There must be connections between the words.
  • The strength of these connections will depend on the words themselves.

Design neural network to encode and process text:

12 of 57

Transformers

  • Motivation
  • Dot-product self-attention
  • Matrix form
  • The transformer
  • NLP pipeline
  • Decoders
  • Large Language models

13 of 57

Dot-product self attention

  • Takes N inputs of size Dx1 and returns N inputs of size Dx1
  • Computes N values (no ReLU)

  • N outputs are weighted sums of these values

  • Weights depend on the inputs themselves

14 of 57

Dot-product self attention

  • Takes N inputs of size Dx1 and returns N inputs of size Dx1
  • Computes N values (no ReLU)

  • N outputs are weighted sums of these values

  • Weights depend on the inputs themselves

15 of 57

Dot-product self attention

  • Takes N inputs of size Dx1 and returns N inputs of size Dx1
  • Computes N values (no ReLU)

  • N outputs are weighted sums of these values

  • Weights depend on the inputs themselves

16 of 57

Attention as routing

17 of 57

Attention as routing

18 of 57

Attention as routing

19 of 57

Attention weights

  • Compute N “queries” and N “keys” from input

  • Calculate similarity and pass through softmax:

  • Weights depend on the inputs themselves

20 of 57

Attention weights

  • Compute N “queries” and N “keys” from input

  • Take dot products and pass through softmax:

  • Weights depend on the inputs themselves

21 of 57

Dot product = measure of similarity

22 of 57

Motivation

Conclusions:�

  • We need a model where parameters don’t increase with input length

  • There must be connections between the words.
  • The strength of these connections will depend on the words themselves.

Design neural network to encode and process text:

23 of 57

Transformers

  • Motivation
  • Dot-product self-attention
  • Matrix form
  • The transformer
  • NLP pipeline
  • Decoders
  • Large Language models

24 of 57

Matrix form

  • Store N input vectors in matrix X

  • Compute values, queries and keys (all dims are (D x 1) x (1 x N) + D x D x (D x N)):�

  • Combine self-attentions (D x N x ColSoftmax( N x D x D x N ))

25 of 57

Matrix form

26 of 57

Positional encoding

Self-attention is equivariant to permuting word order

�But word order is important in language:

The man ate the fish

vs.

The fish ate the man

27 of 57

Positional encoding

28 of 57

Positional encoding

29 of 57

Multi-head self-attention

30 of 57

Why do we learn attention weights?

31 of 57

Transformers

  • Motivation
  • Dot-product self-attention
  • Matrix form
  • The transformer
  • NLP pipeline
  • Encoders, decoders, and encoder-decoders
  • Transformers for vision

32 of 57

Transformers

  • Motivation
  • Dot-product self-attention
  • Matrix form
  • The transformer
  • NLP pipeline
  • Decoders
  • Large Language models

33 of 57

The transformer

34 of 57

Transformers

  • Motivation
  • Dot-product self-attention
  • Matrix form
  • The transformer
  • NLP pipeline
  • Decoders
  • Large Language models

35 of 57

Tokenizer

Goal: Tokenizer chooses input “units”

  • Inevitably, some words (e.g., names) will not be in the vocabulary.
  • It’s not clear how to handle punctuation
  • The vocabulary would need different tokens for versions of the same word with different suffixes (e.g., walk, walks, walked, walking) and there is no way to clarify that these variations are related

Solution: Sub-word tokenization�One particular algorithm: Byte pair encoding

36 of 57

37 of 57

38 of 57

39 of 57

40 of 57

41 of 57

42 of 57

Learning vocabulary

”One hot encoding”

43 of 57

Transformers

  • Motivation
  • Dot-product self-attention
  • Matrix form
  • The transformer
  • NLP pipeline
  • Decoders
  • Large Language models

44 of 57

Three types of transformer layer

  • Encoder (BERT)
  • Decoder (GPT3) will only look at this today
  • Encoder-decoder (Translation)

45 of 57

Decoder model

46 of 57

Decoder model: GPT3

  • One job: predict the next word in a sequence
  • More formally builds an autoregressive probability model

47 of 57

Decoder model: GPT3

  • Builds autoregressive probability model
  • E.g. “It takes great courage to let yourself appear weak”

48 of 57

Predicting all next words simultaneously

49 of 57

Masked self-attention

50 of 57

Transformers

  • Motivation
  • Dot-product self-attention
  • Matrix form
  • The transformer
  • NLP pipeline
  • Decoders
  • Large Language models

51 of 57

GPT3 (Brown et al. 2020)

  • Sequence lengths are 2048 tokens long
  • Batch size is 3.2 million tokens.
  • 96 transformer layers (some of which implement a sparseversion of attention), each of which processes a word embedding of size 12288.
  • 96 heads in the self-attention layers and the value, query, and key dimension is 128.
  • 300 billion tokens
  • 175 billion parameters

52 of 57

try some out at

https://umgpt.umich.edu/

53 of 57

What does it learn?

  • Syntax

“Tomorrow, let’s <VERB>...”

  • General knowledge:

“The train pulled into the …”

54 of 57

Text completion

Understanding Deep Learning is a new textbook from MIT Press by Simon Prince that's designed to offer an accessible, broad introduction to the field. Deep learning is a branch of machine learning that is concerned with algorithms that learn from data that is unstructured or unlabeled. The book is divided into four sections:

  • Introduction to deep learning
  • Deep learning architecture
  • Deep learning algorithms
  • Applications of deep learning

The first section offers an introduction to deep learning, including its history and origins. The second section covers deep learning architecture, discussing various types of neural networks and their applications. The third section dives into deep learning algorithms, including supervised and unsupervised learning, reinforcement learning, and more. The fourth section applies deep learning to various domains, such as computer vision, natural language processing, and robotics.

55 of 57

Few shot learning:

56 of 57

ChatGPT

  • GPT3.5 fine-tuned with human annotations
  • Trained to predict the next word + be “helpful, honest, harmless”

57 of 57

ChatGPT

  • GPT3.5 fine-tuned with human annotations
  • Trained to predict the next word + be “helpful, honest, harmless”