1 of 62

CS 162: Natural Language Processing

Saadia Gabriel

Lecture 8:

Language Modeling Part 5

2 of 62

Announcements

  • Reminder that homework #1 is due this Friday!
  • I’ve posted a sign-up sheet under week 4 for group presentations of final project papers (early sign-ups welcome)

3 of 62

Quiz #3

https://forms.gle/MxUjCaDb3rEeceau5

4 of 62

The weights are the same (same transformation of the input), but the resulting query, word and value vectors are word-specific.

Increasing the complexity of the network usually improves performance because more information can be captured (same rationale for why we use multiple, varying numbers of attention heads).

5 of 62

Yes, efficient NLP is a whole area of research which you can explore for your final project! Often components or entire layers can be removed for certain tasks (pruning), or a smaller model can be trained to mimic the behavior of a larger model (knowledge distillation).

You can also reduce the precision of numbers used to store the model’s parameters (quantization) to save memory.

6 of 62

Computational Limitations of Self-attention

  • The computational cost of self-attention is O(n2d), where n is the sequence length and d is the embedding dimension size.
  • This limits the length of sequences we can feasibly process!
  • In practice, there must be a relatively short fixed context size:

7 of 62

Good question. BPE is used in most modern models, though there are now a number of tokenization variations and some may use a variation.

On a side note, you can use the tiktoken package to see how GPT models will tokenize an input prompt.

Saadia is teaching a class right now.

Sa

adia

is

teaching

a

class

right

now

.

Sparse attention, key-value caching and other modifications allow us to have longer context sizes.

8 of 62

9 of 62

Goals of This Lecture

  • How do we handle limitations of self-attention? (information loss, inefficiency)

  • We’ll motivate model pretraining from word embeddings

  • We’ll dive into pretraining model architectures
    • Unidirectional
    • Independently bidirectional
    • Bidirectional

  • We’ll discuss scaling laws for language models

  • We’ll introduce large language models

Following slides courtesy of Violet Peng and Jia-Chen Gu

10 of 62

Residual Connections

  • Issue of information loss
  • Self-attention can decide not to attend to itself
  • Add input back again in each sublayer!

11 of 62

Residual Connections

12 of 62

Say we want to do movie review classification with our transformer.

Should we train from scratch on our labeled movie review dataset?

13 of 62

Language Model

The Illustrated GPT-2, https://jalammar.github.io/illustrated-gpt2/

14 of 62

Word Embeddings

15 of 62

Contextualized Representations

16 of 62

The Pre-train

Fine-tune Paradigm

Image from Gu et al. (2020)

17 of 62

Why Pre-training? 

18 of 62

The Pretraining Revolution

  • Pretraining has had a significant and tangible impact on how well NLP systems work

19 of 62

Scale Unsupervised Learning and Generalize

Parametric knowledge

acquisition

20 of 62

Pre-training Model Architectures

Jacob Devlin, et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL 2018.

21 of 62

Unidirectional Pre-training

  • Decoder-only and cannot condition on future words: Processes input sequences in a left-to-right manner, so only considers the preceding context when predicting a word (well-suited to text generation)
  • See Improving Language Understanding by Generative Pre-Training (GPT), OpenAI (2018)

22 of 62

GPT Pre-training and Fine-tuning

23 of 62

GPT Architecture

https://jalammar.github.io/illustrated-gpt2/

24 of 62

GPT Technical Details

25 of 62

Fine-tuning GPT for Downstream Tasks

26 of 62

GPT for Natural Language Inference (Entailment)

27 of 62

GPT Results on NLI

28 of 62

Effect of Pre-training in GPT

29 of 62

Effect of Pre-training in GPT

You don’t even have to fine-tune! Tasks can be performed “zero-shot” meaning the task instructions are specified in natural language at test-time in a prompt:

30 of 62

Effect of Pre-training in GPT

31 of 62

Stronger GPT-2

Actually 124M, this was an error in the original release

32 of 62

GPT-2 Results

33 of 62

Stronger GPT-3

https://ourworldindata.org/grapher/artificial-intelligence-training-computation

34 of 62

GPT-3 with In-context Learning

35 of 62

GPT-3 with In-context Learning

36 of 62

Independently Bidirectional Pre-training

37 of 62

ELMo Technical Details and Results

38 of 62

Limitations of ELMo and GPT series

Elmo: Shallow bi-directionality

GPT: no bi-directionality

39 of 62

Bidirectional Pre-training

40 of 62

Masked Language Models

41 of 62

Next Sentence Prediction

BERT was also trained on a next sentence prediction objective to improve its capabilities for capturing long-range dependencies.

42 of 62

Next Sentence Prediction

BERT performance with/without next sentence prediction (NSP)

Devlin et al. (2019)

43 of 62

BERT Input Representation

44 of 62

BERT Technical Details

45 of 62

BERT Fine-tuning

46 of 62

BERT Results

47 of 62

Extension of BERT

48 of 62

Domain-specific BERTs

49 of 62

Extension of Pre-training Encoder-Decoders

50 of 62

Resources for Coding

51 of 62

Scaling Laws

52 of 62

Scaling Laws

53 of 62

*** Discuss ***

Does it matter how we scale our LLMs?

54 of 62

Scaling Laws

55 of 62

Large Language Models and Beyond

Image from https://blogs.cfainstitute.org/investor/2023/05/26/chatgpt-and-large-language-models-six-evolutionary-steps/

Many components of LLM development have been scaling up….

56 of 62

Training LLMs: Data

57 of 62

Training LLMs: Context Size

58 of 62

Next Time…

59 of 62

Post-training

60 of 62

LLM Hallucination

Why does this happen and what can we do about it?

61 of 62

In general, LLM Safety

62 of 62

Lecture Take-aways