1 of 62

CS 162: Natural Language Processing

Saadia Gabriel

Lecture 8:

Language Modeling Part 5

2 of 62

Announcements

  • Reminder that homework #1 is due this Friday!
  • I’ve posted a sign-up sheet under week 4 for group presentations of final project papers (early sign-ups welcome)

3 of 62

Quiz #3

https://forms.gle/MxUjCaDb3rEeceau5

4 of 62

The weights are the same (same transformation of the input), but the resulting query, word and value vectors are word-specific.

Increasing the complexity of the network usually improves performance because more information can be captured (same rationale for why we use multiple, varying numbers of attention heads).

5 of 62

Yes, efficient NLP is a whole area of research which you can explore for your final project! Often components or entire layers can be removed for certain tasks (pruning), or a smaller model can be trained to mimic the behavior of a larger model (knowledge distillation).

You can also reduce the precision of numbers used to store the model’s parameters (quantization) to save memory.

6 of 62

Computational Limitations of Self-attention

  • The computational cost of self-attention is O(n2d), where n is the sequence length and d is the embedding dimension size.
  • This limits the length of sequences we can feasibly process!
  • In practice, there must be a relatively short fixed context size:

7 of 62

Good question. BPE is used in most modern models, though there are now a number of tokenization variations and some may use a variation.

On a side note, you can use the tiktoken package to see how GPT models will tokenize an input prompt.

Saadia is teaching a class right now.

Sa

adia

is

teaching

a

class

right

now

.

Sparse attention, key-value caching and other modifications allow us to have longer context sizes.

8 of 62

9 of 62

Goals of This Lecture

  • How do we handle limitations of self-attention? (information loss, inefficiency)

​

  • We’ll motivate model pretraining from word embeddings

​

  • We’ll dive into pretraining model architectures
    • Unidirectional
    • Independently bidirectional
    • Bidirectional

​

  • We’ll discuss scaling laws for language models

​

  • We’ll introduce large language models

Following slides courtesy of Violet Peng and Jia-Chen Gu

10 of 62

Residual Connections

  • Issue of information loss
  • Self-attention can decide not to attend to itself
  • Add input back again in each sublayer!

11 of 62

Residual Connections

12 of 62

Say we want to do movie review classification with our transformer.

​

Should we train from scratch on our labeled movie review dataset?

13 of 62

Language Model

The Illustrated GPT-2, https://jalammar.github.io/illustrated-gpt2/

14 of 62

Word Embeddings

15 of 62

Contextualized Representations

16 of 62

The Pre-train

Fine-tune Paradigm

Image from Gu et al. (2020)

17 of 62

Why Pre-training? 

18 of 62

The Pretraining Revolution

  • Pretraining has had a significant and tangible impact on how well NLP systems work

19 of 62

Scale Unsupervised Learning and Generalize

Parametric knowledge

acquisition

20 of 62

Pre-training Model Architectures

Jacob Devlin, et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL 2018.

21 of 62

Unidirectional Pre-training

  • Decoder-only and cannot condition on future words: Processes input sequences in a left-to-right manner, so only considers the preceding context when predicting a word (well-suited to text generation)
  • See Improving Language Understanding by Generative Pre-Training (GPT), OpenAI (2018)

22 of 62

GPT Pre-training and Fine-tuning

23 of 62

GPT Architecture

https://jalammar.github.io/illustrated-gpt2/

24 of 62

GPT Technical Details

25 of 62

Fine-tuning GPT for Downstream Tasks

26 of 62

GPT for Natural Language Inference (Entailment)

27 of 62

GPT Results on NLI

28 of 62

Effect of Pre-training in GPT

29 of 62

Effect of Pre-training in GPT

You don’t even have to fine-tune! Tasks can be performed “zero-shot” meaning the task instructions are specified in natural language at test-time in a prompt:

30 of 62

Effect of Pre-training in GPT

31 of 62

Stronger GPT-2

Actually 124M, this was an error in the original release

32 of 62

GPT-2 Results

33 of 62

Stronger GPT-3

https://ourworldindata.org/grapher/artificial-intelligence-training-computation

34 of 62

GPT-3 with In-context Learning

35 of 62

GPT-3 with In-context Learning

36 of 62

Independently Bidirectional Pre-training

37 of 62

ELMo Technical Details and Results

38 of 62

Limitations of ELMo and GPT series

Elmo: Shallow bi-directionality

GPT: no bi-directionality

39 of 62

Bidirectional Pre-training

40 of 62

Masked Language Models

41 of 62

Next Sentence Prediction

BERT was also trained on a next sentence prediction objective to improve its capabilities for capturing long-range dependencies.

42 of 62

Next Sentence Prediction

BERT performance with/without next sentence prediction (NSP)

Devlin et al. (2019)

43 of 62

BERT Input Representation

44 of 62

BERT Technical Details

45 of 62

BERT Fine-tuning

46 of 62

BERT Results

47 of 62

Extension of BERT

48 of 62

Domain-specific BERTs

49 of 62

Extension of Pre-training Encoder-Decoders

50 of 62

Resources for Coding

51 of 62

Scaling Laws

52 of 62

Scaling Laws

53 of 62

*** Discuss ***

Does it matter how we scale our LLMs?

54 of 62

Scaling Laws

55 of 62

Large Language Models and Beyond

Image from https://blogs.cfainstitute.org/investor/2023/05/26/chatgpt-and-large-language-models-six-evolutionary-steps/

Many components of LLM development have been scaling up….

56 of 62

Training LLMs: Data

57 of 62

Training LLMs: Context Size

58 of 62

Next Time…

59 of 62

Post-training

60 of 62

LLM Hallucination

Why does this happen and what can we do about it?

61 of 62

In general, LLM Safety

62 of 62

Lecture Take-aways