1 of 36

CSCI-SHU 376: Natural Language Processing

Hua Shen

2026-02-26

Spring 2026

Lecture 8: Transformers

​

2 of 36

Today’s Plan

  • Transformers

​

3 of 36

Recap: Seq2seq

  • Encoder-decoder Structure
  • The Encoder is a RNN to read the input sentence
  • The Decoder is another RNN to generate output word by word

​

4 of 36

Issues with RNNs

  • RNN is a sequential model: Hard to parallelize
  • Future RNN hidden states cannot be computed before past RNN states

​

​

​

5 of 36

Issues with RNNs

  • It’s still not easy to capture long-context information: due to vanish gradient

​

​

​

6 of 36

Attention is all you need

  • Do we really need RNNs?

​

​

​

7 of 36

Transformers

  • Encoder + Decoder
  • Originally proposed for NMT, and later adapted to all NLP tasks
  • No recurrent structures!
  • Both encoder and decoder have N layers
    • Each encoder layer has 2 sublayers
    • Each decoder layer has 3 sublayers
    • Key point: Multi-head Attention

​

7

8 of 36

Transformers

  • Self-attention
  • Multi-head self-attention
  • Transformer encoder
  • Transformer decoder
  • Other components

​

8

9 of 36

Recap: Attention

10 of 36

Understanding Attention Intuitions

  • Intuition: The weighted sum is a selective summary of the information contained in the values, where the query determines which values to focus on
  • Obtain a fixed-size representation of an arbitrary set of representations, dependent on some other representation
  • Parallelizable

10

11 of 36

Self-Attention

  • From: each state (i.e. input token)
  • To: All other tokens in the sequence

​

  • Attention from the sequence to itself

​

11

12 of 36

Query, Key and Value in Self-Attention

  • Query: asking for information
  • Key: saying it has some information
  • Value: giving the information

​

12

13 of 36

Self-attention in equations

 

 

  • Finally, compute the weighted sum:

​

​

​

​

​

13

14 of 36

Self-attention in Matrix Notations

14

15 of 36

Self-attention in Matrix Notations

15

16 of 36

Multi-Head Attention

16

17 of 36

Multi-Head Attention

  • In practice, we use a reduced dimension for each head

​

​

​

​

  • The total computation cost if similar to single-head attention with full dimensionality

​

​

​

​

17

18 of 36

Multi-Head Attention

18

19 of 36

Positional Encoding

  • Transformer does not have recurrence
  • Include order of tokens!
  • People just use a learnable embedding for every unique position

​

19

20 of 36

Feed-forward Blocks

  • Attention: Gather information from other tokens
  • FFN: Process this information

​

  • There is no elementwise nonlinearities in self-attention; stacking more self-attention just re-average value vectors

​

20

21 of 36

Residual Connections

  • Allow stacking multiple layers

​

21

22 of 36

Layer Norm

  • A trick to help models train faster
  • Normalize vector representation in batch
  • Idea: cut down on uninformative variation in hidden vector values

​

22

23 of 36

Transformers encoder

  • Each encoder layer has two sub-layers:
    • A multi-head self-attention layer
    • A feedforward layer

​

  • Residual connection
  • Layer normalization

​

23

24 of 36

Transformers decoder

  • Each decoder layer has three sub-layers:
    • A masked multi-head self-attention layer
    • A multi-head cross-attention layer
    • A feedforward layer

​

  • Residual connection
  • Layer normalization

​

24

25 of 36

Masked multi-head attention

  • Key point: you can’t see the future words for the decoder!

​

 

25

26 of 36

Multi-head cross-attention

26

27 of 36

Training Transformer

  • Training data: Parallel Corpus
  • Loss: Cross Entropy
  • Back-propagate gradients through both encoder and decode

​

​

27

28 of 36

Summary: Transformer

28

29 of 36

Summary: Transformer

29

30 of 36

Summary: Transformer

30

31 of 36

Transformers: Machine Translation

31

32 of 36

Transformers: document generation

32

33 of 36

Transformers: pros and cons

  • Easier to capture dependencies: attention between every pair of words

​

  • Also easier to parallelize

​

  • Quadratic computation in self-attention
    • Can be quite slow when the sequence length is large

​

33

34 of 36

Work on improving on quadratic self-attention cost

 

34

35 of 36

Still Challenging!

  • Surprisingly, most modifications do not meaningfully improve performance.

​

35

36 of 36

Transformer-XL

  • Relative positional encoding + Segment recurrence for long context

​

36