1 of 43

CSCI-SHU 376: Natural Language Processing

Hua Shen

2026-09-21

Fall 2026

Lecture 6: RNNs and LSTMs

​

2 of 43

Today’s Plan

  • Neural Networks
  • Recurrent neural networks (RNNs)
  • Long Short-Term Memory RNNs (LSTM)
  • Gated Recurrent Units (GRUs)

3 of 43

Feedforward Neural Networks

  • Feedforward NN with a single hidden layer
  • Input vector x, output probability distribution y

​

​

​

​

4 of 43

Recurrent neural networks (RNNs)

  • A family of neural networks that can handle variable length inputs
  • Apply the same weights repeatedly at different positions

​

​

​

5 of 43

Recurrent neural networks (RNNs)

  • Highly effective approach for sequence tasks
  • Better than HMM: no Markov assumption

​

​

​

6 of 43

Recurrent neural networks (RNNs)

 

7 of 43

Recurrent neural networks (RNNs)

8 of 43

RNNs vs Feedforward NNs

9 of 43

Recurrent neural language models

No Markov assumption!

10 of 43

Recurrent neural language models

No Markov assumption!

11 of 43

Weight Tying

No Markov assumption!

12 of 43

LM Performance

  • Perplexity on the Penn Treebank (PTB) dataset

​

13 of 43

RNNs: Pros and Cons

Cons

  • Recurrent computation is slow
  • In practice, difficult to access information from many steps back

​

Pros

  • In theory, can process any length input
  • Model size does not increase for longer input text

​

14 of 43

Training RNNLMs

  • Forward pass + backward pass

​

Correct Next Word

​

Accumulate loss

​

15 of 43

Training RNNLMs - BPTT

  • Backpropagation Through Time

​

16 of 43

Training RNNLMs - BPTT

  • Backpropagation Through Time

​

17 of 43

Today’s Plan

  • Recurrent neural networks (RNNs)
  • Long Short-Term Memory RNNs (LSTM)
  • Gated Recurrent Units (GRUs)

​

​

​

​

​

​

18 of 43

Recurrent neural networks (RNNs)

19 of 43

Multi-layer RNNs

  • RNNs are already ”deep” on one dimension
  • We can also make them “deep” in another direction by applying multiple RNNs
  • 2-4 layers are common (and are better than 1 layer)

​

20 of 43

Bidirectional RNNs

  • Bidirectionality is important in language representations

​

21 of 43

Bidirectional RNNs

  • Bidirectionality is important in language representations

​

22 of 43

Bidirectional RNNs

  • Sequence tagging: yes
  • Text classification: yes
  • Text generation: no

​

23 of 43

Training RNNLMs - BPTT

  • Backpropagation Through Time

​

24 of 43

Training RNNLMs - BPTT

  • Gradient too large (Gradient exploding): difficult to converge
  • Gradient too small (Vanishing Gradient):
    • Can’t capture long-term dependencies
    • May capture a wrong recent dependency

​

25 of 43

Gradient Clipping

  • Gradient clipping: if the norm of gradient is greater than some threshold, scale it down before applying SGD update.

​

26 of 43

Vanishing gradient

  • When some gradients are small, the signal gets smaller and smaller as it backpropagates further.
  • So, model weights are basically updated only with respect to near effects, not long-term effects

27 of 43

Today’s Plan

  • Recurrent neural networks (RNNs)
  • Long Short-Term Memory RNNs (LSTM)
  • Gated Recurrent Units (GRUs)

​

​

​

​

​

​

28 of 43

Long Short-Term Memory RNNs (LSTMs)

  • A type of RNN as a solution to the vanishing gradients problem
  • LSTM does not guarantee that there is no vanishing / exploding gradient, but it does provide an easier way for the model to learn long-distance dependencies

29 of 43

LSTMs: Intuition

 

30 of 43

LSTMs: the formulation

31 of 43

LSTMs: the formulation

  • LSTMs has 4x parameters compared to simple RNNs

​

32 of 43

Visualizing LSTMs

33 of 43

Visualizing LSTMs

34 of 43

Visualizing LSTMs

  • Remove or add information to the cell state

35 of 43

Visualizing LSTMs

Forget gate

36 of 43

Visualizing LSTMs

Input gate

New memory Cell

37 of 43

Visualizing LSTMs

Final memory cell

38 of 43

Visualizing LSTMs

Final hidden cell

39 of 43

Today’s Plan

  • Recurrent neural networks (RNNs)
  • Long Short-Term Memory RNNs (LSTM)
  • Gated Recurrent Units (GRUs)

​

​

​

​

​

​

40 of 43

Gated Recurrent Units (GRUs)

  • Simplified 3 gates to 2 gates: reset gate and update gate, without an explicit cell state

​

41 of 43

Gated Recurrent Units (GRUs)

42 of 43

LSTMs vs GRUs

Music Modeling

43 of 43

LSTMs vs GRUs

Speech Signal modeling