1 of 32

Long Short-Term Memory Networks

Dr. Dinesh K. Vishwakarma

PROFESSOR, DEPARTMENT OF INFORMATION TECHNOLOGY

DELHI TECHNOLOGICAL UNIVERSITY, DELHI.

Webpage: http://www.dtu.ac.in/Web/Departments/InformationTechnology/faculty/dkvishwakarma.php

Email: dinesh@dtu.ac.in

2 of 32

Outlines

    • Introduction of LSTM
    • Structure of LSTM
    • Training of LSTM
    • Summary of LSTM
    • Use cases

Dinesh K. Vishwakarma, Ph.D.

2

4/29/26

Applications of LSTM includes:

  • Time series prediction
  • Speech recognition
  • Rhythm learning
  • Music composition
  • Grammar learning
  • Handwriting recognition
  • Human action recognition
  • Sign language translation
  • Protein homology detection
  • Predicting subcellular localization of proteins
  • Time series anomaly detection
  • Several prediction tasks in the area of business process management
  • Airport passenger management
  • Short-term traffic forecast
  • Market Prediction

3 of 32

What is LSTM?

Dinesh K. Vishwakarma, Ph.D.

3

4/29/26

4 of 32

RNN Cell

Dinesh K. Vishwakarma, Ph.D.

4

4/29/26

5 of 32

Problem with RNN

    • The Problem of Long-Term Dependencies
      • RNN Suffers with Long-Term Dependencies such as:

Dinesh K. Vishwakarma, Ph.D.

5

4/29/26

The colour of tree is “Green”

Performs Good

I live in Kashmir, India….where I like the weather and generally it is “cold”

Performs Poor

Distance between actual context and prediction is very less

6 of 32

LSTM Structure

    • LSTM have chain-like neural network layer. In a standard RNN, the repeating module consists of one single function as shown in the below figure.

    • tanh is squashing function, which makes values between -1 to 1.

Dinesh K. Vishwakarma, Ph.D.

6

4/29/26

7 of 32

LSTM Structure…

    • Each of the functions in the layers has their own structures. The cell state is the horizontal line and it acts like a conveyer belt carrying certain data linearly across the data channel. A simple LSTM consist of 4-gates.

Dinesh K. Vishwakarma, Ph.D.

7

4/29/26

 

Forget Gate

Input Gate

Output Gate

8 of 32

LSTM Structure…

    • The 1st step is to identify, the information which is not required and will be thrown away from the cell state. It is done by a sigmoid layer called as forget gate layer.
    • Forget Gate: After getting the output of previous stateh(t-1), Forget gate helps us to take decisions about what must be removed from h(t-1) state and thus keeping only relevant stuff. It is surrounded by a sigmoid function which helps to crush the input between [0,1].
    • The calculation is done by considering the new input and the previous timestamp which eventually leads to the output of a number between 0 and 1 for each number in that cell state. As typical binary, 1 represents to keep the cell state while 0 represents to trash it.

Dinesh K. Vishwakarma, Ph.D.

8

4/29/26

9 of 32

LSTM Structure…

  •  

Dinesh K. Vishwakarma, Ph.D.

9

4/29/26

10 of 32

LSTM Structure…

  •  

Dinesh K. Vishwakarma, Ph.D.

10

4/29/26

11 of 32

LSTM Structure…

    • In 4th Step, a sigmoid layer is used which decides what parts of the cell state is going to output. It is also called Output Gate.
    • Then, the cell state is kept through tanh (push the values to be between −1 and 1).
    • Later, we multiply it by the output of the sigmoid gate, so that we only output the parts we decided to.
    • The output consists of only the outputs there were decided to be carry forwarded in the previous steps and not all the outputs at once.

Dinesh K. Vishwakarma, Ph.D.

11

4/29/26

12 of 32

LSTM WORKING (STEP BY STEP)

12

https://colah.github.io/posts/2015-08-Understanding-LSTMs/

1

 

What information we’re going to throw away from the cell state

Ex. Language Modelling

Want to add new Gender, then Forget the old one

2

 

 

What new information we’re going to store in the cell state

Ex. Language Modelling

Add the gender of new subject to the cell state

3

 

How much we decided to update each state value

Ex. Language Modelling

Drop old gender and add new Gender

4

 

 

Finally, we need to decide what we’re going to output

Ex. Language Modelling

Based on subject (singular or plural) a relevant verb may be added

13 of 32

LSTM

Dinesh K. Vishwakarma, Ph.D.

13

4/29/26

14 of 32

RNN & LSTM

Dinesh K. Vishwakarma, Ph.D.

14

4/29/26

15 of 32

Summary of LSTM

    • During backpropagation through each LSTM cell, it’s multiplied by different values of forget gate, which makes it less prone to vanishing/exploding gradient.
    • Though, if values of all forget gates are less than 1, it may suffer from vanishing gradient but in practice people tend to initialise the bias terms with some positive number so in the beginning of training f (forget gate) is very close to 1 and as time passes the model can learn these bias terms.
    • In the first step, we found out what was needed to be dropped.
    • The second step consisted of what new inputs are added to the network.
    • The third step was to combine the previously obtained inputs to generate the new cell states.
    • Lastly, we arrived at the output as per requirement.

Dinesh K. Vishwakarma, Ph.D.

15

4/29/26

16 of 32

A case study: LSTM

    • Consider to predict the next word in a sample short story.
    • Start by feeding an LSTM Network with correct sequences from the text of 3 symbols as inputs and 1 labeled symbol.

Dinesh K. Vishwakarma, Ph.D.

16

4/29/26

17 of 32

Dataset

    • The LSTM is trained using a sample short story which consists of 112 unique symbols. Comma and period are also considered as unique symbols in this case.
    • “long ago , the mice had a general council to consider what measures they could take to outwit their common enemy , the cat . some said this , and some said that but at last a young mouse got up and said he had a proposal to make , which he thought would meet the case . you will all agree , said he , that our chief danger consists in the sly and treacherous manner in which the enemy approaches us . now , if we could receive some signal of her approach , we could easily escape from her . i venture , therefore , to propose that a small bell be procured , and attached by a ribbon round the neck of the cat . by this means we should always know when she was about , and could easily retire while she was in the neighborhood . this proposal met with general applause , until an old mouse got up and said that is all very well , but who is to bell the cat ? the mice looked at one another and nobody spoke . then the old mouse said it is easy to propose impossible remedies .”

Dinesh K. Vishwakarma, Ph.D.

17

4/29/26

18 of 32

Training

    • LSTMs can only understand real numbers. So, the first requirement is to convert the unique symbols into unique integer values based on the frequency of occurrence.
    • Doing this will create a customized dictionary that we can make use of later on to map the values.
    • The network will create a 112-element vector consisting of the probability of occurrence of each of these unique integer values.

Dinesh K. Vishwakarma, Ph.D.

18

4/29/26

19 of 32

Variants of LSTM

    • A normal LSTM is discussed in previous topic. But not all LSTMs are the same as we discussed.
    • In fact, most of the paper uses slightly different version of LSTM. The differences are minor, but it’s worth mentioning some of them.
    • One popular LSTM variant, introduced by Gers & Schmidhuber (2000), is adding “peephole connections.” This means that we let the gate layers look at the cell state.
    • The diagram adds peepholes to all the gates, but many papers will give some peepholes and not others.

Dinesh K. Vishwakarma, Ph.D.

19

4/29/26

20 of 32

Variants of LSTM…

    • Another variation is to use coupled forget and input gates.
    • Instead of separately deciding what to forget and what we should add new information to, we make those decisions together.
    • We only forget when we’re going to input something in its place.
    • We only input new values to the state when we forget something older.

Dinesh K. Vishwakarma, Ph.D.

20

4/29/26

21 of 32

Variants of LSTM…GRU

    • A slightly more dramatic variation on the LSTM is the Gated Recurrent Unit, or GRU, introduced by Cho, et al. (2014).
    • Aims to solve the vanishing gradient problem
    • It combines the forget and input gates into a single “update gate.”
    • It also merges the cell state and hidden state, and makes some other changes.
    • The resulting model is simpler than standard LSTM models, and has been growing increasingly popular.

Dinesh K. Vishwakarma, Ph.D.

21

4/29/26

22 of 32

GRU vs LSTM

Dinesh K. Vishwakarma, Ph.D.

22

4/29/26

23 of 32

GRU unit

Dinesh K. Vishwakarma, Ph.D.

23

4/29/26

24 of 32

Update Gate

    • Update gate: the update gate z_t for time step t using the formula:

Dinesh K. Vishwakarma, Ph.D.

24

4/29/26

The update gate helps the model to determine how much of the past information (from previous time steps) needs to be passed along to the future.

25 of 32

#2. Reset gate

    • This gate is used from the model to decide how much of the past information to forget.

Dinesh K. Vishwakarma, Ph.D.

25

4/29/26

26 of 32

#3. Current memory content

Dinesh K. Vishwakarma, Ph.D.

26

4/29/26

A new memory content which will use the reset gate to store the relevant information from the past.

27 of 32

#4. Final memory at current time step

Dinesh K. Vishwakarma, Ph.D.

27

4/29/26

  • z_t — green line is used to calculate 1-z_t which, combined with h’_t — bright green line, produces a result in the dark red line.
  • z_t is also used with h_(t-1) — blue line in an element-wise multiplication.

Finally, h_t — blue line is a result of the summation of the outputs corresponding to the bright and dark red lines

28 of 32

GRU

Dinesh K. Vishwakarma, Ph.D.

28

4/29/26

29 of 32

LSTM and GRU

Dinesh K. Vishwakarma, Ph.D.

29

4/29/26

Amazing! This box of cereal gave me a perfectly balanced breakfast, as all things should be. I only eat half of it but will definitely be buying again!

Amazing! This box of cereal gave me a perfectly balanced breakfast, as all things should be. I only eat half of it but will definitely be buying again!

30 of 32

Summary of LSTM Variants

    • These are only a few of the most notable LSTM variants. There are lots of others, like Depth Gated RNNs by Yao, et al. (2015) and Clockwork RNNs by Koutnik, et al. (2014).
    • A nice comparison is made by Greff, et al. (2015) of popular variants, finding that they’re all about the same. Jozefowicz, et al. (2015) tested more than ten thousand RNN architectures, finding some that worked better than LSTMs on certain tasks.
    • The next step in LSTMs are attention. The idea is to let every step of an RNN pick information to look at from some larger collection of information. For example, if you are using an RNN to create a caption describing an image, it might pick a part of the image to look at for every word it outputs.
    • Attention isn’t the only exciting thread in RNN research. For example, Grid LSTMs by Kalchbrenner, et al. (2015) seem extremely promising. Work using RNNs in generative models – such as Gregor, et al. (2015)Chung, et al. (2015), or Bayer & Osendorfer (2015) – also seems very interesting.
    • The last few years have been an exciting time for recurrent neural networks, and the coming ones promise to only be more so!

Dinesh K. Vishwakarma, Ph.D.

30

4/29/26

31 of 32

References

Dinesh K. Vishwakarma, Ph.D.

31

4/29/26

32 of 32

Thank You�Contact: dinesh@dtu.ac.in �Mobile: +91-9971339840

Dinesh K. Vishwakarma, Ph.D.

32

4/29/26