1 of 99

Machine Learning�Week 10 – Deep Learning (3)

Seungtaek Choi

Division of Language & AI at HUFS

seungtaek.choi@hufs.ac.kr

2 of 99

Announcement

  • Final project presentation: Dec 09

Week

W9

W10

W11

W12

W13

W14

W15

W16

Assign#3 (project proposal)

start

end

Assign#4 (data collection and analysis)

start

end

Assign#5 (model training and evaluation)

start

end

Assign#6 (real usage & final report)

start

end

Assign#7 (presentation)

lecture

Lecture Topic

Unsup. Learning

Unsup.�Learning

Deep Learning

Deep Learning

Advanced�Topics

Advanced�Topics

Final�Project

Final Exam

3 of 99

Deep Learning

Convolutional Neural Networks

4 of 99

What Computers “See”?

  • Images are numbers
    • An image is just a matrix of numbers [0, 255]!
    • i.e., 1080x1080x3 for an RGB image

5 of 99

Tasks in Computer Vision

  • Regression: output variable takes continuous value
  • Classification: output variable takes class label. Can produce probability of belonging to a particular class

Input Image

Pixel Representation

Lincoln

Washington

Jefferson

Obama

Trump

0.8

0.05

0.05

0.01

0.09

classification

6 of 99

“Learning” Feature Representations

  • Motivation
    • The bird occupies a local area and looks the same in different parts of an image.
    • We should construct neural networks which exploit these properties.

7 of 99

Fully Connected Neural Network

  • Input:
    • 2D image
    • Vector of pixel values
  • Fully Connected:
    • Connect neuron in hidden layer to all neurons in input layer
    • No spatial information!
      • Spatial organization of the input is destroyed by flatten.
    • And, many, many parameters!

  • How can we use spatial structure in the input to inform the�architecture of the network?

8 of 99

Fully Connected Layer

9 of 99

Locally Connected Layer

10 of 99

Convolutional Layer

11 of 99

Key Idea

  • A standard neural net applied to images:
    • Scales quadratically with the size of the input
    • Does not leverage stationarity

  • Solution:
    • Connect each hidden unit to a small patch of the input
    • Share the weight across space

  • This is called: convolutional layer.
  • A network with convolutional layers is called convolutional network.

12 of 99

The Convolution Operation

 

Image

Kernel

Feature Map

13 of 99

Feature Extraction with Convolution: A Case Study

  • Features of X

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

1

-1

-1

-1

-1

-1

1

-1

-1

-1

1

-1

-1

-1

1

-1

-1

-1

-1

-1

1

-1

1

-1

-1

-1

-1

-1

-1

-1

1

-1

-1

-1

-1

-1

-1

-1

1

-1

1

-1

-1

-1

-1

-1

1

-1

-1

-1

1

-1

-1

-1

1

-1

-1

-1

-1

-1

1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

1

-1

-1

-1

1

-1

-1

-1

1

-1

-1

-1

-1

-1

1

1

-1

1

-1

-1

-1

-1

-1

-1

-1

1

-1

-1

-1

-1

-1

-1

-1

1

-1

1

1

-1

-1

-1

-1

-1

1

-1

-1

-1

1

-1

-1

-1

1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

14 of 99

Feature Extraction with Convolution: A Case Study

  • Filters to detect X features

1

-1

-1

-1

1

-1

-1

-1

1

1

-1

1

-1

1

-1

1

-1

1

-1

-1

1

-1

1

-1

1

-1

-1

15 of 99

Feature Extraction with Convolution: A Case Study

1

X

1

=

1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

1

-1

-1

-1

-1

-1

1

-1

-1

-1

1

-1

-1

-1

1

-1

-1

-1

-1

-1

1

-1

1

-1

-1

-1

-1

-1

-1

-1

1

-1

-1

-1

-1

-1

-1

-1

1

-1

1

-1

-1

-1

-1

-1

1

-1

-1

-1

1

-1

-1

-1

1

-1

-1

-1

-1

-1

1

-1

-1

-1

-1

-1

-1

-1

-1

-1

-1

1

-1

-1

-1

1

-1

-1

-1

1

1

-1

-1

-1

1

-1

-1

-1

1

=

9

element-wise multiply

add outputs

16 of 99

Producing Feature Maps

Original

Sharpen

Edge Detect

“Strong” Edge Detect

17 of 99

Pooling Layer

18 of 99

Pooling Layer

19 of 99

Spatial Pooling

  • Sum, average, or max
  • Non-overlapping / overlapping regions
  • Role of pooling:
    • Invariance to small transformations
    • Larger receptive fields (see more of input)

20 of 99

Convolutional Neural Networks

  • Feed-forward feature extraction:
    • Convolve input with learned filters: Apply filters to generate feature maps.
    • Non-linearity: Often ReLU.
    • Spatial pooling: Downsampling operation on each feature map.
    • Normalization
  • Supervised training of convolutional filters by back-propagating classification error

21 of 99

Important Concepts in CNN

  • 1. Convolution could have multiple filters.

22 of 99

Important Concepts in CNN

  • 1. Convolution could have multiple filters.
  • 2. For tensor (rank>=3), it still applies element-wise multiplication.

2D Conv on 3D input

3D Conv on 3D input

23 of 99

Important Concepts in CNN

  • 1. Convolution could have multiple filters.
  • 2. For tensor (rank>=3), it still applies element-wise multiplication.
  • 3. Stride is the size of filter step in sliding window.

24 of 99

Important Concepts in CNN

  • 1. Convolution could have multiple filters.
  • 2. For tensor (rank>=3), it still applies element-wise multiplication.
  • 3. Stride is the size of filter step in sliding window.
  • 4. By stacking convolutional layers, we can increase receptive field.

25 of 99

Important Concepts in CNN

  • 1. Convolution could have multiple filters.
  • 2. For tensor (rank>=3), it still applies element-wise multiplication.
  • 3. Stride is the size of filter step in sliding window.
  • 4. By stacking convolutional layers, we can increase receptive field.
  • 5. To include pixels/neurons around the boundary of image, we need padding.

26 of 99

Representation Learning in Deep CNNs

Conv Layer 1

Conv Layer 2

Conv Layer 3

27 of 99

Practice 1: Feature Map Shape

28 of 99

Practice 1: Feature Map Shape

  • Stride = 1 (Default): Moves one pixel at a time

Convolution with 3x3 kernel, zero padding and stride = 1

29 of 99

Practice 1: Feature Map Shape

  • Stride > 1: Moves multiple pixels at a time 🡪 Reduces the output size, leading to downsampling.

Convolution with 3x3 kernel, zero padding and stride = 2

30 of 99

Deep Learning

Recurrent Neural Networks

31 of 99

So Far

  • Regression, Classification, Dimension Reduction, …
  • Based on snapshot-type data

32 of 99

Sequence Matters

  • Given an image of a ball, can you predict where it will go next?

???

33 of 99

Sequence Matters

  • How about this? Can you predict where it will go next?

34 of 99

What is a Sequence?

  • Sentence
    • “This morning I took the dog for a walk.”

  • Medical signals / Speech waveform / Vibration measurement

35 of 99

Sequence Modeling

  • Sequence modeling is the task of predicting what comes next
    • E.g., “This morning I took my dog for a walk.”

    • E.g., given historical air quality, forecast air quality in next couple of hours.

given previous words

predict the next word

36 of 99

A Sequence Modeling Example: Next Word Prediction

  • Idea #1: Use a fixed window

  • Limitation: Cannot model long-term dependencies
    • E.g., “France is where I grew up, but I now live in Boston. I speak fluent ___.”

  • We need information from the distant past to accurately predict the correct word.

“This morning I took my dog for a walk.”

given previous two words

predict the next word

37 of 99

A Sequence Modeling Example: Next Word Prediction

  • Idea #2: Use entire sequence as set of counts

  • Bag-of-words model
    • Define a vocabulary and initialize a zero vector where each element represents for each word
    • Compute word frequency and update the correspond position in the vector

    • Use the vector for prediction

  • Limitation: Counts don’t preserve order
    • “The food was good, not bad at all.” vs. “The food was bad, not good at all.”

  • We need to preserve the information about order.

This morning I took my dog for a walk.”

predict the next word

Here 1 is the count for the word “a

[0 1 0 0 1 0 1 … … 0 0 1 1 0 0 0 1 0]

38 of 99

Sequence Modeling

  • To model sequences, we need to:
    • Handle variable-length sequences
    • Track long-term dependencies
    • Maintain information about order
    • Share parameters across the sequence

  • Solution:
    • Recurrent Neural Networks (RNNs)

39 of 99

A Recurrent Neural Network (RNN)

  • Apply a recurrence relation at every time step to process a sequence:

  • Note: the same function and set of parameters �are used at every time step

output vector

input vector

 

cell state

 

old state

current input

40 of 99

Standard Feed-Forward Neural Network

41 of 99

Recurrent Neural Networks

… and many other architectures and applications

42 of 99

RNN: State Update and Output

  •  

output vector

input vector

 

cell state

 

old state

current input

 

 

43 of 99

RNN: Computational Graph across Time

  • Represent as computational graph unrolled across time

 

44 of 99

RNN: Computational Graph across Time

  • Represent as computational graph unrolled across time

 

45 of 99

RNN: Computational Graph across Time

  • Represent as computational graph unrolled across time

 

46 of 99

RNN: Computational Graph across Time

  • Represent as computational graph unrolled across time

 

47 of 99

RNN: Computational Graph across Time

  • Re-use the same weight matrices at every time step

 

48 of 99

RNN: Computational Graph across Time

  •  

49 of 99

RNN: Computational Graph across Time

  •  

50 of 99

RNN: Backpropagation Through Time

  •  

51 of 99

Standard RNN Gradient Flow

  •  

52 of 99

Standard RNN Gradient Flow: Exploding Gradients

  •  

Many values > 1:

exploding gradients

Gradient clipping to�scale big gradients

Many values < 1:

vanishing gradients

  1. Activation function
  2. Weight initialization
  3. Network architecture

53 of 99

Standard RNN Gradient Flow: Vanishing Gradients

  •  

Many values > 1:

exploding gradients

Gradient clipping to�scale big gradients

Many values < 1:

vanishing gradients

  1. Activation function
  2. Weight initialization
  3. Network architecture

54 of 99

Why is Exploding Gradient a Problem?

  • If the gradient becomes too big, then SGD update step becomes too big:

  • This can be bad updates: we take too large a step and reach a weird and bad parameter configuration (with large loss)
    • You think you’ve found a hill to climb, but suddenly you’re in Iowa.

  • In the worst case, this will result in Inf or NaN �in your network (then you have to restart training from �an earlier checkpoint)

 

55 of 99

Solution: Gradient Clipping

  • Gradient clipping: if the norm of the gradient is greater than some threshold, scale it down before applying SGD update

  • Intuition: take a step in the same direction, but a smaller step

  • In practice: remembering to clip gradients is important, but exploding gradients are an easy problem to solve.

56 of 99

The Problem of Long-Term Dependencies

  • Why are vanishing gradients a problem?

Multiply many small numbers together

Errors due to further back time steps �have smaller and smaller gradients

Bias parameters to capture short-term dependencies

57 of 99

Vanishing Gradient Intuition

 

58 of 99

Vanishing Gradient Intuition

 

Chain rule!

59 of 99

Vanishing Gradient Intuition

 

Chain rule!

 

60 of 99

Vanishing Gradient Intuition

 

Vanishing gradient problem:

When these are small, the gradient signal gets smaller and smaller as it backpropagates further

 

 

61 of 99

Vanishing Gradient Intuition

 

Vanishing gradient problem:

When these are small, the gradient signal gets smaller and smaller as it backpropagates further

 

 

62 of 99

Why is Vanishing Gradient a Problem?

Gradient signal from far away is lost because it’s much smaller than gradient signal from close-by.

So, model weights are basically updated only with respect to near effects, not long-term effects.

63 of 99

Why is Vanishing Gradient a Problem?

  • LM task: When she tried to print her tickets, she found that the printer was out of toner. She went to the stationery store to buy more toner. It was very overpriced. After installing the toner into the printer, she finally printed her _______.

  • To learn from this example, the RNN-LM needs to model the dependency between “tickets” on the 7th step and the target word “tickets” at the end.

  • But if the gradient is small, the model can’t learn this dependency.
    • So, the model is unable to predict similar long-distance dependencies at test time.

  • In practice, a simple RNN will only condition ~7 tokens back.

64 of 99

Solution: Gating Mechanisms in Neurons

  • Use a more complex recurrent unit with gates to control what information is passed through

  • Long Short-Term Memory (LSTM) networks rely on gated cells to track information throughout many time steps.

  • Intuition: could we design an RNN with separate memory which is added to?

65 of 99

Standard RNNs

  • In a standard RNN, recurrent modules contain simple computation

66 of 99

Long Short-Term Memory (LSTM)

  • In an LSTM network, recurrent modules contain gated cells that control the information flow

67 of 99

Long Short-Term Memory (LSTM)

  •  

68 of 99

Long Short-Term Memory (LSTM)

  • Information is added or removed to cell state through structures called gates.

Gates optionally let information through, via a sigmoid layer and pointwise multiplication

69 of 99

LSTM: Forget Irrelevant Information

  •  

 

70 of 99

LSTM: Add New Information

  •  

 

71 of 99

LSTM: Update Cell State

  •  

 

72 of 99

LSTM: Output Filtered Version of Cell State

  •  

 

73 of 99

LSTM: Cell State Matters

  •  

 

74 of 99

LSTM: Cell State Matters

  • You can think of the LSTM equations visually like this:

75 of 99

LSTM: Mitigate Vanishing Gradient

  • Vanilla RNNs

 

 

 

… Vanish!

 

… Explode!

🡪

🡪

 

76 of 99

LSTM: Mitigate Vanishing Gradient

  • Vanilla RNNs

  • LSTM

 

 

 

 

 

So…

We can keep information if we want!�(by adjusting how much we forget)

77 of 99

LSTM: Key Concepts

  • Maintain a separate cell state from what is outputted

  • Use gates to control the flow of information
    • Forget gate gets rid of irrelevant information
    • Selectively updates cell state
    • Output gate returns a filtered version of the cell state

  • LSTM can mitigate vanishing gradient problem

78 of 99

Supplementary: Residual Connection

  • Is vanishing/exploding gradient just an RNN problem?

  • No! It can be a problem for all neural architectures (including feed-forward and convolutional neural networks), especially very deep ones.
    • Due to chain rule / choice of nonlinearity function, gradient can become vanishingly small as it backpropagates
    • Thus, lower layers are learned very slowly (i.e., are hard to train)

  • Another solution: lots of new deep feedforward/convolutional architectures add more direct connections (thus allowing the gradient to flow)

  • For example:
    • Residual connects a.k.a. “ResNet”
    • Also known as skip-connections
    • The identity connection preserves information by default
    • This makes deep networks much easier to train

79 of 99

RNN Applications

  • RNNs can be used for sequence tagging�e.g., part-of-speech tagging, named entity recognition

80 of 99

RNN Applications

  • RNNs can be used as a sentence encoder model�e.g., sentiment classification

How to compute sentence encoding?

81 of 99

RNN Applications

  • RNNs can be used as a sentence encoder model�e.g., sentiment classification

Basic way: use final hidden state

82 of 99

RNN Applications

  • RNNs can be used as a sentence encoder model�e.g., sentiment classification

Usually better: �Take element-wise max or mean of all hidden states

83 of 99

RNN Applications

  • RNNs can be used to generate text based on other information
  • e.g., speech recognition, machine translation, summarization

84 of 99

RNN Applications & Limitations

85 of 99

RNN Applications

86 of 99

Bidirectional and Multi-layer RNNs

  • Motivation

We can regard this hidden state as a representation of the word “terribly” in the context of this sentence. We call this a contextual representation.

element-wise mean/max

element-wise mean/max

the

movie

was

terribly

exciting

!

These contextual representations only contain information about the left context (e.g., “the movie was”).

What about right context?

In this example, “exciting” is in the right context and this modifies the meaning of “terribly” (from negative to positive)

87 of 99

Bidirectional and Multi-layer RNNs

This contextual representation of “terribly” has both left and right context!

Concatenated hidden states

Forward RNN

Backward RNN

88 of 99

Bidirectional RNNs

  •  

This is a general notation to mean “compute one forward step of the RNN” – it could be a simple RNN or LSTM computation.

Concatenated hidden states

Forward RNN

Backward RNN

 

 

 

We regard this as “the hidden state” of a bidirectional RNN.

This is what we pass on the next parts of the network.

Generally, these two RNNs have separate weights

89 of 99

Bidirectional RNNs: Simplified Diagram

The two-way arrows indicate bidirectionality and the depicted hidden states are assumed to be the concatenated forwards+backwards states

90 of 99

Bidirectional RNNs

  • Note: bidirectional RNNs are only applicable if you have access to the entire input sequence
    • They are not applicable to Language Modeling, because in LM you only have left context available.

  • If you do have entire input sequence (e.g., any kind of encoding), bidirectionality is powerful (you should use it by default).

  • For example, BERT (Bidirectional Encoder Representations from Transformers) is a powerful pretrained contextual representation system built on bidirectionality.
    • You will learn more about transformers, including BERT, in a couple of weeks!

91 of 99

Multi-layer RNNs

  • RNNs are already “deep” on one dimension (they unroll over many timesteps)

  • We can also make them “deep” in another dimension by applying multiple RNNs – this is a multi-layer RNN.

  • This allows the network to compute more complex representations
    • The lower RNNs should compute lower-level features and the higher RNNs should compute higher-level features.

  • Multi-layer RNNs are also called stacked RNNs.

92 of 99

Multi-layer RNNs

RNN layer 1

RNN layer 2

RNN layer 3

 

93 of 99

Multi-layer RNNs

  • Multi-layer or stacked RNNs allow a network to compute more complex representations – they work better than just one layer of high-dimensional encodings!
    • The lower RNNs should compute lower-level features and the higher RNNs should compute higher-level features.
  • High-performing RNNs are usually multi-layer (but aren’t as deep as convolutional or feed-forward networks)
  • For example: In a 2017 paper, Britz et al. find that for Neural Machine Translation, 2 to 4 layers is best for encoder RNN, and 4 layers is best for the decoder RNN
    • Often 2 layers is a lot better than 1, and 3 might be a little better than 2
    • Usually, skip-connections/dense-connections are needed to train deeper RNNs (e.g., 8 layers)
  • Transformer-based networks (e.g., BERT) are usually deeper, like 12 or 24 layers.
    • You will learn about Transformers later; they have a lot of skipping-like connections.

94 of 99

RNN Limitations

  • Limitations
    • Encoding bottleneck: Fixed-size hidden state 🡪 information loss
    • Slow, no parallelization: Step-by-step processing 🡪 slow on long sequences
    • Not long memory: Vanishing/exploding gradients 🡪 week long-term dependency

95 of 99

RNN Limitations

  • Limitations
    • Encoding bottleneck: Fixed-size hidden state 🡪 information loss
      • 🡪 Continuous stream
    • Slow, no parallelization: Step-by-step processing 🡪 slow on long sequences
      • 🡪 Parallelization
    • Not long memory: Vanishing/exploding gradients 🡪 week long-term dependency
      • 🡪 Long memory

96 of 99

In Summary

  • 1. LSTMs are powerful

  • 2. Clip your gradients

  • 3. Use bidirectionality when possible

97 of 99

Next

  • Attention and Transformers
  • Large Language Models

98 of 99

  • Language Model: A system that predicts the next word

  • Recurrent Neural Network: A family of neural networks that:
    • Take sequential input of any length; apply the sameweights on each step
    • Can optionally produce output on each step

  • Recurrent Neural Network != Language Model
    • RNNs can be used for many other things

  • Language Modeling is a traditional subcomponent of many NLP tasks, all those involving generating text or estimating the probability of text:
    • Now everything in NLP is being rebuilt upon Language Modeling: GPT-3 is an LM!

99 of 99

Assignment #4