1 of 90

Artificial Neural Networks:�From Perceptron to MLP

2 of 90

Perceptron: Binary Linear Classifier

  • Given weights

2

3 of 90

Perceptron: Geometric Interpretation

  • Given weights

3

4 of 90

Perceptron for New Data

  • Forward propagation with new data

4

5 of 90

Binary Linear Classifier in 2D

5

6 of 90

Binary Linear Classifier in High Dimension

  • More neurons means
    • hyper-plane in a higher dimension

6

7 of 90

From Perceptron to MLP

7

8 of 90

XOR Problem

  • Minsky-Papert Controversy on XOR

  • For not linearly separable

  • Single neuron = one linear classification boundary
    • A perceptron cannot solve due to its linear nature

8

0

0

0

0

1

1

1

0

1

1

1

0

9 of 90

Idea: Nonlinear Curve Approximated by Multiple Lines

  • Nonlinear regression
  • Nonlinear classification

9

10 of 90

XOR Problem

  • At least two lines are required

10

11 of 90

XOR Problem

  • At least two lines are required
  • If two perceptrons are stacked, it represents two hyperplanes.

11

12 of 90

XOR Problem

  • At least two lines are required
  • If two perceptrons are stacked, it represents two hyperplanes.

12

13 of 90

Multiple Perceptrons

  • Multi neurons = multiple linear classification boundaries

13

14 of 90

Multiple Perceptrons

  • Multi neurons = multiple linear classification boundaries

14

15 of 90

Multiple Perceptrons

  • Sigmoid function for nonlinear activation function

15

16 of 90

Multiple Perceptrons

  • In a compact representation

16

17 of 90

Multiple Perceptrons

  • In a compact representation

17

18 of 90

Multiple Perceptrons

  • In a compact representation

18

First layer

with neurons

Second layer

with neurons

19 of 90

Another Perspective:�Hidden Layers as Kernel Learning

19

20 of 90

Nonlinear Classification

20

https://www.youtube.com/watch?v=3liCbRZPrZA

21 of 90

Neuron

  • We can represent this “neuron” as follows:

21

22 of 90

Second Way of Looking at Multiple Perceptrons

  • Can represent nonlinear relationship between input and outputs due to nonlinear activation function

22

23 of 90

Common Activation Functions

23

Source: 6.S191 Intro. to Deep Learning at MIT

Discuss later

24 of 90

XOR Problem in Perceptron

  • The main weakness of linear predictors is their lack of capacity.
  • For classification, the populations have to be linearly separable.

24

25 of 90

Nonlinear Mapping

  • The XOR example can be solved by pre-processing the data to make the two populations linearly separable.

25

Source: Dr. Francois Fleuret at EPFL

26 of 90

Nonlinear Mapping

  •  

26

Source: Dr. Francois Fleuret at EPFL

27 of 90

Nonlinear Mapping

  •  

27

Source: Dr. Francois Fleuret at EPFL

28 of 90

Neuron

  • Suppose that data is not linearly separable

28

29 of 90

Kernel + Neuron

  • Nonlinear mapping + neuron

  • User-defined Kernel

29

 

30 of 90

Neuron + Neuron

  • Nonlinear mapping can be represented by another layer (or neurons)

  • Learnable Kernel
    • Nonlinear activation functions

30

31 of 90

Multi Layer Perceptron (MLP)

  • Nonlinear mapping can be represented by another layer (or neurons)
  • We can generalize an MLP

31

32 of 90

Multi Layer Perceptron (MLP) = Artificial Neural Networks

  • Why do we need multi-layers ?

32

Nonlinear mapping

33 of 90

Multi Layer Perceptron (MLP) = Artificial Neural Networks

  • Why do we need multi-layers ?

33

Nonlinear mapping

34 of 90

Multi Layer Perceptron (MLP) = Artificial Neural Networks

  • Why do we need multi-layers ?

34

Nonlinear mappings

Linearly separable

35 of 90

Multi Layer Perceptron (MLP) = Artificial Neural Networks

  • Why do we need multi-layers ?

35

Nonlinear mappings

Multiple Linear classifiers

Linearly separable

36 of 90

Multi Layer Perceptron (MLP) = Artificial Neural Networks

  • Why do we need multi-layers ?

36

Linear classification

Feature Learning

Nonlinear mappings

Linearly separable

37 of 90

Two Ways of Looking at Artificial Neural Networks

  • Still represent lines

  • Can represent nonlinear relationship between input and outputs due to nonlinear activation function

37

38 of 90

Two Ways of Looking at Artificial Neural Networks

  • Still represent lines

  • Can represent nonlinear relationship between input and outputs due to nonlinear activation function

38

(1)

(2)

39 of 90

Summary

  •  

39

  1. Weight point of view
  2. Value propagation point of view

40 of 90

Deep Artificial Neural Networks

  • Complex/Nonlinear universal function approximator
    • Linearly connected networks
    • Simple nonlinear neurons

40

Feature learning

Classification

Class 1

Class 2

nonlinear

linear

Output

Input

41 of 90

Deep Artificial Neural Networks

  • Complex/Nonlinear universal function approximator
    • Linearly connected networks
    • Simple nonlinear neurons

41

Class 1

Class 2

nonlinear

linear

Feature learning

Classification

Output

Input

42 of 90

Machine Learning vs. Deep Learning

  • Feature engineering

  • Feature learning

42

  • AI refers to the ability of machines to mimic human intelligence without explicit programming

43 of 90

Deep Learning

43

44 of 90

Looking at Parameters in Classification

44

45 of 90

Logistic Regression in a Form of Neural Network

45

46 of 90

Logistic Regression in a Form of Neural Network

  • Neural network convention

46

Do not indicate bias units

47 of 90

Nonlinearly Distributed Data

  • Example to understand network’s behavior
    • Include a hidden layer

47

48 of 90

Nonlinearly Distributed Data

  • Example to understand network’s behavior
    • Include a hidden layer

48

Do not include bias units

49 of 90

Multi Layers

  •  

49

Do not include bias units

50 of 90

Multi Layers

  •  

50

Do not include bias units

51 of 90

Multi Layers

  •  

51

Do not include bias units

52 of 90

Nonlinearly Distributed Data

  • More neurons in hidden layer

52

53 of 90

Nonlinearly Distributed Data

  • More neurons in hidden layer

53

Do not include bias units

54 of 90

More Neurons

  • Multiple linear classification boundaries

54

Do not include bias units

55 of 90

Looking at Parameters in Regression

55

56 of 90

Rectified Linear Unit (ReLU)

  • ReLU activation function

56

57 of 90

Regression in a Form of Neural Network

  • Single Neuron with ReLU

57

58 of 90

Regression in a Form of Neural Network

  • 2 Neurons with ReLU

58

59 of 90

Regression in a Form of Neural Network

  • 4 Neurons with ReLU

59

60 of 90

Regression in a Form of Neural Network

  • 100 Neurons with ReLU

  • Universal Function Approximator

60

61 of 90

Artificial Neural Networks: �Training

61

62 of 90

Training Neural Networks: Optimization

  •  

62

63 of 90

Training Neural Networks: Loss Function

  • Measures error between target values and predictions

  • Example
    • Squared loss (for regression):

    • Cross entropy (for classification):�

63

64 of 90

Training Neural Networks: Gradient Descent

  •  

64

65 of 90

Gradients in ANN

  •  

65

66 of 90

Dynamic Programming

66

67 of 90

Recursive Algorithm

  •  

67

Output

Input

Output

Input

Base Case

68 of 90

Dynamic Programming

  • Dynamic Programming: general, powerful algorithm design technique

  • Fibonacci numbers:

68

69 of 90

Naïve Recursive Algorithm

  • It works. Is it good?

69

 

 

 

 

 

 

 

70 of 90

Memorized Recursive Algorithm

  • Benefit?
    • fib(n) only recurses the first time it’s called

70

 

 

 

 

 

 

 

71 of 90

Dynamic Programming Algorithm

  •  

71

72 of 90

Backpropagation

72

73 of 90

Gradients in ANN

  •  

73

74 of 90

Training Neural Networks: Backpropagation Learning

  • Forward propagation
    • the initial information propagates up to the hidden units at each layer and finally produces output

  • Backpropagation
    • allows the information from the cost to flow backwards through the network in order to compute the gradients

74

75 of 90

Backpropagation

  •  

75

76 of 90

Backpropagation

  •  

76

These are what we need for GD

77 of 90

Backpropagation

  •  

77

These are what we need for GD

78 of 90

Backpropagation

  •  

78

These are what we need for GD

79 of 90

Backpropagation

  •  

79

These are what we need for GD

80 of 90

Backpropagation

  •  

80

These are what we need for GD

81 of 90

Backpropagation

  •  

81

These are what we need for GD

82 of 90

Training Neural Networks with TensorFlow

  •  

82

83 of 90

Core Foundation Review

83

Source: 6.S191 Intro. to Deep Learning at MIT

84 of 90

Artificial Neural Networks with TensorFlow

84

85 of 90

MNIST database

  •  

85

86 of 90

ANN in TensorFlow: MNIST

86

87 of 90

Our Network Model

87

Input layer

(784)

Hidden layer

(100)

Output layer

(10)

Input image

(28 X 28)

Flattened

Digit prediction

in one-hot-encoding

88 of 90

Iterative Optimization

  • We will use
    • Mini-batch gradient descent
    • Adam optimizer

88

89 of 90

Implementation in Python

89

Input layer

(784)

Hidden layer

(100)

Output layer

(10)

 

Flattened

90 of 90

Evaluation

90