Artificial Neural Networks:�From Perceptron to MLP
Perceptron: Binary Linear Classifier
2
Perceptron: Geometric Interpretation
3
Perceptron for New Data
4
Binary Linear Classifier in 2D
5
Binary Linear Classifier in High Dimension
6
From Perceptron to MLP
7
XOR Problem
8
| | |
0 | 0 | 0 |
0 | 1 | 1 |
1 | 0 | 1 |
1 | 1 | 0 |
Idea: Nonlinear Curve Approximated by Multiple Lines
9
XOR Problem
10
XOR Problem
11
XOR Problem
12
Multiple Perceptrons
13
Multiple Perceptrons
14
Multiple Perceptrons
15
Multiple Perceptrons
16
Multiple Perceptrons
17
Multiple Perceptrons
18
First layer
with neurons
Second layer
with neurons
Another Perspective:�Hidden Layers as Kernel Learning
19
Nonlinear Classification
20
https://www.youtube.com/watch?v=3liCbRZPrZA
Neuron
21
Second Way of Looking at Multiple Perceptrons
22
Common Activation Functions
23
Source: 6.S191 Intro. to Deep Learning at MIT
Discuss later
XOR Problem in Perceptron
24
Nonlinear Mapping
25
Source: Dr. Francois Fleuret at EPFL
Nonlinear Mapping
26
Source: Dr. Francois Fleuret at EPFL
Nonlinear Mapping
27
Source: Dr. Francois Fleuret at EPFL
Neuron
28
Kernel + Neuron
29
Neuron + Neuron
30
Multi Layer Perceptron (MLP)
31
Multi Layer Perceptron (MLP) = Artificial Neural Networks
32
Nonlinear mapping
Multi Layer Perceptron (MLP) = Artificial Neural Networks
33
Nonlinear mapping
Multi Layer Perceptron (MLP) = Artificial Neural Networks
34
…
Nonlinear mappings
Linearly separable
Multi Layer Perceptron (MLP) = Artificial Neural Networks
35
…
Nonlinear mappings
Multiple Linear classifiers
Linearly separable
Multi Layer Perceptron (MLP) = Artificial Neural Networks
36
…
Linear classification
Feature Learning
Nonlinear mappings
Linearly separable
Two Ways of Looking at Artificial Neural Networks
37
Two Ways of Looking at Artificial Neural Networks
38
(1)
(2)
Summary
39
Deep Artificial Neural Networks
40
Feature learning
Classification
Class 1
Class 2
nonlinear
linear
…
Output
Input
Deep Artificial Neural Networks
41
Class 1
Class 2
…
…
…
…
…
…
nonlinear
linear
Feature learning
Classification
Output
Input
Machine Learning vs. Deep Learning
42
Deep Learning
43
Looking at Parameters in Classification
44
Logistic Regression in a Form of Neural Network
45
Logistic Regression in a Form of Neural Network
46
Do not indicate bias units
Nonlinearly Distributed Data
47
Nonlinearly Distributed Data
48
Do not include bias units
Multi Layers
49
Do not include bias units
Multi Layers
50
Do not include bias units
Multi Layers
51
Do not include bias units
Nonlinearly Distributed Data
52
Nonlinearly Distributed Data
53
Do not include bias units
More Neurons
54
Do not include bias units
Looking at Parameters in Regression
55
Rectified Linear Unit (ReLU)
56
Regression in a Form of Neural Network
57
Regression in a Form of Neural Network
58
Regression in a Form of Neural Network
59
Regression in a Form of Neural Network
60
Artificial Neural Networks: �Training
61
Training Neural Networks: Optimization
62
Training Neural Networks: Loss Function
63
Training Neural Networks: Gradient Descent
64
Gradients in ANN
65
Dynamic Programming
66
Recursive Algorithm
67
…
Output
Input
…
Output
Input
Base Case
Dynamic Programming
68
Naïve Recursive Algorithm
69
Memorized Recursive Algorithm
70
Dynamic Programming Algorithm
71
Backpropagation
72
Gradients in ANN
73
Training Neural Networks: Backpropagation Learning
74
Backpropagation
75
Backpropagation
76
These are what we need for GD
Backpropagation
77
These are what we need for GD
Backpropagation
78
These are what we need for GD
Backpropagation
79
These are what we need for GD
Backpropagation
80
These are what we need for GD
Backpropagation
81
These are what we need for GD
Training Neural Networks with TensorFlow
82
Core Foundation Review
83
Source: 6.S191 Intro. to Deep Learning at MIT
Artificial Neural Networks with TensorFlow
84
MNIST database
85
ANN in TensorFlow: MNIST
86
Our Network Model
87
Input layer
(784)
Hidden layer
(100)
Output layer
(10)
Input image
(28 X 28)
Flattened
Digit prediction
in one-hot-encoding
Iterative Optimization
88
Implementation in Python
89
Input layer
(784)
Hidden layer
(100)
Output layer
(10)
Flattened
Evaluation
90