Artificial Neural Networks: Deep NN
Prof. Dinesh K. Vishwakarma
DEPARTMENT OF INFORMATION TECHNOLOGY
DELHI TECHNOLOGICAL UNIVERSITY, DELHI.
Webpage: http://www.dtu.ac.in/Web/Departments/InformationTechnology/faculty/dkvishwakarma.php
History of Deep Learning
2
9/9/2025
History of Deep Learning…
3
9/9/2025
History of Deep Learning…
4
9/9/2025
History of Deep Learning…
5
9/9/2025
History of Deep Learning…
6
9/9/2025
History of Deep Learning…
7
9/9/2025
History of Deep Learning…
8
9/9/2025
History of Deep Learning…
9
9/9/2025
History of Deep Learning…
10
9/9/2025
History of Deep Learning…
11
9/9/2025
History of Deep Learning…
12
9/9/2025
Artificial Neural Networks
13
9/9/2025
Brain and Machine
14
9/9/2025
The contrast in architecture
15
9/9/2025
Features of the Brain
16
9/9/2025
The biological inspiration
17
9/9/2025
The biological inspiration…
18
9/9/2025
The Model of Neuron
19
9/9/2025
Synapses vary in strength
Input:DendrIte
Output: axOn
The Model of Neuron
20
9/9/2025
Biological Neuron | Artificial Neuron |
Dendrite | Inputs |
Cell nucleus or Soma | Nodes |
Synapses | Weights |
Axon | Output |
The Artificial Neuron
21
9/9/2025
Output
Input
The Artificial Neuron…
22
9/9/2025
∑
Input
The Artificial Neuron…
23
9/9/2025
∑
Input
Axons
(from other neurons)
Synapses
Dendrites
Cell Body
The Artificial Neuron…
24
9/9/2025
∑
Input
Axons
(from other neurons)
Synapses
Dendrites
Cell Body
+
-
Biasing
The Artificial Neuron…
25
9/9/2025
∑
Input
1
Activation Value
This model is aka Perceptron
Given by Rosenblatt 1958.
Biasing
A Simple Model of a Neuron (Perceptron)…
26
9/9/2025
Example of Perceptron
27
9/9/2025
Example of Perceptron…
28
9/9/2025
Perceptron (s)
29
9/9/2025
Simplified
Multi output
All inputs are connected to all outputs, hence these layers are called as dense layer
Single Layer NN
30
9/9/2025
Single Layer NN
31
9/9/2025
Deep NN
32
9/9/2025
Activation Functions
33
9/9/2025
1
1
-1
Unipolar
Bipolar
Activation Functions…
34
9/9/2025
Unipolar
1
1
-1
Bipolar
Hard Limiting
1
1
-1
Bipolar
Soft Limiting
Unipolar
Activation Functions…
35
9/9/2025
Sigmoid Function
1
Hyperbolic Tangent Function
1
-1
The biological basis of these functions are easily established. Neurons located in different parts of the nervous system have different characteristics. Ocular motor: sigmoid, Visual Cortex: Gaussian
Common Activation Functions
36
9/9/2025
Activation Function
37
9/9/2025
Consider a Example
38
9/9/2025
Quantify Loss
39
9/9/2025
Binary Cross Entropy Loss
40
9/9/2025
Mean Square Error Loss
41
9/9/2025
Training NN
42
9/9/2025
Loss Optimization
43
9/9/2025
Loss Optimization…
44
9/9/2025
Loss Optimization…
45
9/9/2025
Loss Optimization…
46
9/9/2025
Loss Optimization…
47
9/9/2025
Gradient Descent
48
9/9/2025
Computing Gradient
49
9/9/2025
Computing Gradient…
50
9/9/2025
Repeat this for every weight in the network using gradients from later layers
Ex. NN Weight Update
51
9/9/2025
A
B
C
D
E
prediction
Forward Pass Equations
52
9/9/2025
A
B
C
D
E
Prediction
Loss Function
53
9/9/2025
Backpropagation
54
9/9/2025
Backpropagation is the algorithm to change the weights of the neural network in a manner so that the prediction gets closer to the actual output.
Backpropagation…
55
9/9/2025
Using chain rule,
For example,
Backpropagation…
56
9/9/2025
Hence,
Backpropagation…
57
9/9/2025
Similarly,
Backpropagation…
58
9/9/2025
Backpropagation Ex.
59
9/9/2025
A
B
C
D
E
Prediction
Backpropagation Ex.
60
9/9/2025
A
B
C
D
E
Prediction
Backpropagation Ex.
61
9/9/2025
A
B
C
D
E
prediction
Backpropagation Ex.
62
9/9/2025
A
B
C
D
E
prediction
Backpropagation Ex.
63
9/9/2025
A
B
C
D
E
prediction
Backpropagation Ex.
64
9/9/2025
A
B
C
D
E
prediction
Backpropagation Ex.
65
9/9/2025
A
B
C
D
E
prediction
Repeating the entire backpropagation a second time,
Backpropagation Ex.
66
9/9/2025
A
B
C
D
E
prediction
Observations,
Backpropagation Ex.
67
9/9/2025
A
B
C
D
E
prediction
Backpropagation Ex.
68
9/9/2025
A
B
C
D
E
prediction
After second backpropagation round,
Training Perceptrons
69
9/9/2025
t = 0.0
y
x
-1
W1 = ?
W3 = ?
W2 = ?
For AND
A B Output
0 0 0
0 1 0
1 0 0
1 1 1
Training Perceptron's
70
9/9/2025
t = 0.0
y
x
-1
W1 = 0.3
W3 =-0.4
W2 = 0.5
For AND
A B Output
0 0 0
0 1 0
1 0 0
1 1 1
Optimization: In Practice
71
9/9/2025
How to Deal: Learning Rate
72
9/9/2025
Gradient Descent
73
9/9/2025
Entire training dataset is used to compute gradient
Stochastic Gradient Descent
74
9/9/2025
Use single example (or small mini-batch) at each step to compute gradient
B=1, SGD
Mini Batches
75
9/9/2025
Neural Network Problems
76
9/9/2025
Hyperparameters
77
9/9/2025
Hyperparameters…
78
9/9/2025
Hyperparameters are important because they directly control the behaviour of the training algorithm and have a significant impact on the performance of the model is being trained.
Hyperparameters…
79
9/9/2025
Overfittings
80
9/9/2025
Overfittings
81
9/9/2025
Overfittings…
82
9/9/2025
Overfittings…_Underfit
83
9/9/2025
Underfit 1 degree polynomial model on training (left) and testing (right) datasets
Overfittings…_Overfit
84
9/9/2025
Overfit 25 degree polynomial model on training (left) and testing (right) datasets
Overfittings…_Balanced Fit
85
9/9/2025
Balanced Four degree polynomial model on training (left) and testing (right) datasets
Underfit_Overfit_Bestfit
86
9/9/2025
Regularization
87
9/9/2025
Regularization I: Dropout
88
9/9/2025
Regularization I: Dropout
89
9/9/2025
Regularization II: �Early Stopping
90
9/9/2025
Training Time
91
9/9/2025
Regularization II: �Batch Normalization (BN)
92
9/9/2025
Regularization III: �Batch Normalization (BN)
93
9/9/2025
Normalization
94
9/9/2025
Original values
Normalized Value
Offers ‘0’ mean & 1 variance
Normalization…
95
9/9/2025
Features on different scales take longer to reach the minimum
Normalized data helps the network converge faster
Normalization…
96
9/9/2025
A narrow valley causes gradient descent to bounce from one slope to the other
Ex. DNN
97
9/9/2025
The inputs of each hidden layer are the activations from the previous layer, and must also be normalized
Ex. DNN…
98
9/9/2025
The Batch Norm layer normalizes activations from Layer 1 before they reach layer 2
Regularization IV: �Gradient Clipping
99
9/9/2025
Regularization IV: �Gradient Clipping…
100
9/9/2025
Regularization IV: �Gradient Clipping…
101
9/9/2025
Gradient Clipping
By Value
By Norm
Regularization IV: �Gradient Clipping…
102
9/9/2025
Example of Digit Recognition
103
9/9/2025
Machine
“2”
16 x 16 = 256
……
Ink → 1 No ink → 0
……
y1
y2
y10
is 1
is 2
is 0
……
0.1
0.7
0.2
The image is “2”
Example of Neural Network
104
9/9/2025
Sigmoid Function
1
-1
1
-2
1
-1
1
0
4
-2
0.98
0.12
Example of Neural Network
105
9/9/2025
1
-2
1
-1
1
0
4
-2
0.98
0.12
2
-1
-1
-2
3
-1
4
-1
0.86
0.11
0.62
0.83
0
0
-2
2
1
-1
Example of Neural Network
106
9/9/2025
1
-2
1
-1
1
0
0.73
0.5
2
-1
-1
-2
3
-1
4
-1
0.72
0.12
0.51
0.85
0
0
-2
2
Different parameters define different function
0
0
Example of Neural Network
107
9/9/2025
1
-2
1
-1
1
0
4
-2
0.98
0.12
1
-1
Example of Neural Network
108
9/9/2025
……
……
……
……
……
……
……
……
y1
y2
yM
W1
W2
WL
b2
bL
x
a1
a2
y
b1
W1
x
+
b2
W2
a1
+
bL
WL
+
aL-1
b1
Neural Network
109
9/9/2025
……
……
……
……
……
……
……
……
y1
y2
yM
W1
W2
WL
b2
bL
x
a1
a2
y
y
x
b1
W1
x
+
b2
W2
+
bL
WL
+
…
b1
…
Using parallel computing techniques to speed up matrix operation
Softmax
110
9/9/2025
Ordinary Layer
In general, the output of network can be any value.
May not be easy to interpret
Softmax
111
9/9/2025
3
-3
1
2.7
20
0.05
0.88
0.12
≈0
Network Parameters
112
9/9/2025
16 x 16 = 256
……
……
……
……
……
Ink → 1
No ink → 0
……
y1
y2
y10
0.1
0.7
0.2
y1 has the maximum value
Set the network parameters such that ……
Input:
y2 has the maximum value
Input:
is 1
is 2
is 0
Softmax
Visual Information Processing
113
9/9/2025
Enabling Factor of DL
114
9/9/2025
Hierarchical Learning
115
9/9/2025
Low-level features
output
Mid-level features
High-level features
Trainable classifier
Inspired from visual information processing, a representation of Hierarchical Learning is developed, also know as “Deep Learning”
First in 1986 by Rina Dechter
Revolution since 2012
Deep Neural Network
116
9/9/2025
Output Layer
Hidden Layers
Input Layer
Input
Output
Layer 1
……
……
Layer 2
……
Layer L
……
……
……
……
……
y1
y2
yM
Deep means many hidden layers
neuron
Why Deep Network?
117
9/9/2025
Layer X Size | Word Error Rate (%) | Layer X Size | Word Error Rate (%) |
1 X 2k | 24.2 | | |
2 X 2k | 20.4 | | |
3 X 2k | 18.4 | | |
4 X 2k | 17.8 | | |
5 X 2k | 17.2 | 1 X 3772 | 22.5 |
7 X 2k | 17.1 | 1 X 4634 | 22.6 |
| | 1 X 16k | 22.1 |
Seide, Frank, Gang Li, and Dong Yu. "Conversational Speech Transcription Using Context-Dependent Deep Neural Networks." Interspeech. 2011.
Not surprised, more parameters, better performance
Why Deep Network?
118
9/9/2025
Any continuous function f
Can be realized by a network with one hidden layer
(given enough hidden neurons)
Why “Deep” neural network not “Fat” neural network?
119
9/9/2025
Fat + Short v.s. Thin + Tall
……
Deep
……
……
Shallow
Which one is better?
The same number of parameters
Fat + Short v.s. Thin + Tall
120
9/9/2025
Seide, Frank, Gang Li, and Dong Yu. "Conversational Speech Transcription Using Context-Dependent Deep Neural Networks." Interspeech. 2011.
Layer X Size | Word Error Rate (%) | Layer X Size | Word Error Rate (%) |
1 X 2k | 24.2 | | |
2 X 2k | 20.4 | | |
3 X 2k | 18.4 | | |
4 X 2k | 17.8 | | |
5 X 2k | 17.2 | 1 X 3772 | 22.5 |
7 X 2k | 17.1 | 1 X 4634 | 22.6 |
| | 1 X 16k | 22.1 |
When to use Deep Learning?
121
9/9/2025
Fuel of deep learning is the big data by Andrew Ng
Deep
Learning
Machine
Learning
Amount of Data
Performance
Limitations of Deep Learning
122
9/9/2025
Question for Practice
123
9/9/2025
124
9/9/2025
Question for Practice
125
9/9/2025
126
9/9/2025
Reference
127
9/9/2025
Thank you!�dinesh@dtu.ac.in
Slide 128 of 74
9/9/2025