Artificial Neural Networks: Deep NN
Prof. Dinesh K. Vishwakarma
DEPARTMENT OF INFORMATION TECHNOLOGY
DELHI TECHNOLOGICAL UNIVERSITY, DELHI.
Webpage: http://www.dtu.ac.in/Web/Departments/InformationTechnology/faculty/dkvishwakarma.php
History of Deep Learning
Three waves of neural-network research
2
9/9/26
Cybernetics
1940 – 1970 · Golden Age
Connectionism
1980 – 2000 · Dark Age
Deep Learning
2006 – present · Revolution
History of Deep Learning…
Milestones that built the field
3
9/9/26
1943
McCulloch & Pitts
First mathematical model of a neuron — but no way to learn its weights.
1958–1962
Rosenblatt's Perceptron
First learning algorithm for a threshold neuron; convergence proved by Novikoff.
1969
Minsky & Papert
Single-layer perceptrons cannot solve XOR; symbolic AI dominates the 1970s.
1979
Fukushima's Neocognitron
Simple and complex cells introduce convolution and pooling — an early ConvNet.
1986
Backpropagation
Rumelhart, Hinton & Williams: efficient gradients for deep nets; still the workhorse.
1997
LSTM
Hochreiter & Schmidhuber: gated memory fixes vanishing gradients; reshapes NLP.
1998: Convolutional Nets
4
9/9/26
LeCun, Bottou, Bengio & Haffner: Gradient-based learning applied to document recognition, Proc. IEEE, 1998.
Why it matters
LeNet-5 set the convolution + pooling + fully connected pattern that modern vision networks still use.
Artificial Neural Networks
Other terms and names for ANN
5
9/9/26
Connectionist
Behaviour emerges from the connections between many simple units.
Parallel distributed processing
Information is represented and processed across many units at once.
Neural computation
Computation modelled on the signalling of biological neurons.
Adaptive networks
Connection weights are learned from data, not programmed by hand.
Brain and Machine
The Brain
The Machine
6
9/9/26
An ANN borrows the brain's learning and noise tolerance, and runs it on the machine's speed.
Computer vs. Brain
Von Neumann computer
The brain
7
9/9/26
Opposite designs: one fast, exact processor versus many slow, redundant ones working together.
Inside the Brain
10 billion neurons
Each one is a simple processing unit.
8
9/9/26
Thousands of links each
A neuron connects to several thousand others.
Hundreds of ops/second
Individually slow — silicon is millions of times faster.
Neurons die off
They are never replaced, yet performance holds up.
No program
Behaviour comes from the wiring, not from code.
Massive parallelism is the trick
Billions of slow, unreliable units computing at the same time outperform one fast processor.
Fault tolerant
Damage degrades performance gradually, not all at once.
What Biology Teaches Us
What we know
9
9/9/26
What we still cannot copy
So an ANN copies the principles, not the biology
many simple units, working in parallel, with knowledge stored in the connections
How a Neuron Works
10
9/9/26
1 · Dendrites take input
They collect signals from thousands of other neurons.
2 · The soma sums
The cell body adds up all incoming signals.
3 · The axon fires
Past a threshold it sends a spike to the next cells.
Synapses vary in strength
A strong connection passes a large signal, a weak one almost none — these strengths are exactly what an ANN learns.
Neuron: Biology to Model
Threshold firing
A neuron only fires once its summed input crosses a threshold — the same role an activation function plays in an ANN.
11
9/9/26
Biological Neuron | Artificial Neuron |
Dendrite | Inputs |
Cell nucleus or Soma | Nodes |
Synapses | Weights |
Axon | Output |
Synapse strength = weight
A strong synapse passes a large signal, a weak one almost none — this is exactly what a network's weights encode.
This mapping turns a biology diagram into something you can compute
Same vocabulary throughout
This table's terms — inputs, nodes, weights, output — are the ones used for every ANN diagram from here on.
Inside an Artificial Neuron
12
9/9/26
1 Inputs
Signals x₁…xₙ arrive, plus a constant bias.
2 Weights
Each input is multiplied by a learned weight wᵢ.
4 Activation f
f(z) squashes the sum into the output y.
The Artificial Neuron…
13
9/9/26
∑
Input
The Artificial Neuron…
14
9/9/26
∑
Input
Axons
(from other neurons)
Synapses
Dendrites
Cell Body
The Artificial Neuron…
15
9/9/26
∑
Input
Axons
(from other neurons)
Synapses
Dendrites
Cell Body
+
-
Biasing
The Artificial Neuron…
16
9/9/26
∑
Input
1
Activation Value
This model is aka Perceptron
Given by Rosenblatt 1958.
Biasing
A Simple Model of a Neuron
17
9/9/26
1
x₁
x₂
xₘ
Σ
g
ŷ
w₀
w₁
w₂
wₘ
Inputs → Weights → Sum → Activation → Output
The formula
ŷ = g ( w₀ + XᵀW )
Why the activation g?
Without g the neuron is just a straight line. g bends it, so the network can learn non-linear patterns.
1
0
z
Sigmoid: S-shaped, 0 to 1
Example: A Perceptron Is Just a Line
18
9/9/26
1 · Pick the weights
Bias w₀ = 1, weights w = [ 3, −2 ]
3 · It’s just a line!
Setting the output to zero gives a straight line. A single perceptron can only separate data with one straight boundary — it is a linear classifier.
1
x₁
x₂
Σ
g
ŷ
+1
+3
−2
x₁
x₂
Classifying One Point
19
9/9/26
Decision boundary: 1 + 3x₁ − 2x₂ = 0
x₁
x₂
z < 0 → class 0
z > 0 → class 1
z > 0 → class 1
(−1, 2)
1 Take one input
X = [ −1, 2 ] so x₁ = −1, x₂ = 2
2 Compute the sum
z = 1 + 3(−1) − 2(2) = −6
3 Squash it
ŷ = g(−6), about 0.002
Verdict
z is negative, so the point sits on the blue side of the line — the neuron outputs almost 0, i.e. class 0.
Perceptron (s)
20
9/9/26
Single Hidden Layer
21
9/9/26
x₁
x₂
xₘ
z₁
z₂
z₃
zₙ
ŷ₁
ŷ₂
W⁽¹⁾
W⁽²⁾
Inputs
Hidden · n units
Output
Two weight stages
W⁽¹⁾ links inputs to the hidden layer; W⁽²⁾ links hidden to output.
Each unit: sum, then g
Hidden unit i = 1 … n
zᵢ = w₀ + Σⱼ wⱼᵢ xⱼ
Output (sum over i = 1 … n)
ŷ = g( w₀ + Σᵢ wᵢ g(zᵢ) )
Why a hidden layer?
Stacking a layer lets the network bend straight lines into curved boundaries — patterns one perceptron cannot learn.
⋮
Single Layer NN
22
9/9/26
Deep NN
23
9/9/26
Activation Functions
24
9/9/26
1
1
-1
Unipolar
Bipolar
Activation Functions…
25
9/9/26
Unipolar
1
1
-1
Bipolar
Hard Limiting
1
1
-1
Bipolar
Soft Limiting
Unipolar
Activation Functions…
26
9/9/26
Sigmoid Function
1
Hyperbolic Tangent Function
1
-1
The biological basis of these functions are easily established. Neurons located in different parts of the nervous system have different characteristics. Ocular motor: sigmoid, Visual Cortex: Gaussian
Common Activation Functions
27
9/9/26
Activation Function
28
9/9/26
Consider a Example
29
9/9/26
Quantify Loss
30
9/9/26
Binary Cross Entropy Loss
31
9/9/26
Mean Square Error Loss
32
9/9/26
Training NN
33
9/9/26
Loss Optimization
34
9/9/26
Loss Optimization…
35
9/9/26
Loss Optimization…
36
9/9/26
Loss Optimization…
37
9/9/26
Loss Optimization…
38
9/9/26
Gradient Descent
39
9/9/26
01
THE LEARNING PROBLEM
Why backpropagation?
One sweep, all gradients
Backprop computes ∂L/∂w for every weight with a single forward pass and a single backward pass — reusing shared intermediate results.
Gradient descent
Backprop supplies the gradient
the learning rate — how big a step to take
the gradient — which way L rises, and how fast
02
Computing Gradient
42
9/9/26
Computing Gradient…
43
9/9/26
Repeat this for every weight in the network using gradients from later layers
WORKED EXAMPLE — PART 1
Forward pass
w₁ = 0.4
w₂ = 0.6
x
input = 2
ReLU
ŷ
output
target y = 3
Compute left → right
LOSS
3.175
prediction 0.48
vs target 3.0
The backward pass will fix it.
05
WORKED EXAMPLE — PART 2
Backward pass & update
Push the error back with the chain rule (right → left), then step each weight downhill.
Gradients, right → left
w′ = w − η·∂L/∂w
η = 0.1
w₁
0.400 → 0.702
0.4 − 0.1×(−3.024)
w₂
0.600 → 0.802
0.6 − 0.1×(−2.016)
Loss 3.175 → 1.76
after one step
06
Training of NN
46
9/9/26
A
B
C
D
E
prediction
Backpropagation
Forward Pass Equations
47
9/9/26
A
B
C
D
E
Prediction
Loss Function
48
9/9/26
Backpropagation
49
9/9/26
Backpropagation is the algorithm to change the weights of the neural network in a manner so that the prediction gets closer to the actual output.
Backpropagation…
50
9/9/26
Using chain rule,
For example,
Backpropagation…
51
9/9/26
Hence,
Backpropagation…
52
9/9/26
Similarly,
Backpropagation…
53
9/9/26
Backpropagation Ex.
54
9/9/26
A
B
C
D
E
Prediction
Backpropagation Ex.
55
9/9/26
A
B
C
D
E
Prediction
Backpropagation Ex.
56
9/9/26
A
B
C
D
E
prediction
Backpropagation Ex.
57
9/9/26
A
B
C
D
E
prediction
Backpropagation Ex.
58
9/9/26
A
B
C
D
E
prediction
Backpropagation Ex.
59
9/9/26
A
B
C
D
E
prediction
Backpropagation Ex.
60
9/9/26
A
B
C
D
E
prediction
Repeating the entire backpropagation a second time,
Backpropagation Ex.
61
9/9/26
A
B
C
D
E
prediction
Observations,
Backpropagation Ex.
62
9/9/26
A
B
C
D
E
prediction
Backpropagation Ex.
63
9/9/26
A
B
C
D
E
prediction
After second backpropagation round,
Training Perceptrons
64
9/9/26
t = 0.0
y
x
-1
W1 = ?
W3 = ?
W2 = ?
For AND
A B Output
0 0 0
0 1 0
1 0 0
1 1 1
Training Perceptron's
65
9/9/26
t = 0.0
y
x
-1
W1 = 0.3
W3 =-0.4
W2 = 0.5
For AND
A B Output
0 0 0
0 1 0
1 0 0
1 1 1
Optimization: In Practice
66
9/9/26
How to Deal: Learning Rate
67
9/9/26
Approach 1 · Trial and Error
Approach 2 · Adaptive Rate
Train the network several times, each time with a different fixed learning rate, then compare the results.
Aim for the value that is “just right”:
→ Too large: loss oscillates or diverges
→ Too small: training is very slow
Easy to understand, but costly — it needs many full training runs.
The learning rate is no longer fixed — it adjusts automatically during training, often separately for each weight.
It grows or shrinks based on:
Basis of modern optimisers: AdaGrad, RMSProp, Adam.
Gradient Descent
68
9/9/26
Entire training dataset is used to compute gradient
Stochastic Gradient Descent
69
9/9/26
Use single example (or small mini-batch) at each step to compute gradient
B=1, SGD
Mini Batch SGD
70
9/9/26
A mini-batch updates the weights using a small subset of examples at a time — the middle ground between one sample (SGD) and the whole dataset.
More Accurate Gradients
Faster Training
Averaging the gradient over a small batch of examples cancels out the noise of single-sample updates.
This gives:
A whole batch of examples is processed together in one step, instead of one example at a time.
Speed comes from:
The benefit vs. the alternatives — mini-batches sit in the sweet spot:
Full-batch GD
most accurate, slowest
Mini-batch ✓
accurate and fast
Single-sample (SGD)
fastest, noisiest
Neural Network Problems
71
9/9/26
COMMON PROBLEMS
Many parameters to tune
Layers, neurons, learning rate and more must all be chosen carefully.
Risk of overfitting
The model may memorise the training data and generalise poorly.
Long training times
Large networks can take a significant time to train.
KEY PARAMETERS TO SET
Too many neurons increase training time
Keep small — typically around 0.1
Hyperparameters vs Model Parameters
72
9/9/26
Before training, some values are set by us while others are learned by the model.
SET BEFORE TRAINING
Hyperparameters
A parameter of a prior distribution — in statistics it captures prior belief before any data is observed.
Must be initialised before training begins.
LEARNED DURING TRAINING
Model Parameters
Properties the model learns on its own from the training data.
Hyperparameters
73
9/9/26
Model hyperparameters govern the entire training process — the variables you configure before training.
1 Learning Rate
2 Number of Epochs
3 Hidden Layers
4 Hidden Units
5 Activation Functions
WHY THEY MATTER
They directly control the behaviour of the training algorithm and have a significant impact on model performance.
Hyperparameter Optimisation
74
9/9/26
BENEFITS OF GOOD HYPERPARAMETERS
Hyperparameter optimisation — the process of finding the most optimal hyperparameters for a model.
COMMON OPTIMISATION ALGORITHMS
Grid Search
Exhaustively tries every combination in a defined grid.
Random Search
Samples combinations at random across the ranges.
Bayesian Optimisation
Uses past results to pick the next best combination.
Regularization
75
9/9/26
Prevents Overfitting
Adds a penalty for complexity, so the network can’t fit the training data too closely or memorise its noise.
Improves Generalisation
With complexity kept in check, the model performs better on unseen, real-world data — not just the training set.
Handles Multicollinearity
When input features are highly correlated, it stabilises the model by shrinking the influence of redundant features.
Encourages Simpler Models
Penalising large weights favours simpler solutions that are easier to interpret and more robust.
Regularization I: Dropout
76
9/9/26
Regularization I: Dropout
77
9/9/26
Regularization II: �Early Stopping
78
9/9/26
Training Time
79
9/9/26
Regularization II: �Batch Normalization (BN)
80
9/9/26
BN normalises the inputs to each layer during training, countering internal covariate shift — the constant drift in each layer’s input distribution as the network learns.
Why BN Helps
The Problem It Solves
Regularization III: �Batch Normalization (BN)
81
9/9/26
BN standardises a layer’s inputs for every mini-batch, in two steps:
Step 1 · Normalise
Step 2 · Scale & Shift
x̂ = (x − μ) / √(σ² + ε)
y = γ · x̂ + β
For each mini-batch, compute the mean (μ) and variance (σ²) of the inputs, then rescale every value to zero mean and unit variance. ε is a tiny constant that keeps the maths stable.
Apply two learnable parameters — γ (scale) and β (shift). These let the network adjust the normalised values, or even undo the normalisation, whenever that helps learning.
Normalization
82
9/9/26
Original values
Normalized Value
Offers ‘0’ mean & 1 variance
Normalization…
83
9/9/26
Features on different scales take longer to reach the minimum
Normalized data helps the network converge faster
Normalization…
84
9/9/26
Ex. DNN: Inputs Are Previous Activations
85
9/9/26
The inputs of each hidden layer are the activations from the previous layer, and must also be normalized
Ex. DNN…
86
9/9/26
The Batch Norm layer normalizes activations from Layer 1 before they reach layer 2
Regularization IV: �Gradient Clipping
Gradient Clipping keeps training numerically stable by stopping gradients from growing too large.
The problem: exploding gradients
The fix
87
9/9/26
Exploding gradient�‖g‖ = 8.0
Clipping rule (threshold c = 1.0)�if ‖g‖ > c:� g ← c · g / ‖g‖
Stable update�‖g‖ = 1.0
Regularization IV: �Gradient Clipping…
88
9/9/26
Regularization IV: �Gradient Clipping…
89
9/9/26
Gradient Clipping
By Value
By Norm
Regularization IV: �Gradient Clipping…
90
9/9/26
Bias and Variance Tradeoff
91
9/9/26
Model Complexity
92
9/9/26
Under fitting
93
9/9/26
Overfitting
94
9/9/26
Underfitting vs Overfitting
95
9/9/26
What is Bias?
96
9/9/26
What is Variance?
97
9/9/26
Intuitive Difference: Bais vs Variance
98
9/9/26
Bias-Variance Trade-off
99
9/9/26
Example of Digit Recognition
100
9/9/26
Machine
“2”
16 x 16 = 256
……
Ink → 1 No ink → 0
……
y1
y2
y10
is 1
is 2
is 0
……
0.1
0.7
0.2
The image is “2”
Example of Neural Network
101
9/9/26
Sigmoid Function
1
-1
1
-2
1
-1
1
0
4
-2
0.98
0.12
Example of Neural Network
102
9/9/26
1
-2
1
-1
1
0
4
-2
0.98
0.12
2
-1
-1
-2
3
-1
4
-1
0.86
0.11
0.62
0.83
0
0
-2
2
1
-1
Example of Neural Network
103
9/9/26
1
-2
1
-1
1
0
0.73
0.5
2
-1
-1
-2
3
-1
4
-1
0.72
0.12
0.51
0.85
0
0
-2
2
Different parameters define different function
0
0
Example of Neural Network
104
9/9/26
1
-2
1
-1
1
0
4
-2
0.98
0.12
1
-1
Example of Neural Network
105
9/9/26
……
……
……
……
……
……
……
……
y1
y2
yM
W1
W2
WL
b2
bL
x
a1
a2
y
b1
W1
x
+
b2
W2
a1
+
bL
WL
+
aL-1
b1
Neural Network
106
9/9/26
……
……
……
……
……
……
……
……
y1
y2
yM
W1
W2
WL
b2
bL
x
a1
a2
y
y
x
b1
W1
x
+
b2
W2
+
bL
WL
+
…
b1
…
Using parallel computing techniques to speed up matrix operation
Softmax
107
9/9/26
Ordinary Layer
In general, the output of network can be any value.
May not be easy to interpret
Softmax
108
9/9/26
3
-3
1
2.7
20
0.05
0.88
0.12
≈0
Network Parameters
109
9/9/26
16 x 16 = 256
……
……
……
……
……
Ink → 1
No ink → 0
……
y1
y2
y10
0.1
0.7
0.2
y1 has the maximum value
Set the network parameters such that ……
Input:
y2 has the maximum value
Input:
is 1
is 2
is 0
Softmax
Visual Information Processing
110
9/9/26
Enabling Factor of DL
111
9/9/26
Hierarchical Learning
112
9/9/26
Low-level features
output
Mid-level features
High-level features
Trainable classifier
Inspired from visual information processing, a representation of Hierarchical Learning is developed, also know as “Deep Learning”
First in 1986 by Rina Dechter
Revolution since 2012
Deep Neural Network
113
9/9/26
Output Layer
Hidden Layers
Input Layer
Input
Output
Layer 1
……
……
Layer 2
……
Layer L
……
……
……
……
……
y1
y2
yM
Deep means many hidden layers
neuron
Why Deep Network?
114
9/9/26
Layer X Size | Word Error Rate (%) | Layer X Size | Word Error Rate (%) |
1 X 2k | 24.2 | | |
2 X 2k | 20.4 | | |
3 X 2k | 18.4 | | |
4 X 2k | 17.8 | | |
5 X 2k | 17.2 | 1 X 3772 | 22.5 |
7 X 2k | 17.1 | 1 X 4634 | 22.6 |
| | 1 X 16k | 22.1 |
Seide, Frank, Gang Li, and Dong Yu. "Conversational Speech Transcription Using Context-Dependent Deep Neural Networks." Interspeech. 2011.
Not surprised, more parameters, better performance
Why Deep Network?
115
9/9/26
Any continuous function f
Can be realized by a network with one hidden layer
(given enough hidden neurons)
Why “Deep” neural network not “Fat” neural network?
116
9/9/26
Fat + Short v.s. Thin + Tall
……
Deep
……
……
Shallow
Which one is better?
The same number of parameters
Fat + Short v.s. Thin + Tall
117
9/9/26
Seide, Frank, Gang Li, and Dong Yu. "Conversational Speech Transcription Using Context-Dependent Deep Neural Networks." Interspeech. 2011.
Layer X Size | Word Error Rate (%) | Layer X Size | Word Error Rate (%) |
1 X 2k | 24.2 | | |
2 X 2k | 20.4 | | |
3 X 2k | 18.4 | | |
4 X 2k | 17.8 | | |
5 X 2k | 17.2 | 1 X 3772 | 22.5 |
7 X 2k | 17.1 | 1 X 4634 | 22.6 |
| | 1 X 16k | 22.1 |
When to use Deep Learning?
118
9/9/26
Fuel of deep learning is the big data by Andrew Ng
Deep
Learning
Machine
Learning
Amount of Data
Performance
Limitations of Deep Learning
119
9/9/26
Question for Practice
120
9/9/26
121
9/9/26
Question for Practice
122
9/9/26
123
9/9/26
Reference
124
9/9/26
Thank you!�dinesh@dtu.ac.in
Slide 125 of 74
9/9/26