1 of 128

Artificial Neural Networks: Deep NN

Prof. Dinesh K. Vishwakarma

DEPARTMENT OF INFORMATION TECHNOLOGY

DELHI TECHNOLOGICAL UNIVERSITY, DELHI.

Webpage: http://www.dtu.ac.in/Web/Departments/InformationTechnology/faculty/dkvishwakarma.php

2 of 128

History of Deep Learning

  • A Brief History of Deep Learning

2

9/9/2025

3 of 128

History of Deep Learning…

  • A Brief History of Deep Learning

3

9/9/2025

4 of 128

History of Deep Learning…

  • A Brief History of Deep Learning

4

9/9/2025

5 of 128

History of Deep Learning…

  • A Brief History of Deep Learning

5

9/9/2025

6 of 128

History of Deep Learning…

  • A Brief History of Deep Learning

6

9/9/2025

7 of 128

History of Deep Learning…

  • A Brief History of Deep Learning

7

9/9/2025

8 of 128

History of Deep Learning…

  • A Brief History of Deep Learning

8

9/9/2025

9 of 128

History of Deep Learning…

  • A Brief History of Deep Learning

9

9/9/2025

10 of 128

History of Deep Learning…

  • A Brief History of Deep Learning

10

9/9/2025

11 of 128

History of Deep Learning…

  • A Brief History of Deep Learning

11

9/9/2025

12 of 128

History of Deep Learning…

  • A Brief History of Deep Learning

12

9/9/2025

13 of 128

Artificial Neural Networks

  • Other terms/names
      • Connectionist
      • Parallel distributed processing
      • Neural computation
      • Adaptive networks..
  • History
    • 1943-McCulloch & Pitts are generally recognised as the designers of the first neural network.
    • 1949-First learning rule
    • 1969-Minsky & Papert - perceptron limitation - Death of ANN
    • 1980’s - Re-emergence of ANN - multi-layer networks

13

9/9/2025

14 of 128

Brain and Machine

  • The Brain
      • Pattern Recognition
      • Association
      • Complexity
      • Noise Tolerance
  • The Machine
      • Calculation
      • Precision
      • Logic

14

9/9/2025

15 of 128

The contrast in architecture

  • The Von Neumann architecture uses a single processing unit;
      • Tens of millions of operations per second.
      • Absolute arithmetic precision.
  • The brain uses many slow unreliable processors acting in parallel.

15

9/9/2025

16 of 128

Features of the Brain

  • Ten billion (1010) neurons.
  • On average, several thousand connections.
  • Hundreds of operations per second.
  • Die off frequently (never replaced).
  • Compensates for problems by massive parallelism.

16

9/9/2025

17 of 128

The biological inspiration

  • The brain has been extensively studied by scientists.
  • Vast complexity prevents all but rudimentary understanding.
  • Even the behaviour of an individual neuron is extremely complex.

17

9/9/2025

18 of 128

The biological inspiration…

  • Single “percepts” distributed among many neurons
  • Localized parts of the brain are responsible for certain well-defined functions (e.g. vision, motion).

18

9/9/2025

19 of 128

The Model of Neuron

19

9/9/2025

Synapses vary in strength

    • Good connections allowing a large signal
    • Slight connections allow only a weak signal.

Input:DendrIte

Output: axOn

20 of 128

The Model of Neuron

  • Artificial vs Biological Neuron
    • Axons connect to dendrites via synapses.
  • A neuron only fires if its input signal exceeds a certain amount (the threshold) in a short time period.
  • Synapses vary in strength
    • Good connections allowing a large signal
    • Slight connections allow only a weak signal.

20

9/9/2025

Biological Neuron

Artificial Neuron

Dendrite

Inputs

Cell nucleus or Soma

Nodes

Synapses

Weights

Axon

Output

21 of 128

The Artificial Neuron

21

9/9/2025

Output

Input

22 of 128

The Artificial Neuron…

  • Weighting Factors

22

9/9/2025

Input

 

 

 

 

 

 

 

 

 

 

 

 

23 of 128

The Artificial Neuron…

  • Activation Functions

23

9/9/2025

Input

 

 

 

 

 

 

 

 

 

 

Axons

(from other neurons)

Synapses

Dendrites

Cell Body

24 of 128

The Artificial Neuron…

  • Activation Functions

24

9/9/2025

Input

 

 

 

 

 

 

 

 

 

 

Axons

(from other neurons)

Synapses

Dendrites

Cell Body

+

-

 

Biasing

25 of 128

The Artificial Neuron…

  • Activation Functions

25

9/9/2025

Input

 

 

 

 

 

 

 

 

 

 

 

1

Activation Value

 

This model is aka Perceptron

Given by Rosenblatt 1958.

Biasing

26 of 128

A Simple Model of a Neuron (Perceptron)…

26

9/9/2025

27 of 128

Example of Perceptron

27

9/9/2025

28 of 128

Example of Perceptron…

28

9/9/2025

29 of 128

Perceptron (s)

29

9/9/2025

Simplified

Multi output

All inputs are connected to all outputs, hence these layers are called as dense layer

30 of 128

Single Layer NN

30

9/9/2025

31 of 128

Single Layer NN

31

9/9/2025

32 of 128

Deep NN

32

9/9/2025

33 of 128

Activation Functions

  • Mapping the activation value to the output.
  • It can be done using squashing function: unipolar and bipolar.

33

9/9/2025

1

1

-1

Unipolar

Bipolar

34 of 128

Activation Functions…

  • These are also classified as hard limiting activation function and soft limiting activation function.

34

9/9/2025

Unipolar

1

1

-1

Bipolar

Hard Limiting

1

1

-1

Bipolar

Soft Limiting

Unipolar

35 of 128

Activation Functions…

  • Many training algorithms requires the derivative of activation functions. Hence, activation function must be differentiable, E.g. logistic and sigmoid. Soft limiting function meets the requirement.

35

9/9/2025

 

Sigmoid Function

1

 

Hyperbolic Tangent Function

1

 

-1

 

The biological basis of these functions are easily established. Neurons located in different parts of the nervous system have different characteristics. Ocular motor: sigmoid, Visual Cortex: Gaussian

36 of 128

Common Activation Functions

36

9/9/2025

37 of 128

Activation Function

  • Importance: Introduce Non-Linearities

37

9/9/2025

38 of 128

Consider a Example

38

9/9/2025

 

39 of 128

Quantify Loss

  • Loss: Cost incurred from incorrect predictions

  • Empirical Loss: Total Loss over entire dataset

39

9/9/2025

40 of 128

Binary Cross Entropy Loss

  • Cross Entropy Loss can be used with models that output a probability between 0 and 1.

40

9/9/2025

41 of 128

Mean Square Error Loss

  • It is used to measure the regression models that output continuous real numbers.

41

9/9/2025

42 of 128

Training NN

  • Loss Optimization: find n/w weights that achieve lowest loss

42

9/9/2025

43 of 128

Loss Optimization

43

9/9/2025

44 of 128

Loss Optimization…

44

9/9/2025

45 of 128

Loss Optimization…

45

9/9/2025

46 of 128

Loss Optimization…

46

9/9/2025

47 of 128

Loss Optimization…

47

9/9/2025

48 of 128

Gradient Descent

48

9/9/2025

49 of 128

Computing Gradient

  • Backpropagation

  • How does a small change in one weight (ex. W2) affect the final loss J(W)?

49

9/9/2025

50 of 128

Computing Gradient…

  • Chain Rule

50

9/9/2025

Repeat this for every weight in the network using gradients from later layers

51 of 128

Ex. NN Weight Update

51

9/9/2025

A

B

C

D

E

 

 

 

 

 

 

prediction

  • Input Nodes A and B.
  • One hidden layer.
  • Six weights.
  • Prediction E.

52 of 128

Forward Pass Equations

52

9/9/2025

A

B

C

D

E

 

 

 

 

 

 

Prediction

 

 

 

 

53 of 128

Loss Function

53

9/9/2025

 

  • Loss Function calculates the gap between the prediction and actual output.
  • If loss value is very small, this means that the prediction is very close to the actual output which is good.
  • If loss value is large, this means that prediction is far from actual output which is bad.

54 of 128

Backpropagation

54

9/9/2025

 

 

Backpropagation is the algorithm to change the weights of the neural network in a manner so that the prediction gets closer to the actual output.

55 of 128

Backpropagation…

55

9/9/2025

 

Using chain rule,

For example,

 

56 of 128

Backpropagation…

56

9/9/2025

 

 

 

 

 

Hence,

57 of 128

Backpropagation…

57

9/9/2025

 

Similarly,

 

 

 

 

58 of 128

Backpropagation…

58

9/9/2025

 

 

 

 

 

 

59 of 128

Backpropagation Ex.

59

9/9/2025

A

B

C

D

E

 

 

 

 

 

 

Prediction

 

 

 

60 of 128

Backpropagation Ex.

60

9/9/2025

A

B

C

D

E

 

 

 

 

 

 

Prediction

 

 

 

 

61 of 128

Backpropagation Ex.

61

9/9/2025

A

B

C

D

E

 

 

 

 

 

 

prediction

 

 

 

62 of 128

Backpropagation Ex.

62

9/9/2025

A

B

C

D

E

 

 

 

 

 

 

prediction

 

 

 

63 of 128

Backpropagation Ex.

63

9/9/2025

A

B

C

D

E

 

 

 

 

 

 

prediction

 

 

 

64 of 128

Backpropagation Ex.

64

9/9/2025

A

B

C

D

E

 

 

 

 

 

 

prediction

 

65 of 128

Backpropagation Ex.

65

9/9/2025

A

B

C

D

E

 

 

 

 

 

 

prediction

Repeating the entire backpropagation a second time,

  • C = 1.5981 and D = 1.6984
  • Prediction = 1.5946
  • Loss = 2.8928 (2nd backpropagation)

66 of 128

Backpropagation Ex.

66

9/9/2025

A

B

C

D

E

 

 

 

 

 

 

prediction

Observations,

  • The new prediction of 1.5946 is closer to the actual output 4 than the previous prediction 1.07.
  • The second loss of 2.8928 is less than the first loss of 4.2924 signifying that the new weights of the neural network have decreased the gap between prediction and actual output.

67 of 128

Backpropagation Ex.

67

9/9/2025

A

B

C

D

E

 

 

 

 

 

 

prediction

 

68 of 128

Backpropagation Ex.

68

9/9/2025

A

B

C

D

E

 

 

 

 

 

 

prediction

After second backpropagation round,

  • Prediction = 2.1454
  • This prediction is the closest to the actual output 4 as compared to the first prediction of 1.07 and second prediction of 1.5946.
  • Hence, the weights of neural networks are changing to produce an output closer to the value 4.

69 of 128

Training Perceptrons

69

9/9/2025

t = 0.0

y

x

-1

W1 = ?

W3 = ?

W2 = ?

For AND

A B Output

0 0 0

0 1 0

1 0 0

1 1 1

  • What are the weight values?
  • Initialize with random weight values

70 of 128

Training Perceptron's

70

9/9/2025

t = 0.0

y

x

-1

W1 = 0.3

W3 =-0.4

W2 = 0.5

For AND

A B Output

0 0 0

0 1 0

1 0 0

1 1 1

71 of 128

Optimization: In Practice

  • Loss function can be difficult to optimize.
  • Optimization through Gradient Decent

  • Setting Learning Rate

71

9/9/2025

72 of 128

How to Deal: Learning Rate

  • First: Try lots of different learning rates and see what works “just right”.
  • Second: Pick the learning rate adaptively
    • Learning rate no longer fixed
    • Can be made large and small depending upon:
      • How large gradient is
      • How fast learning is happening
      • Size of particular weights
      • Etc.

72

9/9/2025

73 of 128

Gradient Descent

73

9/9/2025

Entire training dataset is used to compute gradient

74 of 128

Stochastic Gradient Descent

74

9/9/2025

Use single example (or small mini-batch) at each step to compute gradient

B=1, SGD

75 of 128

Mini Batches

  • A more accurate estimate of the gradient.
    • Smoother convergences
    • Allow for larger learning rates
  • Mini batches lead to fast training
    • Can parallelize computation
    • Archive significant speed increases on GPUs

75

9/9/2025

76 of 128

Neural Network Problems

  • Many Parameters to be set
  • Overfitting
  • long training times
  • ...
  • Parameter setting
    • Number of layers
    • Number of neurons
        • too many neurons, require more training time
    • Learning rate
      • from experience, value should be small ~0.1
    • Momentum term
    • ..

76

9/9/2025

77 of 128

Hyperparameters

  • In statistics, hyperparameter is a parameter from a prior distribution; it captures the prior belief before data is observed.
  • In any ML algorithm, these parameters need to be initialized before training a model.
  • Model parameters are the properties of training data that will learn on its own during training by the classifier or model.
    • Weights and Biases.
    • Split points in Decision Tree

77

9/9/2025

78 of 128

Hyperparameters…

  • Model Hyperparameters are the properties that govern the entire training process. The below are the variables usually configure before training a model.
    • Learning Rate
    • Number of Epochs
    • Hidden Layers
    • Hidden Units
    • Activations Functions

78

9/9/2025

Hyperparameters are important because they directly control the behaviour of the training algorithm and have a significant impact on the performance of the model is being trained.

79 of 128

Hyperparameters…

  • Choosing good hyperparameters gives two benefits:
    • Efficiently search the space of possible hyperparameters
    • Easy to manage a large set of experiments for hyperparameter tuning.
    • Hyperparameters Optimisation Techniques
      • The process of finding most optimal hyperparameters in machine learning is called hyperparameter optimization.
  • Common algorithms used for optimization is
    • Grid Search
    • Random Search
    • Bayesian Optimisation

79

9/9/2025

80 of 128

Overfittings

    • Problem of Overfitting

80

9/9/2025

81 of 128

Overfittings

    • The problem of Overfitting vs Underfitting, appears when we talk about the polynomial degree.
    • The degree represents how much flexibility is in the model, with a higher power allowing the model freedom to hit as many data points as possible.
    • An underfit model will be less flexible and cannot account for the data.
  • May have poor generalisation ability.
  • Cross-validation with some patterns
    • Typically 30% of training patterns
    • Validation set error is checked each epoch
    • Stop training if validation error goes up.

81

9/9/2025

82 of 128

Overfittings…

82

9/9/2025

83 of 128

Overfittings…_Underfit

83

9/9/2025

Underfit 1 degree polynomial model on training (left) and testing (right) datasets

84 of 128

Overfittings…_Overfit

84

9/9/2025

Overfit 25 degree polynomial model on training (left) and testing (right) datasets

85 of 128

Overfittings…_Balanced Fit

85

9/9/2025

Balanced Four degree polynomial model on training (left) and testing (right) datasets

86 of 128

Underfit_Overfit_Bestfit

86

9/9/2025

87 of 128

Regularization

  • Prevents Overfitting:
    • Regularization introduces a penalty for complexity, discouraging the model from fitting the training data too closely.
  • Improves Generalization:
    • By controlling the model's complexity, regularization helps the model perform better on unseen data, ensuring it generalizes well to real-world scenarios.Improves Generalization
  • Handles Multicollinearity:
    • In datasets with highly correlated features (multicollinearity), regularization can stabilize the model by reducing the impact of redundant features.
  • Encourages Simpler Models:
    • Regularization promotes simpler models by penalizing large coefficients, which often leads to more interpretable and robust solutions. Techniques that constrains our optimization problem to discourage complex models.

87

9/9/2025

88 of 128

Regularization I: Dropout

  • During training set some nodes to 0.

88

9/9/2025

89 of 128

Regularization I: Dropout

  • During training set some nodes to 0.

89

9/9/2025

90 of 128

Regularization II: �Early Stopping

90

9/9/2025

91 of 128

Training Time

  • How many epochs of training?
    • Stop if the error fails to improve (has reached a minimum)
    • Stop if the rate of improvement drops below a certain level
    • Stop if the error reaches an acceptable level
    • Stop when a certain number of epochs have passed

91

9/9/2025

92 of 128

Regularization II: �Batch Normalization (BN)

  •  

92

9/9/2025

93 of 128

Regularization III: �Batch Normalization (BN)

  • BN standardizes the inputs to a layer for each mini-batch of training data. This process involves two main steps:
    • Normalization
    • Scaling and Shifting

93

9/9/2025

94 of 128

Normalization

94

9/9/2025

Original values

Normalized Value

Offers ‘0’ mean & 1 variance

95 of 128

Normalization…

95

9/9/2025

Features on different scales take longer to reach the minimum

Normalized data helps the network converge faster

96 of 128

Normalization…

  • Without normalization, Ex. two features that are on drastically different scales and the network output is a linear combination of each feature vector, this means that the network learns weights for each feature that are also on different scales, which causes the large feature to simply drown out the small feature.
  • In GD, to get minima network would have to make a large update to one weight compared to the other weight. This can cause the GD trajectory to oscillate back and forth along one dimension, thus taking more steps to reach the minima.

96

9/9/2025

A narrow valley causes gradient descent to bounce from one slope to the other

97 of 128

Ex. DNN

  • Consider any of the hidden layers of a network. The activations from the previous layer are simply the inputs to this layer.
  • For instance, from the perspective of Layer 2 in the picture below, if we "blank out" all the previous layers, the activations coming from Layer 1 are no different from the original inputs.

97

9/9/2025

The inputs of each hidden layer are the activations from the previous layer, and must also be normalized

98 of 128

Ex. DNN…

  •  

98

9/9/2025

The Batch Norm layer normalizes activations from Layer 1 before they reach layer 2 

99 of 128

Regularization IV: �Gradient Clipping

  • Gradient Clipping is the process that helps maintain numerical stability by preventing the gradients from growing too large.
  • When training a neural network, the loss gradients are computed through backpropagation.
  • However, if these gradients become too large, the updates to the model weights can also become excessively large, leading to numerical instability.
  • This can result in the model producing NaN (Not a Number) values or overflow errors, which can be problematic. This problem is often referred to as 'gradient exploding', it could be solved by clipping the gradient to the value that we want it to be.

99

9/9/2025

100 of 128

Regularization IV: �Gradient Clipping…

100

9/9/2025

101 of 128

Regularization IV: �Gradient Clipping…

  • Value: A straightforward way to prevent exploding gradients is gradient clipping by value, each gradient component is limited to lie within a specified range [min_value, max_value]. If a gradient component exceeds max_value, it is set to max_value; if it is below min_value, it is set to min_value, otherwise, no change.

101

9/9/2025

Gradient Clipping

By Value

By Norm

102 of 128

Regularization IV: �Gradient Clipping…

  • By Norm
    • Gradient Clipping by Norm is a technique to prevent exploding gradients by rescaling the entire gradient vector if its norm exceeds a specified threshold.

  • The threshold can be decided based on training of neural networks for some epochs, and then look at the statistics of the gradient norms. The average value of gradient norms is a good initial trial.

102

9/9/2025

 

 

103 of 128

Example of Digit Recognition

103

9/9/2025

Machine

“2”

16 x 16 = 256

……

Ink → 1 No ink → 0

……

y1

y2

y10

is 1

is 2

is 0

……

0.1

0.7

0.2

The image is “2”

 

104 of 128

Example of Neural Network

104

9/9/2025

Sigmoid Function

1

-1

1

-2

1

-1

1

0

4

-2

0.98

0.12

105 of 128

Example of Neural Network

105

9/9/2025

1

-2

1

-1

1

0

4

-2

0.98

0.12

2

-1

-1

-2

3

-1

4

-1

0.86

0.11

0.62

0.83

0

0

-2

2

1

-1

106 of 128

Example of Neural Network

106

9/9/2025

1

-2

1

-1

1

0

0.73

0.5

2

-1

-1

-2

3

-1

4

-1

0.72

0.12

0.51

0.85

0

0

-2

2

 

Different parameters define different function

 

 

0

0

107 of 128

Example of Neural Network

107

9/9/2025

 

1

-2

1

-1

1

0

4

-2

0.98

0.12

 

 

 

 

 

 

1

-1

 

108 of 128

Example of Neural Network

108

9/9/2025

……

……

……

……

……

……

……

……

y1

y2

yM

W1

W2

WL

b2

bL

x

a1

a2

y

b1

W1

x

+

 

b2

W2

a1

+

 

bL

WL

+

 

aL-1

b1

109 of 128

Neural Network

109

9/9/2025

 

 

……

……

……

……

……

……

……

……

y1

y2

yM

W1

W2

WL

b2

bL

x

a1

a2

y

y

 

x

b1

W1

x

+

 

b2

W2

+

bL

WL

+

b1

Using parallel computing techniques to speed up matrix operation

110 of 128

Softmax

  • Softmax layer as the output layer

110

9/9/2025

Ordinary Layer

In general, the output of network can be any value.

May not be easy to interpret

111 of 128

Softmax

  • Softmax layer as the output layer

111

9/9/2025

3

-3

1

2.7

20

0.05

0.88

0.12

0

 

112 of 128

Network Parameters

112

9/9/2025

16 x 16 = 256

……

……

……

……

……

Ink → 1

No ink → 0

……

y1

y2

y10

0.1

0.7

0.2

y1 has the maximum value

Set the network parameters such that ……

Input:

y2 has the maximum value

Input:

is 1

is 2

is 0

Softmax

 

113 of 128

Visual Information Processing

  • Visual information processed by our brain is multi-layered.

113

9/9/2025

114 of 128

Enabling Factor of DL

  • Training of deep networks was made computationally feasible by:
      • Faster CPU’s
      • The move to parallel CPU architectures
      • Advent of GPU computing
  • Neural networks are often represented as a matrix of weight vectors.
  • GPU’s are optimized for very fast matrix multiplication
  • 2008 - Nvidia’s CUDA library for GPU computing is released.

114

9/9/2025

115 of 128

Hierarchical Learning

115

9/9/2025

Low-level features

output

Mid-level features

High-level features

Trainable classifier

Inspired from visual information processing, a representation of Hierarchical Learning is developed, also know as “Deep Learning”

First in 1986 by Rina Dechter

Revolution since 2012

116 of 128

Deep Neural Network

116

9/9/2025

Output Layer

Hidden Layers

Input Layer

Input

Output

Layer 1

……

……

Layer 2

……

Layer L

……

……

……

……

……

y1

y2

yM

Deep means many hidden layers

neuron

117 of 128

Why Deep Network?

117

9/9/2025

Layer X Size

Word Error Rate (%)

Layer X Size

Word Error Rate (%)

1 X 2k

24.2

2 X 2k

20.4

3 X 2k

18.4

4 X 2k

17.8

5 X 2k

17.2

1 X 3772

22.5

7 X 2k

17.1

1 X 4634

22.6

1 X 16k

22.1

Seide, Frank, Gang Li, and Dong Yu. "Conversational Speech Transcription Using Context-Dependent Deep Neural Networks." Interspeech. 2011.

Not surprised, more parameters, better performance

118 of 128

Why Deep Network?

  • Universal Theorem

118

9/9/2025

Any continuous function f

Can be realized by a network with one hidden layer

(given enough hidden neurons)

Why “Deep” neural network not “Fat” neural network?

119 of 128

119

9/9/2025

Fat + Short v.s. Thin + Tall

……

Deep

……

……

Shallow

Which one is better?

The same number of parameters

120 of 128

Fat + Short v.s. Thin + Tall

120

9/9/2025

Seide, Frank, Gang Li, and Dong Yu. "Conversational Speech Transcription Using Context-Dependent Deep Neural Networks." Interspeech. 2011.

Layer X Size

Word Error Rate (%)

Layer X Size

Word Error Rate (%)

1 X 2k

24.2

2 X 2k

20.4

3 X 2k

18.4

4 X 2k

17.8

5 X 2k

17.2

1 X 3772

22.5

7 X 2k

17.1

1 X 4634

22.6

1 X 16k

22.1

121 of 128

When to use Deep Learning?

  • Data size is large
  • High end infrastructure
  • Lack of domain understanding
  • Complex problem such as image classification, speech recognition etc.

121

9/9/2025

Fuel of deep learning is the big data by Andrew Ng

Deep

Learning

Machine

Learning

Amount of Data

Performance

122 of 128

Limitations of Deep Learning

  • Very slow to train
  • Models are very complex, with lot of parameters to optimize:
    • Initialization of weights
    • Layer-wise training algorithm
    • Neural architecture
      • Number of layers
      • Size of layers
      • Type – regular, pooling, max pooling, soft max
    • Fine-tuning of weights using back propagation

122

9/9/2025

123 of 128

Question for Practice

123

9/9/2025

124 of 128

124

9/9/2025

125 of 128

Question for Practice

125

9/9/2025

126 of 128

126

9/9/2025

127 of 128

Reference

  • https://towardsdatascience.com/batch-norm-explained-visually-how-it-works-and-why-neural-networks-need-it-b18919692739/

127

9/9/2025

128 of 128

Thank you!�dinesh@dtu.ac.in

Slide 128 of 74

9/9/2025