CS6886: Systems for Deep Learning�Introduction to Deep Learning
Prof. Gopalakrishnan Srinivasan
CSE | BrainSeek Lab | RISE Lab | IITM
https://brainseek-lab.github.io/brainseek-lab/index.html
sgopal@cse.iitm.ac.in
CS6886
2
2011
2016
2022+
IBM Watson defeated human champions in Jeopardy
Google AlphaGo defeated Lee Sedol 4-1 in Go
Deep learning
Statistical machine learning algorithms
Evolution of Machine Learning
Apps built using LLMs (ChatGPT, Perplexity, Gemini) can generate text, images, videos, and solve Olympiad problems …
CS6886
3
Deep Learning Revolution Enabled by Hardware
Image reference: Kendall, J.D. and Kumar, S., 2020. The building blocks of a brain-inspired computer. Applied Physics Reviews, 7(1), p.011305.
AlexNet: Krizhevsky, A., Sutskever, I. and Hinton, G.E., 2012. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25.
CS6886
4
General Principle
Image generated using Perplexity Pro
CS6886
5
Deep Learning Fundamentals
CS6886
6
Fundamental Building Block: Neuron
Biological Neuron
Pre-neuron
Post-neuron
Artificial Neuron
CS6886
7
Common Activation Functions
Image reference: https://www.researchgate.net/figure/Commonly-used-activation-functions-a-Sigmoid-b-Tanh-c-ReLU-and-d-LReLU_fig3_335845675
CS6886
8
Feed-Forward Neural Network
CS6886
9
Biological Neuron
Pre-neuron
Post-neuron
Training Algorithm for Biological Neural Nets
Training Algorithm for Deep Neural Networks: Backpropagation
CS6886
10
CS6886
11
z
y = z
z
y
= 1
ReLU Derivative
y
f’(Z) exists if f(Z) is a continuous function (ReLU, Sigmoid, etc.)
Training Algorithm for Deep Neural Networks: Backpropagation
Loss function is denoted by E
CS6886
12
General Principle
Image generated using Perplexity Pro
CS6886
13
Advances in DNN Architectures
CS6886
14
Car
Airplane
Truck
Dog
Input
Convolutional Layer
Pooling Layer
Fully-Connected Layer
Output
Convolutional layers
Classifier
LeNet: LeCun, Y., Bottou, L., Bengio, Y. and Haffner, P., 2002. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11), pp.2278-2324.
AlexNet: Krizhevsky, A., Sutskever, I. and Hinton, G.E., 2012. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25.
Convolutional Neural Network Architecture
CS6886
15
Convolutional Neural Network Fundamentals
CS6886
16
VGG architecture: Simonyan, K. and Zisserman, A., 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556.
Image reference: https://neurohive.io/en/popular-networks/vgg16/
VGG16 Architecture
CS6886
17
ResNet architecture: He, K., Zhang, X., Ren, S. and Sun, J., 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778).
Image reference: https://www.researchgate.net/figure/Original-ResNet-18-Architecture_fig1_336642248
ResNet Architecture
CS6886
18
Advances in Backpropagation-based Training of DNNs
CS6886
19
Which Loss Function to Use for Classification: Sigmoid + MSE Loss ✕
Image reference: https://medium.com/@hatodi0945/a-comparison-between-mse-cross-entropy-and-hinge-loss-4d4fe63cca12
CS6886
20
Which Loss Function to Use for Classification: Softmax + Cross Entropy Loss ✓
Image reference: https://www.geeksforgeeks.org/deep-learning/the-role-of-softmax-in-neural-networks-detailed-explanation-and-applications/
CS6886
21
Which Loss Function to Use for Classification: Softmax + Cross Entropy Loss ✓
Image reference: https://wikidocs.net/235711
CS6886
22
Advances in Backpropagation based Training
CS6886
23
Weight Initialization is All you Need!
He initialization: He, K., Zhang, X., Ren, S. and Sun, J., 2015. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE international conference on computer vision (pp. 1026-1034).
Xavier initialization: Glorot, X. and Bengio, Y., 2010, March. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics (pp. 249-256). JMLR Workshop and Conference Proceedings.
CS6886
24
Dropout based Regularization
Dropout regularization: Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I. and Salakhutdinov, R., 2014. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1), pp.1929-1958.
CS6886
25
Data Augmentation Pipeline
CS6886
26
Batch Normalization for Intermediate Layers
Batch normalization: Ioffe, S. and Szegedy, C., 2015, June. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning (pp. 448-456). pmlr.
Image reference: https://paperswithcode.com/method/batch-normalization
CS6886
27
Learning Resources