1 of 52

Convolutional Neural Networks Variants

Prof. Dinesh Kumar Vishwakarma,

DEPARTMENT OF INFORMATION TECHNOLOGY

DELHI TECHNOLOGICAL UNIVERSITY, DELHI.

Webpage: http://www.dtu.ac.in/Web/Departments/InformationTechnology/faculty/dkvishwakarma.php

Email: dinesh@dtu.ac.in

2 of 52

Introduction

  • Since 1998, number of CNN models are developed.
    • Recent models are so deep and extremely difficult to visualize. Hence, it becomes like ‘black box’.
    • Starting from LeNet 5 to Latest ResNeXt-50.

3 of 52

Historical Progress Year wise as on 05.04.20

4 of 52

LeNet 5

  • Developed in 1998 by Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner.
    • One of the simplest architecture.
    • It has 02-Convo layer and 03-FC layer. Hence, total 05 layers and very commonly known as LeNet 5.
    • It has 60000 parameters.
    • Implemented on CPU.
    • Novel: Stacking of Conv layer and FC Layer.

T

Tanh

s

SoftMax

5 of 52

Demo of LeNet5

http://yann.lecun.com/exdb/lenet/index.html

6 of 52

Properties of LeNet5/CNN

  • Shift Invariance

Invariance with respect to vertical translations is necessary since the positioning of individual characters in a string is never perfect.

Scale Invariance

Scale invariance is achieved over a wide range of sizes.

http://yann.lecun.com/exdb/lenet/index.html

7 of 52

Properties of LeNet5/CNN…

  • Rotation Invariance

LeNet-5's invariance to small rotations (+-40 degrees).

http://yann.lecun.com/exdb/lenet/index.html

Squeezing Invariance

robustness to variations of the aspect ratio

8 of 52

Properties of LeNet5/CNN…

  • Stroke Width Invariance

http://yann.lecun.com/exdb/lenet/index.html

The robustness to stroke width variation allows LeNet-5 to operate directly on "raw" pixel images without requiring unreliable preprocessing such as line thinning

9 of 52

AlexNet

  • Developed in 2012 by Alex Krizhevsky, Ilya Sutskever, Geoffrey Hinton. University of Toronto, Canada.
    • It has 05-Convo layer and 03-FC layer. Hence, total 08 layers and very commonly known as AlexNet. Novel: First to use ReLU
    • It has 60M parameters and stacked few more layers onto LeNet-5.
    • In 2012, it was “one of the largest CNNs to date on the subsets of ImageNet.” it also won ImageNet challenge in 2012.

R

ReLU

S

SoftMax

10 of 52

Accuracy (Top-1 & 5 Acc.)

    • Measure the performance of model
      • Consider a simple classification problem using deep learning.
      • We gave a input image (blueberry) to the model and get the prediction results (with probability) as follows:
        • cherry: 0.35
        • raspberry: 0.25
        • blueberry: 0.2
        • strawberry: 0.1
        • apple: 0.06
        • orange: 0.04

11 of 52

Accuracy (Top-1 & 5 Acc.)…

    • Using top-1 accuracy
      • We consider the O/P as TRUE prediction as a “cherry”.
    • Using top-5 accuracy
      • We consider the output as FALSE, because blueberry is among the top-5 guesses. A prediction result is as given:
  • cherry: 0.35
  • raspberry: 0.25
  • blueberry: 0.2
  • strawberry: 0.1
  • apple: 0.06
  • orange: 0.04

Model predicted correctly 2 images and the true label turns up 3 times in the top 5 predicted labels

 

Top-1: It measures the proportion of examples for which the predicted label matches the single target label. 2/5=0.4

Top-5: It considers a classification correct if any of the five predictions matches the target label=3/5

12 of 52

AlexNet…

It takes in input a color (RGB) image of dimension 224 X 224.

  • First, a Convolution Layer (CL) of 96 filters of size 11 X 11 and stride 4.
  • Next, a Max-Pooling Layer (M-PL) of filter size 3 X 3 and stride = 2.
  • Again, a CL of 256 filters of size 5 X 5 and stride = 4.
  • Then, a M-PL of filter size 3 X 3 and stride = 2.
  • Again, a CL of 384 filters of size 3 X 3 and stride = 4.
  • Again, a CL of 384 filters of size 3 X 3 and stride = 4.
  • Again, a CL of 256 filters of size 3 X 3 and stride = 4.
  • Then, a M-Pl of filter size 3 X 3 and stride = 2.
  • The output of the last layer, when converted into input-layer like for the Fully Connected Block consists of 9261 nodes, fully connected to a hidden layer with 4096 nodes.
  • The first hidden layer is again fully connected to another hidden layer consisting 4096 nodes.
  • This last hidden layer is fully connected to the output layer implementing “softmax regression” of 1000 nodes.

13 of 52

VGG-16

  • VGG-16 invented by Simonyan et al. (2014) of Visual Geometry Group, CNN gets deeper and deeper with time. They also designed VGG-19.
    • It has 13-Convo layer, 03-FC layer and ReLU Activation. Total 16 layers and that why known as VGG-16. Novel: designing of deep N/W, twice as compared to AlexNet.
    • It has 130M parameters, take 500MB storage space, and stacked more layers onto AlexNet. Uses small filter size 2×2 and 3×3.

R

ReLU

s

SoftMax

Runners up of the ILSVRC-2014 competition

14 of 52

VGG-16

It takes in a color (RGB) image of 224 X 224 dimensions.

    • Convolutional Layer (CL) of 64 filters.
    • CL of 64 filters again.
    • Max-Pooling Layer (M-PL)
    • CL of 128 filters.
    • CL of 128 filters again.
    • M-PL.
    • CL of 256 filters.
    • CL of 256 filters again.
    • CL of 256 filters again.
    • M-PL.
    • CL of 512 filters.

Runners up of the ILSVRC-2014 competition

    • CL of 512 filters again.
    • CL of 512 filters again.
    • M-PL.
    • CL of 512 filters.
    • CL of 512 filters again.
    • CL of 512 filters again.
    • M-PL.
    • The output of the last Pooling Layer is fed into a FC hidden layer consisting of 4096 nodes.
    • This is again FC to another hidden layer consisting again of 4096 nodes.
    • This is FC to an output layer implementing “softmax regression”, classifying among 1000 classes of objects.

15 of 52

GoogleNet or Inception Network V1

  • Developved by Szegedy et al.2014. It has 22-layer architecture with 5M parameters.
    • It uses the concept of N/W in N/W approach. It is done by Inception module.
    • The Inception module is designed by a product of research on approximating sparse structures and each module presents three ideas.
      • a) Having parallel towers of convolutions with different filters, b) followed by concatenation, captures different features at 1×1, 3×3 and 5×5, thereby c) ‘clustering’ them.
      • The above idea is motivated by Arora et al. 2014, where it was suggested that a layer-by layer construction in which one should analyse the correlation statistics of the last layer and cluster them into groups of units with high correlation.
      • 1×1 convolutions are used for dimensionality reduction to remove computational bottlenecks.
      • Due to the activation function from 1×1 convolution, its addition also adds nonlinearity. This idea is based on the Network In Network concept.
      • The authors also introduced two auxiliary classifiers to encourage discrimination in the lower stages of the classifier, to increase the gradient signal that gets propagated back, and to provide additional regularization.

Winner of the ILSVRC 2014

16 of 52

Inception V1 or GoogLeNet…

17 of 52

Inception V1 or GoogLeNet…

  • Challenges involved in image recognition leads to the development of more deeper network with variation of filter size.
    • Salient parts in the image can have extremely large variation in size. Ex. The area occupied by the dog is different in each image.
    • Because of the huge variation in the location of the information, choosing the right kernel size for the convolution operation becomes tough. A larger kernel is preferred for information that is distributed more globally, and a smaller kernel is preferred for information that is distributed more locally.
    • Very deep networks are prone to overfitting and also difficult to pass gradient updates through the entire network.
    • Naively stacking large convolution operations is computationally expensive.

18 of 52

Inception V1 or GoogLeNet…

  • Solution of these challenges are to the development of more deeper network with variation of filter size.
    • To have filters with multiple sizes, which operates on the same level.
    • The network layer would get a bit “wider” rather than “deeper”.
    • Hence, inception module is designed to perform convolution on an I/P with 3 different sizes of filters (1x1, 3x3, 5x5).
    • Additionally, max pooling is also performed.
    • The outputs are concatenated and sent to the next inception module.

19 of 52

Inception V1 or GoogLeNet…

  • Solution to computational expensive of deep network is addressed.
    • To make it cheaper, the authors limit the number of input channels by adding an extra 1x1 convolution before the 3x3 and 5x5 convolutions.
    • Though adding an extra operation may seem counterintuitive, 1x1 convolutions are far more cheaper than 5x5 convolutions, and the reduced number of input channels also help.
    • Do note that however, the 1x1 convolution is introduced after the max pooling layer, rather than before.

20 of 52

Actual Diagram of Network

Inception V1 or GoogLeNet

21 of 52

Inception V1 or GoogLeNet…

  • Architecture Novel: stacking of modules instead of Conv layer
    • GoogLeNet has 9 inception modules stacked linearly. It is 22 layers deep (27, including the pooling layers). It uses global average pooling at the end of the last inception module.
    • Needless to say, it is a pretty deep classifier. As with any very deep network, it is subject to the vanishing gradient problem.
    • To prevent the middle part of the network from “dying out”, the authors introduced two auxiliary classifiers.
    • They essentially applied softmax to the outputs of two of the inception modules, and computed an auxiliary loss over the same labels.
    • The total loss function is a weighted sum of the auxiliary loss and the real loss. Weight value used in the paper was 0.3 for each auxiliary loss and used for training purpose.

22 of 52

Inception V2

  • Inception V2 and V3 both were presented in same work by Szegedy et al. (Dec. 2015).
    • Number of variants are proposed to increase accuracy and reduce the computational cost.
    • Inception V2 explores: Reduce representational bottleneck.
    • The intuition was that, neural networks perform better when convolutions didn’t alter the dimensions of the input drastically.
    • Reducing the dimensions too much may cause loss of information, known as a “representational bottleneck”.
    • Using smart factorization methods, convolutions can be made more efficient in terms of computational complexity.

Szegedy, Christian, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. "Rethinking the inception architecture for computer vision." In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818-2826. 2016.

23 of 52

Inception V2…

  • Solutions of these problems are:
    • Factorize 5x5 convolution to two 3x3 convolution operations to improve computational speed.
    • Although this may seem counterintuitive, a 5x5 convolution is 2.78 times more expensive than a 3x3 convolution.
    • So stacking two 3x3 convolutions infact leads to a boost in performance. This is illustrated in the image.

Szegedy, Christian, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. "Rethinking the inception architecture for computer vision." In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818-2826. 2016.

(A)

24 of 52

Inception V2…

  • Solutions of these problems are:
    • Moreover, they factorize convolutions of filter size nxn to a combination of 1xn and nx1 convolutions.
    • For example, a 3x3 convolution is equivalent to first performing a 1x3 convolution, and then performing a 3x1 convolution on its output.
    • They found this method to be 33% more cheaper than the single 3x3 convolution. This is illustrated in the below image.

Szegedy, Christian, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. "Rethinking the inception architecture for computer vision." In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818-2826. 2016.

(B)

25 of 52

Inception V2…

  • Solutions of representational bottleneck are:
    • The filter banks in the module were expanded (made wider instead of deeper) to remove the representational bottleneck. If the module was made deeper instead, there would be excessive reduction in dimensions, and hence loss of information.
    • The above three principles were used to build three different types of inception modules and these may be denoted as Fig. 5: A, Fig. 6: B and Fig. 7: C. (Original Paper)

Szegedy, Christian, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. "Rethinking the inception architecture for computer vision." In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818-2826. 2016.

(C)

26 of 52

Inception V2 Architecture Parameters

Szegedy, Christian, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. "Rethinking the inception architecture for computer vision." In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818-2826. 2016.

27 of 52

Inception V3

  • In Inception V2, it was noticed that the auxiliary classifiers didn’t contribute much until near the end of the training process, when accuracies were nearing saturation.
  • They argued that the function as regularizes, especially if they have BatchNorm or Dropout operations.
  • Possibilities to improve on the Inception v2 without drastically changing the modules were to be investigated.
  • In Inception V3 uses the following in addition to inception V2.
    • RMSProp Optimizer.
    • Factorized 7x7 convolutions.
    • BatchNorm in the Auxillary Classifiers.
    • Label Smoothing (A type of regularizing component added to the loss formula that prevents the network from becoming too confident about a class. Prevents over fitting).

Szegedy, Christian, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. "Rethinking the inception architecture for computer vision." In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818-2826. 2016.

28 of 52

Layout Diagram of Inception V3

Inception V3

Novel: First Network to have batch normalization. 

Szegedy, Christian, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. "Rethinking the inception architecture for computer vision." In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818-2826. 2016.

29 of 52

Residual Neural Networks: ResNet-50 (2015)

    • Last few CNNs models, It was seen that increasing number of layers in the design, and achieving better performance and also with the increase of network depth, accuracy gets saturated and then slowly degraded.
    • Performance of the model deteriorates both on the training and testing data & reason was not OVERFITTING.
    • Key reasons: Initialization of the network, optimization function, or, more importantly, the problem of vanishing or exploding gradients.
    • Hence, Kaiming et al. from Microsoft Research addressed this problem with ResNet using skip connections (a.k.a. shortcut connections, residuals), while building deeper models.
    • ResNet-50 has 26M parameters.
    • The basic building block for ResNets are the conv and identity blocks.

30 of 52

ResNet-50 (2015)…

Novel: a) Popularised skip connections b) Designing even deeper CNNs (up to 152 layers) without compromising model’s generalization power c) Among the first to use batch normalization

31 of 52

ResNet-50 (2015)…Skip connections works in two ways:

          • Alleviate the issue of vanishing gradient by setting up an alternate shortcut for the gradient to pass through.
          • They enable the model to learn an identity function. This ensures that the higher layers of the model do not perform any worse than the lower layers.
      • The residual blocks make it considerably easier for the layers to learn identity functions. As a result, ResNet improves the efficiency of DNNs with more neural layers while minimizing the percentage of errors.
      • In other words, the skip connections add the outputs from previous layers to the outputs of stacked layers, making it possible to train much deeper networks than previously possible.

FLOPs: floating point operations per seconds

32 of 52

ResNet-50 (2015)…

33 of 52

Xception (2016)

    • Xception is an adaptation from Inception, where the Inception modules have been replaced with depthwise separable convolutions.
    • It has also roughly the same number of parameters as Inception-v1 (23M).
    • Xception takes the Inception hypothesis to an eXtreme.
      • Firstly, cross-channel (or cross-feature map) correlations are captured by 1×1 convolutions.
      • Consequently, spatial correlations within each channel are captured via the regular 3×3 or 5×5 convolutions
    • The idea of an extreme means performing 1×1 to every channel, then performing a 3×3 to each output.
    • This is identical to replacing the Inception module with depthwise separable convolutions.
    • Novel: Introduced CNN based entirely on depthwise separable convolution layers.

34 of 52

Xception (2016)…

  1. Depthwise convolution is the channel-wise n×n spatial convolution. Suppose in the figure above, we have 5 channels, then we will have 5 n×n spatial convolution.
  2. Pointwise convolution actually is the 1×1 convolution to change the dimension.
  3. Compared with conventional convolution: Do not need to perform convolution across all channels and it means the number of connections are fewer and the model is lighter.

Original Depthwise Separable Convolution

35 of 52

Xception (2016)…

The Modified Depthwise Separable Convolution used as an Inception Module in Xception, so called “extreme” version of Inception module (n=3 here)

  • This modification is motivated by the inception module in Inception-v3 that 1×1 convolution is done first before any n×n spatial convolutions.
  • Thus, it is a bit different from the original one. (n=3 here since 3×3 spatial convolutions are used in Inception-v3.)

36 of 52

Xception (2016)

37 of 52

Xception (2016)…

Overall Architecture of Xception (Entry Flow > Middle Flow > Exit Flow)

SeparableConv is the modified depthwise separable convolution. We can see that Separable Convs are treated as Inception Modules and placed throughout the whole deep learning architecture

38 of 52

Inception V4 (2016)

    • The researchers from Google strike again with Inception-v4, 43M parameters.
    • Again, this is an improvement from Inception-v3.
    • The main difference is the Stem group and some minor changes in the Inception-C module.
    • The authors also “made uniform choices for the Inception blocks for each grid size.”
    • They also mentioned that having “residual connections leads to dramatically improved training speed.”
    • Inception-v4 works better because of increased model size.

39 of 52

Inception V4 (2016)…

Improvement from Inception V3

  • Change in Stem module.
  • Adding more Inception modules.
  • Uniform choices of Inception-v3 modules, meaning using the same number of filters for every module.

40 of 52

Inception V4 (2016)…

41 of 52

Inception-ResNet-V2 (2016)

    • In the same paper as Inception-v4, the same authors also introduced Inception-ResNets — a family of Inception-ResNet-v1 and Inception-ResNet-v2.
    • The latter member of the family has 56M parameters.
    • What’s improved from the previous version, Inception-v3?
      • Converting Inception modules to Residual Inception blocks.
      • Adding more Inception modules.
      • Adding a new type of Inception module (Inception-A) after the Stem module.

42 of 52

Inception-ResNet-V2 (2016)…

43 of 52

ResNeXt-50 (2017)

    • ResNeXt-50 has 25M parameters (ResNet-50 has 25.5M).
    • What’s different about ResNeXts is the adding of parallel towers/branches/paths within each module, as seen above indicated by ‘total 32 towers.’
    • The model name, ResNeXt, contains Next.
    • It means the next dimension, on top of the ResNet.
    • This next dimension is called the “cardinality” dimension. And ResNeXt becomes the 1st Runner Up of ILSVRC classification task.
    • What’s novel?
      • Scaling up the number of parallel towers (“cardinality”) within a module (well I mean this has already been explored by the Inception network…)

44 of 52

ResNeXt-50 (2017)…

45 of 52

Network In Network (2014)

    • Recall that in a convolution, the value of a pixel is a linear combination of the weights in a filter and the current sliding window.
    • The authors proposed that instead of this linear combination, let’s have a mini neural network with 1 hidden layer.
    • This is what they coined as Mlpconv.
    • So what we’re dealing with here is a (simple 1 hidden layer) network in a (convolutional neural) network.
    • This idea of Mlpconv is likened to 1×1 convolutions, and became the main feature for Inception architectures.
    • What’s novel?
        • MLP convolutional layers, 1×1 convolutions
        • Global average pooling (taking average of each feature map, and feeding the resulting vector into the softmax layer).

46 of 52

Network In Network (2014)…

  • The overall structure of Network In Network. NINs include the stacking of three mlpconv layers and one global average pooling layer

NINs

47 of 52

Cost Function

  •  

Advertisement

Benefit

 

Company

48 of 52

Cost Function…

  •  

 

 

 

 

 

 

49 of 52

Regularization in ML

    • One of the major aspects of training your machine learning model is avoiding overfitting. 
    • The model will have a low accuracy if it is overfitting. 
    • This happens because model is trying too hard to capture the noise in your training dataset. By noise we mean the data points that don’t really represent the true properties of your data, but random chance.
    • Learning such data points, makes your model more flexible, at the risk of overfitting.

50 of 52

Regularization in ML…

  •  

 

Regularization Term

Regularization Parameter

51 of 52

Regularization in ML…

 

 

 

 

Advertisement

Benefit

52 of 52

Thank You�Contact: dinesh@dtu.ac.in �Mobile: +91-9971339840