1 of 52

ML 101

2 of 52

Deep Learning

3 of 52

What is Deep Learning

  • Some of the most impressive advances in artificial intelligence in recent years have been in the field of deep learning.
  • Deep learning is an approach to machine learning characterized by deep stacks of computations. This depth of computation is what has enabled deep learning models to disentangle the kinds of complex and hierarchical patterns found in the most challenging real-world datasets.
  • Through their power and scalability neural networks have become the defining model of deep learning.

4 of 52

The linear Unit

  • The input is x. Its connection to the neuron has a weight which is w. Whenever a value flows through a connection, you multiply the value by the connection's weight. For the input x, what reaches the neuron is w * x. A neural network "learns" by modifying its weights.
  • The b is a special kind of weight we call the bias. The bias doesn't have any input data associated with it; instead, we put a 1 in the diagram so that the value that reaches the neuron is just b (since 1 * b = b). The bias enables the neuron to modify the output independently of its inputs.
  • The y is the value the neuron ultimately outputs. To get the output, the neuron sums up all the values it receives through its connections. This neuron's activation is y = w * x + b, or as a formula y=wx+b .

5 of 52

Example - The Linear Unit as a Model

Let's think about how this might work on a dataset like 80 Cereals. Training a model with 'sugars' (grams of sugars per serving) as input and 'calories' (calories per serving) as output, we might find the bias is b=90 and the weight is w=2.5. We could estimate the calorie content of a cereal with 5 grams of sugar per serving like this:

6 of 52

Multiple Inputs

  • The formula for this neuron would be y=w0x0+w1x1+w2x2+b . A linear unit with two inputs will fit a plane, and a unit with more inputs than that will fit a hyperplane.

7 of 52

Layers

  • Neural networks typically organize their neurons into layers. When we collect together linear units having a common set of inputs we get a dense layer.
  • You could think of each layer in a neural network as performing some kind of relatively simple transformation.

8 of 52

The Activation Function

  • It turns out, however, that two dense layers with nothing in between are no better than a single dense layer by itself. Dense layers by themselves can never move us out of the world of lines and planes. What we need is something nonlinear. What we need are activation functions.

9 of 52

The Activation Function (2)

  • An activation function is simply some function we apply to each of a layer's outputs (its activations). The most common is the rectifier function max(0,x).

10 of 52

The Activation Function (3)

11 of 52

Stacking Dense Layers

  • Now that we have some nonlinearity, let's see how we can stack layers to get complex data transformations.
  • The layers before the output layer are sometimes called hidden since we never see their outputs directly.

12 of 52

Stochastic Gradient Descent

13 of 52

Stochastic Gradient Descent

  • If we can successfully train a network to do that, its weights must represent in some way the relationship between those features and that target as expressed in the training data.
  • In addition to the training data, we need two more things:
    • A "loss function" that measures how good the network's predictions are.
    • An "optimizer" that can tell the network how to change its weights

​

14 of 52

The Loss Function

  • The loss function measures the disparity between the the target's true value and the value the model predicts.

​

15 of 52

The Optimizer

  • The optimizer is an algorithm that adjusts the weights to minimize the loss.
  • Virtually all of the optimization algorithms used in deep learning belong to a family called stochastic gradient descent.

16 of 52

The Optimizer (2)

  • Each iteration's sample of training data is called a minibatch (or often just "batch"), while a complete round of the training data is called an epoch.

17 of 52

Learning Rate and Batch Size

  • Notice that the line only makes a small shift in the direction of each batch (instead of moving all the way). The size of these shifts is determined by the learning rate. A smaller learning rate means the network needs to see more minibatches before its weights converge to their best values.
  • The learning rate and the size of the minibatches are the two parameters that have the largest effect on how the SGD training proceeds. Their interaction is often subtle and the right choice for these parameters isn't always obvious. (We'll explore these effects in the exercise.)
  • Fortunately, for most work it won't be necessary to do an extensive hyperparameter search to get satisfactory results. Adam is an SGD algorithm that has an adaptive learning rate that makes it suitable for most problems without any parameter tuning (it is "self tuning", in a sense). Adam is a great general-purpose optimizer.

18 of 52

Overfitting and Underfitting

19 of 52

Interpreting the Learning Curves

  • You might think about the information in the training data as being of two kinds: signal and noise.
  • When we train a model we've been plotting the loss on the training set epoch by epoch.
  • These plots we call the learning curves.

20 of 52

Interpreting the Learning Curves (2)

21 of 52

Interpreting the Learning Curves (3)

  • This trade-off indicates that there can be two problems that occur when training a model: not enough signal or too much noise. Underfitting the training set is when the loss is not as low as it could be because the model hasn't learned enough signal. Overfitting the training set is when the loss is not as low as it could be because the model learned too much noise. The trick to training deep learning models is finding the best balance between the two.

22 of 52

Capacity

  • A model's capacity refers to the size and complexity of the patterns it is able to learn. For neural networks, this will largely be determined by how many neurons it has and how they are connected together. If it appears that your network is underfitting the data, you should try increasing its capacity.
  • You can increase the capacity of a network either by making it wider (more units to existing layers) or by making it deeper (adding more layers). Wider networks have an easier time learning more linear relationships, while deeper networks prefer more nonlinear ones. Which is better just depends on the dataset.

23 of 52

Early StoppinG

  • We mentioned that when a model is too eagerly learning noise, the validation loss may start to increase during training. To prevent this, we can simply stop the training whenever it seems the validation loss isn't decreasing anymore. Interrupting the training this way is called early stopping.

24 of 52

Dropout and Batch Normalization

25 of 52

Dropout

  • This is the idea behind dropout. To break up these conspiracies, we randomly drop out some fraction of a layer's input units every step of training, making it much harder for the network to learn those spurious patterns in the training data. Instead, it has to search for broad, general patterns, whose weight patterns tend to be more robust.

26 of 52

DROPout (2)

  • You could also think about dropout as creating a kind of ensemble of networks. The predictions will no longer be made by one big network, but instead by a committee of smaller networks. Individuals in the committee tend to make different kinds of mistakes, but be right at the same time, making the committee as a whole better than any individual. (If you're familiar with random forests as an ensemble of decision trees, it's the same idea.)

27 of 52

Batch Normalization

  • With neural networks, it's generally a good idea to put all of your data on a common scale, perhaps with something like scikit-learn's StandardScaler or MinMaxScaler. The reason is that SGD will shift the network weights in proportion to how large an activation the data produces. Features that tend to produce activations of very different sizes can make for unstable training behavior.

28 of 52

Batch Normalization (2)

  • Now, if it's good to normalize the data before it goes into the network, maybe also normalizing inside the network would be better! In fact, we have a special kind of layer that can do this, the batch normalization layer. A batch normalization layer looks at each batch as it comes in, first normalizing the batch with its own mean and standard deviation, and then also putting the data on a new scale with two trainable rescaling parameters. Batchnorm, in effect, performs a kind of coordinated rescaling of its inputs.

29 of 52

Binary Classification

30 of 52

Binary Classification

  • Accuracy is one of the many metrics in use for measuring success on a classification problem. Accuracy is the ratio of correct predictions to total predictions: accuracy = number_correct / total.
  • The problem with accuracy (and most other classification metrics) is that it can't be used as a loss function. SGD needs a loss function that changes smoothly, but accuracy, being a ratio of counts, changes in "jumps". So, we have to choose a substitute to act as the loss function.

31 of 52

Binary Classification (2)

  • For classification, what we want instead is a distance between probabilities, and this is what cross-entropy provides. Cross-entropy is a sort of measure for the distance from one probability distribution to another.

32 of 52

Making Probabilities with the Sigmoid Function

  • The cross-entropy and accuracy functions both require probabilities as inputs, meaning, numbers from 0 to 1. To covert the real-valued outputs produced by a dense layer into probabilities, we attach a new kind of activation function, the sigmoid activation.

33 of 52

CNN

34 of 52

convolutional neural network (CNN)

  • Convolutional Neural Network (ConvNet/CNN) is a Deep Learning algorithm which can take in an input image, assign importance (learnable weights and biases) to various aspects/objects in the image and be able to differentiate one from the other.

35 of 52

Why ConvNets over Feed-Forward Neural Nets?

  • ConvNet is able to successfully capture the Spatial and Temporal dependencies in an image through the application of relevant filters. The architecture performs a better fitting to the image dataset due to the reduction in the number of parameters involved and reusability of weights. In other words, the network can be trained to understand the sophistication of the image better.

36 of 52

Why ConvNets over Feed-Forward Neural Nets? (2)

37 of 52

Convolution Layer — The Kernel

38 of 52

Convolution Layer — The Kernel (2)

39 of 52

Pooling Layer

  • Similar to the Convolutional Layer, the Pooling layer is responsible for reducing the spatial size of the Convolved Feature. This is to decrease the computational power required to process the data through dimensionality reduction. Furthermore, it is useful for extracting dominant features which are rotational and positional invariant, thus maintaining the process of effectively training of the model.

40 of 52

Pooling Layer

41 of 52

Classification — Fully Connected Layer (FC Layer)

42 of 52

RNN

43 of 52

Recurrent Neural Networks (RNN)

  • A recurrent neural network is a neural network that is specialized for processing a sequence of data x(t)= x(1), . . . , x(τ) with the time step index t ranging from 1 to τ. For tasks that involve sequential inputs, such as speech and language, it is often better to use RNNs. In a NLP problem, if you want to predict the next word in a sentence it is important to know the words before it. RNNs are called recurrent because they perform the same task for every element of a sequence, with the output being depended on the previous computations. Another way to think about RNNs is that they have a “memory” which captures information about what has been calculated so far.

44 of 52

Recurrent Neural Networks (2)

45 of 52

Recurrent Neural Networks (3)

  • Input: x(t)​ is taken as the input to the network at time step t. For example, x1,could be a one-hot vector corresponding to a word of a sentence.
  • Hidden state: h(t)​ represents a hidden state at time t and acts as “memory” of the network. h(t)​ is calculated based on the current input and the previous time step’s hidden state: h(t)​ = f(U x(t)​ + W h(t−1)​). The function f is taken to be a non-linear transformation such as tanh, ReLU.
  • Weights: The RNN has input to hidden connections parameterized by a weight matrix U, hidden-to-hidden recurrent connections parameterized by a weight matrix W, and hidden-to-output connections parameterized by a weight matrix V and all these weights (U,V,W) are shared across time.
  • Output: o(t)​ illustrates the output of the network. In the figure I just put an arrow after o(t) which is also often subjected to non-linearity, especially when the network contains further layers downstream.

46 of 52

RNN

47 of 52

Transfer Learning

48 of 52

Transfer Learning

49 of 52

Transfer Learning (2)

50 of 52

Federated Learning

51 of 52

Federated Learning

52 of 52

Federated Learning (2)