1 of 50

Deep Learning (DEEP-0001)�

6 – Fitting Models

2 of 50

Loss function

  • Training dataset of I pairs of input/output examples:

  • Loss function or cost function measures how bad model is:

or for short:

Returns a scalar that is smaller when model maps inputs to outputs better

3 of 50

Training

  • Loss function:

  • Find the parameters that minimize the loss:

Returns a scalar that is smaller when model maps inputs to outputs better

4 of 50

Example: 1D Linear regression loss function

Loss function:

“Least squares loss function”

5 of 50

Example: 1D Linear regression training

6 of 50

Example: 1D Linear regression training

7 of 50

Example: 1D Linear regression training

8 of 50

Example: 1D Linear regression training

9 of 50

Example: 1D Linear regression training

This technique is known as gradient descent

10 of 50

Fitting models

  • Gradient descent algorithm
  • Linear regression example
  • Gabor model example
  • Stochastic gradient descent
  • Momentum
  • Adam

11 of 50

Gradient descent algorithm

12 of 50

Fitting models

  • Gradient descent algorithm
  • Linear regression example
  • Gabor model example
  • Stochastic gradient descent
  • Momentum
  • Adam

13 of 50

Gradient descent

Step 1: Compute derivatives (slopes of function) with

Respect to the parameters

14 of 50

Gradient descent

Step 1: Compute derivatives (slopes of function) with

Respect to the parameters

15 of 50

Gradient descent

Step 1: Compute derivatives (slopes of function) with

Respect to the parameters

16 of 50

Gradient descent

Step 1: Compute derivatives (slopes of function) with

Respect to the parameters

17 of 50

Gradient descent

Step 1: Compute derivatives (slopes of function) with

Respect to the parameters

Step 2: Update parameters according to rule

 

18 of 50

Gradient descent

Step 1: Compute derivatives (slopes of function) with

Respect to the parameters

Step 2: Update parameters according to rule

 

19 of 50

Gradient descent

20 of 50

Gradient descent

Step 1: Compute derivatives (slopes of function) with

Respect to the parameters

Step 2: Update parameters according to rule

 

21 of 50

Gradient descent

Step 1: Compute derivatives (slopes of function) with

Respect to the parameters

Step 2: Update parameters according to rule

 

22 of 50

Gradient descent

23 of 50

Gradient descent

24 of 50

Gradient descent

25 of 50

Convex problems

Non convex

Non-Convex

Convex

26 of 50

Convex problems

Non convex

Non-Convex

Convex

Test for convexity is that 2nd derivative is positive everywhere

27 of 50

Convexity in higher dimensions

Test for convexity is that determinant of Hessian (2nd derivative matrix) is positive everywhere.

28 of 50

Fitting models

  • Gradient descent algorithm
  • Linear regression example
  • Gabor model example
  • Stochastic gradient descent
  • Momentum
  • Adam

29 of 50

Gabor model

30 of 50

Gabor model

31 of 50

32 of 50

  • Gradient descent gets to the global minimum if we start in the right “valley”
  • Otherwise, descent to a local minimum
  • Or get stuck near a saddle point

33 of 50

Fitting models

  • Gradient descent algorithm
  • Linear regression example
  • Gabor model example
  • Stochastic gradient descent
  • Momentum
  • Adam

34 of 50

IDEA: add noise

  • Stochastic gradient descent
  • Compute gradient based on only a subset of points – a mini-batch
  • Work through dataset sampling without replacement
  • One pass though the data is called an epoch

35 of 50

  •  

36 of 50

37 of 50

38 of 50

Properties of SGD

  • Can escape from local minima
  • Adds noise, but still sensible updates as based on part of data
  • Uses all data equally
  • Less computationally expensive
  • Seems to find better solutions

  • In practice, SGD is used with:
    • Learning rate schedule – decrease learning rate over time

39 of 50

Fitting models

  • Gradient descent algorithm
  • Linear regression example
  • Gabor model example
  • Stochastic gradient descent
  • Momentum
  • Adam

40 of 50

Momentum

Exponential Moving Average:

41 of 50

Momentum

  • Weighted sum of this gradient and previous gradient

42 of 50

43 of 50

Fitting models

  • Gradient descent algorithm
  • Linear regression example
  • Gabor model example
  • Stochastic gradient descent
  • Momentum
  • Adam

44 of 50

Adaptive moment estimation (Adam)

45 of 50

Normalized gradients

  • Measure mean and pointwise squared gradient

  • Normalize:

46 of 50

Normalized gradients

  • Measure mean and pointwise squared gradient

  • Normalize:

47 of 50

Normalized gradients

  • Measure mean and pointwise squared gradient

  • Normalize:

48 of 50

Adaptive moment estimation (Adam)

  • Compute mean and pointwise squared gradients with momentum

  • Moderate near start of the sequence

  • Update the parameters

49 of 50

Adaptive moment estimation (Adam)

50 of 50

Hyperparameters

  • Choice of learning algorithm
  • Learning rate
  • Momentum