1 of 45

Recitation 3: Multiclass, MLP, Backpropagation

Aviv Slobodkin

Partially based on prof. Yoav Goldberg’s slides

And on the slides of Yossi Adi and Felix Kreuk, CS Machine Learning course and the resources [1], [2], [3]

2 of 45

3 of 45

Softmax

  • A vector with a score per class
  • The scores sum to 1 (a probability distribution)

Dog Cat Watermelon toy

0.6 0.2 0.05 0.15

4 of 45

Derivative of softmax

Goal:

A derivative of a quotient:

5 of 45

First, assume i = j

6 of 45

Now, assume i j

In summary:

7 of 45

Gradient descent for multiclass logistic regression

8 of 45

Logistic regression: gradient calculation

We have one example (xi, yi) and calculate gradient with respect to row k in W

9 of 45

Logistic regression: gradient calculation

10 of 45

Beyond linear classifiers

11 of 45

12 of 45

13 of 45

14 of 45

15 of 45

16 of 45

17 of 45

18 of 45

19 of 45

20 of 45

21 of 45

22 of 45

23 of 45

24 of 45

25 of 45

26 of 45

Reminder: Chain Rule

27 of 45

Visualization

28 of 45

Backpropagation

We will calculate parameter updates for the following MLP model:

29 of 45

Backpropagation

30 of 45

We start by calculating ∂l/∂W2. This is easy! Just ignore the rest of the network.

31 of 45

We will use the chain rule to calculate the gradient ∂l/∂W1. We will start by calculating ∂l/∂h1

32 of 45

Now that we have ∂l/∂h1, we will continue by calculating ∂h1[p]/∂zi[j]. What is the dimensionality of the gradient ∂h1/∂z?

33 of 45

Now that we have both ∂l/∂h1[p] and ∂h1[p]/∂zi[j] we can calculate ∂l/∂z1[j] with the chain rule:

34 of 45

With ∂l/∂z1[j] we can finally calculate ∂l/∂w1[j] (!!!)

Completely analogous to the formula for ∂l/∂w2[j] !

35 of 45

What if we had another layer before xt?

36 of 45

Conclusions

  • We know how to calculate the error signal for the last layer:
    • δout = Ŷ - Yt (the “base case” of the dynamic programming)
  • For any other layer i, let Wi be the layer’s weight matrix, its activation function gi, its activation vector hi, and let the input to the activation be zi:
    • zi = Wi * hi-1
    • hi = g(zi)

*

37 of 45

Backpropagation Algorithm for multiclass MLP

  1. Perform a forward pass on the network with input xt
    1. Store: Ŷ; zi, hi for each layer i = 1...n
  2. Initialize δn = δout ← Ŷ - Yt
  3. Update Wn according to logistic regression
  4. For each layer i = n-1 … 1 (backwards) do:
    • δi ← (Wi+1δi+1) ∘gi’(zi) > error signal = zi
    • Wi hi-1Tδi > gradient calculation
    • Wi ← Wi -η Wi > update layer’s parameters

38 of 45

In practice: automatic differentiation

39 of 45

40 of 45

The computation graph

41 of 45

42 of 45

In practice: computation graphs

43 of 45

44 of 45

Forward pass

45 of 45

Backward pass