1 of 144

AI Experience Lab

[Loss functions]

Ue-Hwan, Kim

2 of 144

Where are we now?

  • Types of Learning
  • Overall Workflow
  • Theoretical Perspective
    • 1. Data Preparation
    • 2. Building Models: Neural Networks
    • 3. Training Models
      • Loss functions
      • Backpropagation
      • Optimization
    • 4. Evaluating Models
    • 5. Improving Performance
  • Linear Algebra and Numpy
  • Summary

2

3 of 144

Contents

  • Example problem
  • Linear classifier
  • Softmax classifier
  • More on loss functions
  • Regularization
  • Summary

3

4 of 144

Example problem

4

5 of 144

5

credit: Stanford CS231n

6 of 144

6

credit: Stanford CS231n

7 of 144

7

credit: Stanford CS231n

8 of 144

8

credit: Stanford CS231n

9 of 144

9

credit: Stanford CS231n

10 of 144

10

credit: Stanford CS231n

11 of 144

11

credit: Stanford CS231n

12 of 144

12

credit: Stanford CS231n

13 of 144

13

credit: Stanford CS231n

14 of 144

14

credit: Stanford CS231n

15 of 144

Let’s implement

15

credit: Stanford CS231n

16 of 144

16

credit: Stanford CS231n

17 of 144

17

credit: Stanford CS231n

18 of 144

18

credit: Stanford CS231n

19 of 144

19

credit: Stanford CS231n

20 of 144

Poll

Which of the following is true for image classification

  • It is a task of extracting semantics from an array of numbers
  • Hard-coded algorithms show satisfactory performance

20

21 of 144

Poll

Which of the following is true for image classification

  • It is a task of extracting semantics from an array of numbers
  • Hard-coded algorithms show satisfactory performance

21

22 of 144

Linear classifiers

22

23 of 144

23

credit: Stanford CS231n

24 of 144

24

credit: Stanford CS231n

25 of 144

25

credit: Stanford CS231n

26 of 144

26

credit: Stanford CS231n

27 of 144

27

credit: Stanford CS231n

28 of 144

28

credit: Stanford CS231n

29 of 144

29

credit: Stanford CS231n

30 of 144

30

credit: Stanford CS231n

31 of 144

31

credit: Stanford CS231n

32 of 144

32

credit: Stanford CS231n

33 of 144

33

credit: Stanford CS231n

34 of 144

34

credit: Stanford CS231n

Today!

Next class!

35 of 144

35

credit: Stanford CS231n

36 of 144

36

credit: Stanford CS231n

37 of 144

37

credit: Stanford CS231n

38 of 144

38

credit: Stanford CS231n

39 of 144

39

credit: Prof. Dr. Andreas Geiger

40 of 144

40

credit: Prof. Dr. Andreas Geiger

41 of 144

41

credit: Stanford CS231n

42 of 144

42

credit: Stanford CS231n

43 of 144

43

credit: Stanford CS231n

44 of 144

44

credit: Stanford CS231n

45 of 144

45

credit: Stanford CS231n

Calculate!

46 of 144

46

credit: Stanford CS231n

47 of 144

47

credit: Stanford CS231n

48 of 144

48

credit: Stanford CS231n

49 of 144

49

credit: Stanford CS231n

50 of 144

50

credit: Stanford CS231n

51 of 144

51

credit: Stanford CS231n

No change

0 / inf

(C - 1)

52 of 144

52

credit: Stanford CS231n

53 of 144

53

credit: Stanford CS231n

Just adding 1

54 of 144

54

credit: Stanford CS231n

55 of 144

55

credit: Stanford CS231n

Just rescaling

56 of 144

56

credit: Stanford CS231n

57 of 144

57

credit: Stanford CS231n

Squared hinge loss

58 of 144

58

credit: Stanford CS231n

Squared hinge loss

59 of 144

59

credit: Stanford CS231n

60 of 144

Let’s implement!

60

61 of 144

Softmax classifier

61

62 of 144

62

credit: Stanford CS231n

63 of 144

63

credit: Stanford CS231n

64 of 144

64

credit: Stanford CS231n

65 of 144

65

credit: Stanford CS231n

66 of 144

66

credit: Stanford CS231n

67 of 144

67

credit: Stanford CS231n

68 of 144

68

credit: Stanford CS231n

69 of 144

69

credit: Stanford CS231n

70 of 144

70

credit: Stanford CS231n

71 of 144

71

credit: Stanford CS231n

72 of 144

72

credit: Stanford CS231n

73 of 144

73

credit: Stanford CS231n

0 / inf

74 of 144

74

credit: Stanford CS231n

75 of 144

75

credit: Stanford CS231n

76 of 144

76

credit: Stanford CS231n

77 of 144

77

credit: Stanford CS231n

78 of 144

78

credit: Stanford CS231n

79 of 144

79

credit: Stanford CS231n

SVM: nothing much

Softmax: sensitive

80 of 144

Let’s implement!

80

81 of 144

More on loss functions

81

82 of 144

82

credit: Prof. Dr. Andreas Geiger

83 of 144

83

credit: Prof. Dr. Andreas Geiger

84 of 144

84

credit: Prof. Dr. Andreas Geiger

85 of 144

85

credit: Prof. Dr. Andreas Geiger

86 of 144

More on loss functions

  • regression -

86

87 of 144

87

credit: Prof. Dr. Andreas Geiger

88 of 144

88

credit: Prof. Dr. Andreas Geiger

89 of 144

89

credit: Prof. Dr. Andreas Geiger

90 of 144

90

credit: Prof. Dr. Andreas Geiger

91 of 144

91

credit: Prof. Dr. Andreas Geiger

92 of 144

92

credit: Prof. Dr. Andreas Geiger

93 of 144

Poll

손실함수(Loss functions)에 대하여 옳은 설명은?

  • 손실 함수는 어떤 함수라도 될 수 있습니다.
  • 손실 함수를 계산하여 도출하는 대신 우도(likelihood)를 최대화합니다.
  • 제곱 손실(L2)은 절대 손실(L1)보다 특이값에 더 강건(robust)합니다.
  • 훈련 샘플이 i.i.d일 필요는 없다고 가정합니다. (독립적이고 동일하게 분산됨)

93

94 of 144

Poll

손실함수(Loss functions)에 대하여 옳은 설명은?

  • 손실 함수는 어떤 함수라도 될 수 있습니다.
  • 손실 함수를 계산하여 도출하는 대신 우도(likelihood)를 최대화합니다.
  • 제곱 손실(L2)은 절대 손실(L1)보다 특이값에 더 강건(robust)합니다.
  • 훈련 샘플이 i.i.d일 필요는 없다고 가정합니다. (독립적이고 동일하게 분산됨)

94

95 of 144

Poll

Which is correct about loss functions?

  • A loss function can be any functions
  • Rather than manually derive loss functions, we maximize the likelihood
  • The squared loss (L2) is more robust to outliers than the absolute loss (L1) is
  • We assume training samples do not have to be i.i.d. (independent and identically distributed)

95

96 of 144

Poll

Which is correct about loss functions?

  • A loss function can be any functions
  • Rather than manually derive loss functions, we maximize the likelihood
  • The squared loss (L2) is more robust to outliers than the absolute loss (L1) is
  • We assume training samples do not have to be i.i.d. (independent and identically distributed)

96

97 of 144

More on loss functions

  • classification -

97

98 of 144

98

credit: Prof. Dr. Andreas Geiger

99 of 144

99

credit: Prof. Dr. Andreas Geiger

100 of 144

100

credit: Prof. Dr. Andreas Geiger

101 of 144

101

credit: Prof. Dr. Andreas Geiger

102 of 144

What about the logistic regression?

102

103 of 144

What about the logistic regression?

103

104 of 144

What about the logistic regression?

104

105 of 144

What about the logistic regression?

105

106 of 144

106

107 of 144

107

108 of 144

108

109 of 144

109

credit: Prof. Dr. Andreas Geiger

110 of 144

110

credit: Prof. Dr. Andreas Geiger

111 of 144

111

112 of 144

112

credit: Prof. Dr. Andreas Geiger

113 of 144

113

114 of 144

114

credit: Prof. Dr. Andreas Geiger

115 of 144

115

credit: Prof. Dr. Andreas Geiger

116 of 144

116

117 of 144

117

118 of 144

118

119 of 144

119

120 of 144

120

credit: Prof. Dr. Andreas Geiger

121 of 144

121

credit: Prof. Dr. Andreas Geiger

122 of 144

122

credit: Prof. Dr. Andreas Geiger

123 of 144

Let’s implement

123

124 of 144

Regularization

124

125 of 144

125

credit: Stanford CS231n

126 of 144

126

credit: Stanford CS231n

127 of 144

127

credit: Stanford CS231n

128 of 144

128

credit: Stanford CS231n

129 of 144

129

credit: Stanford CS231n

130 of 144

130

credit: Stanford CS231n

131 of 144

131

credit: Stanford CS231n

132 of 144

132

credit: Stanford CS231n

133 of 144

133

credit: Stanford CS231n

134 of 144

134

credit: Stanford CS231n

135 of 144

135

credit: Stanford CS231n

136 of 144

136

credit: Stanford CS231n

137 of 144

137

credit: Stanford CS231n

138 of 144

138

credit: Stanford CS231n

139 of 144

139

credit: Stanford CS231n

140 of 144

140

credit: Stanford CS231n

141 of 144

141

credit: Stanford CS231n

142 of 144

142

credit: Stanford CS231n

143 of 144

143

credit: Stanford CS231n

Next class!

144 of 144

Summary

Today we

  • talked about the image classification task
  • discussed linear classifiers
  • had a look at the SVM and softmax losses
  • introduced a probabilistic view on loss functions
  • considered regularization to weigh between weights

144