Loss Function
Mean Squared Error (MSE)
Mean Absolute Error (MAE)
Softmax Function
Let’s understand this with an example. Let’s say the models (such as those trained using algorithms such as multi-class LDA, and multinomial logistic regression) output three different values such as 5.0, 2.5, and 0.5 for a particular input. In order to convert these numbers into probabilities, these numbers are fed into the ure.softmax function as shown in fig.
Softmax Layer
Introduction to Neural Networks
Slide credit: Hung-yi Lee – Deep Learning Tutorial
A Layer with Sigmoid Activations
3
-3
1
0.95
0.05
0.73
11
CS 404/504, Fall 2021
Softmax Layer
Introduction to Neural Networks
A Softmax Layer
3
-3
1
2.7
20
0.05
0.88
0.12
≈0
Slide credit: Hung-yi Lee – Deep Learning Tutorial
12
CS 404/504, Fall 2021
Example for Softmax Function
Entrophy
Cross Entrophy
Binary Cross Entrophy
Binary Cross Entropy is the negative average of the log of corrected predicted probabilities.
.
Predicted Probabilities
Here in the table, we have three columns
ID: It represents a unique instance.
Actual: It is the class the object originally belongs to.
Predicted_probabilities.: The is output given by the model that tells, the probability object belongs to class 1.
Corrected Probabilities- It is the probability that a particular observation belongs to its original class.
Predicted Prob=prob object belongs to class 1
Log(Corrected probabilities)
Now we will calculate the log value for each of the corrected probabilities. The reason behind using the log value is, the log value offers less penalty for small differences between predicted probability and corrected probability. when the difference is large the penalty will be higher.
Here we have calculated log values for all the corrected probabilities. Since all the corrected probabilities lie between 0 and 1, all the log values are negative.
In order to compensate for this negative value, we will use a negative average of the values
The value of the negative average of corrected probabilities we calculate comes to be 0.214 which is our Log loss or Binary cross-entropy for this particular example.
Further, instead of calculating corrected probabilities, we can calculate the Log loss using the formula given below.
Here, pi is the probability of class 1, and (1-pi) is the probability of class 0.
When the observation belongs to class 1 the first part of the formula becomes active and the second part vanishes and vice versa in the case observation’s actual class are 0. This is how we calculate the Binary cross-entropy.
Binary Cross Entropy for Multi-Class classification
If you are dealing with a multi-class classification problem you can calculate the Log loss in the same way. Just use the formula given below.