1 of 26

2 of 26

Loss Function

  • A loss function is a function that compares the target and predicted output values; measures how well the neural network models the training data. When training, we aim to minimize this loss between the predicted and target outputs.

​

  • The hyperparameters are adjusted to minimize the average loss — we find the weights, wT, and biases, b, that minimize the value of J (average loss).

​

3 of 26

Mean Squared Error (MSE)

  • One of the most popular loss functions, MSE finds the average of the squared differences between the target and the predicted outputs.

4 of 26

  • This function has numerous properties that make it especially suited for calculating loss.

​

  • The difference is squared, which means it does not matter whether the predicted value is above or below the target value; however, values with a large error are penalized.

5 of 26

Mean Absolute Error (MAE)

  • MAE finds the average of the absolute differences between the target and the predicted outputs.

6 of 26

  • This loss function is used as an alternative to MSE in some cases. As mentioned previously, MSE is highly sensitive to outliers, which can dramatically affect the loss because the distance is squared.

​

  • MAE is used in cases when the training data has a large number of outliers to mitigate this.

​

  • It also has some disadvantages; as the average distance approaches 0, gradient descent optimization will not work, as the function's derivative at 0 is undefined (which will result in an error, as it is impossible to divide by 0).

7 of 26

Softmax Function

  • The softmax function is a function that turns a vector of K real values into a vector of K real values that sum to 1. The input values can be positive, negative, zero, or greater than one, but the softmax transforms them into values between 0 and 1, so that they can be interpreted as probabilities.

​

  • If one of the inputs is small or negative, the softmax turns it into a small probability, and if an input is large, then it turns it into a large probability, but it will always remain between 0 and 1.

​

  • It is most commonly used as an activation function for the last layer of the neural network in the case of multi-class classification. 

​

​

8 of 26

  • the Softmax function is described as a combination of multiple sigmoids. �
  • It calculates the relative probabilities. Similar to the sigmoid/logistic activation function, the SoftMax function returns the probability of each class. 

​

9 of 26

Let’s understand this with an example. Let’s say the models (such as those trained using algorithms such as multi-class LDA, and multinomial logistic regression) output three different values such as 5.0, 2.5, and 0.5 for a particular input. In order to convert these numbers into probabilities, these numbers are fed into the ure.softmax function as shown in fig.

10 of 26

  • Notice that the softmax outputs are less than 1. And, the outputs of the softmax function sum up to 1. Owing to this property, the Softmax function is considered an activation function in neural networks and algorithms such as multinomial logistic regression. Note that for binary logistic regression, the activation function used is the sigmoid function.

​

  • Based on the above, it could be understood that the output of the softmax function maps to a [0, 1] range. And, it maps outputs in a way that the total sum of all the output values is 1. Thus, it could be said that the output of the softmax function is a probability distribution.

​

11 of 26

Softmax Layer

  • In multi-class classification tasks, the output layer is typically a softmax layer
    • I.e., it employs a softmax activation function
    • If a layer with a sigmoid activation function is used as the output layer instead, the predictions by the NN may not be easy to interpret
      • Note that an output layer with sigmoid activations can still be used for binary classification

​

Introduction to Neural Networks

​

Slide credit: Hung-yi Lee – Deep Learning Tutorial

A Layer with Sigmoid Activations

3

-3

1

0.95

0.05

0.73

11

CS 404/504, Fall 2021

12 of 26

Softmax Layer

  • The softmax layer applies softmax activations to output a probability value in the range [0, 1]
    • The values z inputted to the softmax layer are referred to as logits

Introduction to Neural Networks

​

A Softmax Layer

3

-3

1

2.7

20

0.05

0.88

0.12

≈0

 

Slide credit: Hung-yi Lee – Deep Learning Tutorial

12

CS 404/504, Fall 2021

13 of 26

Example for Softmax Function

14 of 26

Entrophy

  • Entropy of a random variable X is the level of uncertainty inherent in the variables possible outcome.
  • For p(x) — probability distribution and a random variable X, entropy is defined as follows

​

15 of 26

  • Reason for negative sign: log(p(x))<0 for all p(x) in (0,1) . p(x) is a probability distribution and therefore the values must range between 0 and 1.

16 of 26

17 of 26

18 of 26

  • As expected the entropy for the first and third container is smaller than the second one. This is because probability of picking a given shape is more certain in container 1 and 3 than in 2.

19 of 26

Cross Entrophy

  • Cross-entropy builds upon the idea of information theory entropy and measures the difference between two probability distributions for a given random variable/set of events.

​

  • Cross-entropy loss refers to the contrast between two random variables; it measures them in order to extract the difference in the information they contain, showcasing the results.

​

20 of 26

  • We use this type of loss function to calculate how accurate our machine learning or deep learning model is by defining the difference between the estimated probability with our desired outcome.

​

  • Cross entropy can be applied in both binary and multi-class classification problems. 

21 of 26

Binary Cross Entrophy

  • Let’s consider the earlier example, where we answer whether a student will pass the SAT exams. In this case, we work with four students. 

​

  • We have two models, A and B, that predict the likelihood of the four students passing the exam, as shown in the figure in the next slide.

22 of 26

  • What is Binary Cross Entropy Or Logs Loss?
  • Binary cross entropy compares each of the predicted probabilities to actual class output which can be either 0 or 1. It then calculates the score that penalizes the probabilities based on the distance from the expected value. That means how close or far from the actual value.

Binary Cross Entropy is the negative average of the log of corrected predicted probabilities.

.

Predicted Probabilities

       

Here in the table, we have three columns

ID: It represents a unique instance.

Actual: It is the class the object originally belongs to.

Predicted_probabilities.: The is output given by the model that tells, the probability object belongs to class 1.

Corrected Probabilities- It is the probability that a particular observation belongs to its original class. 

      

Predicted Prob=prob object belongs to class 1

23 of 26

Log(Corrected probabilities)

Now we will calculate the log value for each of the corrected probabilities. The reason behind using the log value is, the log value offers less penalty for small differences between predicted probability and corrected probability. when the difference is large the penalty will be higher.

        

Here we have calculated log values for all the corrected probabilities. Since all the corrected probabilities lie between 0 and 1, all the log values are negative.

24 of 26

In order to compensate for this negative value, we will use a negative average of the values

              

The value of the negative average of corrected probabilities we calculate comes to be 0.214 which is our Log loss or Binary cross-entropy for this particular example.

25 of 26

Further, instead of calculating corrected probabilities, we can calculate the Log loss using the formula given below.

                                     

Here, pi is the probability of class 1, and (1-pi) is the probability of class 0.

When the observation belongs to class 1 the first part of the formula becomes active and the second part vanishes and vice versa in the case observation’s actual class are 0. This is how we calculate the Binary cross-entropy.

26 of 26

Binary Cross Entropy for Multi-Class classification

If you are dealing with a multi-class classification problem you can calculate the Log loss in the same way. Just use the formula given below.