1 of 29

Introduction to Neural Networks

PHYS591000 2023.03.23

​

2 of 29

Outline

  • Here we come: The most famous AI algorithms�
  • How neurons work: Activation function�
  • The goal of the neural network: Minimize Loss function�
  • Make the machine learn faster: Optimizers�
  • A word on Regularization�

2

3 of 29

Warming up

  • As usual, take 3 mins to introduce yourself to your teammates for this week!�– “What do you know about neural networks? Have you played with them before?”

​

3

4 of 29

Neural Network

  • Artificial Neural Network (ANN), or just neural network (NN), is a model which tries to simulate the functions of biological neural networks.�
  • Each circle on the plot is a ‘neuron’. Can have multiple neurons for a layer. Can have many hidden layers.

4

Source: Wikipedia

5 of 29

How does a neuron work?

  • A neuron works like a ‘switch’: its output is determined by summing over inputs with different weights (importance of each input), plus a bias (a ‘threshold’ value for this neuron)

5

X1 * W1

X2 * W2

​

+ b

W1, W2: weights

b: bias

6 of 29

How does a neuron work?

  • When training the model we want to change the output ‘little by little’, so a simple 0/1 binary output is not ideal. ��– We’d like a smooth output which acts like a switch (0/1)

6

7 of 29

Activation Function

  • Idea: Introduce an activation function that smooths the output

7

‘Off’

‘On’

Courtesy of Prof. Kai-Feng Chen (NTU)

8 of 29

Activation Function

  • Two common activation functions are the rectified linear unit (ReLU) function and sigmoid function

8

Sigmoid

9 of 29

Network Architecture

9

Courtesy of Prof. Kai-Feng Chen (NTU)

10 of 29

Network Architecture

  • If all neurons are connected to every neuron in the next layer, it is called a fully-connected network.�
  • If the information always flows from left to right (input → hidden layer 1 → … → output layer), this is called a feedforward network.

10

11 of 29

Review: Goal of training a model

  • Recall in Regression the way to find the line (parameters) is to minimize the function Σi(h(xi)-yi)2 → method of least squares

​

​

11

y (Energy)

x (Number of shower particles)

y = h(x) = θ0 + θ1 x

θ0 , θ1 : parameters

Intercept θ0

Slope θ1

12 of 29

Review: Goal of training a model

  • Σi(h(xi)-yi)2 is loss function. The goal of training is to find the optimal parameters which minimize the loss function. �

​

12

y (Energy)

x (Number of shower particles)

y = h(x) = θ0 + θ1 x

θ0 , θ1 : parameters

Intercept θ0

Slope θ1

13 of 29

How to Train Neural Network

  • Same for NN: The goal of training is to find optimal weights (wi) and biases (bj) which minimize the loss function. A typical choice of loss function is the mean square error (MSE):

13

Number of data points

Current output given input x

True value y for input x

14 of 29

How to Train Neural Network

  • If we need to optimize a lot of wi and bj we need a smart algorithm to do so, e.g.,�the method of gradient descent

14

15 of 29

How to minimize the loss function

  • Method of gradient descent: The next step is proportional to the negative of the local gradient

15

Gradient of Loss (average over the input data)

a tunable hyperparameter

Courtesy of Prof. Kai-Feng Chen (NTU)

16 of 29

Optimizer/Solver for minimizing loss function

  • When the input data size is large it will take a lot of time calculating the gradients, and make the NN too slow. There are several optimizer/solver to speed up the learning process.�
  • A common solution is to use a random subset of input data instead:�

The method is called stochastic gradient descent (SGD).

16

Gradient of Loss approximated by average over a subset of input data

17 of 29

Optimizer choice

  • SGD is an easy and useful algorithm for minimizing the loss function, but there can be problems like slow training.�
  • Many other choices of optimizers. E.g.
    • Adam: Increase the ‘step size’ (learning rate) when gradient becomes smaller
    • Adagrad: Regularized learning rate – Larger/smaller gradient would give smaller/larger learning rate.
    • Adadelta : extended Adagrad with reduced dependence to the global learning rate.

17

18 of 29

Training Process hyperparameters

  • So we use a subset of data to calculate the gradient of loss.�
  • This subset is called mini-batch. We can specify how many samples to process at one time (‘batch size’).�
  • Epoch: One cycle through the full training dataset (all subsets in batches). Training a NN usually requires O(10-1000) epochs.

18

19 of 29

Neural Network Parameters

  • When you train a NN you have to decide �
    • Number of hidden layers and �number of neurons for each hidden layer
    • Which activation function to use
    • Which loss function to use
    • How to minimize the loss function (optimizer)�...

19

20 of 29

Choice of Loss Function

  • Training with MSE can be slow since the gradient depends on the slope of the activation function σ(z):

20

Courtesy of Prof. Kai-Feng Chen (NTU)

21 of 29

Choice of Loss Function

  • Another choice of loss functions is the cross entropy function:�The form of binary (two-class) cross entropy looks like

21

Courtesy of Prof. Kai-Feng Chen (NTU)

22 of 29

Choice of Loss Function and Output Layer

  • For multi-classification we can use categorical cross entropy:���
  • Which often combines with an output layer using Softmax:�

22

Courtesy of Prof. Kai-Feng Chen (NTU)

23 of 29

Choice of Loss Function and Output Layer

  • Categorical cross entropy loss + softmax output again makes the gradient independent of the first derivative of the activation function, and speeds up the learning.�
  • Furthermore, the softmax output is not just 0 or 1 (yes or no for one class) – it’s continuous between 0 to 1, and thus can be interpreted as the probability of being in this class.

23

24 of 29

Regularization

  • As usual we need to be careful about overtraining (overfitting)�
  • L1/L2 regularization: Add extra term λΣi|Wi| (L1) or λΣiWi2 (L2) to the loss function�

​

24

Number of epochs

25 of 29

In-class exercise

  • We’ll use Keras with Tensorflow backend on the MNIST data!��– Tensorflow: an open source software library for high performance numerical computation originally developed by Google.��– Keras: a kind of “wrapper” or package which can help to build most of the conventional NN models, and the real calculation can be carried out by TensorFlow, as one of the supported backends.

25

26 of 29

In-class exercise

  • We’ll build a 784-30-10 NN:

26

27 of 29

Lab for this week

  • For the Lab this week, we’ll apply what we’ve learned about NN to a binary classification task: classify jets from �1) W/Z (a kind of heavy gauge boson) bosons �2) light quark or gluons (‘QCD’) �
  • You’ll compare results with different hyperparameters and see what you can learn about hyperparameter tuning.

27

28 of 29

Backup

28

29 of 29

Neural Network Information Flow

  • If the NN processes the information only in one way (input → output without any loops) it is called feedforward.�
  • Focus on feedforward NN now. Later we’ll talk about RNN (recurrent NN) which is non-feedforward.

29

Source: Wikipedia

Flow of information