Bayesian neural networks
Philip Smolenski-Jensen�Working with Marek Cygan, Piotr Tempczyk, Ksawery Smoczyński and Rafał Michaluk
Bayesian neural networks vs standard neural networks
Components of BNN
Math behind BNN’s
Predictions in Bayesian Neural Nets
Advantages of BNN’s
Finding posterior distribution
Variational Inference: general idea
VI intro: KL divergence
Variational inference: ELBO
We have:
after substitution we get:
Variational Inference: ELBO
Stochastic Variational Inference
Elbo has the form of:
We can compute the gradient of E[f(z)] in general case:
The estimator of ELBO gradient tends to have high variance.
Alternative: Reparameterization trick
Suppose:
Then we can write:
Where
This allows us to produce more stable ELBO gradient estimator. Such a function g exists for multiple distributions. This is called the reparameterization trick.
What about non-Bayesian parameters?
Variational Inverence: Full algorithm
What families of variational distributions can one use?
Motivation behind the project
Log Likelihood distribution for ReLU vs Leaky ReLU
Expected Calibration Error
Experiments
Sample LeakyReLU results: MLEClassify, FashionMNIST
Sample LeakyReLU results: ConvClassify, MNIST
Some observations:
Sample Results: DeepMLEClassify, MNIST
Sample Results: ConvClassify, FashionMNIST
Full Results
Decalibration with ReLU
No decalibration with LeakyReLU
Conclusions
Next steps
Thank you for the attention!