A Variational Bayesian Method for Autoregressive and Recurrent Models
Max Welling
CuspAI, University of Amsterdam
”all I know is that I know nothing” - Socrates
Dario Coscia
International School of Advanced Studies
University of Amsterdam
From Experiment to In Silico Simulation
Sören von Bülow, Mateusz Sikora, Gerhard Hummer. MPI of Biophysics
data
COMPUTATIONAL COMPLEXITY
TIME
Era 3: In-silico design
Era 2: Data-driven modelling
Era 1: Trial-and-error
From Simulation to Emulation
slow
fast
Simulate
Emulate
Simulate Train NN Surrogate Emulate
Amortization
Surrogate
Key Question
(ensembles are accurate but slow)
Sequence Models: A Brief Overview
Radford, Alec. "Improving language understanding by generative pre-training."
Van Den Oord, Aaron, et al. "WaveNet: A generative model for raw audio."
Keller, T. Anderson, et al. "Traveling waves encode the recent past and enhance sequence learning."
Sequence models are a class of machine learning models designed to process and generate data that is inherently sequential, e.g. text or audio
Autoregressive and Recurrent Models in AI4Science
Autoregressive models can be used to build Neural Emulators to solve PDEs or perform weather forecast by iteratively applying the emulator on its output given an initial state
Bodnar, Cristian, et al. "Aurora: A foundation model of the atmosphere"
Recurrent models can be used to discover novel Molecules by generating chemical structures as sequence of tokens, Chemical Language Models (CLMs)
Özçelik, Rıza, et al. "Chemical language modeling with structured state space sequence models."
Can we trust Autoregressive and Recurrent Models?
Large Language Models Hallucination
Diverging Solutions between Emulator and PDE Solver
time
diverging
Bayesian Methods in AI
Uncertainty Quantification is fundamental in AI4Science to have reliable emulators
Bayesian methods are key for search, e.g. molecular design/ language models alignment
Korbak, Tomasz, et al. "RL with KL penalties is better viewed as Bayesian inference."
Variational Bayes for Bayesian Deep Learning
Variational Inference optimizes the parameters of some parameterized posterior, e.g. Neural Networks, such that it is a close approximation of the true intractable posterior
Variational Dropout and Local Reparametrization Trick: Scaling to Large Networks
Variational Dropout generalizes the concept of Dropout to learnable dropout rates
The Local Reparameterization trick translates uncertainty about global parameters into local noise that is independent across datapoints in the minibatch
Sampling activation allows scaling
Bayesianize Sequence Models: BARNN
BARNN is a practical, scalable, calibrated and accurate methodology to turn any autoregressive or recurrent model into its Bayesian version
Joint states-weight Sequence Model
Jointly sample in an alternating way weights and states
Variational Bayes Training
Optimize via Variational Inference the Evidence Lower BOund
Temporal Variational Dropout and Temporal VAMP
The posterior is obtained by extending Variational Dropout to variable in time dropout rates
The best prior maximazing the Temporal ELBO is given by the aggregated posterior in time
The optimizable weights are decoupled
=
scalability
(Results in Tractable ELBO)
BARNN in AI4Science: PDEs and Molecules
Choosing a gaussian state update with small variance is equivalent to update with a standard Autoregressive PDE Solvers
Molecules can be represented as text using SMILES and a Language Model can be used to generate new molecules
BARNN solves PDEs with Calibrated Uncertainties
SOTA results on solving PDEs
Having a small root mean-square error (RMSE) between CFD solution and Neural Emulator solution does not mean to have calibrated uncertainties! BARNN achieves both
BARNN dynamically adapts uncertainties over time
MAP provides convergent metrics when uncertainties are not needed
Confidence intervals are tighter in regions of small uncertainties and larger otherwise in the Kuramoto Sivashinsky PDE
BARNN excels in molecule generation and long-range dependency modelling
The molecules generated by BARNN are statistically richer
BARNN generates more complex molecules (many rings) showing great long-dependency modelling abilities
Total generated structures
Total valid generated structures
BARNN excels in molecular properties prediction
The molecules generated by BARNN better match the training dataset properties, showing higher density matching ability
Conclusions
BARNN combines Bayesian and sequence models, enabling UQ and achieving state-of-the-art predictive results while remaining scalable to large nets.
AI4Science highlights the necessity of Uncertainty Quantification, without UQ surrogate models can not be safely used in applications
Extending these approaches to Large Language Models will be impactful for chatbots, allowing them to recognize when they lack knowledge