1 of 21

A Variational Bayesian Method for Autoregressive and Recurrent Models

Max Welling

CuspAI, University of Amsterdam

”all I know is that I know nothing” - Socrates

2 of 21

Dario Coscia

International School of Advanced Studies

University of Amsterdam

3 of 21

From Experiment to In Silico Simulation

Sören von Bülow, Mateusz Sikora, Gerhard Hummer. MPI of Biophysics

data

COMPUTATIONAL COMPLEXITY

TIME

Era 3: In-silico design

Era 2: Data-driven modelling

Era 1: Trial-and-error

4 of 21

From Simulation to Emulation

slow

fast

Simulate

Emulate

Simulate Train NN Surrogate Emulate

Amortization

Surrogate

5 of 21

Key Question

  • How can we trust a surrogate when it operates outside of its training regime?

  • Answer: trustworthy and computationally efficient uncertainty prediction!

(ensembles are accurate but slow)

6 of 21

Sequence Models: A Brief Overview

Radford, Alec. "Improving language understanding by generative pre-training."

Van Den Oord, Aaron, et al. "WaveNet: A generative model for raw audio."

Keller, T. Anderson, et al. "Traveling waves encode the recent past and enhance sequence learning."

Sequence models are a class of machine learning models designed to process and generate data that is inherently sequential, e.g. text or audio

7 of 21

Autoregressive and Recurrent Models in AI4Science

Autoregressive models can be used to build Neural Emulators to solve PDEs or perform weather forecast by iteratively applying the emulator on its output given an initial state

Bodnar, Cristian, et al. "Aurora: A foundation model of the atmosphere"

Recurrent models can be used to discover novel Molecules by generating chemical structures as sequence of tokens, Chemical Language Models (CLMs)

Özçelik, Rıza, et al. "Chemical language modeling with structured state space sequence models."

8 of 21

Can we trust Autoregressive and Recurrent Models?

Large Language Models Hallucination

Diverging Solutions between Emulator and PDE Solver

time

diverging

9 of 21

Bayesian Methods in AI

Uncertainty Quantification is fundamental in AI4Science to have reliable emulators

Bayesian methods are key for search, e.g. molecular design/ language models alignment

Korbak, Tomasz, et al. "RL with KL penalties is better viewed as Bayesian inference."

10 of 21

Variational Bayes for Bayesian Deep Learning

Variational Inference optimizes the parameters of some parameterized posterior, e.g. Neural Networks, such that it is a close approximation of the true intractable posterior

11 of 21

Variational Dropout and Local Reparametrization Trick: Scaling to Large Networks

Variational Dropout generalizes the concept of Dropout to learnable dropout rates

The Local Reparameterization trick translates uncertainty about global parameters into local noise that is independent across datapoints in the minibatch

Sampling activation allows scaling

12 of 21

Bayesianize Sequence Models: BARNN

BARNN is a practical, scalable, calibrated and accurate methodology to turn any autoregressive or recurrent model into its Bayesian version

13 of 21

Joint states-weight Sequence Model

Jointly sample in an alternating way weights and states

14 of 21

Variational Bayes Training

Optimize via Variational Inference the Evidence Lower BOund

15 of 21

Temporal Variational Dropout and Temporal VAMP

The posterior is obtained by extending Variational Dropout to variable in time dropout rates

The best prior maximazing the Temporal ELBO is given by the aggregated posterior in time

The optimizable weights are decoupled

=

scalability

(Results in Tractable ELBO)

16 of 21

BARNN in AI4Science: PDEs and Molecules

Choosing a gaussian state update with small variance is equivalent to update with a standard Autoregressive PDE Solvers

Molecules can be represented as text using SMILES and a Language Model can be used to generate new molecules

17 of 21

BARNN solves PDEs with Calibrated Uncertainties

SOTA results on solving PDEs

Having a small root mean-square error (RMSE) between CFD solution and Neural Emulator solution does not mean to have calibrated uncertainties! BARNN achieves both

18 of 21

BARNN dynamically adapts uncertainties over time

MAP provides convergent metrics when uncertainties are not needed

Confidence intervals are tighter in regions of small uncertainties and larger otherwise in the Kuramoto Sivashinsky PDE

19 of 21

BARNN excels in molecule generation and long-range dependency modelling

The molecules generated by BARNN are statistically richer

BARNN generates more complex molecules (many rings) showing great long-dependency modelling abilities

Total generated structures

Total valid generated structures

20 of 21

BARNN excels in molecular properties prediction

The molecules generated by BARNN better match the training dataset properties, showing higher density matching ability

21 of 21

Conclusions

BARNN combines Bayesian and sequence models, enabling UQ and achieving state-of-the-art predictive results while remaining scalable to large nets.

AI4Science highlights the necessity of Uncertainty Quantification, without UQ surrogate models can not be safely used in applications

Extending these approaches to Large Language Models will be impactful for chatbots, allowing them to recognize when they lack knowledge