1 of 66

CS294-158 Deep Unsupervised Learning

Lecture 3 Likelihood Models: Flow Models

Pieter Abbeel, Wilson Yan, Kevin Frans, Philipp Wu

2 of 66

Overview

  • What do we want from a generative model?
    • Good fit to the training data (really, the underlying distribution!)
    • For new x, ability to evaluate
    • Ability to sample from
    • A latent representation / embedding space that’s meaningful

  • Recall L2 Autoregressive Models?
    • Check many boxes except
      • Sampling is serial (hence slow)
      • Lacks latent representation / embedding space
      • Limited to discrete data (at least in terms of where experimental success has been)

  • Flow Models will check all boxes (but performance not as good as other models…)

2

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

3 of 66

Flow Models

  • Goal: Fit a density model with continuous

  • What do we want from this model?
    • Good fit to the training data (really, the underlying distribution!)
    • For new x, ability to evaluate
    • Ability to sample from
    • A latent representation / embedding space that’s meaningful

3

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

4 of 66

Outline

  • Foundations of Flows (1-D)
  • 2-D Flows
  • N-D Flows
  • Dequantization

4

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

5 of 66

Quick Refresher: Probability Density Models

5

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

6 of 66

How to fit a density model?

Continuous data

0.22159854, 0.84525919, 0.09121633, 0.364252 , 0.30738086,

0.32240615, 0.24371194, 0.22400792, 0.39181847, 0.16407012,

0.84685229, 0.15944969, 0.79142357, 0.6505366 , 0.33123603,

0.81409325, 0.74042126, 0.67950372, 0.74073271, 0.37091554,

0.83476616, 0.38346571, 0.33561352, 0.74100048, 0.32061713,

0.09172335, 0.39037131, 0.80496586, 0.80301971, 0.32048452,

0.79428266, 0.6961708 , 0.20183965, 0.82621227, 0.367292 ,

0.76095756, 0.10125199, 0.41495427, 0.85999877, 0.23004346,

0.28881973, 0.41211802, 0.24764836, 0.72743029, 0.20749136,

0.29877091, 0.75781455, 0.29219608, 0.79681589, 0.86823823,

0.29936483, 0.02948181, 0.78528968, 0.84015573, 0.40391632,

0.77816356, 0.75039186, 0.84709016, 0.76950307, 0.29772759,

0.41163966, 0.24862007, 0.34249207, 0.74363912, 0.38303383, …

6

Maximum Likelihood:

Equivalently:

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

7 of 66

Example Density Model: Mixtures of Gaussians

Parameters: means and variances of components, mixture weights

7

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

8 of 66

Aside on Mixtures of Gaussians

Do mixtures of Gaussians work for high-dimensional data?

Not really. The sampling process is:

  1. Pick a cluster center
  2. Add Gaussian noise

Imagine this for modeling natural images! The only way a realistic image can be generated is if it is a cluster center, i.e. if it is already stored directly in the parameters.

8

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

9 of 66

Flow Models: Invertible transform from data x to embedding z

9

...

Original data space, complex distribution to model

x

Embedding space, we enforce a simpler distribution

Note: VAE has very similar diagram, but learns both directions rather than imposing invertibility (see L4)

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

10 of 66

Change of Variables Formula

10

Note: for this simple formula to work, �it requires invertible & differentiable

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

11 of 66

Flow Models: Training

11

Assuming we have an expression for ,

this can be optimized with Stochastic Gradient Descent

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

12 of 66

Flow Models: Examples

12

Two choices to make:

Invertible function class

Embedding space density

I.e., in 1-D this is any monotonically increasing (or decreasing) function

E.g.:

“normalizing flow”

mixture of Gaussians

-> make sure maps to [0,1]

E.g.:

- a x + b (for positive a)�- polynomials with positive coefficients and only odd powers�- exp (theta x)�- sigmoid (a x + b)�- cumulative density functions, e.g. CDF of mixture of Gaussians or weighted sum of logistics�- composition of flows = flow

Ideally an easy distribution to sample from

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

13 of 66

Example: Flow to Gaussian z

13

Before training

After training

True distribution of x

Flow x → z

Empirical distribution of z

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

14 of 66

Example: Flow to Uniform z

14

Before training

After training

True distribution of x

Flow x → z

Empirical distribution of z

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

15 of 66

Example: Flow to Beta(5,5) z

15

Before training

After training

True distribution of x

Flow x → z

Empirical distribution of z

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

16 of 66

1-D Flow Models Summary

16

Training:

Inference: evaluate training objective

Sampling: first sample z from then compute

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

17 of 66

Aside: special case of a CDF and p(z) U[0,1]

E.g. Gaussian CDF, mixture CDF, logit CDF, multiple CDFs

17

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

18 of 66

Refresher: Cumulative Density Function (CDF)

18

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

19 of 66

Sampling via inverse CDF

19

Sampling from the model:

The CDF is an invertible, differentiable map from data to [0, 1]

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

20 of 66

CDF Flows: special case of a CDF and p(z) U[0,1]

If we use a flow defined as a parameterized CDF, we recover the original objective for fitting the corresponding parameterized PDF.

20

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

21 of 66

CDF Flows: special case of a CDF and p(z) U[0,1]

If we use a flow defined as a parameterized CDF, we recover the original objective for fitting the corresponding parameterized PDF.

21

this term constant per p(z) uniform

cdf’(x) = pdf(x)

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

22 of 66

How general are flows?

  • Can every (smooth) distribution be represented by a (normalizing) flow? [considering 1-D for now]

22

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

23 of 66

How general are flows?

  • CDF turns any density into uniform
  • Inverse flow is flow

23

→ can turn any (smooth) p(x) into any (smooth) p(z)

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

24 of 66

Recap of Flow Models in 1D

Flow: a differentiable, invertible mapping from x (data) to z (noise)

  • Train so that it turns the data distribution into a base distribution p(z)
    • Common choices: uniform, standard normal

  • This way, mapping z ~ p(z) through the flow’s inverse will yield good samples

24

x (data)

z (noise)

f (inference)

f-1 (sampling)

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

25 of 66

Outline

  • Foundations of Flows (1-D)
  • 2-D Flows
  • N-D Flows
  • Dequantization

25

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

26 of 66

2-D Autoregressive Flow

26

Note that the dependence on x1 in f_theta2 can be arbitrarily complex (including any neural net), no invertibility requirements

Why?

x1 is also given when inverting from z2 to x2

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

27 of 66

Example 2-D Autoregressive Flow

27

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

28 of 66

Training Objective

28

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

29 of 66

2-D Autoregressive Flow: Two Moons

Architecture:

  • Base distribution: Uniform[0,1]^2
  • x1: mixture of 5 Gaussians
    • i.e. z1 = cdf of Mo5G
  • x2: mixture of 5 Gaussians, conditioned on x1
    • I.e. z2 = cdf of Mo5G with mixture weights w, means mu, variances sigma conditioned on x1. I.e. arbitrary NN can be trained to map from x1 to w, mu, sigma

29

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

30 of 66

2-D Autoregressive Flow: Face

Architecture:

  • Base distribution: Uniform[0,1]^2
  • x1: mixture of 5 Gaussians
  • x2: mixture of 5 Gaussians, conditioned on x1

30

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

31 of 66

Outline

  • Foundations of Flows (1-D)
  • 2-D Flows
  • N-D Flows
  • Dequantization

31

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

32 of 66

High-dimensional data

32

f (inference)

f-1 (sampling)

x and z must have the same dimension

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

33 of 66

Outline

  • Foundations of Flows (1-D)
  • 2-D Flows
  • N-D Flows
    • Autoregressive Flows and Inverse Autoregressive Flows
    • RealNVP (like) architectures
    • Glow, Flow++, FFJORD
  • Dequantization

33

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

34 of 66

Autoregressive flows

  • The sampling process of a Bayes net is a flow
    • If autoregressive, this flow is called an autoregressive flow

  • Sampling is an invertible mapping from z to x

34

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

35 of 66

Autoregressive flows

  • How to fit autoregressive flows?
    • Map x to z
    • Fully parallelizable
  • Notice
    • x → z has the same structure as the log likelihood computation of an autoregressive model
    • z → x has the same structure as the sampling procedure of an autoregressive model

35

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

36 of 66

Autoregressive flows

36

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

37 of 66

Inverse autoregressive flows

  • The inverse of an autoregressive flow is also a flow, called the inverse autoregressive flow (IAF)

    • x → z has the same structure as the sampling in an autoregressive model

    • z → x has the same structure as log likelihood computation of an autoregressive model. So, IAF sampling is fast

37

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

38 of 66

AF

IAF

38

Training: long serial chain�Sampling: parallel / fast

Training: parallel / fast�Sampling: long serial chain

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

39 of 66

AF vs IAF

  • Autoregressive flow
    • Fast evaluation of p(x) for arbitrary x
    • Slow sampling
  • Inverse autoregressive flow
    • Slow evaluation of p(x) for arbitrary x, so training directly by maximum likelihood is slow.
    • Fast sampling
    • Fast evaluation of p(x) if x is a sample
  • There are models (Parallel WaveNet, IAF-VAE) that exploit IAF’s fast sampling

39

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

40 of 66

Best of both AF and IAF

Key idea:

  • Train AF -> FAST
  • Distillation:
    • Sample from AF, keep activitations (bit slow..)
    • Train IAF on the samples using access to activations -> FAST

40

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

41 of 66

AF and IAF

Naively, both end up being as deep as the number of variables!

  • E.g. 1MP image → 1M layers…

Can do parameter sharing as in Autoregressive Models from lecture 2 [e.g. RNN, masking]

41

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

42 of 66

Outline

  • Foundations of Flows (1-D)
  • 2-D Flows
  • N-D Flows
    • Autoregressive Flows and Inverse Autoregressive Flows
    • RealNVP (like) architectures
    • Glow, Flow++, FFJORD
  • Dequantization

42

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

43 of 66

Change of MANY variables

For z ~ p(z), sampling process f-1 linearly transforms a small cube dz to a small parallelepiped dx. Probability is conserved:

Intuition: x is likely if it maps to a “large” region in z space

43

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

44 of 66

Flow models: training

Change-of-variables formula lets us compute the density over x:

Train with maximum likelihood:

New key requirement: the Jacobian determinant must be easy to calculate and differentiate!

44

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

45 of 66

Constructing flows: composition

  • Flows can be composed

x → f1 → f2 → … fk → z

  • Easy way to increase expressiveness

45

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

46 of 66

Affine flows

  • Another name for affine flow: multivariate Gaussian.
    • Parameters: an invertible matrix A and a vector b
    • f(x) = A-1(x - b)
  • Sampling: x = Az + b, where z ~ N(0, I)
  • Log likelihood is expensive when dimension is large.
    • The Jacobian of f is A-1
    • Log likelihood involves calculating det(A)

46

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

47 of 66

Elementwise flows

  • Lots of freedom in elementwise flow
    • Can use elementwise affine functions or CDF flows.
  • The Jacobian is diagonal, so the determinant is easy to evaluate.

47

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

48 of 66

NICE/RealNVP

Affine coupling layer

  • Split variables in half: x1:d/2, xd/2+1:d

  • Invertible! Note that sθ and tθ can be arbitrary neural nets with no restrictions.
    • Think of them as data-parameterized elementwise flows.

48

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

49 of 66

NICE/RealNVP

  • It also has a tractable Jacobian determinant

  • The Jacobian is triangular, so its determinant is the product of diagonal entries.

49

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

50 of 66

RealNVP

  • Takeaway: coupling layers allow unrestricted neural nets to be used in flows, while preserving invertibility and tractability

50

[Dinh et al. Density estimation using Real NVP. ICLR 2017]

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

51 of 66

RealNVP Architecture

Input x: 32x32xc image

  • Layer 1: (Checkerboard x3, channel squeeze, channel x3)
    • Split result to get x1: 16x16x2c and z1: 16x16x2c (fine-grained latents)
  • Layer 2: (Checkerboard x3, channel squeeze, channel x3)
    • Split result to get x2: 8x8x4c and z2: 8x8x4c (coarser latents)
  • Layer 3: (Checkerboard x3, channel squeeze, channel x3)
    • Get z3: 4x4x16c (latents for highest-level details)

51

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

52 of 66

RealNVP: How to partition variables?

52

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

53 of 66

Good vs Bad Partitioning

Checkerboard x4; channel squeeze; channel x3; channel unsqueeze; checkerboard x3

(Mask top half; mask bottom half; mask left half; mask right half) x2

53

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

54 of 66

Outline

  • Foundations of Flows (1-D)
  • 2-D Flows
  • N-D Flows
    • Autoregressive Flows and Inverse Autoregressive Flows
    • RealNVP (like) architectures
    • Glow, Flow++, FFJORD
  • Dequantization

54

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

55 of 66

Choice of coupling transformation

  • A Bayes net defines coupling dependency, but what invertible transformation f to use is a design question

  • Affine transformation is the most commonly used one (NICE, RealNVP, IAF-VAE, …)

  • More complex, nonlinear transformations -> better performance
    • CDFs and inverse CDFs for Mixture of Gaussians or Logistics (Flow++)
    • Piecewise linear/quadratic functions (Neural Importance Sampling)

55

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

56 of 66

NN architecture also matters

  • Flow++ = MoL transformation + self-attention in NN
    • Bayes net (coupling dependency), transformation function class, NN architecture all play a role in a flow’s performance. Still an

56

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

57 of 66

Other classes of flows

  • Glow (link)
    • Invertible 1x1 convolutions
    • Large-scale training

  • Continuous time flows (FFJORD)
    • Allows for unrestricted architectures. Invertibility and fast log probability computation guaranteed.

57

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

58 of 66

Glow: Interpolation

58

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

59 of 66

Glow: Attribute Control

59

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

60 of 66

Outline

  • Foundations of Flows (1-D)
  • 2-D Flows
  • N-D Flows
  • Dequantization

60

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

61 of 66

Flow on Discrete Data Without Dequantization...

61

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

62 of 66

Continuous flows for discrete data

  • A problem arises when fitting continuous density models to discrete data: degeneracy
    • When the data are 3-bit pixel values,
    • What density does a model assign to values between bins like 0.4, 0.42…?
  • Correct semantics: we want the integral of probability density within a discrete interval to approximate discrete probability mass

62

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

63 of 66

Continuous flows for discrete data

  • Solution: Dequantization. Add noise to data.
    • We draw noise u uniformly from

63

[Theis, Oord, Bethge, 2016]

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

64 of 66

Flow on Discrete Data With Dequantization

64

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

65 of 66

Future directions

  • The ultimate goal: a likelihood-based model with
    • fast sampling
    • fast inference
    • fast training
    • good samples
    • good compression
  • Flows seem to let us achieve some of these criteria.
  • But how exactly do we design and compose flows for great performance? That’s an open question.
  • Some requirements that might pose permanent challenges:
    • Dimensionality preserving
    • Invertibility
    • Cheap determinant

65

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models

66 of 66

Bibliography

NICE: Dinh, Laurent, David Krueger, and Yoshua Bengio. "NICE: Non-linear independent components estimation." arXiv preprint arXiv:1410.8516 (2014).

RealNVP: Dinh, Laurent, Jascha Sohl-Dickstein, and Samy Bengio. "Density estimation using Real NVP." arXiv preprint arXiv:1605.08803 (2016).

AF: Chen, Xi, et al. "Variational lossy autoencoder." arXiv preprint arXiv:1611.02731 (2016).;

Masked autoregressive flow for density estimation, Papamakarios, George, Theo Pavlakou, and Iain Murray. Advances in Neural Information Processing Systems. 2017.

IAF: Improved variational inference with inverse autoregressive flow. Kingma, Durk P., et al. Advances in neural information processing systems. 2016.

Neural Importance Sampling: Müller, Thomas, et al. arXiv preprint arXiv:1808.03

Glow: Kingma, Durk P., and Prafulla Dhariwal. "Glow: Generative flow with invertible 1x1 convolutions." Advances in Neural Information Processing Systems. 2018.

FFJORD: Grathwohl, Will, et al. "Ffjord: Free-form continuous dynamics for scalable reversible generative models." arXiv preprint arXiv:1810.01367 (2018).

Neural Autoregressive Flow: Huang, Chin-Wei, et al. arXiv:1804.00779 (2018)

Residual Flows for invertible generative modeling, Ricky T. Q. Chen, Jens Behrmann, David Duvenaud, Jörn-Henrik Jacobsen, arXiv: 1906.02735

Flow++: Improving Flow-Based Generative Models with Variational Dequantization and Architecture Design. Ho, Jonathan, Chen, Srinivas, Duan, Abbeel ICML 2019, arXiv:1902.00275 (2019).

MintNet: Building Invertible Neural Networks with Masked Convolutions, Song, Meng, Ermon, NeurIPS 2019, arXiv: 1907.07945

Normalizing Flows for Probabilistic Modeling and Inference, George Papamakarios, Eric Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, Balaji Lakshminarayanan https://arxiv.org/abs/1912.02762, JMLR [survey/tutorial]

Normalizing Flows: An Introduction and Review of Methods, Ivan Kobyzev, Simon JD Prince, Marcus A Brubaker, IEEE PAMI 2021 [survey/tutorial]

66

UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models