CS294-158 Deep Unsupervised Learning
Lecture 3 Likelihood Models: Flow Models
Pieter Abbeel, Wilson Yan, Kevin Frans, Philipp Wu
Overview
2
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Flow Models
3
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Outline
4
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Quick Refresher: Probability Density Models
5
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
How to fit a density model?
Continuous data
0.22159854, 0.84525919, 0.09121633, 0.364252 , 0.30738086,
0.32240615, 0.24371194, 0.22400792, 0.39181847, 0.16407012,
0.84685229, 0.15944969, 0.79142357, 0.6505366 , 0.33123603,
0.81409325, 0.74042126, 0.67950372, 0.74073271, 0.37091554,
0.83476616, 0.38346571, 0.33561352, 0.74100048, 0.32061713,
0.09172335, 0.39037131, 0.80496586, 0.80301971, 0.32048452,
0.79428266, 0.6961708 , 0.20183965, 0.82621227, 0.367292 ,
0.76095756, 0.10125199, 0.41495427, 0.85999877, 0.23004346,
0.28881973, 0.41211802, 0.24764836, 0.72743029, 0.20749136,
0.29877091, 0.75781455, 0.29219608, 0.79681589, 0.86823823,
0.29936483, 0.02948181, 0.78528968, 0.84015573, 0.40391632,
0.77816356, 0.75039186, 0.84709016, 0.76950307, 0.29772759,
0.41163966, 0.24862007, 0.34249207, 0.74363912, 0.38303383, …
6
Maximum Likelihood:
Equivalently:
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Example Density Model: Mixtures of Gaussians
Parameters: means and variances of components, mixture weights
7
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Aside on Mixtures of Gaussians
Do mixtures of Gaussians work for high-dimensional data?
Not really. The sampling process is:
Imagine this for modeling natural images! The only way a realistic image can be generated is if it is a cluster center, i.e. if it is already stored directly in the parameters.
8
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Flow Models: Invertible transform from data x to embedding z
9
...
Original data space, complex distribution to model
x
Embedding space, we enforce a simpler distribution
Note: VAE has very similar diagram, but learns both directions rather than imposing invertibility (see L4)
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Change of Variables Formula
10
Note: for this simple formula to work, �it requires invertible & differentiable
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Flow Models: Training
11
Assuming we have an expression for ,
this can be optimized with Stochastic Gradient Descent
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Flow Models: Examples
12
Two choices to make:
Invertible function class
Embedding space density
I.e., in 1-D this is any monotonically increasing (or decreasing) function
E.g.:
“normalizing flow”
mixture of Gaussians
-> make sure maps to [0,1]
E.g.:
- a x + b (for positive a)�- polynomials with positive coefficients and only odd powers�- exp (theta x)�- sigmoid (a x + b)�- cumulative density functions, e.g. CDF of mixture of Gaussians or weighted sum of logistics�- composition of flows = flow
Ideally an easy distribution to sample from
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Example: Flow to Gaussian z
13
Before training
After training
True distribution of x
Flow x → z
Empirical distribution of z
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Example: Flow to Uniform z
14
Before training
After training
True distribution of x
Flow x → z
Empirical distribution of z
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Example: Flow to Beta(5,5) z
15
Before training
After training
True distribution of x
Flow x → z
Empirical distribution of z
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
1-D Flow Models Summary
16
Training:
Inference: evaluate training objective
Sampling: first sample z from then compute
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Aside: special case of a CDF and p(z) U[0,1]
E.g. Gaussian CDF, mixture CDF, logit CDF, multiple CDFs
17
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Refresher: Cumulative Density Function (CDF)
18
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Sampling via inverse CDF
19
Sampling from the model:
The CDF is an invertible, differentiable map from data to [0, 1]
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
CDF Flows: special case of a CDF and p(z) U[0,1]
If we use a flow defined as a parameterized CDF, we recover the original objective for fitting the corresponding parameterized PDF.
20
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
CDF Flows: special case of a CDF and p(z) U[0,1]
If we use a flow defined as a parameterized CDF, we recover the original objective for fitting the corresponding parameterized PDF.
21
this term constant per p(z) uniform
cdf’(x) = pdf(x)
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
How general are flows?
22
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
How general are flows?
23
→ can turn any (smooth) p(x) into any (smooth) p(z)
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Recap of Flow Models in 1D
Flow: a differentiable, invertible mapping from x (data) to z (noise)
24
x (data)
z (noise)
f (inference)
f-1 (sampling)
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Outline
25
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
2-D Autoregressive Flow
26
Note that the dependence on x1 in f_theta2 can be arbitrarily complex (including any neural net), no invertibility requirements
Why?
x1 is also given when inverting from z2 to x2
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Example 2-D Autoregressive Flow
27
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Training Objective
28
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
2-D Autoregressive Flow: Two Moons
Architecture:
29
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
2-D Autoregressive Flow: Face
Architecture:
30
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Outline
31
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
High-dimensional data
32
f (inference)
f-1 (sampling)
x and z must have the same dimension
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Outline
33
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Autoregressive flows
34
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Autoregressive flows
35
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Autoregressive flows
36
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Inverse autoregressive flows
37
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
AF
IAF
38
Training: long serial chain�Sampling: parallel / fast
Training: parallel / fast�Sampling: long serial chain
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
AF vs IAF
39
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Best of both AF and IAF
Key idea:
40
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
AF and IAF
Naively, both end up being as deep as the number of variables!
Can do parameter sharing as in Autoregressive Models from lecture 2 [e.g. RNN, masking]
41
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Outline
42
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Change of MANY variables
For z ~ p(z), sampling process f-1 linearly transforms a small cube dz to a small parallelepiped dx. Probability is conserved:
Intuition: x is likely if it maps to a “large” region in z space
43
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Flow models: training
Change-of-variables formula lets us compute the density over x:
Train with maximum likelihood:
New key requirement: the Jacobian determinant must be easy to calculate and differentiate!
44
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Constructing flows: composition
x → f1 → f2 → … fk → z
45
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Affine flows
46
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Elementwise flows
47
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
NICE/RealNVP
Affine coupling layer
48
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
NICE/RealNVP
49
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
RealNVP
50
[Dinh et al. Density estimation using Real NVP. ICLR 2017]
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
RealNVP Architecture
Input x: 32x32xc image
51
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
RealNVP: How to partition variables?
52
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Good vs Bad Partitioning
Checkerboard x4; channel squeeze; channel x3; channel unsqueeze; checkerboard x3
(Mask top half; mask bottom half; mask left half; mask right half) x2
53
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Outline
54
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Choice of coupling transformation
55
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
NN architecture also matters
56
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Other classes of flows
57
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Glow: Interpolation
58
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Glow: Attribute Control
59
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Outline
60
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Flow on Discrete Data Without Dequantization...
61
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Continuous flows for discrete data
62
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Continuous flows for discrete data
63
[Theis, Oord, Bethge, 2016]
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Flow on Discrete Data With Dequantization
64
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Future directions
65
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models
Bibliography
NICE: Dinh, Laurent, David Krueger, and Yoshua Bengio. "NICE: Non-linear independent components estimation." arXiv preprint arXiv:1410.8516 (2014).
RealNVP: Dinh, Laurent, Jascha Sohl-Dickstein, and Samy Bengio. "Density estimation using Real NVP." arXiv preprint arXiv:1605.08803 (2016).
AF: Chen, Xi, et al. "Variational lossy autoencoder." arXiv preprint arXiv:1611.02731 (2016).;
Masked autoregressive flow for density estimation, Papamakarios, George, Theo Pavlakou, and Iain Murray. Advances in Neural Information Processing Systems. 2017.
IAF: Improved variational inference with inverse autoregressive flow. Kingma, Durk P., et al. Advances in neural information processing systems. 2016.
Neural Importance Sampling: Müller, Thomas, et al. arXiv preprint arXiv:1808.03
Glow: Kingma, Durk P., and Prafulla Dhariwal. "Glow: Generative flow with invertible 1x1 convolutions." Advances in Neural Information Processing Systems. 2018.
FFJORD: Grathwohl, Will, et al. "Ffjord: Free-form continuous dynamics for scalable reversible generative models." arXiv preprint arXiv:1810.01367 (2018).
Neural Autoregressive Flow: Huang, Chin-Wei, et al. arXiv:1804.00779 (2018)
Residual Flows for invertible generative modeling, Ricky T. Q. Chen, Jens Behrmann, David Duvenaud, Jörn-Henrik Jacobsen, arXiv: 1906.02735
Flow++: Improving Flow-Based Generative Models with Variational Dequantization and Architecture Design. Ho, Jonathan, Chen, Srinivas, Duan, Abbeel ICML 2019, arXiv:1902.00275 (2019).
MintNet: Building Invertible Neural Networks with Masked Convolutions, Song, Meng, Ermon, NeurIPS 2019, arXiv: 1907.07945
Normalizing Flows for Probabilistic Modeling and Inference, George Papamakarios, Eric Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, Balaji Lakshminarayanan https://arxiv.org/abs/1912.02762, JMLR [survey/tutorial]
Normalizing Flows: An Introduction and Review of Methods, Ivan Kobyzev, Simon JD Prince, Marcus A Brubaker, IEEE PAMI 2021 [survey/tutorial]
66
UC Berkeley -- Spring 2024 -- Deep Unsupervised Learning -- Pieter Abbeel, Kevin Frans, Philipp Wu, Wilson Yan -- L3 Flow Models