CSE 5524: �Generative models - 2
1
HW 3 & HW 4 & quizzes
Final project (30%)
Today (33; skip 34)
4
Recap: Recognition models vs. generative models
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Generative models achieve “one-to-many” mapping by making the generator a “stochastic function”
latent variables
Recap: Unconditional generative models
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Gray: visible; White: latent
Recap: What is the objective?
Recap: Direct and indirect approaches
Direct “generator”
Indirect “generator”
Recap: What we have seen or will see
Recap: Learning density function
Maximum likelihood estimation (MLE)
= Minimum KL divergence
Recap: Maximum likelihood estimation (MLE)
similar
How likely the model can generate the true data?
Recap: Autoregressive density model
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Recap: Diffusion model
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Recap: Generative adversarial net (GAN)
14
Generator
Discriminator
REAL
FAKE
[Credits: Mengdi Fan and Xinyu Zhou, CSE 5539 course presentation]
Today (33; skip 34)
15
Representation learning vs. generative modeling
“Unlabeled” data
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Representation learning vs. generative modeling
“Unlabeled” data
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Map data to simple underlying representation (embedding)
Map simple base distribution (noise) to data
Representation learning vs. generative modeling
“Unlabeled” data
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Representation learning vs. generative modeling
“Unlabeled” data
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Latent variables are representations of the data
Technical setting for connecting them
Representation
Generation
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Today (33; skip 34)
21
Question: difference from auto-encoder (AE)
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Now we assume we know or force Pz: Gaussian
Auto-encoder (AE)
Variational
Auto-encoder (VAE)
Trick from AE to VAE
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Auto-encoder (AE) vs. Variational Auto-encoder (VAE)
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Auto-encoder (AE)
Variational
Auto-encoder (VAE)
Today (33; skip 34)
25
Objective function & Hypothesis space
Mixture model
Gaussian
Conditional Gaussian
Infinite mixture of Gaussian
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Infinite mixture of Gaussian
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Today (33; skip 34)
29
Optimization (1)
[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]
Optimization (2)
Today (33; skip 34)
35
Optimization (3)
Optimization (3)
Optimization (3)
Optimization (3)
“specific” distribution for sampling z
Optimization (3)
Optimization (3)
Today (33; skip 34)
42
Optimization (4)
Put everything together
Put everything together
Put everything together
Put everything together
Today (33; skip 34)
48
Overall training
Overall training
AE vs. VAE
Auto-encoder (AE)
Variational
Auto-encoder (VAE)
Today (33; skip 34)
55
Examples
Gaussian encourages “disentanglement”
Recap: Generative models with “z” as Gaussian
[Credits: What are Diffusion Models?]