1 of 60

CSE 5524: �Generative models - 2

1

2 of 60

HW 3 & HW 4 & quizzes

  • HW 3
    • Caution: please re-download the data
    • Due: 4/11/2025

  • HW 4
    • Plan: A lighter homework
    • Planned release: 4/7/2025
    • Due: 4/21/2025

  • Quizzes:
    • Two quizzes to be released in the next two weeks --- True/False, multiple choices, unlimited tries

3 of 60

Final project (30%)

  • Project proposal: 4/3 (3%)
  • Instructions are on Carmen

  • Baseline:
    • Something “simple” that you can implement in a day or two
    • Get basic idea of the computation
    • Get basic idea of the performance (e.g., classification accuracy)

4 of 60

Today (33; skip 34)

  • Recap
  • Variational Auto-encoder (VAE)

4

5 of 60

Recap: Recognition models vs. generative models

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

Generative models achieve “one-to-many” mapping by making the generator a “stochastic function”

latent variables

6 of 60

Recap: Unconditional generative models

  • Given a collection of images x without label y

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

Gray: visible; White: latent

7 of 60

Recap: What is the objective?

  • High-level: the output of the generator looks like real data

  • Different realization:
    • Match certain statistics: mean color, color variance, etc.
    • Synthetic data have high probability under a density model fit to real data
    • Synthetic data and real data are indistinguishable

8 of 60

Recap: Direct and indirect approaches

Direct “generator”

Indirect “generator”

    • Pros: easy to sample image (in testing)
    • Cons: hard to compute the density (in training)
    • Pros: easy to compute the density (in training) Cons: hard to sample image (in testing)

9 of 60

Recap: What we have seen or will see

  • Density model
  • Autoregressive density model
  • Diffusion model
  • Generative adversarial net (GAN)
  • Variational auto encoder (VAE)

10 of 60

Recap: Learning density function

  • Minimize the Kullback–Leibler (KL) divergence between the model and data

Maximum likelihood estimation (MLE)

= Minimum KL divergence

11 of 60

Recap: Maximum likelihood estimation (MLE)

similar

How likely the model can generate the true data?

12 of 60

Recap: Autoregressive density model

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

13 of 60

Recap: Diffusion model

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

14 of 60

Recap: Generative adversarial net (GAN)

14

Generator

Discriminator

REAL

FAKE

[Credits: Mengdi Fan and Xinyu Zhou, CSE 5539 course presentation]

 

 

 

15 of 60

Today (33; skip 34)

  • Recap
  • Variational Auto-encoder (VAE)
    • Concepts
    • AE vs. VAE
    • Math of VAE

15

16 of 60

Representation learning vs. generative modeling

“Unlabeled” data

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

17 of 60

Representation learning vs. generative modeling

“Unlabeled” data

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

Map data to simple underlying representation (embedding)

Map simple base distribution (noise) to data

18 of 60

Representation learning vs. generative modeling

“Unlabeled” data

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

19 of 60

Representation learning vs. generative modeling

“Unlabeled” data

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

Latent variables are representations of the data

20 of 60

Technical setting for connecting them

  • Notation:

  • Goal:

Representation

Generation

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

21 of 60

Today (33; skip 34)

  • Recap
  • Variational Auto-encoder (VAE)
    • Concepts
    • AE vs. VAE
    • Math of VAE

21

22 of 60

Question: difference from auto-encoder (AE)

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

Now we assume we know or force Pz: Gaussian

Auto-encoder (AE)

Variational

Auto-encoder (VAE)

23 of 60

Trick from AE to VAE

  • Regularize the latent distribution to squish into a Gaussian

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

24 of 60

Auto-encoder (AE) vs. Variational Auto-encoder (VAE)

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

Auto-encoder (AE)

Variational

Auto-encoder (VAE)

25 of 60

Today (33; skip 34)

  • Recap
  • Variational Auto-encoder (VAE)
    • Concepts
    • AE vs. VAE
    • Math of VAE

25

26 of 60

Objective function & Hypothesis space

  • Data log likelihood:

  • What is ?

Mixture model

Gaussian

Conditional Gaussian

27 of 60

Infinite mixture of Gaussian

  • VAE hypothesis space:

  • Standard Mixture of Gaussian:

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

28 of 60

Infinite mixture of Gaussian

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

29 of 60

Today (33; skip 34)

  • Recap
  • Variational Auto-encoder (VAE)
    • Concepts
    • AE vs. VAE
    • Math of VAE

29

30 of 60

Optimization (1)

  • Maximize the log likelihood

[Figure credit: A. Torralba, P. Isola, and W. T. Freeman, Foundations of Computer Vision.]

31 of 60

Optimization (2)

  • Integration approximated by Monte Carlo estimate (sampling + average)

32 of 60

33 of 60

34 of 60

35 of 60

Today (33; skip 34)

  • Recap
  • Variational Auto-encoder (VAE)
    • Concepts
    • AE vs. VAE
    • Math of VAE

35

36 of 60

Optimization (3)

  • Problem: inefficient usage of “z”

  • Decomposition

    • Only some of “z” will lead to high probability

37 of 60

Optimization (3)

  • Problem: inefficient usage of “z”

  • Decomposition

    • Only some of “z” will lead to high probability

38 of 60

Optimization (3)

  • Importance sampling

39 of 60

Optimization (3)

  • Importance sampling

“specific” distribution for sampling z

40 of 60

Optimization (3)

  • Importance sampling

41 of 60

Optimization (3)

  • Importance sampling

  • Optimal q:

42 of 60

Today (33; skip 34)

  • Recap
  • Variational Auto-encoder (VAE)
    • Concepts
    • AE vs. VAE
    • Math of VAE

42

43 of 60

Optimization (4)

  • It could be hard

  • Variational inference: approximate by another conditional Gaussian

44 of 60

Put everything together

45 of 60

Put everything together

46 of 60

Put everything together

47 of 60

Put everything together

48 of 60

Today (33; skip 34)

  • Recap
  • Variational Auto-encoder (VAE)
    • Concepts
    • AE vs. VAE
    • Math of VAE

48

49 of 60

Overall training

50 of 60

Overall training

51 of 60

AE vs. VAE

Auto-encoder (AE)

Variational

Auto-encoder (VAE)

52 of 60

53 of 60

54 of 60

55 of 60

Today (33; skip 34)

  • Recap
  • Variational Auto-encoder (VAE)
    • Concepts
    • AE vs. VAE
    • Math of VAE
    • Examples

55

56 of 60

Examples

57 of 60

58 of 60

Gaussian encourages “disentanglement”

59 of 60

Recap: Generative models with “z” as Gaussian

[Credits: What are Diffusion Models?]

60 of 60