1 of 56

Diffusion For Babies

A Tutorial

Christian Simon

4

2 of 56

Meta -Emu

Runway

Imagen

Stable Image

3 of 56

Diffusion in Physics

Can we borrow this concept for machine learning problems?

4 of 56

How to get to this point?

5 of 56

Neural Network Training

  1. What is exactly neural network trying to do?
  2. How is neural network trained?

6 of 56

Neural Network Training

  • What is exactly neural network trying to do?
  • How is neural network trained?

x

y

f(.;θ)

7 of 56

Gaussian Distribution

Love Distribution

Decoder θ

pθ(x)

Training a function to map from Gaussian dist. to ♡ dist.

8 of 56

Gaussian Distribution

Love Distribution

Decoder θ

pθ(x)

Dataset: {x1, x2, x3, …, xN} ~ pdata(x)

9 of 56

How to estimate θ?

Decoder θ

10 of 56

Gaussian Distribution

Love Distribution

Diffusion

Decoder θ

pθ(x)

Dataset: {x1, x2, x3, …, xN} ~ pdata(x)

11 of 56

Maximum Likelihood Estimation (MLE)

  • What is MLE? - a method of estimating the parameters of an assumed probability distribution, given some observed data.
  • Take for example, when fitting a Gaussian to our dataset, we take the sample mean and sample variance, and use it as the parameter of our Gaussian.

12 of 56

Gaussian Distribution

Love Distribution

Diffusion

Decoder θ

pθ(x)

Dataset: {x1, x2, x3, …, xN} ~ pdata(x)

13 of 56

Let’s make it easier to compute!

14 of 56

Let’s make it easier to compute!

15 of 56

Remember expectation to integral

16 of 56

Remember expectation to integral

17 of 56

Similarity Between 2 Distributions

18 of 56

Matching Two Distributions

Gaussian Distribution

Love Distribution

Diffusion

Estimation

True/Data

Decoder θ

pθ(x)

Dataset: {x1, x2, x3, …, xN} ~ pdata(x)

pdata(x)

KL

Div.

p(z)

19 of 56

Matching Two Distributions

Gaussian Distribution

Love Distribution

Decoder θ

pθ(x)

p(z)

Remember: p(x,z) = p(z|x)p(x)

20 of 56

Simplifying the Formulation

Gaussian Distribution

Love Distribution

Decoder θ

p(z)

z

q(z | x)

Dataset: {x1, x2, x3, …, xN} ~ pdata(x)

Encoder

pθ(x|z)

We can use this integral to 1, depending on preference.

21 of 56

Remember integral to expectation

22 of 56

Derive

Remember:

23 of 56

Derive

24 of 56

Derive

25 of 56

Derive

REMEMBER THIS?????

26 of 56

Variational AutoEncoder

27 of 56

How do we connect these two algorithms?

Encoding is a function to add noise

Decoding is (trainable) parameterized function to denoise

Layer is represented as denoising iterations, while VAE is not

28 of 56

29 of 56

q

p

Remember:

30 of 56

q

31 of 56

q

32 of 56

Back to Distribution Matching using KL Divergence

33 of 56

34 of 56

Recursive

35 of 56

36 of 56

37 of 56

38 of 56

39 of 56

40 of 56

41 of 56

42 of 56

43 of 56

44 of 56

Simplify

45 of 56

46 of 56

47 of 56

48 of 56

49 of 56

50 of 56

51 of 56

Inference DDIM

DDIM >>>> DDPM (Sampling in standard diffusion models)

Denoising Diffusion Implicit Models - DDIM

200 steps >>>> 1000 steps

52 of 56

Stable Diffusion

53 of 56

Diffusion for Classification

54 of 56

Diffusion - Interrupted Optimization

55 of 56

Drawbacks

  • Diffusion is hard to control. The randomness in diffusion process makes the generation hard to predict and controlled.

  • Slow generation. DDPM might take 20 minutes on a low end GPU.

  • Generation is hard to quantify with existing metrics. Most of the time, evaluation requires human in the loop to check the fidelity and quality of generated samples

56 of 56

Some Open Topics

  • Fast diffusion model inference -1 step?
  • Diffusion for classification purposes.
  • Diffusion that can be combined with existing large models (e.g., LLMs).
  • Diffusion models that can work on multimodalities