1 of 34

GANime

VIDEO GENERATION OF ANIME CONTENT CONDITIONED ON�TWO FRAMES

Farid Abdalla – 12/09/2022

2 of 34

Table of contents

  • Context
  • Introduction
  • Related work
  • Approach
  • Results
  • Conclusion
  • References

2

3 of 34

Context

  • Master Thesis in partnership with Osaka Prefecture University, Japan
    • Duration: 1st of April - 26th of August 2022
  • Supervised by
    • Stefano Carrino of Haute École Arc
    • Motoi Iwata of Osaka Prefecture University

3

4 of 34

Introduction

  • Anime production is expensive
    • In terms of cost
      • ¥ 9-10 million ($70k) for one episode of Kimetsu no Yaiba [1]
      • ¥ 2 billion ($15mio) for the Kimetsu no Yaiba movie
    • But also in time
      • Around 2000 drawn frames for a slow anime [2]
      • Up to 10’000 drawn frames when lots of action
      • Around 9 months to make a 12 episode anime [3]

4

5 of 34

Introduction

  • But it is also very profitable
    • The Kimetsu no Yaiba movie grossed around $500 million [3]
    • It helps to sell the manga or derived products
      • In 2020, Kimetsu no Yaiba franchise generated ¥ 1 trillion ($ 7 billion) in estimated revenue sales
  • Reducing the number of drawn frames would significantly reduce the production time and costs

5

6 of 34

Introduction

  • Image generation has never been so popular
    • DALL-E (available in beta now [4])
    • Imagen
  • Video generation still lacking
    • More complex
    • Coherence between frames
    • More processing power required
  • Generally conditioned on one frame
  • Idea:
    • Use image generation capabilities to generate intermediate frames based on the first and last frame of anime content

6

VideoGPT: generation conditioned on a single frame

7 of 34

Related work

  • Autoencoders
  • Variational autoencoders (VAE)
  • Vector Quantized Variational Autoencoders (VQ-VAE)
  • Vector-Quantized Generative Adversarial Network (VQ-GAN)
  • VideoGPT
  • Point-to-Point Video Generation

7

8 of 34

Autoencoders

  • Encoder
    • Encode data into a latent space
  • Decoder
    • Reconstruct the data
  • Loss:
    • Reconstruction loss (MSE)
  • Problem
    • Irregular latent space

8

Architecture of an autoencoder

9 of 34

Variational Autoencoders

  • Changes
    • Continuous latent space
      • Mean, variance
    • Loss
      • Reconstruction loss
      • Similarity loss (KL Divergence)

9

Architecture of a variational autoencoder

Regularized latent space

10 of 34

Vector-Quantized Variational Autoencoders

  • Changes
    • Discrete latent space
      • Real world data can be discrete
        • Shape, color, size
      • Useful to use with transformers

10

Architecture of a variational autoencoder

  • Loss:
    • Reconstruction (MSE)
    • Codebook
    • Commitment
  • Two stage training

11 of 34

Two stage training

  • First stage
    • Learn a codebook of image constituents
    • Any image can be represented as a sequence of codebook indices
  • Second stage
    • Use another model (PixelCNN, Transformer…) to predict the next index in a sequence
    • E.g. when given [1, 42, 3, 3] the prediction would be 94

11

Architecture of the VQ-GAN

12 of 34

Vector-Quantized GAN

  • Changes
    • Add perceptual loss (VGG16) to the reconstruction loss
    • Add adversarial loss
    • Use transformer as the second stage
      • A transformer is good at predicting the next token (e.g. GPT model)

12

Architecture of the VQ-GAN

13 of 34

Perceptual loss

  • Use an intermediate pretrained model (VGG16)
  • Used for example in style transfer
    • An image can be perceptually similar, but pixel wise very different

13

Example of style transfer

14 of 34

Adversarial loss

  • Min-max game between generator and discriminator
  • The generator try to generate images to fool the discriminator
  • The discriminator try to distinguish real from fake pictures
  • PatchGAN
    • Real / Fake applied on image patches and not on full image

14

Generator

Discriminator

Real / Fake

15 of 34

VideoGPT

  • 3D VQ-VAE
    • Generate video in one shot (not frame by frame)
    • Number of frame not customizable
  • Transformer for second stage
  • Conditioned on a single frame

15

Example of video generation with VideoGPT

16 of 34

Point-to-point video generation

  • Generate video frame by frame between two frames
  • Can specify the length
  • Image generation
    • Variational Autoencoder
  • Temporal consistency
    • LSTM

16

17 of 34

Approach

  • Possible approaches
    • Start from VideoGPT and add VQ-GAN improvements
    • Start from VQ-GAN and add temporal dimension
  • VQ-GAN code was simpler
  • VideoGPT frame number is not customizable
  • Decided to start from VQ-GAN

17

18 of 34

First stage

  • Converted the VQ-GAN code from PyTorch to Tensorflow
  • Trained on datasets
    • Moving MNIST
      • 10’000 videos of 20 frames
      • Size: 200’000 x 64 x 64 x 3
    • Kimetsu no Yaiba
      • 44 episodes of ~34’000 frames
      • 1.5 mio frames
      • Reduced to 300’000 frames because of duplicates
      • Size: 300’000 x 64 x 128 x 3

18

Example of the Moving MNIST dataset

Example of the Kimetsu no Yaiba anime dataset

19 of 34

First stage

  • Replaced VGG16 with VGG19
    • Better results

  • Trained without the discriminator
    • Strange artifacts appeared

19

Artifacts appearing when using the adversarial loss

20 of 34

Second stage

  • Dataset
    • Used PySceneDetect to split the episodes into scenes
    • ~300-500 scenes per episode
    • After postprocessing: ~35’000 scenes

20

Example of scene splitting

21 of 34

Second stage

  • Trained a transformer to generate a video frame by frame
    • First trained a GPT like transformer from the VQ-GAN paper
    • Switched to the Huggingface GPT2 transformer
      • Better results
      • Faster to train
    • Available models
      • gpt2: 117M parameters
      • gpt2-medium: 345M parameters
      • gpt2-large: 774M parameters
      • gpt2-xl: 1558M parameters

21

22 of 34

Frame by frame generation

22

23 of 34

Frame by frame generation - example

23

first frame

last frame

VQ-GAN encoder

[35, 28, 16, 58, 32, …]

indices first frame

[62, 15, 57, 58, 32, …]

indices last frame

GPT2

Number of remaining frames

frame to predict

[38, 28, 51, 58, 32, …]

indices frame to predict

[38, 29, 50, 58, 32, …]

prediction

Loss

VQ-GAN decoder

prediction

24 of 34

M previous frames

24

Input frames

Generated frames

0, last

1

0, 1, last

2

0, 1, 2, last

3

1, 2, 3, last

4

2, 3, 4, last

5

With M=3

25 of 34

Using ground truth during N epochs

25

Ground truth

Frame 0

Frame 1

Frame 2

Frame 3

Frame N

Ground truth

Frame 0

Frame 1

Frame 2

Frame 3

Frame N

Generated

Frame 1

Generated

Frame 2

Ground truth is used until the model starts to produce meaningful results

26 of 34

Using ground truth during N epochs

26

Ground truth

Frame 0

Frame 1

Frame 2

Frame 3

Frame N

Generated

Frame 1

Generated

Frame 2

Generated

Frame 3

27 of 34

Results

27

Generated

Ground truth

On the test set:

28 of 34

Results on the test set

Good results:

28

29 of 34

Results on the test set

Surprising results:

29

30 of 34

Conclusion

  • Generating videos with VQ-GAN + transformer seems promising
    • Some results are coherent and close to the ground truth
    • Others are more creative
  • Some improvements to the VQ-GAN could be done
    • Correct the problem with the adversarial loss
      • Some artifacts are visible when activating the discriminator
  • Using larger transformer could be beneficial
    • The GPT-2 medium used has 345 million parameters but the GPT-2 xl could be used with 1.5 billion parameters

30

31 of 34

Other improvements

  • Improve the dataset
    • Use multiple anime to have more diversity
    • Remove static scenes
  • Use PyTorch to take advantage of memory reduction
    • Gradient Checkpointing
    • Mixed Precision
    • Model parallelism

31

32 of 34

Self-evaluation

I am happy with the results considering that 5 months is a rather short period. It was interesting to work on this project and also be able to travel to Japan to realise it. Even though there are some possible improvements, I hope it can encourage future works on video synthesis

Source code available here: https://github.com/Kurokabe/GANime

Updates will come with a Google colab to test the model (if it fits on the available memory)

32

33 of 34

Questions

33

34 of 34

References

34