Language Model Beats Diffusion�– Tokenizer is Key to Visual Generation
Lijun Yu
04/2024
lijun@lj-y.com
Motivation
Here LMs refer to transformer models that learn discrete token sequences
Background: LMs in Visual Generation
| | Tokenizer | LM Type |
Image | ImageGPT | Color clustering | AR-LM & MLM |
DALL·E | dVAE | AR-LM | |
Taming transformer | VQGAN | AR-LM | |
Parti | ViT-VQGAN | AR-LM | |
MaskGIT & Muse | VQGAN | MLM | |
Video | Phenaki | CViViT VQGAN | MLM |
MAGVIT | 3D VQGAN | MLM |
Preliminary: Image Tokenization
Video Tokenization
Issues with MAGVIT Tokenizer
Introducing MAGVIT-v2 Tokenizer
---- Previously ----�SPAE: Semantic Pyramid AutoEncoder
SPAE Tokenization Example
Text to Image �with frozen LLMs
The first time a frozen LLM generates images�without relying on external models, e.g., stable diffusion
With SPAE, we can transform any LLM into a Gemini-style multimodal model even without tuning.
Lookup-Free Quantization
Lookup-Free Quantization
Joint Image-Video Tokenization
Joint Image-Video Tokenization
Exploring two new designs:
Joint Image-Video Tokenization
| # Params | FID↓ | FVD↓ |
MAGVIT | 39M | n/a | 107.15 |
C-ViViT | 90M | 28.02 | 437.54 |
C-ViViT + MAGVIT | 67M | 13.52 | 316.70 |
MAGVITv2 | 58M | 7.06 | 96.33 |
Comparing joint/causal tokenization architectures on UCF-101.
FID is calculated on the first frame.
Architecture Ablations
Image Reconstruction
Original
VQGAN
(ImageNet)
MAGVIT-v2
(ImageNet)
MAGVIT-v2
(Web Images)
Video Compression
MAGVIT-v2 is preferred over MAGVIT, HEVC (H.265), and VVC (H.266) in subjective rater study.
MAGVIT: Masked Generative Video Transformer
Squeezing something
Frame
Prediction
Frame
Interpolation
Out painting
Inpainting
Class-conditional �generation
The first multi-task masked transformer for video generation,�with state-of-the-art generation quality and efficiency.
Masked Video Synthesis
Bidirectional Transformer
Masked tokens at each step: �Mask Condition Prediction
Input video
Generated videos
Sampling progress
Here we decode intermediate states for visualization, which does not happen in standard sampling.
Token Factorization
Image Generation
Type | Method | FID↓ | Guided�FID↓ | # Params | # Steps | Latent |
Diff. + VAE* | DiT-XL/2 | 12.03 | 3.04 | 675M | 250 | 642 |
Diffusion | RIN | 3.95 | | 320M | 1000 | |
Diffusion | VDM++ | 2.99 | 2.65 | 2B | 512 | |
MLM + VQ | MaskGIT | 7.32 | | 227M | 12 | 322 |
MLM + VQ | DPC | 3.62 | | 619M | 72 | 322 |
MLM + LFQ | MAGVIT-v2 | 4.61 | | 307M | 12 | 162 |
3.07 | 1.91 | 64 |
Image Generation
Type | Method | FID↓ | Guided�FID↓ | # Params | # Steps |
Diffusion + VAE* | MDT | 6.23 | 1.79 | 676M | 250 |
Diffusion | RIN | 3.42 | | 410M | 1000 |
Diffusion | VDM++ | 2.40 | 2.12 | 2B | 512 |
MLM + VQ | Contextual RQ | 3.41 | | 1.4B | 72 |
MLM + VQ | DPC | 4.45 | | 454M | 180 |
MLM + LFQ | MAGVIT-v2 | 3.65 | 1.78 | 307M | 64 |
Video Generation
Type | Method | K600 FVD↓ | UCF FVD↓ | # Params | # Steps |
Diffusion | VDM | 16.2±0.3 | | 1.1B | 256 |
Diffusion | RIN | 10.8 | | 411M | 1000 |
MLM + VQ | MAGVIT | 9.9±0.3 | 76±2 | 306M | 12 |
MLM + LFQ | MAGVIT-v2 | 5.2±0.2 | | 307M | 12 |
4.3±0.1 | 58±3 | 24 |
Video Generation: Frame prediction on Kinetics-600 and class-conditional generation on UCF-101
Video Generation
MAGVIT-v2 enables remarkable video generation quality with transformers using various objectives
MAGVIT-v2 Tokenizer
AR-LM
LDM
MLM
VideoPoet: �A Large Language Model for Zero-Shot Video Generation
VideoPoet
Training
Preliminary Scaling
Video Generation
Audio Generation
Text-to-Video
Model | CLIPSIM | FVD |
Video LDM | 0.2929 | |
Make-A-Video | 0.3049 | - |
Show-1 | 0.3072 | 538 |
VideoPoet (pretrain) | 0.3049 | 213 |
VideoPoet (task adapt) | 0.3123 | - |
Zero-shot text-to-video evaluation on MSR-VTT
See more: https://sites.research.google/videopoet
Text-to-Video
Human Evaluation
Image-to-Video
Source: https://twitter.com/AIsonesone
Video-to-Audio
On generated videos
W.A.L.T
W.A.L.T
Related: NaViT
Dehghani et al. Patch n’ Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution. In NeurIPS 2023.
Related: Sora
Our best guesses – it’s (almost) all about execution
Brooks et al. Video generation models as world simulators. 2024.
Related: Genie
Bruce et al. Genie: Generative Interactive Environments. 2024.
Takeaways
We have covered a lot:
MAGVIT, SPAE, MAGVIT-v2, VideoPoet, W.A.L.T, …�
One point that may be noteworthy:
Language models are at least as good as diffusion models on visual synthesis, if a good tokenizer/representation is available.
And language models are much more general (and scalable).
Just generating good-looking videos will never be our end goal.
Multi-modal native, interactive world models, learn from raw signals
Super intelligence?
Thank you!