1 of 13

EG3D: Efficient Geometry-aware 3D Generative Adversarial Networks

Present by Yihang Liu

Eric Ryan Chan, Connor Zhizhen Lin, Matthew Aaron Chan, Koki Nagano,

Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas Guibas,

Jonathan Tremblay, Sameh Khamis, Tero Karras, Gordon Wetzstein

Stanford University & NVIDIA

2 of 13

One Sentence Summary

Use an efficient tri-plane representation to achieve high-quality, real-time 3D-aware image synthesis

2

3 of 13

The Problem - Quality v.s 3D Awareness

3D GANs Quality <<<<< 2D GAN Quality

3D rely on

Neural implicit representations

–Too slow for high-resolution training

Voxel grids

–Memory inefficient, hard to scale, low quality

2D does not have view consistency and impairs 3D geometry quality

But some hybrid methods look good…

3

4 of 13

Method - Tri-plane hybrid 3D representation

Three axis-aligned orthogonal feature planes (NxNxC)

4

Query process:

  • Project x ∈ R³ onto each plane
  • Retrieve features via bilinear interpolation
  • Aggregate three feature vectors
  • MLP interprets for color/density
  • Volume rendering for RGB images

Advantages:

  • Efficient
    • Keeps decoder small
    • Use explicit features
  • Maintains expressiveness

5 of 13

Method - 3D GAN Architecture Based on 2D

Generator Backbone

StyleGAN2 generate the features for tri-plane

Latent Code and camera parameters are fed into the Mapping Network (mod. latent)

Output 256x256x96 feature map, then reshaped

Rendering

Sample features from the tri-planes

Aggregate by summation, and fed a lightweight decoder (→scalar density σ, 32-channel feature)

Neural volume renderer → 2D feature image

Upsamples low-resolution neural rendering (2 StyleGAN2 Conv.)

Dual Discrimination

Modified StyleGAN2 discriminator

dual discrimination - avoid multi-view inconsistency

Use camera pose matrices as conditioning labels

6 of 13

Method - Dual Discrimination

Purpose:

Previous methods had multi-view inconsistency issues when using 2D CNN upsampling.

This method want to ensures consistency between:

Raw neural rendering (I_RGB)

Super-resolved final output (I⁺_RGB)

It makes final output to match real image distribution

For Generated Images:

  • Take I_RGB
  • Bilinearly upsample it to same resolution as I⁺_RGB
  • Concatenate with I⁺_RGB
  • Feed this 6-channel image to discriminator

For Real Images:

  • Take original real image
  • Create a blurred copy
  • Concatenate them into 6-channel image

7 of 13

Method - Modeling Pose-correlated Attributes

  • Real datasets (like FFHQ) have inherent pose-appearance correlations
    • People looking at camera are more likely to be smiling than those at angles
  • So, naive handling leads to view-inconsistent results
  • Previous works trade off view consistency v.s. ignore pose correlations

Generator Pose Conditioning:

  • Feed both latent code z and camera parameters P to backbone mapping network
  • Allows target view to influence scene synthesis
  • Training Strategy: generator can model pose-dependent biases in dataset and reproduces real image distributions
  • Fix the conditioning pose when rendering from moving camera
  • Prevents scene from shifting with camera pose
  • Randomly swap conditioning pose with another random pose with 50% probability during training

8 of 13

Experiments

Datasets

  • FFHQ
    • Real-world human face dataset
    • Uses off-the-shelf pose estimators
    • Augmented with horizontal flips

  • AFHQv2 Cats
    • Small, real-world cat face dataset
    • Apply transfer learning from FFHQ checkpoints
    • Use adaptive data augmentation

9 of 13

Experiments - Result

10 of 13

Experiments - Ablation

11 of 13

Applications

Style Mixing

Single-view 3D reconstruction

12 of 13

Key Contributions

  1. Tri-plane Hybrid Representation
  2. Dual Discrimination
  3. Generator Pose Conditioning (Random Pose Swapping)

13 of 13

Thank you!

And … Questions