1 of 34

Video Autoencoder: self-supervised disentanglement

of static 3D structure and motion

ICCV 2021

1

Sifei Liu

NVIDIA

Alexei A. Efros

UC Berkeley

Xiaolong Wang

UC San Diego

Zihang Lai

CMU

2 of 34

Disentangle the visual world

2

Movement

Depth

Structure and viewpoint

3 of 34

Prior work

From the very beginnings of computer vision, …

3

Barrow and Tenenbaum, Comput. Vis. Syst., 1978

Tomasi and Kanade, IJCV, 1992

4 of 34

Prior work

Learning disentangled visual representation from auto-encoders

4

Park et al., NeurIPS, 2020

Liu et al., ECCV, 2020

Kulkarni et al., NIPS, 2015

Kim et al., ICML, 2018

5 of 34

Prior work

5

Tung et al., CVPR, 2019

Nguyen-Phuoc et al., ICCV, 2019

Wiles et al., CVPR, 2020

Learning 3D from 2D supervisions

6 of 34

Objective

In this work, we learn to separate 3D structure from Camera Motion without any human annotations

6

Input raw video

ENCODER

7 of 34

Test time

The features obtained (3D structure, Camera Motion) can be used for several downstream tasks:

7

Novel view synthesis

Pose estimation in video

Video following

7

Single image

Raw video

Single image + Raw video

*Actual results

8 of 34

Test time

Novel view synthesis

8

ENCODER

Pre-defined

trajectory

Single still image

DECODER

Novel view synthesis

9 of 34

Learning from temporal continuity of videos

9

Spatio-temporal continuity

No spatio-temporal continuity

10 of 34

Learning from temporal continuity of videos

Assume that a local snippet of video is capturing a static scene

10

Video snippet

Frame 1

Frame 2

Frame 3

Frame 4

Frame 5

Static scene

11 of 34

Learning from temporal continuity of videos

Assume that a local snippet of video is capturing a static scene

11

Video snippet

Frame 1

Frame 2

Frame 3

Frame 4

Frame 5

3D scene structure

Trajectory

Encode

[ ]

R1, t1

R2, t2

12 of 34

Learning from temporal continuity of videos

Assume that a local snippet of video is capturing a static scene

12

Video snippet

Frame 1

Frame 2

Frame 3

Frame 4

Frame 5

Decode

13 of 34

Model architecture

13

ENCODER

3D Structure

(Deep voxels)

3D Trajectory

Structure + Motion

3D

Transform

Video

DECODER

Reconstruction

3D Encoder

Traj. Encoder

Decoder

Consistency

loss

Camera

transformation

Joint

optimisation

14 of 34

Model architecture

14

3D Encoder

1

Traj. Encoder

2

3

Decoder

4

ResNet

50

Reshape

3D conv

Input image

3D deep voxels

15 of 34

Model architecture

15

conv

Stacked image pair

relu

conv

relu

conv

relu

 

6D Pose

3D Encoder

1

Traj. Encoder

2

3

Decoder

4

16 of 34

Model architecture

16

3D Encoder

1

Traj. Encoder

2

3

Decoder

4

3D conv

Reshape

2D conv

Output image

3D deep voxels

3D Transform

Camera pose

17 of 34

Training loss

  • Reconstruction loss:

  • GAN loss:

  • Consistency loss between deep voxels:

17

18 of 34

Video Autoencoder

18

Results

19 of 34

Datasets

19

RealEstate10K

Matterport3D

Replica

20 of 34

Results

Novel view synthesis

20

Single input image

Output video

RealEstate10K dataset

21 of 34

21

A Japanese living room (out-of-distribution)

Single input image

Output video

Novel view synthesis (Out-of-domain results)

Results

22 of 34

22

Novel view synthesis (Out-of-domain results)

Results

Single input image

Output video

“Spirited Away”

23 of 34

23

Bedroom in Arles, Vincent van Gogh

Single input image

Output video

Novel view synthesis (Out-of-domain results)

Results

24 of 34

Comparison with previous methods

24

Input image

RealEstate10K dataset

Mustikovela et al.

Tung et al.

Yu et al.

Wiles et al.

Ours

Ground truth

More artefacts

Blank area

Results

Tung et al.

Mustikovela et al.

Yu et al.

Wiles et al.

25 of 34

25

Input image

RealEstate10K dataset

Comparison with previous methods

Results

Mustikovela et al.

Tung et al.

Yu et al.

Wiles et al.

Ours

Ground truth

More artefacts

Blank area

26 of 34

Results

Novel view synthesis (RealEstate10K)

26

 

27 of 34

Results

Novel view synthesis (Matterport3D & Replica)

27

Matterport 3D

MP3D → Replica

Method

Pose

PSNR ↑

SSIM ↑

PSNR ↑

SSIM ↑

Methods without any camera supervision

Ours

×

20.58

0.64

21.72

0.77

Methods with full camera supervision

Dosovisky et al.

14.79

0.57

14.36

0.68

Appearance Flow

15.87

0.53

17.42

0.66

SynSin (w/ voxel)

20.62

0.70

19.77

0.75

SynSin (w/ point cloud)

20.91

0.72

21.94

0.81

28 of 34

Results

Input video

Trajectory Prediction

Camera pose estimation

29 of 34

Results

Camera pose estimation

29

Input video

Trajectory Prediction

30 of 34

Results

Camera pose estimation

30

Comparisons on 30-frame videos

 

 

 

31 of 34

31

Input image

Followed Video

Result

Camera Shaking

Video following

Results

32 of 34

32

Input image

Followed Video

Result

Rotating Right

Video following

Results

33 of 34

Future work

Towards training on completely unstructured datasets

33

Posed videos (RealEstate10K)

Video length: 3.85 days

Raw videos (YouTube)

Video length: ~2M years

34 of 34

Thanks!

34

Please see paper and website for details

https://zlai0.github.io/VideoAutoencoder