Video Autoencoder: self-supervised disentanglement
of static 3D structure and motion
ICCV 2021
1
Sifei Liu
NVIDIA
Alexei A. Efros
UC Berkeley
Xiaolong Wang
UC San Diego
Zihang Lai
CMU
Disentangle the visual world
2
Movement
Depth
Structure and viewpoint
Prior work
From the very beginnings of computer vision, …
3
Barrow and Tenenbaum, Comput. Vis. Syst., 1978
Tomasi and Kanade, IJCV, 1992
Prior work
Learning disentangled visual representation from auto-encoders
4
Park et al., NeurIPS, 2020
Liu et al., ECCV, 2020
Kulkarni et al., NIPS, 2015
Kim et al., ICML, 2018
Prior work
5
Tung et al., CVPR, 2019
Nguyen-Phuoc et al., ICCV, 2019
Wiles et al., CVPR, 2020
Learning 3D from 2D supervisions
Objective
In this work, we learn to separate 3D structure from Camera Motion without any human annotations
6
Input raw video
ENCODER
Test time
The features obtained (3D structure, Camera Motion) can be used for several downstream tasks:
7
Novel view synthesis
Pose estimation in video
Video following
7
Single image
Raw video
Single image + Raw video
*Actual results
Test time
Novel view synthesis
8
ENCODER
Pre-defined
trajectory
Single still image
DECODER
Novel view synthesis
Learning from temporal continuity of videos
9
Spatio-temporal continuity
No spatio-temporal continuity
Learning from temporal continuity of videos
Assume that a local snippet of video is capturing a static scene
10
Video snippet
Frame 1
Frame 2
Frame 3
Frame 4
Frame 5
Static scene
Learning from temporal continuity of videos
Assume that a local snippet of video is capturing a static scene
11
Video snippet
Frame 1
Frame 2
Frame 3
Frame 4
Frame 5
3D scene structure
Trajectory
Encode
[ ]
R1, t1
R2, t2
…
Learning from temporal continuity of videos
Assume that a local snippet of video is capturing a static scene
12
Video snippet
Frame 1
Frame 2
Frame 3
Frame 4
Frame 5
Decode
Model architecture
13
ENCODER
3D Structure
(Deep voxels)
3D Trajectory
Structure + Motion
3D
Transform
Video
DECODER
Reconstruction
3D Encoder
Traj. Encoder
Decoder
Consistency
loss
Camera
transformation
Joint
optimisation
Model architecture
14
3D Encoder
1
Traj. Encoder
2
3
Decoder
4
ResNet
50
Reshape
3D conv
Input image
3D deep voxels
Model architecture
15
conv
Stacked image pair
relu
conv
relu
conv
relu
…
6D Pose
3D Encoder
1
Traj. Encoder
2
3
Decoder
4
Model architecture
16
3D Encoder
1
Traj. Encoder
2
3
Decoder
4
3D conv
Reshape
2D conv
Output image
3D deep voxels
3D Transform
Camera pose
Training loss
17
Video Autoencoder
18
Results
Datasets
19
RealEstate10K
Matterport3D
Replica
Results
Novel view synthesis
20
Single input image
Output video
RealEstate10K dataset
21
A Japanese living room (out-of-distribution)
Single input image
Output video
Novel view synthesis (Out-of-domain results)
Results
22
Novel view synthesis (Out-of-domain results)
Results
Single input image
Output video
“Spirited Away”
23
Bedroom in Arles, Vincent van Gogh
Single input image
Output video
Novel view synthesis (Out-of-domain results)
Results
Comparison with previous methods
24
Input image
RealEstate10K dataset
Mustikovela et al.
Tung et al.
Yu et al.
Wiles et al.
Ours
Ground truth
More artefacts
Blank area
Results
Tung et al.
Mustikovela et al.
Yu et al.
Wiles et al.
25
Input image
RealEstate10K dataset
Comparison with previous methods
Results
Mustikovela et al.
Tung et al.
Yu et al.
Wiles et al.
Ours
Ground truth
More artefacts
Blank area
Results
Novel view synthesis (RealEstate10K)
26
Results
Novel view synthesis (Matterport3D & Replica)
27
| | Matterport 3D | MP3D → Replica | ||
Method | Pose | PSNR ↑ | SSIM ↑ | PSNR ↑ | SSIM ↑ |
Methods without any camera supervision | |||||
Ours | × | 20.58 | 0.64 | 21.72 | 0.77 |
Methods with full camera supervision | |||||
Dosovisky et al. | √ | 14.79 | 0.57 | 14.36 | 0.68 |
Appearance Flow | √ | 15.87 | 0.53 | 17.42 | 0.66 |
SynSin (w/ voxel) | √ | 20.62 | 0.70 | 19.77 | 0.75 |
SynSin (w/ point cloud) | √ | 20.91 | 0.72 | 21.94 | 0.81 |
Results
Input video
Trajectory Prediction
Camera pose estimation
Results
Camera pose estimation
29
Input video
Trajectory Prediction
Results
Camera pose estimation
30
Comparisons on 30-frame videos
31
Input image
Followed Video
Result
Camera Shaking
Video following
Results
32
Input image
Followed Video
Result
Rotating Right
Video following
Results
Future work
Towards training on completely unstructured datasets
33
Posed videos (RealEstate10K)
Video length: 3.85 days
Raw videos (YouTube)
Video length: ~2M years
Thanks!
34
Please see paper and website for details
https://zlai0.github.io/VideoAutoencoder