1 of 31

1

Amanda Adkins

April 9, 2021

Unsupervised Monocular Depth Learning in Dynamic Scenes

2 of 31

Motivation

  • Understanding 3D geometry is important for robots and autonomous vehicles
  • Using monocular cameras is desirable due to cost

2

3 of 31

Monocular Depth Estimation

  • Depth and motion from monocular images is an ill-posed problem
  • Existing works leverage prior knowledge
    • Deep networks can learn these priors
    • Some works use other sensors, such as LiDAR, to provide ground truth
    • Self-supervision using monocular video avoids cost of labeled data

3

4 of 31

Predicting Depth in Dynamic Scenes

  • Traditional Structure from Motion approaches and deep networks struggle when there are moving objects in the scene

4

5 of 31

Predicting Depth in Dynamic Scenes

  • Existing approaches leverage prior knowledge to deal with moving objects
    • Optical flow prediction used as an intermediate step; object motion detected from optical flow
    • Stereo data
    • Semantic segmentation as a prior for what could move
    • Detect when object is at same pose relative to camera and mark it as moving

5

6 of 31

Goal and Assumptions

  • Learn depth, ego-motion, and per-pixel object motion from monocular video only
    • No semantics, stereo, or ground truth
  • Models object motion feasible by translation only
    • Does not model rotating objects

6

7 of 31

Key Insights

  • Object motion should be sparse
    • Most pixels should belong to the background
  • Per-pixel object translation estimates should be consistent throughout a rigid moving object

7

8 of 31

High Level Approach

  • Depth, ego-motion, object motion learned simultaneously from monocular videos using self-supervision
  • Use pairs of adjacent video frames as training data
  • Predict depth map D(u, v) for each of the images in the pair
  • Add depth as another channel to the images and feed into motion prediction network
  • Motion prediction network predicts 3D translation map at original resolution and 6D ego-motion vector
  • TODO add equations

8

9 of 31

High Level Approach

9

10 of 31

High Level Approach

10

11 of 31

High Level Approach

11

12 of 31

High Level Approach

12

13 of 31

High Level Approach

  • At inference time
    • For a single frame, get depth per pixel
    • For pairs of frames, get object translation field (per pixel) and camera motion

13

14 of 31

Training Losses

  • Total loss is sum of:
    • Depth regularization loss
    • Motion regularization loss
    • Consistency regularization loss
  • Depth regularization applied to each image
  • Motion and consistency regularization are applied twice to the pair of images, using both possible orderings for the pair of images

14

15 of 31

Depth Regularization

  • Encourages depth to be consistent for neighboring pixels

  • Loss is weaker for pixels with large color variation

15

16 of 31

Motion Regularization

  • Group smoothness loss minimizes changes within moving areas

16

17 of 31

Motion Regularization

  • Sparsity loss encourages pixels to have same translation as background

17

18 of 31

Motion Regularization

  • Composed of group smoothness loss and sparsity loss

18

19 of 31

Consistency Regularization

  • Consistency loss is composed of the motion cycle consistency loss and the photometric consistency loss
  • Motion cycle consistency loss encourages motion from frame 1 to frame 2 to be the inverse of motion from frame 2 to frame 1

19

20 of 31

Consistency Regularization

  • Occlusion aware photometric consistency loss encourages photometric consistency of corresponding areas of the two frames
  • Composed of L1 loss on intensity and structural similarity loss

  • A mask is applied to mitigate the effect of occlusions

20

21 of 31

Experiments

  • Test depth prediction on 4 datasets
    • Cityscapes - highly dynamic urban driving
    • KITTI - urban environment, some dynamic elements
    • Waymo Open Dataset - autonomous driving dataset incorporating dynamics, nighttime driving, diverse weather conditions
    • YouTube dataset - randomly selected videos from Youtube taken with handheld cameras while walking
  • Achieve depth inference at ~5.3 ms per frame

21

22 of 31

Experiments

  • Training Set Sizes
    • Cityscapes - 22,973 image pairs
    • KITTI - 22,600 image pairs
    • Waymo Open Dataset - 100,000 image pairs (1000 videos)

22

23 of 31

Results - CityScapes

23

24 of 31

Results - Cityscapes Ablation

24

25 of 31

Results - Kitti

25

26 of 31

Results - Waymo

26

27 of 31

Results - Youtube

27

28 of 31

Takeaways

  • Unsupervised depth learning in dynamic scenes that outperforms prior approaches
  • Doesn’t require semantic cues, though incorporating a mask does yield modest improvements

28

29 of 31

Limitations / Results Shortcomings

  • Don’t report metrics on motion prediction
  • Only report metrics on autonomous driving datasets
  • Don’t account for objects that are deforming or rotating in their model

29

30 of 31

Discussion

  • Should networks like this try to leverage prior knowledge (such as semantics) or learn independently?

  • How can traditional model-based approaches (ex. feature matching) be incorporated into deep learned networks to improve performance?

  • Is fully self-supervised the best way to go? Opportunities for including (sparse) ground truth?

  • Other thoughts?

30

31 of 31

Other Resources

31