1 of 52

Copy-and-Paste Networks for�Deep Video Inpainting

연세대학교 Computational Intelligence & Photography Lab

1

May 23, 2019

Computational Intelligence & Photography Lab, Yonsei University

May 27, 2019 / NAVER Tech Talk

2 of 52

Speaker Biography

연세대학교 Computational Intelligence & Photography Lab

2

May 23, 2019

  • Research Interest
    • Image / video enhancement based on deep learning
    • Video Inpainting
    • Computer Vision / Machine Learning

  • Education
    • M.S. student, CS, Yonsei University (Advisor: Seon Joo Kim)
    • B.S. in CS, Yonsei University

  • Project Experience
    • Video enhancement / Inpainting @ Hyundai mnsoft
    • Partial object classification @ Ministry of Trade, Industry and Energy

3 of 52

Contents 

  1. Introduction to Video Inpainting
    • Video Inpainting? / Single Image Inpainting VS Video Inpainting
  2. Proposed Methods
  3. Experiments: Video Object Removal / Video Restoration
  4. Analysis
  5. Applications: Under/over-exposed video enhancement
  6. Conclusion & Future works

연세대학교 Computational Intelligence & Photography Lab

3

May 23, 2019

4 of 52

1. Introduction - Video inpainting?

  • Video inpainting is the process of completing corrupted or missing regions in videos.

연세대학교 Computational Intelligence & Photography Lab

4

May 23, 2019

5 of 52

1. Introduction - Video Inpainting?

  • Video inpainting for video editing

연세대학교 Computational Intelligence & Photography Lab

5

May 23, 2019

Input

Segmentation Mask

Output

+

=

6 of 52

1. Introduction - Video Inpainting?

  • Video inpainting for autonomous driving simulation

연세대학교 Computational Intelligence & Photography Lab

6

May 23, 2019

Li, Wei, et al. "AADS: Augmented autonomous driving simulation using data-driven algorithms.“, arXiv 2019

7 of 52

1. Introduction - Image inpainting VS Video inpainting

연세대학교 Computational Intelligence & Photography Lab

7

May 23, 2019

GT

Output

Input

Single image inpainting

Input

Output

GT

8 of 52

1. Introduction - Image inpainting VS Video inpainting

연세대학교 Computational Intelligence & Photography Lab

8

May 23, 2019

Video inpainting

Input

Output

Temporal Consistency

Input

Output

9 of 52

1. Introduction - Related works

  • Optimization based approach : Long execution time(> 90 frames / 15min)
    • Alignment based on the homography

    • Optical flow optimization method(state-of-the-art)

연세대학교 Computational Intelligence & Photography Lab

9

May 23, 2019

M.Granados. et al, “Background Inpainting for Videos with Dynamic Objects and a Free-moving Camera”, ECCV 2012

Jia-Bin Huang. et al, “Temporally Coherent Completion of Dynamic Video”, SIGGRAPH ASIA 2016

10 of 52

1. Introduction - Related works

  • Deep learning based approach
    • 3D-2D encoder-decoder based deep video inpainting

    • Optical flow based deep video inpainting (state-of-the-art)

연세대학교 Computational Intelligence & Photography Lab

10

May 23, 2019

C.Wang. et al, “Video Inpainting by Jointly Learning Temporal Structure and Spatial Details”, AAAI 2019

D. Kim. et al, “Deep Video Inpainting”, CVPR 2019

11 of 52

2. Proposed methods – Overview

연세대학교 Computational Intelligence & Photography Lab

11

May 23, 2019

Target frame

Reference frames

Aligned Reference frames

2. Copy

1. Alignment

4. Update

3. Paste

Input

Video

12 of 52

2. Proposed methods - Overview

연세대학교 Computational Intelligence & Photography Lab

12

May 23, 2019

13 of 52

2. Proposed methods – Alignment Network

연세대학교 Computational Intelligence & Photography Lab

13

May 23, 2019

14 of 52

2. Proposed methods – Alignment Network

  • There are some frames that cannot be aligned through affine transformation.
  • Therefore, it is necessary to combine reference frames based on context matching with the target frame.

연세대학교 Computational Intelligence & Photography Lab

14

May 23, 2019

Target

frame

Reference

frames

Aligned

Reference

frames

15 of 52

2. Proposed methods – Copy-and-Paste Network

연세대학교 Computational Intelligence & Photography Lab

15

May 23, 2019

16 of 52

2. Proposed methods – Context Matching Module

연세대학교 Computational Intelligence & Photography Lab

16

May 23, 2019

17 of 52

2. Proposed methods – Masked Softmax example

연세대학교 Computational Intelligence & Photography Lab

17

May 23, 2019

18 of 52

2. Proposed methods – Context Matching Module

연세대학교 Computational Intelligence & Photography Lab

18

May 23, 2019

19 of 52

2. Proposed methods – Copy-and-paste network

연세대학교 Computational Intelligence & Photography Lab

19

May 23, 2019

20 of 52

2. Proposed methods – Reference update

연세대학교 Computational Intelligence & Photography Lab

20

May 23, 2019

21 of 52

2. Proposed methods – Summary

연세대학교 Computational Intelligence & Photography Lab

21

May 23, 2019

  • Self-supervised deep alignment networks

  • Copy-and-Paste Networks with context matching module

Target

frame

Reference

frame

Aligned

reference

frame

Overlap

Target

frame

Aligned

reference

frame

Restored output

22 of 52

2. Proposed methods – Summary

연세대학교 Computational Intelligence & Photography Lab

22

May 23, 2019

As a result,

Input + Segmentation Mask

Output

23 of 52

3. Experiments – Training dataset

  • We synthesized the videos by compositing background image sequences with object masks.
  • Background image sequences

- Places dataset(image, 1.8M) + Youtube dataset(Video, 7.3K)

  • Foreground mask

- MIT Saliency Benchmark(mask image, 11K), Pascal VOC 2012 (mask image, 14.3K)

연세대학교 Computational Intelligence & Photography Lab

23

May 23, 2019

24 of 52

3. Experiments – DAVIS video dataset

  • Video object segmentation dataset.
  • 60 videos (70.3 frames per videos)

연세대학교 Computational Intelligence & Photography Lab

24

May 23, 2019

25 of 52

3. Experiments – Video Object Removal

  • Test dataset : Randomly selected 30 video sequences in DAVIS dataset(240p)
  • Goal: Video object removal task
  • Evaluation: User study(Amazon Mechanical Turk, rank the video completion results)

연세대학교 Computational Intelligence & Photography Lab

25

May 23, 2019

26 of 52

3. Experiments – Video Object Removal

연세대학교 Computational Intelligence & Photography Lab

26

May 23, 2019

Input Video

VINet

Huang et al.

Ours

27 of 52

3. Experiments – Video Object Removal

연세대학교 Computational Intelligence & Photography Lab

27

May 23, 2019

Input Video

VINet

Huang et al.

Ours

28 of 52

3. Experiments – Video Object Removal

연세대학교 Computational Intelligence & Photography Lab

28

May 23, 2019

Input Video

VINet

Huang et al.

Ours

29 of 52

3. Experiments – Video Object Removal

연세대학교 Computational Intelligence & Photography Lab

29

May 23, 2019

Input Video

VINet

Huang et al.

Ours

30 of 52

3. Experiments – Video Object Removal

  • User study results (Lower is better)

연세대학교 Computational Intelligence & Photography Lab

30

May 23, 2019

Average execution time

Huang et al.

952s

Ours

27.14s

31 of 52

3. Experiments – Video Restoration

  • Test dataset : 25 video synthesized by compositing randomly selected background video and mask sequences in DAVIS dataset

  • Goal : Video restoration to synthesized mask region
  • Evaluation : PSNR and SSIM measures
  • Results

연세대학교 Computational Intelligence & Photography Lab

31

May 23, 2019

PSNR

SSIM

Huang et al.

28.14

0.859

Ours

28.37

0.851

Background video(GT)

Foreground Mask

Synthesized video

+

=

*note: VINet is excluded in this experiment because official code has not published yet.

32 of 52

4. Analysis – Masked softmax (ablation study)

연세대학교 Computational Intelligence & Photography Lab

32

May 23, 2019

Result using softmax

Result using masked softmax

Input

33 of 52

4. Analysis – Reference update (ablation study)

연세대학교 Computational Intelligence & Photography Lab

33

May 23, 2019

Without update

With update

Input

34 of 52

4. Analysis - Limitation

  • We observed the synthesized blurry results when there is a large invisible region in a video.

연세대학교 Computational Intelligence & Photography Lab

34

May 23, 2019

Input Video

VINet

Huang et al.

Ours

35 of 52

5. Application

  • We extend our method for restoring under/over-exposed image sequences.
  • Goal: Under / over-exposed video restoration on road scene.

연세대학교 Computational Intelligence & Photography Lab

35

May 23, 2019

36 of 52

5. Application - Limitation

연세대학교 Computational Intelligence & Photography Lab

36

May 23, 2019

Enhanced Brightness

Dark Image

Enhanced Brightness

Oversaturated Image

  • Limitation – There is no information about saturated parts.

37 of 52

5. Application - Idea

연세대학교 Computational Intelligence & Photography Lab

37

May 23, 2019

  • Ideas: Multi-frame based approach

- Under/over-exposed parts gradually become adaptable to the next frame.

- Each near frame has enough redundant information in different form.

Frame 1

Frame 2

Frame 3

Frame 4

Propagation

38 of 52

5. Application – Training dataset

  • Training dataset : Simulate the rapid exposure change situation from 20K frames of normal videos.

연세대학교 Computational Intelligence & Photography Lab

38

May 23, 2019

Over-exposed frames example

Under-exposed frames example

Original

Synthesized

frames

39 of 52

5. Application – approach

연세대학교 Computational Intelligence & Photography Lab

39

May 23, 2019

Target frame / Mask

Reference frame / Mask 1,2 …

Our model

Enhanced output

40 of 52

5. Application – Results

  • Test dataset : 469 frames videos that contains rapid exposure changes due to tunnels in/out.
  • Evaluation: color histogram-based lane detection accuracy measure.
  • Results

연세대학교 Computational Intelligence & Photography Lab

40

May 23, 2019

Lane detection accuracy

Over/under-exposed image

46.69%

Restored input by our model

83.00%

Under-exposed frame example

Over-exposed frame example

41 of 52

6. Conclusion & Future works

  • The proposed method inpaints the missing information by copy-and-pasting contents from the reference frames.
  • The reference update process ensure the temporal consistency.
  • We extended our framework to restore over/under-exposed videos and were able to significantly increase the lane detection accuracy.
  • Limitation
    • Blurry artifacts when there is a large invisible region in a video.
  • Future works
    • Improve the results of invisible region based on generative approach.

연세대학교 Computational Intelligence & Photography Lab

41

May 23, 2019

42 of 52

연세대학교 Computational Intelligence & Photography Lab

42

May 23, 2019

Q&A

43 of 52

Thank you for your attention.�

연세대학교 Computational Intelligence & Photography Lab

43

May 23, 2019

If you have a question, please contact to me

l.sh@yonsei.ac.kr

44 of 52

2. Methods – Alignment network

연세대학교 Computational Intelligence & Photography Lab

44

May 23, 2019

45 of 52

2. Methods – Alignment network

연세대학교 Computational Intelligence & Photography Lab

45

May 23, 2019

46 of 52

2. Methods – Loss function

연세대학교 Computational Intelligence & Photography Lab

46

May 23, 2019

 

 

 

 

 

 

 

47 of 52

2. Methods – Copy-and-Paste networks

연세대학교 Computational Intelligence & Photography Lab

47

May 23, 2019

48 of 52

2. Methods – Copy-and-paste networks

연세대학교 Computational Intelligence & Photography Lab

48

May 23, 2019

49 of 52

2. Methods – Loss function

연세대학교 Computational Intelligence & Photography Lab

49

May 23, 2019

 

 

 

 

 

 

 

50 of 52

4. Analysis – Reference update

연세대학교 Computational Intelligence & Photography Lab

50

May 23, 2019

 

 

Input

 

 

good

good

Time

51 of 52

4. Analysis – Reference update

연세대학교 Computational Intelligence & Photography Lab

51

May 23, 2019

(c) Without update

(d) With update

(a) sample frame

(b) Input + Mask

52 of 52

4. Analysis – Masked softmax

연세대학교 Computational Intelligence & Photography Lab

52

May 23, 2019

Input frame

Time

(a) Result using softmax

(b) Result using masked softmax