Copy-and-Paste Networks for�Deep Video Inpainting
연세대학교 Computational Intelligence & Photography Lab
1
May 23, 2019
Computational Intelligence & Photography Lab, Yonsei University
May 27, 2019 / NAVER Tech Talk
Sungho Lee
Speaker Biography
연세대학교 Computational Intelligence & Photography Lab
2
May 23, 2019
Sungho Lee
Contents
연세대학교 Computational Intelligence & Photography Lab
3
May 23, 2019
1. Introduction - Video inpainting?
연세대학교 Computational Intelligence & Photography Lab
4
May 23, 2019
1. Introduction - Video Inpainting?
연세대학교 Computational Intelligence & Photography Lab
5
May 23, 2019
Input
Segmentation Mask
Output
+
=
1. Introduction - Video Inpainting?
연세대학교 Computational Intelligence & Photography Lab
6
May 23, 2019
Li, Wei, et al. "AADS: Augmented autonomous driving simulation using data-driven algorithms.“, arXiv 2019
1. Introduction - Image inpainting VS Video inpainting
연세대학교 Computational Intelligence & Photography Lab
7
May 23, 2019
GT
Output
Input
Single image inpainting
Input
Output
GT
1. Introduction - Image inpainting VS Video inpainting
연세대학교 Computational Intelligence & Photography Lab
8
May 23, 2019
Video inpainting
Input
Output
Temporal Consistency
Input
Output
1. Introduction - Related works
연세대학교 Computational Intelligence & Photography Lab
9
May 23, 2019
M.Granados. et al, “Background Inpainting for Videos with Dynamic Objects and a Free-moving Camera”, ECCV 2012
Jia-Bin Huang. et al, “Temporally Coherent Completion of Dynamic Video”, SIGGRAPH ASIA 2016
1. Introduction - Related works
연세대학교 Computational Intelligence & Photography Lab
10
May 23, 2019
C.Wang. et al, “Video Inpainting by Jointly Learning Temporal Structure and Spatial Details”, AAAI 2019
D. Kim. et al, “Deep Video Inpainting”, CVPR 2019
2. Proposed methods – Overview
연세대학교 Computational Intelligence & Photography Lab
11
May 23, 2019
…
Target frame
Reference frames
Aligned Reference frames
…
2. Copy
1. Alignment
…
4. Update
3. Paste
Input
Video
2. Proposed methods - Overview
연세대학교 Computational Intelligence & Photography Lab
12
May 23, 2019
2. Proposed methods – Alignment Network
연세대학교 Computational Intelligence & Photography Lab
13
May 23, 2019
2. Proposed methods – Alignment Network
연세대학교 Computational Intelligence & Photography Lab
14
May 23, 2019
Target
frame
Reference
frames
Aligned
Reference
frames
…
…
2. Proposed methods – Copy-and-Paste Network
연세대학교 Computational Intelligence & Photography Lab
15
May 23, 2019
2. Proposed methods – Context Matching Module
연세대학교 Computational Intelligence & Photography Lab
16
May 23, 2019
2. Proposed methods – Masked Softmax example
연세대학교 Computational Intelligence & Photography Lab
17
May 23, 2019
2. Proposed methods – Context Matching Module
연세대학교 Computational Intelligence & Photography Lab
18
May 23, 2019
2. Proposed methods – Copy-and-paste network
연세대학교 Computational Intelligence & Photography Lab
19
May 23, 2019
2. Proposed methods – Reference update
연세대학교 Computational Intelligence & Photography Lab
20
May 23, 2019
2. Proposed methods – Summary
연세대학교 Computational Intelligence & Photography Lab
21
May 23, 2019
Target
frame
Reference
frame
Aligned
reference
frame
Overlap
…
Target
frame
Aligned
reference
frame
Restored output
2. Proposed methods – Summary
연세대학교 Computational Intelligence & Photography Lab
22
May 23, 2019
As a result,
Input + Segmentation Mask
Output
3. Experiments – Training dataset
- Places dataset(image, 1.8M) + Youtube dataset(Video, 7.3K)
- MIT Saliency Benchmark(mask image, 11K), Pascal VOC 2012 (mask image, 14.3K)
연세대학교 Computational Intelligence & Photography Lab
23
May 23, 2019
3. Experiments – DAVIS video dataset
연세대학교 Computational Intelligence & Photography Lab
24
May 23, 2019
3. Experiments – Video Object Removal
연세대학교 Computational Intelligence & Photography Lab
25
May 23, 2019
3. Experiments – Video Object Removal
연세대학교 Computational Intelligence & Photography Lab
26
May 23, 2019
Input Video
VINet
Huang et al.
Ours
3. Experiments – Video Object Removal
연세대학교 Computational Intelligence & Photography Lab
27
May 23, 2019
Input Video
VINet
Huang et al.
Ours
3. Experiments – Video Object Removal
연세대학교 Computational Intelligence & Photography Lab
28
May 23, 2019
Input Video
VINet
Huang et al.
Ours
3. Experiments – Video Object Removal
연세대학교 Computational Intelligence & Photography Lab
29
May 23, 2019
Input Video
VINet
Huang et al.
Ours
3. Experiments – Video Object Removal
연세대학교 Computational Intelligence & Photography Lab
30
May 23, 2019
| Average execution time |
Huang et al. | 952s |
Ours | 27.14s |
3. Experiments – Video Restoration
연세대학교 Computational Intelligence & Photography Lab
31
May 23, 2019
| PSNR | SSIM |
Huang et al. | 28.14 | 0.859 |
Ours | 28.37 | 0.851 |
Background video(GT)
Foreground Mask
Synthesized video
+
=
*note: VINet is excluded in this experiment because official code has not published yet.
4. Analysis – Masked softmax (ablation study)
연세대학교 Computational Intelligence & Photography Lab
32
May 23, 2019
Result using softmax
Result using masked softmax
Input
4. Analysis – Reference update (ablation study)
연세대학교 Computational Intelligence & Photography Lab
33
May 23, 2019
Without update
With update
Input
4. Analysis - Limitation
연세대학교 Computational Intelligence & Photography Lab
34
May 23, 2019
Input Video
VINet
Huang et al.
Ours
5. Application
연세대학교 Computational Intelligence & Photography Lab
35
May 23, 2019
5. Application - Limitation
연세대학교 Computational Intelligence & Photography Lab
36
May 23, 2019
Enhanced Brightness
Dark Image
Enhanced Brightness
Oversaturated Image
5. Application - Idea
연세대학교 Computational Intelligence & Photography Lab
37
May 23, 2019
- Under/over-exposed parts gradually become adaptable to the next frame.
- Each near frame has enough redundant information in different form.
Frame 1
Frame 2
Frame 3
Frame 4
Propagation
5. Application – Training dataset
연세대학교 Computational Intelligence & Photography Lab
38
May 23, 2019
Over-exposed frames example
Under-exposed frames example
Original
Synthesized
frames
5. Application – approach
연세대학교 Computational Intelligence & Photography Lab
39
May 23, 2019
Target frame / Mask
Reference frame / Mask 1,2 …
Our model
Enhanced output
…
5. Application – Results
연세대학교 Computational Intelligence & Photography Lab
40
May 23, 2019
| Lane detection accuracy |
Over/under-exposed image | 46.69% |
Restored input by our model | 83.00% |
Under-exposed frame example
Over-exposed frame example
6. Conclusion & Future works
연세대학교 Computational Intelligence & Photography Lab
41
May 23, 2019
연세대학교 Computational Intelligence & Photography Lab
42
May 23, 2019
Q&A
Thank you for your attention.�
연세대학교 Computational Intelligence & Photography Lab
43
May 23, 2019
If you have a question, please contact to me
2. Methods – Alignment network
연세대학교 Computational Intelligence & Photography Lab
44
May 23, 2019
2. Methods – Alignment network
연세대학교 Computational Intelligence & Photography Lab
45
May 23, 2019
2. Methods – Loss function
연세대학교 Computational Intelligence & Photography Lab
46
May 23, 2019
2. Methods – Copy-and-Paste networks
연세대학교 Computational Intelligence & Photography Lab
47
May 23, 2019
2. Methods – Copy-and-paste networks
연세대학교 Computational Intelligence & Photography Lab
48
May 23, 2019
2. Methods – Loss function
연세대학교 Computational Intelligence & Photography Lab
49
May 23, 2019
4. Analysis – Reference update
연세대학교 Computational Intelligence & Photography Lab
50
May 23, 2019
Input
…
…
…
good
good
Time
4. Analysis – Reference update
연세대학교 Computational Intelligence & Photography Lab
51
May 23, 2019
(c) Without update
(d) With update
(a) sample frame
(b) Input + Mask
4. Analysis – Masked softmax
연세대학교 Computational Intelligence & Photography Lab
52
May 23, 2019
Input frame
Time
(a) Result using softmax
(b) Result using masked softmax