Blind Image Inpainting via Omni-dimensional Gated Attention and Wavelet Queries
Shruti S. Phutke, Ashutosh Kulkarni, Santosh Kumar Vipparthi, and Subrahmanyam Murala
Computer Vision and Pattern Recognition Lab
Indian Institute of Technology Ropar, Rupnagar, Punjab
Motivation
The contributions of the proposed work are:
Contributions
Proposed architecture for blind image inpainting.
Transformer Block
Transformer Block
Feed-forward Network
Downsample with stride 2
Upsample with stride 2
Subtraction
Multiplication
GELU Activation
Transformer Block
Transformer Block
Transformer Block
R
Reshape
1x1 Conv
Sigmoid Activation
Concat🡪1x1 Conv
Conv3
ODConv
Omni-dimensional Gated Attention (OGA)
Wavelet Query Multi-head Attention (WQMA)
R
R
R
K
QW
V
R
WCP
DC3
DC3
DC3
3x3 Depth-wise separable Conv
Wavelet Coefficient Processing
FFN
Transformer Block
FFN
WQMA
DWT
LL
LH
HL
HH
DC3
DC3
DC3
DC3
IDWT
ODConv
Omnidimensional Convolution
Transformer Block
OGA
OGA
OGA
Proposed Method
Proposed Method
Experimental Dataset
CelebA-HQ [1]
FFHQ [2]
[1] Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. In ICLR, 2018.
[2] Tero Karras, Samuli Laine, and Timo Aila.. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF CVPR, PP 4401–4410, 2019.
Experimental Dataset
Paris Street View [1]
Places2 [2]
[1] Carl Doersch, Saurabh Singh, Abhinav Gupta, Josef Sivic, and Alexei Efros. What makes paris look like paris? ACM Transactions on Graphics, 31(4), 2012.
[2] Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE TPAMI, 40(6):1452–1464, 2017
Network Configuration | PSNR↑ | SSIM↑ | L1↓ | FID↓ |
(a) TransCNNHAE [1] | 26.72 | 0.896 | 0.0352 | 41.50 |
| 26.89 | 0.885 | 0.0347 | 46.63 |
| 27.50 | 0.901 | 0.0324 | 43.11 |
| 27.05 | 0.898 | 0.0328 | 44.32 |
| 27.81 | 0.905 | 0.0301 | 40.64 |
Ablation study on different configurations of the proposed network on ParisSV dataset for blind image inpainting (Note: ↑- Higher is better, ↓- Lower is better).
[1]
[1] Xiefan Guo, Hongyu Yang, and Di Huang. Image inpainting via conditional texture and structure dual generation. IEEE/CVF International Conference on Computer Vision, pages 14134–14143, 2021.
Qualitative result analysis of ablation study on different configurations of the proposed network for blind image inpainting.
Ablation Study
[1] Yi Wang, Ying-Cong Chen, Xin Tao, and Jiaya Jia. Vcnet: A robust approach to blind image inpainting. In European Conference on Computer Vision, pages 752–768. Springer,2020.
[2] Haoru Zhao, Zhaorui Gu, Bing Zheng, and Haiyong Zheng. Transcnn-hae: Transformer-cnn hybrid autoencoder for blind image inpainting ACM International Conference on Multimedia, pages 6813–6821,2022.
[3] Xiefan Guo, Hongyu Yang, and Di Huang. Image inpainting via conditional texture and structure dual generation. IEEE/CVF International Conference on Computer Vision, pages 14134–14143, 2021.
Comparison of the proposed method (ours) and existing state-of-the-art methods for blind image inpainting.
Result Analysis
Metric | Dataset | VCNet [1] | CTSDG [2] | TransCNNHAE [3] | Ours |
PSNR | CelebA-HQ | 25.59 | 26.94 | 27.71 | 28.21 |
FFHQ | 23.62 | 24.62 | 27.05 | 28.19 | |
ParisSV | 23.62 | 26.08 | 26.72 | 27.81 | |
Places2 | 24.09 | 26.05 | 26.87 | 27.55 | |
SSIM | CelebA-HQ | 0.874 | 0.934 | 0.949 | 0.951 |
FFHQ | 0.861 | 0.935 | 0.941 | 0.952 | |
ParisSV | 0.824 | 0.861 | 0.869 | 0.905 | |
Places2 | 0.869 | 0.905 | 0.910 | 0.918 | |
| CelebA-HQ | 0.0396 | 0.0318 | 0.0250 | 0.0221 |
FFHQ | 0.0482 | 0.0392 | 0.0281 | 0.0234 | |
ParisSV | 0.0527 | 0.0412 | 0.0352 | 0.0301 | |
Places2 | 0.0429 | 0.0308 | 0.0261 | 0.0231 | |
FID | CelebA-HQ | 9.275 | 8.561 | 7.251 | 7.235 |
FFHQ | 10.148 | 9.586 | 9.424 | 8.639 | |
ParisSV | 64.215 | 43.015 | 41.505 | 40.646 | |
Places2 | 28.821 | 18.685 | 17.640 | 17.521 |
Qualitative results comparison of the proposed method (Ours) with existing state-of-the-art methods (VCNet [1], CTSDG [2], TransCNNHAE [3]) on Celeb (first two rows) and FFHQ (last two rows) dataset for blind image inpainting.
Qualitative results comparison of the proposed method (Ours) with existing state-of-the-art methods (VCNet [1], CTSDG [3], TransCNNHAE [2]) on ParisSV (first two rows) and Places2 (last two rows) dataset for blind image inpainting.
Qualitative results comparison of the proposed method (Ours) with existing state-of-the-art method (TransCNNHAE [2]) on unseen patterns.
[1] Yi Wang, Ying-Cong Chen, Xin Tao, and Jiaya Jia. Vcnet: A robust approach to blind image inpainting. In European Conference on Computer Vision, pages 752–768. Springer,2020.
[2] Xiefan Guo, Hongyu Yang, and Di Huang. Image inpainting via conditional texture and structure dual generation. IEEE/CVF International Conference on Computer Vision, pages 14134–14143, 2021.
[3] Haoru Zhao, Zhaorui Gu, Bing Zheng, and Haiyong Zheng. Transcnn-hae: Transformer-cnn hybrid autoencoder for blind image inpainting ACM International Conference on Multimedia, pages 6813–6821,2022.
Result Analysis
Summary
Thank You
If you are interested in our work, please visit the repository at: