Progress Presentation�on
Development of Deep Learning Models for Image Super-Resolution
Research & Research Progress Colloquium - II
Presented By:-
Neeraj Baghel
Roll No:-RSI2021003
03-06-2023
May, 2023
Computer Vision and Biometrics Lab (CVBL)
Department of Information Technology (IT)
Indian Institute of Information Technology Allahabad (IIIT-A)
Devghat, Jhalwa, Prayagraj-211015, U. P. INDIA
Supervisor:-
Dr. Shiv Ram Dubey
(Assistant Professor)
Solution?
Image Super-Resolution
2
Recognition in Low-Dimensional
Storage, Transmission High-Resolution
2. Motivation
Motivation:
3
Crop and Zoom
High-Resolution image
Low-Resolution Image
Applications:
Low-Resolution
High-Resolution
3. Problem Statement
4
Development of Deep Learning Models for Image Super-Resolution
Objective
Develop model for improving the quality of low-resolution images in various applications
4. Related Work
Interpolation Techniques
5
[24] Phillip Isola, “Image-to-image translation with conditional adversarial networks”, CVPR 2017.
[25] Ronneberger, Olaf, Philipp Fischer, and Thomas Brox. "U-net: Convolutional networks for biomedical image segmentation." Medical Image Computing and Computer-Assisted Intervention–MICCAI, 2015.
Fig: Training a conditional GAN to map edges→photo [24]
4. Related Work
Image-to-image translation with conditional adversarial networks [24]
6
Fig: Training a conditional GAN to map edges→photo [24]
[24] Phillip Isola, “Image-to-image translation with conditional adversarial networks”, CVPR 2017.
[25] Ronneberger, Olaf, Philipp Fischer, and Thomas Brox. "U-net: Convolutional networks for biomedical image segmentation." Medical Image Computing and Computer-Assisted Intervention–MICCAI, 2015.
Vision Transformer (ViT) [26]
4. Related Work
7
Method 2: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (VIT) [1]
[26] Dosovitskiy, Alexey, et al. "An image is worth 16x16 words: Transformers for image recognition at scale." arXiv (2020).
ViTGAN: Training gans with vision transformers [27]
4. Related Work
8
[27] Lee, K., Chang, H., Jiang, L., Zhang, H., Tu, Z., Liu, C.: Vitgan: Training gans with vision transformers. arXiv (2021).
Transgan: Two pure transformers can make one strong gan, and that can scale up [28]
4. Related Work
9
Grid Transformer Block
Fig: TransGAN Framework [28]
[28] Jiang, "Transgan: Two pure transformers can make one strong gan, and that can scale up." NIPS (2021).
TTSR: Learning Texture Transformer Network for Image Super-Resolution [29]
4. Related Work
10
Author proposed
1) A Texture Transformer,
2) Cross-scale feature integration module (CSFI)
Texture Transformer
Learnable Texture Extractor:
Texture extraction for reference images.
Whose parameters will be updated during end-to-end training.
Relevance Embedding:
Unfold Q and K into n patches.
Calculate relevance b/w them.
Then transfer from most relevant position in V.
[29] Yang, Fuzhi, et al. "Learning texture transformer network for image super-resolution." CVPR (2020).
4. Related Work
11
ESRT: Efficient Transformer for Single Image Super-Resolution [8]
4. Related Work
12
The overall architecture of the proposed Efficient SR Transformer
[8] Zhisheng Lu and Hong Liu and Juncheng Li and Linlin Zhang, “Efficient Transformer for Single Image Super-Resolution”, arXiv preprint arXiv: 2108.11084 (2021).
The pre and post-processing for the Efficient Transformer (ET).
The schematic diagram of the propose HFM
4. Related Work
13
State of Art
Zhisheng Lu and Hong Liu and Juncheng Li and Linlin Zhang, “Efficient Transformer for Single Image Super-Resolution”, arXiv preprint arXiv: 2108.11084 (2021).
4. Related Work
14
Fig: The diagram of the proposed image processing transformer (IPT)
[9] Chen, Hanting, et al. "Pre-trained image processing transformer." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021.
IPT: Pre-Trained Image Processing Transformer [9]
4. Related Work
15
Fig: Pipeline Overview. MaskGIT follows a two-stage design, with 1) a tokenizer that tokenizes images into visual tokens, and 2) a bidirectional tranformer model that performs MVTM
[30] Chang, H., Zhang, H., Jiang, L., Liu, C., & Freeman, W. T. (2022). MaskGIT: Masked Generative Image Transformer. arXiv preprint arXiv:2202.04200.
MaskGIT: Masked Generative Image Transformer [30]
Fig: Comparison between sequential decoding and MaskGIT’s scheduled parallel decoding
4. Related Work
16
Restormer: Efficient Transformer for High-Resolution Image Restoration [31]
[31] Zamir, Syed Waqas, et al. "Restormer: Efficient Transformer for High-Resolution Image Restoration." arXiv preprint arXiv:2111.09881 (2021).
17
5. Research Gaps
SRCNN [1] | DRCN [14] VDSR[3] FSRCNN[2] | SRGAN[5] DRRN[16] MemNet[17] EDSR[4] | RCAN[10] ESRGAN[6] LapSRN[15] SRMDNF[18] CARN[11] RDN[20] | SAN[12] IMDN[19] OSIR-RK3[21] RNAN[22] | GMGAN[13] IGNN[23] | IPT[9] DUS-GAN[7] | ESRT[8] |
|
2015 | 2016 | 2017 | 2018 | 2019 | 2020 | 2021 | 2022 | |
Timeline
Proposed Patch Translator for Image Super-Resolution
6. PTSR: Patch Translator for Image Super-Resolution [Proposed]
18
Fig. Proposed patch translator for image translation based on multi-head attention driven transformer. It divides image I into vector patches Vi with Positional embedding PE. Then the proposed transformer module is used to convert it back into vector patches leading to the generation of the image with same shape.
Patch Translator based Generator
6. PTSR: Patch Translator for Image Super-Resolution [Proposed]
19
Fig. The proposed PTSR generator framework GR2S uses a patch translator for image super-resolution. It takes XR and learns the features (i.e., the difference between ↑XR and YR ).
Proposed Transformer based Discriminator
6. PTSR
20
Fig. The Vision Transformer based Discriminator Network for image super-resolution.
Experimental setup: Dataset
6. PTSR
21
DIV2K | Set5 | Set14 | Bsd100 |
Experimental setup: Loss Function
6. PTSR
22
Fig. Proposed features are further utilised in calculating reconstruction loss for image super-resolution. HR image shows the original ground truth image, Lln×2 and Lln×4 image shows the features that model should learn for 2× and 4× scale respectively. Lld×2 and Lld×4 image shows the features that model have learned for 2× and 4× scale, respectively
The overall Generator loss can be interpreted as:
Adversarial loss
Image Super-resolution Reconstruction loss
weight coefficients of Adversarial loss
weight coefficients of Reconstruction loss
Experimental setup: Visual Activation Map
6. PTSR
23
Fig. Visual Activation Maps for proposed PTSR highlight the regions in input images where the model concentrates more. The blue colour represents the least essential regions, and the red colour represents the most important regions in terms of the image super-resolution.
Where, ∇ is the cost difference between the YR and YS for the gradient at XR image and normalised η in the range of (0,1). The VMap is shown in Fig. for XR image, which shows the model focuses on high-frequency regions for better YS image generation.
Experimental Results:
6. PTSR
24
Table: PSNR AND SSIM COMPARISON AMONG SR METHOD AT 2× and 4× SCALE .
HERE , * DENOTES THE REPRODUCED RESULTS .
Quantitative Evaluation
Experimental Results:
6. PTSR
25
Qualitative Evaluation
Fig. Visualization results for 2× super-resolution (a) High-Resolution image, (b) EDSR results, (c) IPT results, (d) ESRT results and (e) PTSR(Ours) results on different dataset images (i) Baby image in Set5, (ii) Face image in Set14 and (iii) 69020 image in BSD100 dataset
Experimental Results:
6. PTSR
26
Ablation Study
Table: IMPACT OF DIFFERENT LOSS FUNCTION AND TRANSFORMER STACK.
COMPARISON IS BASED ON PSNR AND SSIM.
HERE, # DENOTES THE SELECTED PARAMETER FOR PROPOSED MODEL.
Conclusion
6. PTSR
27
7. SRTransGAN: Image Super-Resolution using Transformer based
Generative Adversarial Network [Proposed]
28
Fig. Proposed SRTransGAN framework for image super-resolution
Key Features:
SRTransG as Generator Network
SRTransD as Discriminator Network
Proposed Method
Proposed method: Transformer Generator
7. SRTransGAN
29
Fig. Proposed SRTransG framework for image super-resolution
GDFF Block
MDTA Block
Proposed method: Transformer Discriminator
7. SRTransGAN
30
Fig. Proposed SRTransD framework for image super-resolution
Experimental setup: Dataset for SISR problem
7. SRTransGAN
31
Set5 | Set14 | Bsd100 | Urban100 |
Experimental setup: Dataset for FSR problem
7. SRTransGAN
Train Set: 17943 Female and 100057 Male, Val Set: 1000 Female and 1000 Male. Test 100 images
Test Set: 100
Train Set: 51901 images, Test Set: 100 images
32
CelebA-HQ | CelebA | FFHQ |
Fig. The proposed Different
Experimental Results:
7. SRTransGAN
33
Table: PSNR AND SSIM COMPARISON AMONG SR METHOD AT 2× SCALE .
HERE , * DENOTES THE REPRODUCED RESULTS .
Quantitative Evaluation
(SISR)
Experimental Results:
7. SRTransGAN
34
Table: PSNR AND SSIM COMPARISON AMONG SR METHOD AT 4× SCALE .
HERE , * DENOTES THE REPRODUCED RESULTS .
Quantitative Evaluation
(SISR)
Experimental Results:
7. SRTransGAN
35
Table: PSNR AND SSIM COMPARISON AMONG SR METHOD AT 4× SCALE .
HERE , * DENOTES THE REPRODUCED RESULTS .
Quantitative Evaluation
(FSR)
Experimental Results: Qualitative Evaluation (FSR)
7. SRTransGAN
36
4× super-resolution
2× super-resolution
Experimental Results:
Ablation Study
7. SRTransGAN
37
Table: IMPACT OF DIFFERENT LOSS FUNCTION AND TRANSFORMER STACK.
COMPARISON IS BASED ON PSNR AND SSIM.
HERE, # DENOTES THE SELECTED PARAMETER FOR PROPOSED MODEL.
Observations: The number of levels in the proposed model is directly proportional to the performance of model.
Observations: The number of Stack in the proposed model is directly proportional to the performance of model.
Observations: As the training set is increased model will perform better.
Contribution
7. SRTransGAN
38
UpCrossViT based Generator for Face SR
8. UpCrossViT: Face Super-Resolution with Up-Cross Attention based Vision Transformer [Proposed]
39
Fig. The proposed UpCrossViT generator framework
UpCrossViT based Generator for Face SR
8. UpCrossViT
40
Fig. The proposed Different Proposed Transformer
Experimental setup: Dataset
8. UpCrossViT
Train Set: 17943 Female and 100057 Male, Val Set: 1000 Female and 1000 Male. Test 100 images
Test Set: 100
Train Set: 51901 images, Test Set: 100 images
41
CelebA-HQ | CelebA | FFHQ |
Fig. The proposed Different
Experimental setup: Loss Function
8. UpCrossViT
42
The overall Generator loss can be interpreted as:
Adversarial loss
Image Super-resolution Reconstruction loss
weight coefficients of Adversarial loss
weight coefficients of Reconstruction loss
Lld Lln
Fig. Proposed features are further utilised in calculating reconstruction loss for image super-resolution. Lld image shows the features that model have learned for 4× scale. Lln image shows the features that model should learn for 4× scale.
Results Comparison with different settings and State-of-the-Art at 4x
8. UpCrossViT
43
Face4x0314attloss0 Test module: 17/03/2023 11:31:52 FFHQtest: psnr/ssim: 35.724/0.577 celeba: psnr/ssim: 42.705/0.747 CelebATest: psnr/ssim: 37.792/0.645 | face4x0314twogen1 Test module: 22/03/2023 13:24:03 FFHQtest: psnr/ssim: 36.803/0.612 celeba: psnr/ssim: 44.557/0.762 CelebATest: psnr/ssim: 39.422/0.690 | upatt+trans+adblock Test module: 15/03/2023 15:33:48 FFHQtest: psnr/ssim: 35.818/0.581 celeba: psnr/ssim: 43.307/0.750 CelebATest: psnr/ssim: 37.970/0.650 1000 male psnr,ssim 37.73 0.64 1000 female psnr,ssim 37.30 0.636 |
Results Comparison among SR Method at 4× Scale.
8. UpCrossViT
44
Fig. 4. Visualization results for 4× super-resolution (a) Low-Resolution, (b)Super-Resolution Image, (c) High-Resolution Image.
Qualitative Evaluation
(a) Low-Resolution (b)Super-Resolution Image (c) High-Resolution Image
Proposed method: Transformer Generator
9. CrossTrans: Cross Attention based Transformer for image related task [Proposed]
45
Fig. Proposed SRTransG framework for image super-resolution
GDFF Block
MDTA Block
UpCrossViT based Generator for Face SR
9. CrossTrans:
46
Fig. The proposed Different Proposed Transformer
Experimental setup: Dataset
9. CrossTrans:
Train Set: 17943 Female and 100057 Male, Val Set: 1000 Female and 1000 Male. Test 100 images
Test Set: 100
Train Set: 51901 images, Test Set: 100 images
47
CelebA-HQ | CelebA | FFHQ |
Fig. The proposed Different
Experimental setup: Loss Function
9. CrossTrans:
48
The overall Generator loss can be interpreted as:
Adversarial loss
Image Super-resolution Reconstruction loss
weight coefficients of Adversarial loss
weight coefficients of Reconstruction loss
Lld Lln
Fig. Proposed features are further utilised in calculating reconstruction loss for image super-resolution. Lld image shows the features that model have learned for 4× scale. Lln image shows the features that model should learn for 4× scale.
Results Comparison with different settings and State-of-the-Art at 4x
9. CrossTrans:
49
Face4x0314attloss0 Test module: 17/03/2023 11:31:52 FFHQtest: psnr/ssim: 35.724/0.577 celeba: psnr/ssim: 42.705/0.747 CelebATest: psnr/ssim: 37.792/0.645 | face4x0314twogen1 Test module: 22/03/2023 13:24:03 FFHQtest: psnr/ssim: 36.803/0.612 celeba: psnr/ssim: 44.557/0.762 CelebATest: psnr/ssim: 39.422/0.690 | upatt+trans+adblock Test module: 15/03/2023 15:33:48 FFHQtest: psnr/ssim: 35.818/0.581 celeba: psnr/ssim: 43.307/0.750 CelebATest: psnr/ssim: 37.970/0.650 1000 male psnr,ssim 37.73 0.64 1000 female psnr,ssim 37.30 0.636 |
Test module: 29/10/2023 17:24:38
FFHQtest: x2 psnr/ssim: 27.608/0.794
celeba: x2 psnr/ssim: 31.876/0.897
CelebATest: x2 psnr/ssim: 28.939/0.830
: x2 psnr/ssim: 29.359/0.823
9. CrossTrans:
50
tip
spl
IEEE Transactions on Neural Networks and Learning Systems 10.4
IEEE Transactions on Pattern Analysis and Machine Intelligence 23.6
Neurocomputing 6.0
Neural Networks 7.8
Image and Vision Computing 4.7
The Visual Computer 3.5
Results Comparison with different settings and State-of-the-Art at 4x
9. CrossTrans:
51
cross+up Test module: 29/10/2023 16:33:52 100h: x2 psnr/ssim: 21.134/0.688 100l: x2 psnr/ssim: 26.346/0.831 test100: x2 psnr/ssim: 22.733/0.777 test1200: x2 psnr/ssim: 26.974/0.794 test2800: x2 psnr/ssim: 26.970/0.814 | cross: Test module: 29/10/2023 22:27:26 input: x2 psnr/ssim: 22.761/0.764 input: x2 psnr/ssim: 28.681/0.891 input: x2 psnr/ssim: 24.743/0.849 input: x2 psnr/ssim: 31.612/0.901 input: x2 psnr/ssim: 30.726/0.903 | Test module: 30/10/2023 18:00:37 input: x2 psnr/ssim: 25.441/0.814 input: x2 psnr/ssim: 32.138/0.937 input: x2 psnr/ssim: 25.960/0.890 input: x2 psnr/ssim: 31.696/0.931 |
Results Comparison among SR Method at 4× Scale.
9. CrossTrans:
52
Fig. 4. Visualization results for 4× super-resolution (a) Low-Resolution, (b)Super-Resolution Image, (c) High-Resolution Image.
Qualitative Evaluation
(a) Low-Resolution (b)Super-Resolution Image (c) High-Resolution Image
References
53
References
54
References
55
THANK YOU