1 of 55

Progress Presentation�on

Development of Deep Learning Models for Image Super-Resolution

Research & Research Progress Colloquium - II

Presented By:-

Neeraj Baghel

Roll No:-RSI2021003

03-06-2023

May, 2023

Computer Vision and Biometrics Lab (CVBL)

Department of Information Technology (IT)

Indian Institute of Information Technology Allahabad (IIIT-A)

Devghat, Jhalwa, Prayagraj-211015, U. P. INDIA

Supervisor:-

Dr. Shiv Ram Dubey

(Assistant Professor)

2 of 55

  1. Introduction

Solution?

Image Super-Resolution

2

Recognition in Low-Dimensional

Storage, Transmission High-Resolution

3 of 55

2. Motivation

Motivation:

  • Improved image quality
  • Better image analysis
  • Increased efficiency
  • Compatibility
  • Cost-effective

3

Crop and Zoom

High-Resolution image

Low-Resolution Image

Applications:

  • Image Enhancement.
  • Face Enhancement
  • Medical Imaging.
  • Satellite Imaging.
  • Digital Zoom in Camera.

Low-Resolution

High-Resolution

4 of 55

3. Problem Statement

4

Development of Deep Learning Models for Image Super-Resolution

  • High-quality
  • Visually pleasing images
  • Improved quantitative
  • Improved qualitative

Objective

  • Various deep learning architectures
  • Optimisation methods
  • Performance with benchmark datasets.

Develop model for improving the quality of low-resolution images in various applications

5 of 55

4. Related Work

Interpolation Techniques

5

[24] Phillip Isola, “Image-to-image translation with conditional adversarial networks”, CVPR 2017.

[25] Ronneberger, Olaf, Philipp Fischer, and Thomas Brox. "U-net: Convolutional networks for biomedical image segmentation." Medical Image Computing and Computer-Assisted Intervention–MICCAI, 2015.

Fig: Training a conditional GAN to map edges→photo [24]

6 of 55

4. Related Work

Image-to-image translation with conditional adversarial networks [24]

6

Fig: Training a conditional GAN to map edges→photo [24]

[24] Phillip Isola, “Image-to-image translation with conditional adversarial networks”, CVPR 2017.

[25] Ronneberger, Olaf, Philipp Fischer, and Thomas Brox. "U-net: Convolutional networks for biomedical image segmentation." Medical Image Computing and Computer-Assisted Intervention–MICCAI, 2015.

7 of 55

Vision Transformer (ViT) [26]

4. Related Work

7

Method 2: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (VIT) [1]

[26] Dosovitskiy, Alexey, et al. "An image is worth 16x16 words: Transformers for image recognition at scale." arXiv (2020).

8 of 55

ViTGAN: Training gans with vision transformers [27]

4. Related Work

8

[27] Lee, K., Chang, H., Jiang, L., Zhang, H., Tu, Z., Liu, C.: Vitgan: Training gans with vision transformers. arXiv (2021).

9 of 55

Transgan: Two pure transformers can make one strong gan, and that can scale up [28]

4. Related Work

9

Grid Transformer Block

Fig: TransGAN Framework [28]

[28] Jiang, "Transgan: Two pure transformers can make one strong gan, and that can scale up." NIPS (2021).

10 of 55

TTSR: Learning Texture Transformer Network for Image Super-Resolution [29]

4. Related Work

10

Author proposed

1) A Texture Transformer,

2) Cross-scale feature integration module (CSFI)

Texture Transformer

Learnable Texture Extractor:

Texture extraction for reference images.

Whose parameters will be updated during end-to-end training.

Relevance Embedding:

Unfold Q and K into n patches.

Calculate relevance b/w them.

Then transfer from most relevant position in V.

[29] Yang, Fuzhi, et al. "Learning texture transformer network for image super-resolution." CVPR (2020).

11 of 55

4. Related Work

11

12 of 55

ESRT: Efficient Transformer for Single Image Super-Resolution [8]

4. Related Work

12

The overall architecture of the proposed Efficient SR Transformer

[8] Zhisheng Lu and Hong Liu and Juncheng Li and Linlin Zhang, “Efficient Transformer for Single Image Super-Resolution”, arXiv preprint arXiv: 2108.11084 (2021).

The pre and post-processing for the Efficient Transformer (ET).

The schematic diagram of the propose HFM

13 of 55

4. Related Work

13

State of Art

Zhisheng Lu and Hong Liu and Juncheng Li and Linlin Zhang, “Efficient Transformer for Single Image Super-Resolution”, arXiv preprint arXiv: 2108.11084 (2021).

14 of 55

4. Related Work

14

Fig: The diagram of the proposed image processing transformer (IPT)

[9] Chen, Hanting, et al. "Pre-trained image processing transformer." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021.

IPT: Pre-Trained Image Processing Transformer [9]

15 of 55

4. Related Work

15

Fig: Pipeline Overview. MaskGIT follows a two-stage design, with 1) a tokenizer that tokenizes images into visual tokens, and 2) a bidirectional tranformer model that performs MVTM

[30] Chang, H., Zhang, H., Jiang, L., Liu, C., & Freeman, W. T. (2022). MaskGIT: Masked Generative Image Transformer. arXiv preprint arXiv:2202.04200.

MaskGIT: Masked Generative Image Transformer [30]

Fig: Comparison between sequential decoding and MaskGIT’s scheduled parallel decoding

16 of 55

4. Related Work

16

Restormer: Efficient Transformer for High-Resolution Image Restoration [31]

[31] Zamir, Syed Waqas, et al. "Restormer: Efficient Transformer for High-Resolution Image Restoration." arXiv preprint arXiv:2111.09881 (2021).

17 of 55

17

  1. Convolution neural networks (CNN) based methods such as
    • SRCNN [1], FSRCNN [2], VDSR [3], EDSR [4].
    • limited performance: lack of objective functions,
    • unable to focus on the important regions.
  2. Generative adversarial networks (GANs) such as
    • SRGAN [5], ESRGAN [6], DUS-GAN [7].
    • limited performance by the learning capability of generator and discriminator networks.
  3. Transformer based networks such as
    • ESRT [8], IPT [9].
    • Their performance is limited as they have not exploited the adversarial training.

5. Research Gaps

SRCNN [1]

DRCN [14]

VDSR[3]

FSRCNN[2]

SRGAN[5]

DRRN[16]

MemNet[17]

EDSR[4]

RCAN[10]

ESRGAN[6]

LapSRN[15]

SRMDNF[18]

CARN[11]

RDN[20]

SAN[12]

IMDN[19]

OSIR-RK3[21]

RNAN[22]

GMGAN[13]

IGNN[23]

IPT[9]

DUS-GAN[7]

ESRT[8]

  • CNN’s
  • GAN’s
  • Transformer’s

2015

2016

2017

2018

2019

2020

2021

2022

Timeline

18 of 55

Proposed Patch Translator for Image Super-Resolution

6. PTSR: Patch Translator for Image Super-Resolution [Proposed]

18

Fig. Proposed patch translator for image translation based on multi-head attention driven transformer. It divides image I into vector patches Vi with Positional embedding PE. Then the proposed transformer module is used to convert it back into vector patches leading to the generation of the image with same shape.

19 of 55

Patch Translator based Generator

6. PTSR: Patch Translator for Image Super-Resolution [Proposed]

19

Fig. The proposed PTSR generator framework GR2S uses a patch translator for image super-resolution. It takes XR and learns the features (i.e., the difference between ↑XR and YR ).

20 of 55

Proposed Transformer based Discriminator

6. PTSR

20

Fig. The Vision Transformer based Discriminator Network for image super-resolution.

21 of 55

Experimental setup: Dataset

6. PTSR

  • DIV2K Dataset : 800 RGB training images and 100 validation images, Rich textures (2K resolution: 2040,1404)
  • Set5, Set14: Evaluation dataset for Super Resolution (5,14 images respectively) (buildings, animal,etc.) (High-Resolution: (512,512),(768,512) respectively)
  • Bsd100: Comparison for 4 scale factor (100 images) (Resolution: 481,321)

21

DIV2K

Set5

Set14

Bsd100

22 of 55

Experimental setup: Loss Function

6. PTSR

22

Fig. Proposed features are further utilised in calculating reconstruction loss for image super-resolution. HR image shows the original ground truth image, Lln×2 and Lln×4 image shows the features that model should learn for 2× and 4× scale respectively. Lld×2 and Lld×4 image shows the features that model have learned for 2× and 4× scale, respectively

The overall Generator loss can be interpreted as:

Adversarial loss

Image Super-resolution Reconstruction loss

weight coefficients of Adversarial loss

weight coefficients of Reconstruction loss

23 of 55

Experimental setup: Visual Activation Map

6. PTSR

23

Fig. Visual Activation Maps for proposed PTSR highlight the regions in input images where the model concentrates more. The blue colour represents the least essential regions, and the red colour represents the most important regions in terms of the image super-resolution.

Where, is the cost difference between the YR and YS for the gradient at XR image and normalised η in the range of (0,1). The VMap is shown in Fig. for XR image, which shows the model focuses on high-frequency regions for better YS image generation.

24 of 55

Experimental Results:

6. PTSR

24

Table: PSNR AND SSIM COMPARISON AMONG SR METHOD AT 2× and 4× SCALE .

HERE , * DENOTES THE REPRODUCED RESULTS .

Quantitative Evaluation

25 of 55

Experimental Results:

6. PTSR

25

Qualitative Evaluation

Fig. Visualization results for 2× super-resolution (a) High-Resolution image, (b) EDSR results, (c) IPT results, (d) ESRT results and (e) PTSR(Ours) results on different dataset images (i) Baby image in Set5, (ii) Face image in Set14 and (iii) 69020 image in BSD100 dataset

26 of 55

Experimental Results:

6. PTSR

26

Ablation Study

Table: IMPACT OF DIFFERENT LOSS FUNCTION AND TRANSFORMER STACK.

COMPARISON IS BASED ON PSNR AND SSIM.

HERE, # DENOTES THE SELECTED PARAMETER FOR PROPOSED MODEL.

27 of 55

Conclusion

6. PTSR

  • Proposed: Novel patch translator based GAN architecture (PTSR) for image super-resolution.
  • The generator contains a patch translator module capable of super-resolution image patches with the help of a transformer.
  • As the basic transformer is not suitable at patch level, we introduce the patch translator based on a vision transformer which can be used for any image-to-image translation task.
  • The proposed patch translator based generator network transforms patches into embeddings, passes through transformer layer and converts back into patches.
  • We have achieved very promising results for 2× resolution and 4x resolution.
  • The patch translator module can be used for other image-to-image translation tasks also.

27

28 of 55

7. SRTransGAN: Image Super-Resolution using Transformer based

Generative Adversarial Network [Proposed]

28

Fig. Proposed SRTransGAN framework for image super-resolution

Key Features:

SRTransG as Generator Network

SRTransD as Discriminator Network

Proposed Method

29 of 55

Proposed method: Transformer Generator

7. SRTransGAN

29

Fig. Proposed SRTransG framework for image super-resolution

GDFF Block

MDTA Block

30 of 55

Proposed method: Transformer Discriminator

7. SRTransGAN

30

Fig. Proposed SRTransD framework for image super-resolution

31 of 55

Experimental setup: Dataset for SISR problem

7. SRTransGAN

  • CUFED Dataset: 11,871 pair images
  • CUFED5 Dataset: 126 testing images with 5 reference set
  • DIV2K Dataset : 800 RGB training images and 100 validation images, Rich textures (2K resolution)
  • Set5, Set14: Evaluation dataset for Super Resolution (5,14 images) (buildings, animal,etc.)
  • Bsd100, Urban100: Comparison for 4 scale factor (100 images)

31

Set5

Set14

Bsd100

Urban100

32 of 55

Experimental setup: Dataset for FSR problem

7. SRTransGAN

  • CelebA-HQ: 30,000 high-resolution face images selected from the CelebA dataset (1024*1024).

Train Set: 17943 Female and 100057 Male, Val Set: 1000 Female and 1000 Male. Test 100 images

  • CelebA: 202,599 number of face images.

Test Set: 100

  • FFHQ: 52,000 high-quality having 512×512 resolution with variation in terms of age, ethnicity and image background with coverage of accessories such as eyeglasses, sunglasses, hats, etc.

Train Set: 51901 images, Test Set: 100 images

32

CelebA-HQ

CelebA

FFHQ

Fig. The proposed Different

33 of 55

Experimental Results:

7. SRTransGAN

33

Table: PSNR AND SSIM COMPARISON AMONG SR METHOD AT 2× SCALE .

HERE , * DENOTES THE REPRODUCED RESULTS .

Quantitative Evaluation

(SISR)

34 of 55

Experimental Results:

7. SRTransGAN

34

Table: PSNR AND SSIM COMPARISON AMONG SR METHOD AT 4× SCALE .

HERE , * DENOTES THE REPRODUCED RESULTS .

Quantitative Evaluation

(SISR)

35 of 55

Experimental Results:

7. SRTransGAN

35

Table: PSNR AND SSIM COMPARISON AMONG SR METHOD AT 4× SCALE .

HERE , * DENOTES THE REPRODUCED RESULTS .

Quantitative Evaluation

(FSR)

36 of 55

Experimental Results: Qualitative Evaluation (FSR)

7. SRTransGAN

36

4× super-resolution

2× super-resolution

37 of 55

Experimental Results:

Ablation Study

7. SRTransGAN

37

Table: IMPACT OF DIFFERENT LOSS FUNCTION AND TRANSFORMER STACK.

COMPARISON IS BASED ON PSNR AND SSIM.

HERE, # DENOTES THE SELECTED PARAMETER FOR PROPOSED MODEL.

Observations: The number of levels in the proposed model is directly proportional to the performance of model.

Observations: The number of Stack in the proposed model is directly proportional to the performance of model.

Observations: As the training set is increased model will perform better.

38 of 55

Contribution

7. SRTransGAN

  1. We propose PTSR for image super-resolution, which is a complete transformer network with no CNN backbone network in generator and discriminator network.
  2. As the basic transformer is not suitable at patch level translation, we introduce the patch translator based on a vision transformer which can be used for any image-to-image translation task.
  3. We introduce the transformer module, which contains the adaptive embedding for learning the distribution while preserving the patch location. It is beneficial for image-to-image translation tasks and more beneficial for super-resolution tasks.
  4. We introduce a new Image Super-resolution Reconstruction loss based on the difference in residuals.
  5. We conduct the extensive experiments and observe superior performance using the proposed PTSR models as compared to the existing models.

38

39 of 55

UpCrossViT based Generator for Face SR

8. UpCrossViT: Face Super-Resolution with Up-Cross Attention based Vision Transformer [Proposed]

39

Fig. The proposed UpCrossViT generator framework

40 of 55

UpCrossViT based Generator for Face SR

8. UpCrossViT

40

Fig. The proposed Different Proposed Transformer

41 of 55

Experimental setup: Dataset

8. UpCrossViT

  • CelebA-HQ: 30,000 high-resolution face images selected from the CelebA dataset (1024*1024).

Train Set: 17943 Female and 100057 Male, Val Set: 1000 Female and 1000 Male. Test 100 images

  • CelebA: 202,599 number of face images.

Test Set: 100

  • FFHQ: 52,000 high-quality having 512×512 resolution with variation in terms of age, ethnicity and image background with coverage of accessories such as eyeglasses, sunglasses, hats, etc.

Train Set: 51901 images, Test Set: 100 images

41

CelebA-HQ

CelebA

FFHQ

Fig. The proposed Different

42 of 55

Experimental setup: Loss Function

8. UpCrossViT

42

The overall Generator loss can be interpreted as:

Adversarial loss

Image Super-resolution Reconstruction loss

weight coefficients of Adversarial loss

weight coefficients of Reconstruction loss

Lld Lln

Fig. Proposed features are further utilised in calculating reconstruction loss for image super-resolution. Lld image shows the features that model have learned for 4× scale. Lln image shows the features that model should learn for 4× scale.

43 of 55

Results Comparison with different settings and State-of-the-Art at 4x

8. UpCrossViT

43

Face4x0314attloss0

Test module: 17/03/2023 11:31:52

FFHQtest: psnr/ssim: 35.724/0.577

celeba: psnr/ssim: 42.705/0.747

CelebATest: psnr/ssim: 37.792/0.645

face4x0314twogen1

Test module: 22/03/2023 13:24:03

FFHQtest: psnr/ssim: 36.803/0.612

celeba: psnr/ssim: 44.557/0.762

CelebATest: psnr/ssim: 39.422/0.690

upatt+trans+adblock

Test module: 15/03/2023 15:33:48

FFHQtest: psnr/ssim: 35.818/0.581

celeba: psnr/ssim: 43.307/0.750

CelebATest: psnr/ssim: 37.970/0.650

1000 male psnr,ssim 37.73 0.64

1000 female psnr,ssim 37.30 0.636

44 of 55

Results Comparison among SR Method at 4× Scale.

8. UpCrossViT

44

Fig. 4. Visualization results for 4× super-resolution (a) Low-Resolution, (b)Super-Resolution Image, (c) High-Resolution Image.

Qualitative Evaluation

(a) Low-Resolution (b)Super-Resolution Image (c) High-Resolution Image

45 of 55

Proposed method: Transformer Generator

9. CrossTrans: Cross Attention based Transformer for image related task [Proposed]

45

Fig. Proposed SRTransG framework for image super-resolution

GDFF Block

MDTA Block

46 of 55

UpCrossViT based Generator for Face SR

9. CrossTrans:

46

Fig. The proposed Different Proposed Transformer

47 of 55

Experimental setup: Dataset

9. CrossTrans:

  • CelebA-HQ: 30,000 high-resolution face images selected from the CelebA dataset (1024*1024).

Train Set: 17943 Female and 100057 Male, Val Set: 1000 Female and 1000 Male. Test 100 images

  • CelebA: 202,599 number of face images.

Test Set: 100

  • FFHQ: 52,000 high-quality having 512×512 resolution with variation in terms of age, ethnicity and image background with coverage of accessories such as eyeglasses, sunglasses, hats, etc.

Train Set: 51901 images, Test Set: 100 images

47

CelebA-HQ

CelebA

FFHQ

Fig. The proposed Different

48 of 55

Experimental setup: Loss Function

9. CrossTrans:

48

The overall Generator loss can be interpreted as:

Adversarial loss

Image Super-resolution Reconstruction loss

weight coefficients of Adversarial loss

weight coefficients of Reconstruction loss

Lld Lln

Fig. Proposed features are further utilised in calculating reconstruction loss for image super-resolution. Lld image shows the features that model have learned for 4× scale. Lln image shows the features that model should learn for 4× scale.

49 of 55

Results Comparison with different settings and State-of-the-Art at 4x

9. CrossTrans:

49

Face4x0314attloss0

Test module: 17/03/2023 11:31:52

FFHQtest: psnr/ssim: 35.724/0.577

celeba: psnr/ssim: 42.705/0.747

CelebATest: psnr/ssim: 37.792/0.645

face4x0314twogen1

Test module: 22/03/2023 13:24:03

FFHQtest: psnr/ssim: 36.803/0.612

celeba: psnr/ssim: 44.557/0.762

CelebATest: psnr/ssim: 39.422/0.690

upatt+trans+adblock

Test module: 15/03/2023 15:33:48

FFHQtest: psnr/ssim: 35.818/0.581

celeba: psnr/ssim: 43.307/0.750

CelebATest: psnr/ssim: 37.970/0.650

1000 male psnr,ssim 37.73 0.64

1000 female psnr,ssim 37.30 0.636

Test module: 29/10/2023 17:24:38

FFHQtest: x2 psnr/ssim: 27.608/0.794

celeba: x2 psnr/ssim: 31.876/0.897

CelebATest: x2 psnr/ssim: 28.939/0.830

: x2 psnr/ssim: 29.359/0.823

50 of 55

9. CrossTrans:

50

51 of 55

Results Comparison with different settings and State-of-the-Art at 4x

9. CrossTrans:

51

cross+up

Test module: 29/10/2023 16:33:52

100h: x2 psnr/ssim: 21.134/0.688

100l: x2 psnr/ssim: 26.346/0.831

test100: x2 psnr/ssim: 22.733/0.777

test1200: x2 psnr/ssim: 26.974/0.794

test2800: x2 psnr/ssim: 26.970/0.814

cross:

Test module: 29/10/2023 22:27:26

input: x2 psnr/ssim: 22.761/0.764

input: x2 psnr/ssim: 28.681/0.891

input: x2 psnr/ssim: 24.743/0.849

input: x2 psnr/ssim: 31.612/0.901

input: x2 psnr/ssim: 30.726/0.903

Test module: 30/10/2023 18:00:37

input: x2 psnr/ssim: 25.441/0.814

input: x2 psnr/ssim: 32.138/0.937

input: x2 psnr/ssim: 25.960/0.890

input: x2 psnr/ssim: 31.696/0.931

52 of 55

Results Comparison among SR Method at 4× Scale.

9. CrossTrans:

52

Fig. 4. Visualization results for 4× super-resolution (a) Low-Resolution, (b)Super-Resolution Image, (c) High-Resolution Image.

Qualitative Evaluation

(a) Low-Resolution (b)Super-Resolution Image (c) High-Resolution Image

53 of 55

References

  1. Dong, C., Loy, C.C., He, K., Tang, X.: Image super-resolution using deep convolutional networks. IEEE transactions on pattern analysis and machine intelligence 38(2), 295–307 (2015)
  2. Dong, C., Loy, C.C., Tang, X.: Accelerating the super-resolution convolutional neural network. In: European conference on computer vision, pp. 391–407. Springer (2016)
  3. Kim, J., Lee, J.K., Lee, K.M.: Accurate image super-resolution using very deep convolutional networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1646–1654 (2016)
  4. Lim, B., Son, S., Kim, H., Nah, S., Mu Lee, K.: Enhanced deep residual networks for single image super-resolution. In: Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pp. 136–144 (2017)
  5. C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang et al., “Photo-realistic single image super-resolution using a generative adversarial network,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017.
  6. X. Wang, K. Yu, S. Wu, J. Gu, Y. Liu, C. Dong, Y. Qiao, and C. Change Loy, “Esrgan: Enhanced super-resolution generative adversarial networks,” in Proceedings of the European conference on computer vision (ECCV) workshops, 2018.
  7. K. Prajapati, V. Chudasama, H. Patel, K. Upla, K. Raja, R. Ramachandra, and C. Busch, “Direct unsupervised super-resolution using generative adversarial network (dus-gan) for real-world data,” IEEE Transactions on Image Processing, vol. 30, pp. 8251–8264, 2021.
  8. Z. Lu, H. Liu, J. Li, and L. Zhang, “Efficient transformer for single image super-resolution,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshop, 2022.
  9. H. Chen, Y. Wang, T. Guo, C. Xu, Y. Deng, Z. Liu, S. Ma, C. Xu, C. Xu, and W. Gao, “Pre-trained image processing transformer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
  10. Y. Zhang, K. Li, K. Li, L. Wang, B. Zhong, and Y. Fu, “Image super-resolution using very deep residual channel attention networks,” in Proceedings of the European conference on computer vision (ECCV),2018.
  11. N. Ahn, B. Kang, and K.-A. Sohn, “Fast, accurate, and lightweight super-resolution with cascading residual network,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018.
  12. T. Dai, J. Cai, Y. Zhang, S.-T. Xia, and L. Zhang, “Second-order attention network for single image super-resolution,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019.
  13. X. Zhu, L. Zhang, L. Zhang, X. Liu, Y. Shen, and S. Zhao, “Gan-based image super-resolution with a novel quality loss,” Mathematical Problems in Engineering, 2020.
  14. J. Kim, J. K. Lee, and K. M. Lee, “Deeply-recursive convolutional network for image super-resolution,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016.

53

54 of 55

References

  1. W.-S. Lai, J.-B. Huang, N. Ahuja, and M.-H. Yang, “Fast and accurate image super-resolution with deep laplacian pyramid networks,” IEEE transactions on pattern analysis and machine intelligence, vol. 41, no. 11, pp. 2599–2613, 2018.
  2. Y. Tai, J. Yang, and X. Liu, “Image super-resolution via deep recursive residual network,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017.
  3. Y. Tai, J. Yang, X. Liu, and C. Xu, “Memnet: A persistent memory network for image restoration,” in Proceedings of the IEEE international conference on computer vision, 2017.
  4. K. Zhang, W. Zuo, and L. Zhang, “Learning a single convolutional super-resolution network for multiple degradations,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018,
  5. Z. Hui, X. Gao, Y. Yang, and X. Wang, “Lightweight image super-resolution with information multi-distillation network,” in Proceedings of the 27th ACM International Conference on Multimedia, 2019.
  6. Y. Zhang, Y. Tian, Y. Kong, B. Zhong, and Y. Fu, “Residual dense network for image super-resolution,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018.
  7. X. He, Z. Mo, P. Wang, Y. Liu, M. Yang, and J. Cheng, “Ode-inspired network design for single image super-resolution,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019
  8. Y. Zhang, K. Li, K. Li, B. Zhong, and Y. Fu, “Residual non-local attention networks for image restoration,” arXiv preprint arXiv:1903.10082,2019
  9. S. Zhou, J. Zhang, W. Zuo, and C. C. Loy, “Cross-scale internal graph neural network for image super-resolution,” Advances in neural information processing systems, vol. 33, pp. 3499–3509, 2020
  10. Isola, P., Zhu, J.Y., Zhou, T., Efros, A.A.: Image-to-image translation with conditional adversarial networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1125–1134 (2017)
  11. Ronneberger, Olaf, Philipp Fischer, and Thomas Brox. "U-net: Convolutional networks for biomedical image segmentation." Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18. Springer International Publishing, 2015.
  12. Dosovitskiy, Alexey, et al. "An image is worth 16x16 words: Transformers for image recognition at scale." arXiv preprint arXiv:2010.11929 (2020).
  13. Lee, K., Chang, H., Jiang, L., Zhang, H., Tu, Z., Liu, C.: Vitgan: Training gans with vision transformers. arXiv preprint arXiv:2107.04589 (2021).
  14. Jiang, Yifan, Shiyu Chang, and Zhangyang Wang. "Transgan: Two pure transformers can make one strong gan, and that can scale up." Advances in Neural Information Processing Systems 34 (2021): 14745-14758.Yang, Fuzhi, et al. "Learning texture transformer network for image super-resolution." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020.

54

55 of 55

References

  1. Yang, Fuzhi, et al. "Learning texture transformer network for image super-resolution." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020.
  2. Chang, H., Zhang, H., Jiang, L., Liu, C., & Freeman, W. T. (2022). MaskGIT: Masked Generative Image Transformer. arXiv preprint arXiv:2202.04200.
  3. Zamir, Syed Waqas, et al. "Restormer: Efficient Transformer for High-Resolution Image Restoration." arXiv preprint arXiv:2111.09881 (2021).
  4. Esmaeilzehi, Alireza, M. Omair Ahmad, and M. N. S. Swamy. "SRNSSI: a deep light-weight network for single image super resolution using spatial and spectral information." IEEE Transactions on Computational Imaging 7 (2021): 409-421.

55

THANK YOU