1 of 16

DRCT: Saving Image Super-resolution

away from Information Bottleneck

Advanced Computer Vision Laboratory, National Cheng Kung University

Project page: https://github.com/ming053l/DRCT

Chih-Chung Hsu, Chia-Ming Lee, Yi-Shiuan Chou

NTIRE2024

2 of 16

Task Description

SwinIR: Image Restoration Using Swin Transformer, Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, Radu Timofte, 2021

Image-super Resolution is to enlarge given LR image into HR image.

HR

LR

SR

Bicubic

Downsampling

Recover

Lost spatial

Information

3 of 16

Task Description

SwinIR: Image Restoration Using Swin Transformer, Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, Radu Timofte, 2021

Image-super Resolution is to enlarge given LR image into HR image.

For achieving stronger performance, the mainstream is to use SwinIR-based model by its shifting-window attention mechanism to capture long-range spatial Information.

4 of 16

Many works design a sophisticated shifting-window mechanism to enhance performance.

(a) HAT, Overlapping Cross-Attention

Cross Aggregation Transformer for Image Restoration, Zheng Chen, Yulun Zhang, Jinjin Gu, Yongbing Zhang, Linghe Kong, and Xin Yuan, NeurIPS 2022

Activating More Pixels in Image Super-Resolution Transformer, Xiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao and Chao Dong, CVPR 2023

(b) CAT, Rectangle-Window Self-Attention

Task Description

5 of 16

Activating More Pixels in Image Super-Resolution Transformer, Xiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao and Chao Dong, CVPR 2023

Channel-Attention to activate more pixel have been regarded as powerful method.

Task Description

6 of 16

Motivation

Activating More Pixels in Image Super-Resolution Transformer, Xiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao and Chao Dong, CVPR 2023

These methods make SR model too complex. Can we use simpler network to achieve better/same results?

PSNR in Set5

#params

SwinIR*

32.92

11.9M

HAT*

33.04

20.7M

DRCT [proposed]

33.11

14.1M

PSNR in Set5

#params

SwinIR –XL**

-

23.1M

HAT-L*

33.30

40.8M

DRCT-L [proposed]

33.37

27.5M

*Default training setting in the paper

SwinIR: Image Restoration Using Swin Transformer, Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, Radu Timofte, 2021

** window_size=16, emb_dim=180, depth=12 (equally to HAT-L, DRCT-L )

7 of 16

Observation (Information Bottleneck in SR)

The maximum and minimum values of the feature map intensities in the SR network increase with the network depth but decrease to smaller values in the final layer.

This indicates that information flow in the network cannot be maintained as the depth increases.

Deep learning and the information bottleneck principle, Naftali Tishby and Noga Zaslavsky, IEEE Information Theory Workshop 2015

8 of 16

Observation (Information Bottleneck in SR)

SwinIR (Unstable)

HAT (Unstable)

DRCT (Stable)

9 of 16

Proposed Method (adding Dense-connection)

Adding dense-connection to stabilize information flow and enhance receptive field.

Shifting-window to adaptively capture long-range dependency.

* x6 RDGs in DRCT, x12 RDGs for DRCT-L

10 of 16

Proposed Method (adding Dense-connection)

Inspired by RRDB, which stabilize information flow in CNN-based SR models.

ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks, Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Chen Change Loy, Yu Qiao, Xiaoou Tang, ECCVW 2018

11 of 16

Experiment Results (Visual)

DRCT can recover a details very well.

12 of 16

Experiment Results (Numeric)

Method/PSNR

Set5

Set14

BSD100

Urban100

Manga109

SwinIR

32.92

29.09

27.92

27.45

32.03

HAT

33.04

29.23

28.00

27.97

32.48

DRCT [proposed]

33.11

29.35

28.18

28.06

32.59

LAM and DI (To observe the model’s receptive field)

Complexity comparison

Performance comparison on the benchmark datasets

13 of 16

Limitation

Training Speed is slow, especially with adding more component.

Pretraining a DRCT on ImageNet need 6~8 days

Pretraining a HAT on ImageNet need 5~7 days

Pretraining a DRCT+ CAB [Channel Attention] on ImageNet need 12~15 days

* On the 4 NVIDIA-RTX-3090s

14 of 16

Conclusion and Contribution

1. We observe that there is a ‘information bottleneck’ which is in the final layer of SR model.

* The feature map intensive shrinkage to smaller values.

15 of 16

Conclusion and Contribution

2. We import dense-connection in SwinIR-based SR model, achieving ideal performance with fewer computational burden.

1. We observe that there is a ‘information bottleneck’ which is in the final layer of SR model.

* The feature map intensive shrinkage to smaller values.

* Without a sophisticated shifting window attention.

* Without a parameter-dense channel attention block (CAB).

16 of 16

Thanks

for

Your Attention

Contact: Chia-Ming Lee, zuw408421476@gmail.com

Yi-Shiuan Chou, nelly910421@gmail.com

https://github.com/ming053l/DRCT

GitHub Page