DRCT: Saving Image Super-resolution
away from Information Bottleneck
Advanced Computer Vision Laboratory, National Cheng Kung University
Project page: https://github.com/ming053l/DRCT
Chih-Chung Hsu, Chia-Ming Lee, Yi-Shiuan Chou
NTIRE2024
Task Description
SwinIR: Image Restoration Using Swin Transformer, Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, Radu Timofte, 2021
Image-super Resolution is to enlarge given LR image into HR image.
HR
LR
SR
Bicubic
Downsampling
Recover
Lost spatial
Information
Task Description
SwinIR: Image Restoration Using Swin Transformer, Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, Radu Timofte, 2021
Image-super Resolution is to enlarge given LR image into HR image.
For achieving stronger performance, the mainstream is to use SwinIR-based model by its shifting-window attention mechanism to capture long-range spatial Information.
Many works design a sophisticated shifting-window mechanism to enhance performance.
(a) HAT, Overlapping Cross-Attention
Cross Aggregation Transformer for Image Restoration, Zheng Chen, Yulun Zhang, Jinjin Gu, Yongbing Zhang, Linghe Kong, and Xin Yuan, NeurIPS 2022
Activating More Pixels in Image Super-Resolution Transformer, Xiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao and Chao Dong, CVPR 2023
(b) CAT, Rectangle-Window Self-Attention
Task Description
Activating More Pixels in Image Super-Resolution Transformer, Xiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao and Chao Dong, CVPR 2023
Channel-Attention to activate more pixel have been regarded as powerful method.
Task Description
Motivation
Activating More Pixels in Image Super-Resolution Transformer, Xiangyu Chen, Xintao Wang, Jiantao Zhou, Yu Qiao and Chao Dong, CVPR 2023
These methods make SR model too complex. Can we use simpler network to achieve better/same results?
| PSNR in Set5 | #params |
SwinIR* | 32.92 | 11.9M |
HAT* | 33.04 | 20.7M |
DRCT [proposed] | 33.11 | 14.1M |
| PSNR in Set5 | #params |
SwinIR –XL** | - | 23.1M |
HAT-L* | 33.30 | 40.8M |
DRCT-L [proposed] | 33.37 | 27.5M |
*Default training setting in the paper
SwinIR: Image Restoration Using Swin Transformer, Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, Radu Timofte, 2021
** window_size=16, emb_dim=180, depth=12 (equally to HAT-L, DRCT-L )
Observation (Information Bottleneck in SR)
The maximum and minimum values of the feature map intensities in the SR network increase with the network depth but decrease to smaller values in the final layer.
This indicates that information flow in the network cannot be maintained as the depth increases.
Deep learning and the information bottleneck principle, Naftali Tishby and Noga Zaslavsky, IEEE Information Theory Workshop 2015
Observation (Information Bottleneck in SR)
SwinIR (Unstable)
HAT (Unstable)
DRCT (Stable)
Proposed Method (adding Dense-connection)
Adding dense-connection to stabilize information flow and enhance receptive field.
Shifting-window to adaptively capture long-range dependency.
* x6 RDGs in DRCT, x12 RDGs for DRCT-L
Proposed Method (adding Dense-connection)
Inspired by RRDB, which stabilize information flow in CNN-based SR models.
ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks, Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Chen Change Loy, Yu Qiao, Xiaoou Tang, ECCVW 2018
Experiment Results (Visual)
DRCT can recover a details very well.
Experiment Results (Numeric)
Method/PSNR | Set5 | Set14 | BSD100 | Urban100 | Manga109 |
SwinIR | 32.92 | 29.09 | 27.92 | 27.45 | 32.03 |
HAT | 33.04 | 29.23 | 28.00 | 27.97 | 32.48 |
DRCT [proposed] | 33.11 | 29.35 | 28.18 | 28.06 | 32.59 |
LAM and DI (To observe the model’s receptive field)
Complexity comparison
Performance comparison on the benchmark datasets
Limitation
Training Speed is slow, especially with adding more component.
Pretraining a DRCT on ImageNet need 6~8 days
Pretraining a HAT on ImageNet need 5~7 days
Pretraining a DRCT+ CAB [Channel Attention] on ImageNet need 12~15 days
* On the 4 NVIDIA-RTX-3090s
Conclusion and Contribution
1. We observe that there is a ‘information bottleneck’ which is in the final layer of SR model.
* The feature map intensive shrinkage to smaller values.
Conclusion and Contribution
2. We import dense-connection in SwinIR-based SR model, achieving ideal performance with fewer computational burden.
1. We observe that there is a ‘information bottleneck’ which is in the final layer of SR model.
* The feature map intensive shrinkage to smaller values.
* Without a sophisticated shifting window attention.
* Without a parameter-dense channel attention block (CAB).
Thanks
for
Your Attention
Contact: Chia-Ming Lee, zuw408421476@gmail.com
Yi-Shiuan Chou, nelly910421@gmail.com
https://github.com/ming053l/DRCT
GitHub Page