1 of 15

OPDN: Omnidirectional Position-aware Deformable Network for Omnidirectional Image Super-Resolution

Xiaopeng Sun*1, Weiqi Li*1,2, Zhenyu Zhang1,2, Qiufang Ma1, Xuhan Sheng2, Ming Cheng1, Haoyu Ma1, Shijie Zhao+1, Jian Zhang2 , Junlin Li1, Li Zhang1

2 of 15

Introduction

3 of 15

Main Contributions

  • Proposing an Omnidirectional Positionaware Deformable Network (OPDN) for 360-degree image super resolution
  • introducing a new module named OPDB, which incorporates a frequency block and Fourier upsampling to enhance the final performance of image improvement
  • a favorable balance between enhancement performance and model complexity

4 of 15

OverView�

[1] Chen X, Wang X, Zhou J, et al. Activating more pixels in image super-resolution transformer[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023: 22367-22377.

  • Fig (a): The overall pipeline of proposed two-stage framework, in which two ×4 SR models are employed in stage 1, while stage 2 performs a same-resolution enhancement.

  • Fig (b): The network architecture of our proposed model A. Model B incorporates the SFF model on top of Model A to learn frequency domain information.

5 of 15

Architectures�

  • OPDB:
    • Combines dimensional information and position encoding information
    • The spherical coordinate information be aggregated through deformable convolution, achieving a larger receptive field and superior reconstruction

positional encoding can be linearly represented by position, reflecting its relative position relationship

6 of 15

Architectures�

  • OPDB:
    • is capable of adapting to dimensional changes in 360-degree images and leveraging a wider range of information to reconstruct SR results achieving a larger receptive field and superior reconstruction

Fig. Results of LAM visualization. From left to right, (a) and (b) show the LAM contribution, area of contribution and SR results

Fig. Visualizations of offset maps in OPDB. Reference and deformed points are depicted in green and red, respectively

7 of 15

Architectures�

  • SFF:
    • achieve the image-wide receptive field for exploring global contextual information.

Man Zhou, Jie Huang, Keyu Yan, Hu Yu, Xueyang Fu, Aiping Liu, Xian Wei, and Feng Zhao. Spatial-frequency domain information integration for pan-sharpening. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XVIII, pages 274–291. Springer, 2022. 5

8 of 15

Architectures�

  • Fourier Upsamping
    • spatial upsampling operators heavily depend on local pixel attention, incapably exploring the global dependency.
    • the Fourier domain obeys the nature of global modeling according to the spectral convolution theorem

Zhou, Hu Yu, Jie Huang, Feng Zhao, Jinwei Gu, Chen Change Loy, Deyu Meng, and Chongyi Li. Deep fourier up-sampling. arXiv preprint arXiv:2210.05171, 2022. 5

9 of 15

Architectures�

  • Tricks
    • Simulating Data Degradation
      • Increase the amount of training data for better performance of transformer-based models

Yanze Wu, Xintao Wang, Gen Li, and Ying Shan. Animesr: Learning real-world super-resolution models for animation videos. In Advances in Neural Information Processing Systems, 2022. 4

Fig. Visualization comparisons of simulated LR and the groudtruth

10 of 15

Architectures�

  • Tricks
    • self-ensemble x8 strategy
      • Increase the amount of training data for better performance of transformer-based models

Fig. Ensemble results in Flickr360 dataset

Fig. Self-ensemble strategy for ERP images

11 of 15

Results�

  • Ablation Study

12 of 15

Results�

  • Quantitative Results

13 of 15

Results�

  • Quanlitative Results

14 of 15

Results�

  • Qualitative Results

15 of 15

Thanks for watching!