1 of 17

Semantic Guidance Learning for High-Resolution Non-homogeneous Dehazing�New Trends in Image Restoration and Enhancement workshop (NTIRE 2023)

1

Hao-Hsiang Yang2, I-Hsiang Chen2, Chia-Hsuan Hsieh4, Hua-En Chang2, Yuan-Chun Chiang2, Yi-Chung Chen3, Zhi-Kai Huang2, Wei-Ting Chen1, Sy-Yen Kuo2

1GIEE, 2 EE, 3 GICE, National Taiwan University, Taiwan

4 ServiceNow, USA

2 of 17

Outline

2

1

INTRODUCTION

High-Resolution Non-homogeneous Dehazing

3

EXPERIMENTAL RESULTS

Ablation Experiments

Comparison with State-of-the-art Methods

2

METHODOLOGY

Semantic Guided Loss Functions

Overall Neural Network

Post-Processing for Inference

4

CONCLUSION

3 of 17

INTRODUCTION

3

2021/3/19

1

4 of 17

High-Resolution Non-homogeneous Dehazing

4

    • Reconstruct clear images from non-homogeneous hazed images
      • Cannot follow traditional haze model.
        • Consider the end-to-end deep learning solution.
      • The image size is very huge. (e.g. 4000 x 6000)
        • Design the efficient training/testing pipeline.
      • The number of images is limited. (e.g. 40 pics)
        • Need use extra information as guidance.

5 of 17

Motivations

5

    • Use semantic information as guidance
      • Details on dehazed images are blurry.
      • Color tones are shifted.
        • Semantic segmentation maps can distinguish tiny objects.
        • Provide color clues on roads, vegetation and sky.
    • Efficient training/testing pipeline.
      • Training: crop images to pass the model.
      • Testing:
        • Crop images to pass the model
          • Unnatural boundaries.
        • Directly pass whole images to model
          • Degraded results.
          • Gap between different train/test size.
          • Apply post-processing to handle it.

6 of 17

Contributions

6

(a): Non-homogeneous haze image

(b): Baseline dehazed results.

(c): Proposed results.

(d): Dehazed image by feeding whole images.

(e): Dehazed image by feeding whole images + post-processing.

Cropped images causes unnatural boundaries.

7 of 17

METHODOLOGY

7

2021/3/19

2

8 of 17

Semantical Guidance Losses

8

    • Two loss functions based on semantic information
      • Semantic corresponding loss
        • Directly align with semantic features.
        • Measure the similarity of two semantic maps.
      • Semantic color tone consistency loss
        • Semantic maps also provide color clues.
          • Sky: blue, vegetation: green, road: gray & brown…
          • Calculate the averaged color values of different classes.

9 of 17

Overall Neural Network

9

    • Two-branch neural network

Discrete wavelet transform (DWT) branch

    • Encoder-decoder structure based on U-Net
    • Extract various frequency and color tone features.

Res2Net branch

    • Extract high level features.
    • Refine features by attention modules.

Combine two features to reconstruct clear images

Use large kernels to increase receptive fields.

10 of 17

Post-Processing

10

    • Two strategies for post-processing

    • Test time augmentation (TTA)
      • Rotate and flip images and passed through the neural network.
      • 8 different orientation images for TTA.

    • Test-time Local Converter (TLC)
      • Different image size between training/testing phase.
      • Converts global operation to local one.
      • Extract representations based on local spatial region of features as in training phase.

11 of 17

EXPERIMENTAL RESULTS

11

2021/3/19

3

12 of 17

Datasets and Training Details

LAFFNet: A Lightweight Adaptive Feature Fusion Network for Underwater Image Enhancement

12

2021/3/19

    • Datasets
      • 40 images from 2023 Challenge
      • Also use hazy datasets from NTIRE Challenge 2019 - 2021.
    • Pretrained segmentation model
      • DeepLab v3 trained on the Cityscape datasets
      • Select fence, sky, terrain, vegetation, sidewalk and road.
    • Data augmentation
      • Crop as 512 × 512, random flip, rotation.

13 of 17

Ablation Experiments

13

    • Test on our validation dataset
    • Baseline
      • Use original DW-GAN
    • Use large kernels and semantic guidance
      • PSNR + 0.60, SSIM + 0.0165
    • Use extra post-processing
      • PSNR + 0.38, SSIM + 0.0077

14 of 17

Comparison with State-of-the-art Methods

14

Challenge results

Compare with other methods

15 of 17

CONCLUSIONS

15

2021/3/19

4

16 of 17

Conclusions

16

Conclusions

    • Propose two loss functions based on semantic features.
    • Achieve competitive performance in NTIRE 2023 Dehazing Challenge.

Future works

    • Use stronger semantic segmentation models for guidance. (e.g. SAM)
    • Use other information like language or text-image embedding

17 of 17

LAFFNet: A Lightweight Adaptive Feature Fusion Network for Underwater Image Enhancement azing

17

2021/3/19

THANKS