1 of 25

Enhancing Satellite Object Localization with Dilated Convolutions and Attention-aided Spatial Pooling

Authors: Seraj Al Mahmud Mostafa, Chenxi Wang, Jia Yue, Yuta Hozumi, and Jianwu Wang

Paper id: S3542

Presenter: Jianwu Wang

Department of Information Systems, University of Maryland, Baltimore County, MD, USA

July 20th, 2025

2025

2 of 25

2

Outline

  • Introducing the datasets
  • Research problems
  • Proposed solution
  • Related works
  • Methodology
  • Results

3 of 25

3

Gravity Wave Data

Gravity waves are buoyancy acts as the restoring force, typically caused by disturbances such as airflow over mountains or convection.

Data Source and format

  • VIIRS Day/Night Band (DNB), particularly night band.
  • Total 50 scenes from upper atmosphere in the HDF format.

Preprocessing

  • Negative values (like 1.0e-9) clipped to small positive integer and any value greater than 1 is clipped to 1.
  • Any out of range values set to either black or white.

Augmentation & Labels

  • Rotated (90°, 180°, 270°) → 15,000 total patches of 200x200
  • Classes: gw (gravity wave) vs. ngw (non-GW)
  • Train/Val/test split: 70:20:10 (balanced).

Fig: Gravity Waves

4 of 25

4

Bore Data

Mesospheric bores are sharp airglow fronts, marked by sudden brightness changes, caused by steepened gravity waves moving through temperature or wind ducts.

Data Source and format

  • VISI sensor (ISS), observing airglow (~95 km altitude) [10].
  • Total 306 mesospheric bore events, night side only

Preprocessing

  • Manual inspection sharp brightness fronts with trailing waves
  • Bright: emission below duct → brighter airglow
  • Dark: emission above duct → darker airglow

Augmentation & Labels

  • Events manually labeled: 133 bright bores, 173 dark bores
  • Classes: only ‘bore’ class.

Fig: Mesospheric bores

5 of 25

5

Ocean Eddy Data

Ocean eddy is a swirling, circulating current of water that breaks off from larger currents and moves through the ocean like a slow-moving whirlpool.

Data Source and format

  • Synthetic Aperture Radar (SAR) satellite images if GeoTIFF format
  • 100 eddy and 400 non-eddy labeled images (manually annotated)

Preprocessing

  • Converted using GeoSpatial (GDAL) library to PNG files

Augmentation & Labels

  • Eddy images rotated (90°, 180°, 270°) → 400 total eddy images
  • Final dataset: 800 grayscale images (50% eddy / 50% non-eddy)
  • Train/Val/test split: 70:20:10.

Fig: Ocean Eddy

6 of 25

6

Datasets

Aspect

Gravity Wave

Mesospheric Bore

Ocean Eddy

Definition

Buoyancy-driven wave from airflow or convection

Sharp airglow front from steepened gravity waves in ducts

Swirling current breaking off from larger currents

Sensor / Source

VIIRS DNB (night band), HDF format

VISI sensor (ISS), airglow at ~95 km

SAR satellite imagery, GeoTIFF

Data Size

50 HDF files → 15,000 patches (200×200)

306 bore events (night only)

100 eddy + 400 non-eddy images

Preprocessing

Negative clipped to 1e-9, >1 clipped to 1, B/W for out-of-range

Manual inspection of sharp fronts, brightness labeled by duct position

GeoTIFF → PNG using GDAL

Augmentation

Rotation (90°, 180°, 270°)

Labels / Classes

‘gw’ vs. ‘ngw’ (gravity wave / non-GW)

‘bore’ (ony bore class)

‘eddy’ vs. ‘non-eddy’

Split

70:20:10 (train/val/test)

7 of 25

7

Research Challenges

  • Significant scale variability makes consistent object localization difficult.
  • Accurate localization requires understanding how fragmented or distorted parts of an object relate within the global context, especially under challenging conditions like city lights, cloud cover, or sensor noise.

Fig: Gravity Waves

Fig: Ocean Eddy

Fig: Mesospheric bores

8 of 25

8

Proposed Solution

We proposed a hybrid localization approach that addresses scale variability while capturing the global structure of target objects.

  • First, we apply multi-dilation convolution to extract features across multiple scales.
  • Second, we enhance the SPP (Spatial Pyramid Pooling) layer with an attention mechanism to better detect distorted or occluded regions of the object.

9 of 25

9

Related Works

Multi Dilation related works:

  • Liu et al. improved road area extraction in semantic segmentation by combining dilated convolutions with residual learning [11].
  • Zhang et al. utilized dilated convolutions in atrous CNNs to capture more semantic information for ultrasound image segmentation [12].
  • Chen et al. investigated optimal dilation rates for aggregating multiscale features and enlarging the receptive field [13].
  • Wang et al. developed smoothing techniques to address gridding artifacts in dense predictions [14].

Attention related works:

  • Zhou et al. proposed the Scale Aware Spatial Pyramid Pooling (SSPP) module, alongside Encoder Mask and Scale-Attention modules, addressing challenges in scale-awareness, boundary sharpness, and long-range dependency modeling [15].
  • Cao et al. designed a network incorporating a Context Extraction Module and an Attention Guided Module to enhance object localization and recognition by leveraging contextual information and adaptive attention mechanisms [16].
  • Feng et al. introduced AttSPPnet, combining a soft attention mechanism with Spatial Pyramid Pooling (SPP) for action recognition, allowing the network to focus on relevant regions and improving robustness to action deformation [17].

Limitations from the related works: Most existing methods focus on small object detection and overlook challenges like diverse object patterns, occlusion, and overlap, issues particularly in remote sensing data that remain largely unexplored.

10 of 25

10

Methodology

  • We proposed ‘YOLO-DCAP’, which can localize objects with varying scales and extents, even when mixed with interference like occlusion or overlap, making localalization more challenging.

YOLO-DCAP Backbone

Ref: Multi dilations [11-14]

Fig: YOLOv5 architecture

Fig: YOLOv5 Backbone

Fig: YOLO-DCAP Backbone

11 of 25

11

Methodology … cont.

Feature pyramyd Network (FPN)

+

Path Aggregation Network (PANet)

YOLOv5 Backbone Neck Head

YOLO-DCAP Backbone

Feature pyramyd Network (FPN)

+

Path Aggregation Network (PANet)

YOLO-DCAP Backbone Neck Head

12 of 25

12

Methodology … MDRC

Fig: Proposed Multi Dilated Residual Convolution (MDRC) for YOLO-DCAP.

MDRC addresses scale variation by employing parallel dilated convolutions with dilation rates, typically set to [2, 3], to capture multi-scale features. The MDRC in fact addresses the scale variability challenge.

Ref: Multi dilations [11-14]

Input

Conv

BatchNorm

SiLU Activation

Output

Fig: The CONV layer in YOLOv5.

SiLU= Sigmoid weighted Linear Unit.

Fig: Regular convolution vs dilated convolution (increased receptive fields).

13 of 25

13

Methodology … AaSP

Fig: Proposed Attention-aided Spatial Pooling (AaSP) with sequential MaxPool layers and Attention mechanism to captures local and global features at scale.

  • Stage 1: the input is reduced along the channel dimension, where the first pooling focuses on fine-grained features, and the second captures global context, reducing complexity unlike SPPF (uses 3 MaxPool layers).
  • Stage 2: 1) in squeeze phase, the global spatial information is collected using the global average pooling; 2) in excitation phase, channel interdependencies are captured using two fully connected layers with ReLU and Sigmoid activations.
  • In the end, the final output is determined, emphasizing important ones while suppressing less relevant ones.

Ref: SPP [18], Squeeze and Excitation network [19]

Fig: SPPF (spatial pyramid pooling fast) working principle, with 3 Maxpool layers to captures features [ref].

Input

MaxPool

(5x5)

MaxPool

(5x5)

MaxPool

(5x5)

concat

Output

multi-scale-feature

14 of 25

14

Methodology … AaSP

Fig: Proposed Attention-aided Spatial Pooling (AaSP) with sequential MaxPool layers and Attention mechanism to captures local and global features at scale.

  • AaSP in Sage 1, the input is reduced along the channel dimension, where the first pooling focuses on fine-grained features, and the second captures global context, reducing complexity unlike SPP (uses 3 parallel MaxPool layers).
  • AaSP in Stage 2, in the squeeze phase, the global spatial information is collected using the global average pooling and the excitation phase, channel interdependencies are captured using two fully connected layers with ReLU and sigmoid activations.
  • Finally, the final output is determined, emphasizing the important ones while suppressing less relevant ones.

Ref: SPP [18], Squeeze and Excitation network [19]

Fig: SPP working principle, with 3 parallel Maxpool layers to captures features at scale.

Input

MaxPool

(5x5)

MaxPool

(9x9)

MaxPool

(13x13)

concat

Output

multi-scale-feature

15 of 25

15

Comparisons with state-of-the-arts approaches

Table: Compares the state-of-the-art baseline methods with proposed MDRC, AaSP and enhanced models across GW, Bore, and OE datasets.

YOLO-DCAP = MDRC + AaSP

16 of 25

16

Mean and Std. Dev. Comparisons

Table: Mean and Standard Deviation comparison between baselines and the proposed YOLO-DCAP approaches.

The mean and standard deviation are calculated based on 5 runs.

17 of 25

17

Fig 1: Localization comparison across baseline models and proposed YOLO-DCAP on GW, Bore, and OE datasets.

Comparisons with state-of-the-arts

18 of 25

18

Ablation studies

  • We evaluated a modified approach having MDRC+Simplified Attention together without replacing SPPF layer.

YOLO-DCAP Backbone

Fig: YOLOv5 Backbone

Fig: YOLO-DCAP Backbone

Fig: YOLO modified Backbone (MDRC+CCSA)

19 of 25

19

Ablation studies

Table 3: Performance comparison of SSCA and AaSP, both combining MDRC with state of the arts.

20 of 25

20

Ablation studies

Table: Effects of multi-scale dilation impact on different layers of YOLO-DCAP backbone across all datasets.

21 of 25

21

Key Contribution and Conclusion

  • Designed an architecture (YOLO-DCAP) leveraging MDRC and AaSP modules to improve object localization under occlusions, scale variability, and complex spatial patterns.
  • Demonstrated superior performance and robustness, with consistent IoU gains over baseline YOLO and other state of the art attention-based models, validating effectiveness under real-world conditions.
  • Our codes are open sourced at https://github.com/AI-4-atmosphere-remote-sensing/satellite-object-localization

22 of 25

22

References

  1. Sreekanth, V.S., Raghunath, K., Mishra, D.: Deep kernel dictionary learning for detection of wave breaking features in atmospheric gravity waves. Computers & Geosciences p. 105361 (2023)
  2. González, J.L., Chapman, T., Chen, K., Nguyen, H., Chambers, L., Mostafa, S.A., Wang, J., Purushotham, S., Wang, C., Yue, J.: Atmospheric gravity wave detection using transfer learning techniques. In: 2022 IEEE/ACM International Conference on Big Data Computing, Applications and Technologies (BDCAT). pp. 128–137. IEEE (2022)
  3. Bazi, Y., Bashmal, L., Rahhal, M.M.A., Dayil, R.A., Ajlan, N.A.: Vision transformers for remote sensing image classification. Remote Sensing 13(3), 516 (2021)
  4. Hozumi, Yuta, Akinori Saito, Takeshi Sakanoi, Atsushi Yamazaki, Keisuke Hosokawa, and Takuji Nakamura. "Geographical and seasonal variability of mesospheric bores observed from the International Space Station." Journal of Geophysical Research: Space Physics 124, no. 5 (2019): 3775-3785.
  5. Z. Liu, R. Feng, L. Wang, Y. Zhong, and L. Cao, “D-resunet: Resunet and dilated convolution for high resolution satellite imagery road extraction,” in IGARSS 2019-2019 IEEE International Geoscience and Remote Sensing Symposium, pp. 3927–3930, IEEE, 2019.
  6. L. Zhang, J. Zhang, Z. Li, and Y. Song, “A multiple-channel and atrous convolution network for ultrasound image segmentation,” Medical Physics, vol. 47, no. 12, pp. 6270–6285, 2020.
  7. H. Chen and H. Lin, “An effective hybrid atrous convolutional network for pixel-level crack detection,” IEEE Transactions on Instrumentation and Measurement, vol. 70, pp. 1–12, 2021.
  8. Z. Wang and S. Ji, “Smoothed dilated convolutions for improved dense prediction,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 2486–2495, 2018.
  9. F. Zhou, Y. Hu, and X. Shen, “Scale-aware spatial pyramid pooling with both encoder-mask and scale-attention for semantic segmentation,” Neurocomputing, vol. 383, pp. 174–182, 2020.
  10. J. Cao, Q. Chen, J. Guo, and R. Shi, “Attention-guided context feature pyramid network for object detection,” arXiv preprint arXiv:2005.11475, 2020.
  11. W. Feng, X. Zhang, X. Huang, and Z. Luo, “Attention focused spatial pyramid pooling for boxless action recognition in still images,” in Artificial Neural Networks and Machine Learning–ICANN 2017: 26th International Conference on Artificial Neural Networks, Alghero, Italy, September 11-14, 2017, Proceedings, Part II 26, pp. 574–581, Springer, 2017.
  12. K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” IEEE transactions on pattern analysis and machine intelligence, vol. 37, no. 9, pp. 1904– 1916, 2015.
  13. N. Vosco, A. Shenkler, and M. Grobman, “Tiled squeeze-and-excite: Channel attention with local spatial context,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 345–353, 2021.
  14. Z. Chen, J. Hu, G. Min, C. Luo, and T. El-Ghazawi, “Adaptive and efficient resource allocation in cloud datacenters using actor-critic deep reinforcement learning,” IEEE Transactions on Parallel and Distributed Systems, vol. 33, no. 8, pp. 1911–1923, 2022.

23 of 25

23

Thank You

Contact: jianwu@umbc.edu

24 of 25

24

25 of 25

  • Hello
  • world

25