1 of 15

Comprehensive Transformer-based Model Architecture for Real-World Storm Prediction

1University of Delaware, 2University of Louisiana at Lafayette

3Intel Corporation, 4Tulane University

Fudong Lin1, Xu Yuan1, Yihe Zhang2, Purushottam Sigdel3,

Li Chen2, Lu Peng4, Nian-Feng Tzeng2

2 of 15

Significance of Storm Predictions

Timely and precise storm prediction can provide an early alert for preparation, avoiding potential damage to property and human safety.

  • Unexpected storm events (e.g., hurricanes) cause significant damage to human property and health every year.

  • Hurricane Katrina resulted in 1,392 fatalities and caused property damage estimated between $97.4 billion to $145.5 billion in late August 2005.

Hurricane Katrina (August 2005)

3 of 15

Existing Solutions for Storm Predictions

Conventional Physical Models

In this paper, our goal is to develop a DL-based model, specifically using Transformers, to enhance storm prediction performance while minimizing computational resource requirements.

  • Superb performance but incurring excessive computational overhead

Deep Learning (DL)-based Models

  • Computationally efficient but achieving unsatisfactory performance

4 of 15

Background: Vision Transformers (ViT)

Advantage

  • Scalability to large model and data size
  • Global representation
  • Connection of visual and textual data

Limitation

  • Computational Overhead
  • Overfitting

Figure Credit: Alexey, Dosovitskiy, et. al. “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”, ICCV 2021

ViT Model Architecture

5 of 15

Background: Masked Auto-Encoder (MAE)

MAE is effective for learning visual representation without human-supervision by reconstructing images from the masked inputs.

Figure Credit: Kaiming, He, et. al. “Masked Autoencoders Are Scalable Vision Learners”, CVPR 2022

Encoder:

  • Consider a small subset of visible image patches
  • No positional embeddings required

Decoder:

  • Consider a full set of image patches, including both visible patches and masked tokens
  • Require positional embeddings

MAE Model Architecture

6 of 15

Background: SEVIR Dataset

Details of SEVIR Dataset

  • It contains a collection of sensor images captured by satellite and radar, characterizing weather events during 2017-2019.
  • It consists of 10180 normal events and 2559 storm events.
  • Four types of images with different resolutions are considered.

Illustration of Four Types of Sensor Data

Description of the SEVIR Dataset

Mark Veillette, et. al. “SEVIR : A Storm Event Imagery Dataset for Deep Learning Applications in Radar and Satellite Meteorology”, NeurIPS 2020

7 of 15

Challenge

Three challenges prevent researchers from using the SEVIR for storm predictions:

  • Limited Observational Samples: The storm events only account for 20% of total events (2559 storm events v.s. 10180 normal events).
  • Intangible Pattern: Weather images usually include erratic and intangible shapes.
  • Multi-scale Data: The resolutions of weather images vary from 192x192 pixels to 768x768 pixels.

8 of 15

Our Design: Model Overview

  • Three MAE encoders with different scales to learn high-quality visual representations from multi-scale solutions
  • Temporal representation for incorporating domain knowledge into deep learning model for storm predictions
  • Representation concatenation for addressing the limited observational storm events by considering visual and temporal representation simultaneously

Our Model Architecture

Insights Underlying Our Design

9 of 15

Our Design: Content Embedding

Our Model Architecture

Content Embedding

Intuitions Underlying Our Content Embedding

  • Devised for differentiating the membership of data sources
  • MAE Encoders: Only the small-scale MAE with the content embedding
  • ViT Encoder: Including Five types of content embeddings (4 for sensor data and 1 for temporal data)

10 of 15

Experiment: Overall Performance

  • Settings: Binary Classification, i.e., normal events or storm events
  • Baselines: ResNet-50 and ViT-Base
  • Metrics: Precision, Recall, F1-Score, and Accuracy

Our approach outperforms two baselines under all scenarios, with an overall accuracy of 94.4 % and an F1-Score of 85.0% on storm events.

Overall Performance for Storm Predictions

11 of 15

Experiment: Ablation Studies

Temporal Representation w/ Different Time Intervals

Removal of Different Components on Our Model

12 of 15

Experiment: Ablation Studies

Content Embeddings on the MAE Encoder

Positional Embeddings on the ViT Encoder

13 of 15

Experiment: A Failure Scenario

  • Settings: Multi-class Classification, i.e., predicting specific event types
  • Baselines: ResNet-50 and ViT-Base
  • Metrics: Precision, Recall, F1-Score, Accuracy

Our approach still outperforms two baseline models but all methods achieve a very poor performance. This is because the number of training data for some storm types is extremely limited.

Overall Performance for Storm Type Predictions

14 of 15

Conclusion and Future Work

Conclusion

  • This work has developed a comprehensive Transformer-based model architecture for real-world storm prediction, utilizing both ViT and MAE as key components.
  • With a collection of novel designs (such as image representation concatenation, temporal representation, and content embedding), our approach can achieve state-of-the-art storm prediction performance.
  • Although we conduct experiments on the curated SEVIR dataset, our model architecture can be generalized to effectively handle any type of real-world satellite and radar image data.

Future Work

  • Both our approach and the baseline models exhibit poor performance in accurately predicting storm types due to the extremely limited number of events for certain storm types, calling for the creation of multi-modal storm datasets to provide a more comprehensive and holistic view of storm events.

15 of 15

Q & A

Thank you!