Comprehensive Transformer-based Model Architecture for Real-World Storm Prediction
1University of Delaware, 2University of Louisiana at Lafayette
3Intel Corporation, 4Tulane University
Fudong Lin1, Xu Yuan1, Yihe Zhang2, Purushottam Sigdel3,
Li Chen2, Lu Peng4, Nian-Feng Tzeng2
Significance of Storm Predictions
Timely and precise storm prediction can provide an early alert for preparation, avoiding potential damage to property and human safety.
Hurricane Katrina (August 2005)
Existing Solutions for Storm Predictions
Conventional Physical Models
In this paper, our goal is to develop a DL-based model, specifically using Transformers, to enhance storm prediction performance while minimizing computational resource requirements.
Deep Learning (DL)-based Models
Background: Vision Transformers (ViT)
Advantage
Limitation
Figure Credit: Alexey, Dosovitskiy, et. al. “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”, ICCV 2021
ViT Model Architecture
Background: Masked Auto-Encoder (MAE)
MAE is effective for learning visual representation without human-supervision by reconstructing images from the masked inputs.
Figure Credit: Kaiming, He, et. al. “Masked Autoencoders Are Scalable Vision Learners”, CVPR 2022
Encoder:
Decoder:
MAE Model Architecture
Background: SEVIR Dataset
Details of SEVIR Dataset
Illustration of Four Types of Sensor Data
Description of the SEVIR Dataset
Mark Veillette, et. al. “SEVIR : A Storm Event Imagery Dataset for Deep Learning Applications in Radar and Satellite Meteorology”, NeurIPS 2020
Challenge
Three challenges prevent researchers from using the SEVIR for storm predictions:
Our Design: Model Overview
Our Model Architecture
Insights Underlying Our Design
Our Design: Content Embedding
Our Model Architecture
Content Embedding
Intuitions Underlying Our Content Embedding
Experiment: Overall Performance
Our approach outperforms two baselines under all scenarios, with an overall accuracy of 94.4 % and an F1-Score of 85.0% on storm events.
Overall Performance for Storm Predictions
Experiment: Ablation Studies
Temporal Representation w/ Different Time Intervals
Removal of Different Components on Our Model
Experiment: Ablation Studies
Content Embeddings on the MAE Encoder
Positional Embeddings on the ViT Encoder
Experiment: A Failure Scenario
Our approach still outperforms two baseline models but all methods achieve a very poor performance. This is because the number of training data for some storm types is extremely limited.
Overall Performance for Storm Type Predictions
Conclusion and Future Work
Conclusion
Future Work
Q & A
Thank you!