1 of 13

Scaling Machine Learning for Remote Sensing on Cloud Computing Environment

Summer School on High-Performance and Disruptive Computing in Remote Sensing

IEEE Geoscience and Remote Sensing Society

Earth Science Informatics Technical Committee

Manil Maskey, Iksha Gurung, Muthukumaran Ramasubramanian, Shubhankar Gahlot, Drew Bollinger

NASA IMPACT

June 3, 2021

2 of 13

NASA IMPACT

3 of 13

Summer School Goals

  • Provide technical guidance on performing end-to-end machine learning use case for remote sensing
  • Utilize cloud computing for machine learning on remote sensing
  • Promote open science via collaboration
  • Develop machine learning expertise for RS
  • Provide platform for sharing experiences and lessons learned
  • Promote collaboration amongst machine learning experts, domain experts, and software developers

4 of 13

Expected Outcome

Participants are expected to:

  • Learn the fundamentals of end to end machine learning life cycle
  • Design, implement, and deploy deep learning models on the cloud
  • Gain insights with cloud computing for machine learning

Everyone is expected to:

  • Exchange ideas
  • Foster collaboration

5 of 13

Machine Learning

Rapid adoption of ML due to:

Large data volumes

Advanced algorithms

Networks

Cloud computing

Hardware

6 of 13

Cloud Computing

Big data close to compute

Data storage

Scalable compute

Cloud native

7 of 13

Usecase

Image from AGU poster presentation https://ntrs.nasa.gov/citations/20190030822.

Dust storm off Alaska - Earth Observatory

8 of 13

Course Chapters

  • Chapter-0: Setup
  • End-to-end ML Lifecycle
    • Chapter-1: Data labeling
      • Identify the problem type (classification, segmentation, localization)
      • Label data
    • Chapter-2: Data Preprocessing
      • Convert raw labels and examples into machine learning trainable form
        • Eg. Change shapefiles to bitmaps
      • Prepare splits (train/validation/test)
    • Chapter-3: Model Training
      • Choose the appropriate base model architecture or create one
      • Use train and validation splits
      • Tune hyper-parameters
    • Chapter-4: Evaluation & Deployment
      • Use test data split for evaluation
      • Retrain if needed
      • Make model available for public usage.
      • Make predictions on new cases.

9 of 13

Machine learning lifecycle

10 of 13

Overview

11 of 13

References

12 of 13

Prerequisites

13 of 13

Meeting Link