1 of 44

Introducing the Second HDR ML Challenge:

Scientific Modeling out of Distribution

2 of 44

Outline

  • ML Challenge Year One in Review
  • Motivation: FAIR-IFYING ML Challenges
  • Overview: ML Challenge Year Two
  • Timeline
  • Sponsorships (AWS, AMD, NVIDIA)
  • Looking Forward
  • How to Get Involved – Hackathons (interactive)

3 of 44

Success of Year One: Anomaly Detection (AAAI 2025 workshop)

171 Participants

404 Submissions

264 Participants

2466 Submissions

163 Participants

390 Submissions

We had more than 550 teams participating!

Looking forward to this year’s

4 of 44

Motivation

Why FAIR ML Challenges?

FAIR principles make for better, more reproducible science!

ML Challenges help raise awareness about—and enthusiasm to solve—scientific problems

5 of 44

FAIR-ifying ML Challenges (Year One - 2024)

  • Goal: FAIR-ify both data and submissions
    • Submissions need to be fully reproducible
    • Software workflow description is embedded into submission
    • Ensures fully FAIR-ified workflow

5

Requirements.txt

Lists packages

Needed for inference

Whitelist

Checks against

Allowed packages for NERSC

Job Submission

Scripts will install missing packages and run inference on dataset

Start with base container w/standard builds

(temporary) human checking scripts

6 of 44

FAIR-ifying ML Challenges (Year TWo - 2025)

  • Goal: FAIR-ify both data and submissions
    • Submissions need to be fully reproducible
    • Software workflow description is embedded into submission
    • Ensures fully FAIR-ified workflow

6

Requirements.txt

Lists packages

Needed for inference

Whitelist

Checks against

Allowed packages for Openness & Reproducibility

Job Submission

Scripts will install missing packages and run inference (on NRP) on dataset

Start with base container w/standard builds

Automated script checks

7 of 44

Running on NRP

  • Previously hosted challenge at NERSC
    • Lawrence Berkeley National Lab in concert with FairUniverse project
    • Several restrictions coming from the fact that this was at LBL
      • Major restriction to 1 submission/day
  • Hosting server on the National Research Platform (NRP)
    • NSF funded initiative centered at the San Diego SuperComputer Center
      • https://nrp.ai/
    • National Kubernetes based system with GPUs/FPGAs/CPUs
    • Designed for usage of Heterogeneous computing
  • Currently hosting this on A30s within the NRP
    • We have the ability to scale up the resources by demand
    • Limiting submissions to 10/day

8 of 44

FAIR Workflows

  • Full reproducibility has been a major effort for this project
    • Workflows are all publicly available on github, datasets in other repos.
      • Scoring, submission, data preparation, example models & training code.
    • We adhere to a common, public, docker container across all institutes.
      • Additional packages (if needed) are installed at submission time.
        • A whitelist is enforced to avoid corrupt/exploitive software and ensure open-source dependencies.
  • Our level of reproducibility/FAIRness was not present in other challenges

8

A3D3

Imageomics

iHARP

9 of 44

Access to Platforms supporting FAIRness

9

10 of 44

Year two: Scientific Modeling out of distribution

Beetles as Sentinel Taxa

Forecasting Monkey neurons

PREDICTING COASTAL FLOODING EVENTS

11 of 44

Beetles as Sentinel Taxa

An Imageomics FAIR ML Challenge

12 of 44

National Ecological Observatory Network (NEON)

Ecological Monitoring Observatory across 20 U.S. Eco-Climatic Domains and 81 Field Sites

13 of 44

65

Sample Types

NEON Biorepository

>100,000

specimens added to the Biorepository each year

  • Small mammals
  • Fishes
  • Ground beetles
  • Mosquitos
  • Ticks
  • Zooplankton
  • Vascular plants & algae
  • Microbes
  • Soil
  • Dust
  • Precipitation
  • …and more!

14 of 44

Beetles as Sentinel Taxa: the Dataset

  • Beetles collected across NEON sites and pinned: Imaged and segmented
  • Sites selected for both in and out of domain subjects
  • Scientific name for each�Individual
  • NEON Eco-climatic domain �Where collection occurred

15 of 44

Time

(not to scale)

Trap deployment

Trap collection

SPEI_30d

SPEI_1y

SPEI_2y

01 July 2024

01 July 2023

01 July 2022

2 week duration

Trained model

Inputs: images

Outputs: SPEI predictions

Images of trapped beetle specimens

Objective: Predict drought status from characteristics of ground beetles collected during a sampling event.

Target variable: Standardized Precipitation Evapotranspiration Index (SPEI)

  • Derived from remote sensing data for a given location (e.g., NEON Site)
  • Calculated over a given time window (e.g., 30 days, 1 year, 2 years)
  • Range is [-3.0, 3.0]
  • < -1 indicates dry conditions
  • > 1 indicates wet conditions

Predict SPEI values

16 of 44

The Team

Elizabeth G. Campolongo

Wei-Lun Chao

Chandra Earl (NEON)

Hilmar Lapp

Kayla Perry

Sydne Record

Eric Sokol (NEON)

David E. Carlyn

Alyson East

Connor Kilrain

Fangxun Liu

Zheda Mai

S M Rayeed

Jiaman Wu

With our amazing students:

17 of 44

Forecasting Monkey Motor Neuron behavior

An A3D3 FAIR ML Challenge

18 of 44

Motivating A3D3 year 2 ChaLLenGE

  • A3D3 aims to target real-time AI applications for science
  • Our focus is not just on AI algorithms, but on efficienct fast implementation

For critical technical problems: Scientific challenges are leading the world

More data, More algorithms, less latency big challenge!

19 of 44

Motivating A3D3 anomaly detection

Real-time AI pipeline is critical for doing brain to body controls!

This is one of our focuses in A3D3!

20 of 44

Motivating A3D3 anomaly detection

Real-time AI pipeline is critical for doing brain to body controls!

This is one of our focuses in A3D3!

21 of 44

neural forecasting: why this task?

  • Most AI in neuroscience models:
    • “behavior decoding”: neurons(t) → behavior(t)
    • “neural dynamics”:
      • neurons(t) → neurons(t+1)
      • neuron_A(t) → neuron_B(t)�
  • Neural forecasting: neurons(t) → neurons(t+N)�
  • Why?
    • Challenging generalization problem (non-stationarity)
    • Potential benefits to discover/model relationships in data (e.g., neuron connections)
    • Potential benefits for real-time applications

22 of 44

The Data and the Team

  • micro-electrocorticography (uECoG) from monkeys performing reaching tasks
    • Collected by L. Scholl, P. Rajeswaran
    • Related paper: Ouchi, et al., J Neurosci 2025�
  • Voltage time-series from frontal motor cortices
    • 239 electrodes (monkey A)
    • 89 electrodes (monkey B)
  • No behavioral data/alignment

Pavi Rajeswaran

Leo Scholl

23 of 44

The A3D3 idea for the challenge

  • Consider forecasting Neuron singles from previous data
    • Neuron data consists of time series of 100+ neurons of Monkeys doing activities
  • Predicting a timeseries in the future would be the challenge

23

Can be used to do artificial limb control

24 of 44

The Challenge

  • Learning the Neural Dynamics through Prediction:
    • Propose methods to measure the changes in neural dynamics from recorded neural activity
    • trained model should predict future activities given past neural activities�
  • Generalization to Unseen Sessions:
    • Unseen datasets from other recording sessions (different day/sensors)
    • Used to validate trained model

25 of 44

What we learn

  • From the learned inputs, we also extract the correlation of the neurons
    • Allows us to understand which series of neurons are firing to trigger the signal of interest
    • The example code could be provided
    • Additional synthetic dataset could be used for evaluation of discovering underlying connection

This leads to correlation between the different channels

We can compute this from the resulting channels

26 of 44

Come and Try it

27 of 44

PREDICTING COASTAL FLOODING EVENTS

Two Weeks in advance

An iHARP FAIR ML Challenge

28 of 44

U.S EAST COAST SEA LEVEL:Motivation & The Dataset

  • Source: National Data Buoy Center (NDBC)
  • Temporal Data of Tidal and Sea Level Variation
  • 70 Years of Historical Sea Level (1950 - 2020)
  • Hourly Data
  • U.S. East Coastal Stations
  • Encompassing a wide range of oceanographic and geomorphological settings e.g differing tidal regimes, storm surge susceptibility, rates of sea level rise, geographic diversity, regional-specific challenges etc

29 of 44

The TASK

  • Training Dataset:
    • 70 Years
    • 12 Coastal Stations
    • 12 Flooding Thresholds (1 for each station)
  • Given Seed time windows:
    • 15 predefined historical time windows of 7-day intervals each e.g. 1/1/1990 to 1/7/1990, 3/1/2000 to 3/7/2000, etc,
  • Objective: Participants should predict minor flooding events in a set predicted time window of 14 days after each historical time window.
    • Number of coastal flooding days
    • Actual date(s) associated to coastal flooding event(s)
    • Flooding level value

30 of 44

The TASK

  • Model Evaluation:
    • Hidden Test Set: 4 Coastal Stations over a 70 year period (Out-of-Domain)
    • Groundtruth Evaluation associated with flooding for each coastal station
    • Metrics include: TP, TN, FP, FN, Accuracy, F-Score & Matthew’s Correlation Coefficient

31 of 44

The team

Faculty

  • Dr. Aneesh Subramanian, CUB
  • Dr. Bayu Tama, UMBC
  • Dr. Ratnaksha Lele, UMBC
  • Dr. Vandana Janeja, UMBC
  • Dr. Josephine Namayanja, UMBC

PhD Students

  • Emam Hossain, UMBC
  • Sai Vikas Amaraneni, UMBC
  • Maloy Devnath Kumar, UMBC
  • Subhankar Ghosh, UMN

32 of 44

Timeline

  • The challenge starts now! Come and Join!
  • Challenge runs through January 31st, 2026.
    • We reserve the right to extend the challenge beyond this date.
  • Awards ceremony will be Spring 2026 (exact date pending).
    • Conference to highlight broadly scientific modeling out of distribution.
  • Learn more about the challenge program from our white paper:
    • Major contribution is an end-to-end FAIR framework.
    • Our framework can work on the NERSC & NRP clusters.
    • Drives a path towards something more sophisticated.

33 of 44

SPONSORSHIPS

  • We are still seeking sponsorships for the challenge.
    • We acknowledge the following:
      • Cloud credits from AWS and NVIDIA for the winners.
      • $5000 Cloud Credits from NVIDIA(Lambda) for winners.
      • Lambda is also providing fixed credits to teams for training.
        • $400 Cloud credits on Lambda
      • Prize from AMD (TBD) (negotiating GPUs + cash)

34 of 44

Looking Forward

  • We are hosting a spin-off workshop in Spring 2026!
    • Highlight broadly scientific modeling out of distribution.
  • This work involved organizing and FAIR-ifying both data and algorithms across many domains.
    • Communication was crucial.
    • Dataset and organization were essential.
  • The Final ML Challenge Year Three (2026)
    • Building on the momentum of the first two challenges to address uncertainty quantification.
    • Come and join the effort!

35 of 44

Looking Forward

  • We are hosting a spin-off workshop in Spring 2026
  • This work involved organizing and FAIR-ifying data/algorithms across many domains
    • Communication took time
    • Dataset/organization was critical
  • The third and final challenge next year
    • Building on the momentum of the first two challenges
    • Come and join effort!

Placeholder – FARR workshop (April/May)

36 of 44

Thank you!

Join the 2nd HDR ML Challenge Today!

https://www.nsfhdr.org/mlchallenge-y2

37 of 44

Get Involved:

HAckathons!

38 of 44

First Hackathon of Year 2: October (Gathertown)

  1. Learn about running your own hackathon:

We provide the MATERIALS!

You bring the SPACE & ENTHUSIASM!

  1. Learn about the challenges:

What are the questions?

What data is available?

  1. Form a team!

39 of 44

Join our First Hackathon

October on Gathertown!

Don’t just try the challenge, learn how to run your own hackathon too!

40 of 44

Extra slides follow this one

41 of 44

First Hackathon: October XX (Virtual)

Start with How-to, then actual hackathon itself

We’re running a hackathon, would you like to too?

  • Encourage sign-up
  • Encourage running more
  • Participate in hackathon planning
  • Can be part day to multi-day (range of getting people started on the challenge(s) to actually digging into working on them)

How many run last year—table, pictures

42 of 44

What we have vs What you bring

We have MATERIALS

You bring the SPACE & ENTHUSIASM

43 of 44

Accessing All details

43

44 of 44