1 of 31

CONVERGENT INTELLIGENCE WORKSHOP · DISCOVERY PARTNERS INSTITUTE, CHICAGO

Real-Time AI

for Science

The Accelerated AI Algorithms for Data-Driven Discovery Institute (A3D3)

Shih-Chieh Hsu

University of Washington · A3D3 Director · September 29, 2026

a3d3.ai · fastmachinelearning.org · NSF award PHY-2117997

AI-assisted slide

2 of 31

2

AI EXPANDING SCIENTIFIC HORIZONS

AI lets us ask bigger questions, from the smallest to the largest scales

The small

10⁻¹⁵ m · particles

Credit: Keiko Murano

The medium

10⁻⁴ m · neurons

Credit: Joseph Caputo

The big

10¹⁵ m · black holes

Credit: SXS Lensing

2013 Physics �François Englert and �Peter W. Higgs

2014 Physiology or Medicine �John O’Keefe, May-Britt Moser �and Edvard Moser

2017 Physics�Rainer Weiss, Barry C. Barish and �Kip S. Thorne

3 of 31

3

WHAT DO THESE DOMAINS SHARE?

Ultra-low latency at extreme throughput

High energy physics

Proton collisions every 25 ns (40 MHz); future LHC data rates above 1 Pb/s.

Multi-messenger astrophysics

 

Neuroscience

Closed-loop experiments need low-latency, causal inference.

4 of 31

4

WHERE A3D3 FOCUSES

Training builds the model; real-time inference is where the hardware limits bite

Training

Very large datasets and memory; CPU/GPU/TPU farms; floating point. Done in an ML framework such as TensorFlow or PyTorch.

Minimal compute, often integer quantization. Real-time performance and power limits that require custom hardware: FPGA, ASIC, edge devices.

Diagram credit: Marzieh Vaez Torshizi

Inference · focus of A3D3

5 of 31

5

THE INSTITUTE

An NSF Harnessing the Data Revolution Institute, since 2021

Mission

To advance scientific discovery through real-time AI applications at scale.

Vision

To empower researchers with the knowledge and tools for effective real-time AI use across scientific fields.

$15M

NSF award, plus a $1.2M supplement

2021 → 2027

Started 2021; extended to Sep 30, 2027

NSF award PHY-2117997 · a3d3.ai

6 of 31

6

THE PEOPLE

21 institutions, three continents

21

institutions

180

members

75%

of members are trainees

7 of 31

7

MULTI-DISCIPLINARY BY DESIGN

Physicists, astronomers, neuroscientists and engineers side by side

National recognition

Academy members / Fellows: Chen, Scholberg

Early career awards

Cremonesi, Li, Gonski, Aarestad, Rankin, Orsborn

Alumni → faculty

Rankin, Chou, Chen, Liu became assistant professors; Chen is a Google Fellow

8 of 31

8

THE PROGRAM

Training the next generation, and opening problems to everyone

Flagship postbac program

Applications per year, holistic review with an equity-minded rubric

13

Y1

55

Y2

87

Y3

103

Y4

1063

Y5

80% of alumni pursue STEM graduate degrees

HDR ML Challenges

553

participants · Y1 anomaly detection (2,656 submissions)

391

participants · Y2 out-of-distribution modeling �(4,827 submissions)

9 of 31

9

HARDWARE–ALGORITHM CO-DESIGN

Convergence: where science, AI and compute meet

Scientific domain applications

Compute elements & systems

AI algorithms

Science data pipelines

Domain-inspired ML

ML-specific systems

A3D3

Algorithms

Handle irregular data, scarce labels, and the need for interpretable models.

Hardware

Dedicated platforms for low latency and high throughput, within power and memory limits.

Design tools

Automation so domain experts can deploy their own models on hardware.

10 of 31

10

FROM MODEL TO SILICON

hls4ml turns trained neural networks into FPGA and ASIC firmware

Enables domain experts to implement and deploy AI directly on hardware, and streamlines the path from algorithm development to hardware execution. Open source: github.com/fastmachinelearning/hls4ml

Research productivity

11 of 31

TopLevelSynth: Try hls4ml by Chatting, Design by Prompting

Cloud platform �Running Multiple FastML Workflows Online in a Hardware-Agnostic Environment

Agentic workflow �Learn by Chatting, Design by Prompting

Demo video: �tinyurl.com/THeavy duty, production level expectations, and the structuralop-Level-Synth ��Platform (preview registration open): toplevel.h125.net ��Join slack #Top-Level-Synth of FastML server for More information! ��Developer: Qibin Liu

11

12 of 31

13 of 31

13

AI ON THE EDGE AT SCALE

SuperSonic: GPU and FPGA inference as a service

Data sent to GPU/FPGA servers via gRPC; results returned synchronously or asynchronously. Kubernetes + NVIDIA Triton, with autoscaling.

Used by

14 of 31

14

Scaling up: a shared multi-facility inference platform across experiments, part of the DOE Genesis Mission’s American Science Cloud (Bhattacharya et al.).

https://www.youtubeeducation.com/watch?v=tsRtYQaBNwY

15 of 31

15

16 of 31

16

THREE DOMAIN SCIENCES

Three very different sciences, one real-time problem

HEP

LHC at CERN

40 MHz

collisions: a new bunch crossing every 25 ns

Trigger decisions in microseconds; up to 200 simultaneous collisions per crossing at the HL-LHC.

MMA

Multi-messenger astronomy

20–2000 Hz

LIGO Filtering waveforms;

Detections and directions must reach fast; rare signals hide in noise.

NEUROSCIENCE

Brain–computer interfaces

Real time

causal inference for closed-loop experiments

Neurons drift across days; labels are scarce; decoders must never look into the future.

Sources: A3D3 HEP, MMA and Neuroscience team overviews

17 of 31

HEP HIGHLIGHT · REAL-TIME TRIGGER

Anomaly detection on every LHC collision, in the Level-1 trigger

A variational autoencoder learns to compress normal collisions; the reconstruction error scores how unusual an event is. Unsupervised, so no signal model is needed.

Unique

A large fraction of AXOL1TL-selected events would otherwise have been rejected by existing triggers.

CMS-DP-2024-059 (CMS Preliminary, 2024) · slide material: Kaito Sugizaki

18 of 31

HEP HIGHLIGHT · INTELLIGENCE AT THE SENSOR EDGE

A neural network inside the pixel readout chip

50 ns

latency, a new inference every 25 ns

0.29 mm²

smallest HLS area, 28 nm CMOS

~10×

less bandwidth: 72.2 → 5.77 Gbps

Co-design in practice

2-bit charge per pixel, 4–8 bit fixed-point weights, ADC thresholds learned jointly with the network. Toolchain: QKeras → hls4ml → Catapult HLS.

Single-layer position precision comparable to offline multi-layer reconstruction, opening pixel data to Level-1 track triggers.

Fig. 5, x/y residuals vs. non-ML reconstruction

Dickinson et al., “On-chip probabilistic inference for charged-particle tracking at the sensor edge,” arXiv:2602.15946

19 of 31

HEP HIGHLIGHT · REAL-TIME RECONSTRUCTION

One transformer from raw detector hits to full particle tracks

98.6%

efficiency at 0.8% fake rate

15.1 ms

per event on one A100 GPU

52×

faster than the ACORN-GNN

Miao, Govil et al., “HEPTv2: End-to-End Efficient Point Transformer for Charged Particle Reconstruction,” arXiv:2606.20437 · Fig. 2 reproduced unmodified

20 of 31

MMA HIGHLIGHT · GRAVITATIONAL WAVES

Aframe: a machine-learning search for black-hole mergers

38

catalog events recovered at p_astro > 0.5

3

new candidates beyond GWTC-3, confirmed

27

events louder than 100 years of background

A ResNet-34 scans 1.5 s windows of both LIGO detectors at 4 Hz. First end-to-end ML search of all of O3 with standard false-alarm rates, sensitive volume and p_astro — competitive with matched filtering at high mass.

E. Marx, W. Benoit et al., “A machine learning-enabled search for binary black hole mergers in LVK O3,” arXiv:2505.21261 (2025)

21 of 31

MMA HIGHLIGHT · OPTICAL TRANSIENTS

AppleCiDEr: multimodal early classification of sky alerts

Light curves, image cutouts, metadata and spectra fused in one classifier; being deployed in SkyPortal and the BOOM broker, with Rubin LSST next.

88%

photometry accuracy

87%

spectra accuracy

90–98%

test accuracy for SN I, SN II, CV and AGN

Also in MMA: ML event tagging for supernova-neutrino pointing in DUNE — 78% eES vs. νₑCC (G. Qin for DUNE, DPF 2026, in progress).

A. Junell, A. Sasli et al., “AppleCiDEr I: Data set, methods, and infrastructure,” arXiv:2507.16088 (2025) · code: github.com/skyportal/applecider

22 of 31

NEUROSCIENCE HIGHLIGHT · AUTONOMOUS LAB · MORE IN HAO FANG’S TALK

A real-time, closed-loop brain–computer interface on an FPGA

Record

A Neuropixel probe records ~50–100 neurons simultaneously.

Decode on FPGA

Real-time AI (Hauck group) tracks a latent signal of movement planning and detects when it crosses threshold.

Act

Change the task or stimulate the brain, then test the impact on behavior.

23 of 31

NEUROSCIENCE HIGHLIGHTS

Decoders that adapt without daily recalibration

STABLE IBCI DECODING

SPINT

0.66 R²

on held-out sessions with 0 test-time labels or gradient updates

Latency ratio 0.13: real-time capable. One unlabeled calibration trial suffices.

Le et al., NeurIPS 2025

NEURAL FOUNDATION MODEL

RPNT

0.8778 R²

cross-subject + cross-task, vs. 0.7717 best baseline

Pretrained with no behavior labels, then adapted few-shot.

Fang et al., arXiv:2601.17641 (2026)

SPEECH BCI: BRAIN-TO-TEXT

DCoND

5.77% WER

vs. 8.93% for the previous leading method

Context-aware neural decoding fused with LLMs; state of the art on Brain-to-Text 2024.

Li et al., J. Neural Eng. 22 (2025) 056026

24 of 31

LOOKING AHEAD · AGENTIC AI

AI agents that plan, run, review and write up a physics analysis

~10 h

per end-to-end analysis, vs. �~1 year today

8 / 9

results within |pull| < 2 of published values

A first: the primary Lund jet plane density in e⁺e⁻, from archived ALEPH data. Orchestrator plus sub-agents, six reviewer agents and a human gate at unblinding. All results still require independent expert verification.

Moreno, Bright-Thonney, Novak et al., “AI Agents Can Already Autonomously Perform Experimental High Energy Physics,” arXiv:2603.20179

25 of 31

COMMUNITY AND ECOSYSTEM

FastML: an open ecosystem across academia, national labs and industry

High-performance data systems with low latency, high-throughput processing, real-time control modules and custom processing elements.

26 of 31

COMMUNITY BUILDING

Where the community meets, learns and builds together

Summer school / Workshops

HDR ML Challenges / FPGA Hackthons

FastML for Science 2026 · UC San Diego

FastML Server on Slack for everyday exchange

AD, FAIR, Agentic AI define benchmark dataset for sceintifc community

This workshop, with POSE HLS4ML and DPI

27 of 31

ACADEMIC–INDUSTRY CONNECTION

Open tools flow out; industry needs and platforms flow in

FastML Foundation, a non-profit with A3D3 leadership: Phil Harris (President), Thea Aarrestad, Alex Tapper, Ryan Kastner, Javier Duarte, Nhan Tran.

NSF POSE Phase II · HLS4ML

An open-source ecosystem for collaborative rapid design of edge-AI hardware accelerators, from science to closed-loop real-time control.

fastmachinelearning.org · fastmachinelearning.org/pose

28 of 31

INTERNATIONAL AND INDUSTRY PARTNERSHIP

The Cross-Pacific AI Initiative: �UW × Tsukuba, with Amazon and NVIDIA

Advancing AI research to drive societal change and foster long-term collaboration that accelerates innovation.

X-PAI: AI-powered sleep-disorder assessment

Sleep apnea affects 1 billion+ people, 90% undiagnosed. Multimodal explainable AI on face, speech and response patterns reaches >90% accuracy. Tsukuba clinical data × UW real-time AI; 2025 data → 2027 smartphone app.

Kei Muroi, H. Kitagawa, S. Maeda, H. Noma, A. Ishii, H. Tanaka, H. Yamamoto, H. Hsu, S. Hauck, E. Shlizerman, A. Orsborn, S. Nakajima

29 of 31

FUTURE OF A3D3

Next: agents, foundation models and national-scale platforms

DOE Genesis Mission

A $5B national initiative. A3D3 researchers lead 2 proposals and co-lead 7 more, e.g.:

  • AI-accelerated supernova burst response with DUNE (Scholberg)
  • Real-time AI for quantum-enabled rare-event discovery (Li)
  • AIDA-Scout: real-time anomaly detection in CMS L1 scouting (Duarte, Diaz)
  • Foundation models for particle and nuclear physics (Cremonesi)

HDR Challenge Y3

Train a foundation model that encodes HEP simulation for AI agents, so participants can fine-tune it for analysis.

Yulei Zhang (UW), Philip Harris (MIT), Yuan-Tang Chou (NTHU), Oz Amram (Fermilab). Target: early next year (TBD).

Community-driven autonomous lab for HEP

Extending agentic AI from single analyses to community-scale discovery.

arXiv:2602.22248 · arXiv:2602.17582

energy.gov/undersecretaryforscience/genesis-mission

30 of 31

We warmly welcome the Convergence Intelligent community to join CPAD 2026 �and help shape the next generation of detector, trigger, readout, and computing systems where fast, reliable, hardware-aware AI can have transformative impact.

​

Coordinating Panel for Advanced Detectors

31 of 31

SUMMARY

Real-time AI for science is co-design, end to end

Enabling real-time discovery

Advanced AI on specialized hardware (FPGAs, ASICs) delivers the rapid, high-throughput processing experimental science needs.

Driving scalable AI ecosystems

Hardware–software co-design and AI-as-a-service platforms give flexible, efficient, scalable solutions for complex scientific workflows.

Bridging multi-scale phenomena

AI-powered tools let us explore from microscopic particles to cosmic events, and connect them.

Thank you · Partner with Us

a3d3.ai · fastmachinelearning.org