1 of 16

GAIA HazLab

LECTURE 1 | SUMMER 2026

Agents for GAIA

From research question to checked action

Agent anatomy | scientific workflows | human checkpoints

Delegate the work.

Keep the scientific decisions.

GAIA HazLab | Agents for GAIA | Summer 2026

1/16

2 of 16

GAIA HazLab

START WITH THE WORK, NOT THE PRODUCT

What work do we actually want to delegate?

Prepare data

RESEARCH QUESTION

analysis-ready sample + QC

Build software

ISSUE OR SPECIFICATION

tested pull request

Run models

EXPERIMENT BRIEF

comparison + model card

Translate

CROSS-DOMAIN QUESTION

cited workflow

Coordinate

NOTES AND STATUS

draft actions + digest

The design question is not only what the agent can do. It is which decisions remain ours.

GAIA HazLab | What to delegate | Summer 2026

2/16

3 of 16

GAIA HazLab

PART 1 OF 4

Agent anatomy: a model embedded in a system

HARNESS / CONTROLLER

LANGUAGE MODEL

proposes the

next step

GOAL +

CONTEXT

HUMAN

SCIENTIST

TOOLS

ENVIRONMENT

STATE + HISTORY

PERMISSIONS

FIVE PIECES

1 goal + context

​

2 model

​

3 tools

​

4 controller + permissions

​

5 environment + state

The model proposes.

The system executes.

GAIA HazLab | Agent anatomy | Summer 2026

3/16

4 of 16

GAIA HazLab

PART 2 OF 4

Agent anatomy: one cycle

OBSERVE

read context

and results

DECIDE

select the

next step

ACT

call a tool

INSPECT

read what

happened

ONE

CYCLE

EXAMPLE: DATA PREPARATION

1

Observe

Read the request and approved catalog metadata.

2

Decide

Query coverage before downloading any data.

3

Act

Call the catalog tool with region and dates.

4

Inspect

Coverage is incomplete; revise the source choice.

The result of one action changes the next action.

GAIA HazLab | Agent anatomy | Summer 2026

4/16

5 of 16

GAIA HazLab

PART 3 OF 4

Correctness is checked at every transition

Different steps require different kinds of checking.

GOAL -> PLAN

All requested variables, dates, limits, and outputs are represented.

Does the plan reflect the scientific intent?

PLAN -> ACTION

Tool exists; arguments are valid; action is permitted.

Is this the right scientific or operational step?

ACTION -> RESULT

Command completed; expected files or state changes exist.

Does the output mean what the agent says it means?

RESULT -> CONTINUE / STOP

Named completion criteria and thresholds are satisfied.

Is this enough to proceed, scale, publish, or stop?

Automatic checks establish state. Human checks establish meaning and acceptability.

GAIA HazLab | Agent anatomy | Summer 2026

5/16

6 of 16

GAIA HazLab

PART 4 OF 4

Scientists choose where to enter the loop

MORE HUMAN ATTENTION

MORE BOUNDED AUTONOMY

Human closes

every loop

CHAT / MANUAL

Approve the

plan

PLAN FIRST

Approve consequential

actions

PERMISSION GATES

Review milestones

and final product

BOUNDED RUN

Unattended inside

strict limits

HEADLESS / SCHEDULED

BEFORE THE RUN

Goal, sources of truth, allowed tools, limits, and stop condition.

AFTER THE PLAN

Interpretation, assumptions, sequence of work, and proposed checks.

BEFORE CONSEQUENCE

Expensive, external, irreversible, or outward-facing actions.

AT MILESTONES

Plots, diffs, sampled data, validation report, and final artifact.

Human-in-the-loop is not one setting. It is a separate design choice for each action.

GAIA HazLab | Agent anatomy | Summer 2026

6/16

7 of 16

GAIA HazLab

CHOOSE THE SIMPLEST FORM THAT FITS

Chat, script, workflow, or agent?

CHAT

Human chooses every next step

BEST FOR

Questions, synthesis, drafting, reflection

SCRIPT

Every branch is written in advance

BEST FOR

Stable, repeated, deterministic procedures

FIXED WORKFLOW

Stages are predefined; a model works inside them

BEST FOR

Repeatable pipelines with bounded language tasks

AGENT

The model chooses the next step from observations

BEST FOR

Open-ended, multi-step, checkable work

USE AN

AGENT WHEN

the task is multi-step

​

the next step depends on what it finds

​

tools can expose the state

​

success can be checked

AUDIENCE PAUSE

summarize one paper

convert 500 files

debug a failing pipeline

select data sources

GAIA HazLab | Choosing the mode | Summer 2026

7/16

8 of 16

GAIA HazLab

SCIENTISTS FRAME THE WORK AND SIGN OFF

Agents sit inside a scientific workflow

Frame the science

Interpret and release

01

Research

question

SCIENTIST

02

Assumptions

SCIENTIST

03

Goals +

constraints

SCIENTIST

04

Specification

CO-DESIGN

05

Experiment

CO-DESIGN

06

Implement

AGENT + RSE

07

Validate

CO-DESIGN

08

Handoff +

reuse

SCIENTIST

question is answerable

assumptions named

success defined

inputs + outputs

comparison is meaningful

tests + artifacts

automated + manual checks

ready for community use

CHECK AT EACH PHASE

Scientists own the framing, the meaning of the checks, the interpretation, and the decision to release.

GAIA HazLab | Scientific workflow | Summer 2026

8/16

9 of 16

GAIA HazLab

ILLUSTRATIVE GAIA CASE

Running case: a small analysis-ready hazard dataset

RUNNING CASE

7-day sample cube

precipitation | soil moisture | terrain

Stop before scaling beyond the sample.

GAIA AGENT TASK CARD

GOAL

Prepare a small, inspectable dataset for a western Washington landslide-nowcast pilot.

AUTHORITATIVE CONTEXT

Approved catalog records, variable definitions, and GAIA data recipes.

ALLOWED TOOLS

Catalog search, metadata reads, Python/Xarray, local plots, and file writes.

CONSTRAINTS

Seven days only. Sample under 100 MB. No full cloud job. No publishing.

DONE WHEN

The sample opens, variables and units are documented, and a QC figure is produced.

HUMAN GATE

A scientist approves data choice and sample quality before any scale-up.

This task is bounded by an artifact, a size limit, named checks, and an approval gate.

GAIA HazLab | Running case | Summer 2026

9/16

10 of 16

GAIA HazLab

RUNNING CASE

Step 1: interpret the scientific request

1

INTERPRET

2

PLAN

3

SAMPLE

4

DELIVER

SCIENTIST REQUEST

"Prepare a seven-day sample for a western Washington landslide-nowcast pilot. I need precipitation, soil moisture, terrain, a quick-look plot, and provenance. Do not scale yet."

Natural language contains both intent and ambiguity.

STRUCTURED INTERPRETATION

REGION

western Washington

PERIOD

7 days

VARIABLES

precipitation; soil moisture; terrain

OUTPUT

sample cube + QC + provenance

SCALE LIMIT

< 100 MB; no full run

TARGET RESOLUTION

not specified

TIME RESOLUTION

hourly or daily?

Pause: clarify resolution before planning.

AGENT DOES

Extract the region, period, variables, output, constraints, and unresolved questions.

AUTOMATIC CHECK

Required fields exist; coordinates and dates parse; scale limit is explicit.

HUMAN CHECKPOINT

Confirm variable definitions, target resolution, time resolution, and intended use.

GAIA HazLab | Running case | Summer 2026

10/16

11 of 16

GAIA HazLab

RUNNING CASE

Step 2: choose data and propose a plan

1

INTERPRET

2

PLAN

3

SAMPLE

4

DELIVER

ILLUSTRATIVE CATALOG ENTRIES

precip-grid

VARIABLE

precipitation

RESOLUTION

1 km | hourly

UNITS

mm h-1

ACCESS

coverage: complete

soil-moisture

VARIABLE

volumetric water content

RESOLUTION

9 km | daily

UNITS

m3 m-3

ACCESS

coverage: complete

terrain-dem

SELECTED

VARIABLE

elevation + slope

RESOLUTION

30 m | static

UNITS

m; degrees

ACCESS

coverage: complete

PROPOSED PLAN

1 fetch metadata only

2 align the sample region

3 stage seven days under 100 MB

4 harmonize coordinates and units

5 write quick-look plot and provenance

6 stop for approval

AGENT DOES

Read catalog metadata, select one source per variable, and produce a stepwise plan.

AUTOMATIC CHECK

Coverage intersects the request; variables and units exist; access and license are known.

HUMAN CHECKPOINT

Approve scientific suitability, scale assumptions, and the plan before data movement.

GAIA HazLab | Running case | Summer 2026

11/16

SELECTED

SELECTED

12 of 16

GAIA HazLab

RUNNING CASE

Step 3: stage a small sample and inspect it

1

INTERPRET

2

PLAN

3

SAMPLE

4

DELIVER

QUICK-LOOK MAP

sample only | synthetic preview

XARRAY SUMMARY

DIMENSIONS

time: 168 | y: 240 | x: 320

SIZE

61.8 MB

COORDINATES

time UTC | latitude | longitude

PRECIPITATION

float32 | mm h-1

SOIL_MOISTURE

float32 | m3 m-3

SLOPE

float32 | degrees

MISSING

1.8% | concentrated at boundary

Decision: inspect the pattern before any scale-up.

AGENT DOES

Stage only the bounded sample; harmonize coordinates and units; create a quick-look.

AUTOMATIC CHECK

Files open; dimensions, coordinates, units, missingness, and size are reported.

HUMAN CHECKPOINT

Inspect spatial and temporal patterns. Decide whether the sample is scientifically plausible.

GAIA HazLab | Running case | Summer 2026

12/16

13 of 16

GAIA HazLab

RUNNING CASE

Step 4: deliver the artifact and request approval to scale

1

INTERPRET

2

PLAN

3

SAMPLE

4

DELIVER

DELIVERABLE BUNDLE

plan.md

interpretation, assumptions, selected sources

data/sample.zarr

bounded, analysis-ready sample

qc/quicklook.png

visual inspection of the sample

provenance.yaml

source records, versions, transformations

validation.md

automatic checks and required manual checks

trace.json

tool calls, results, limits, and stop reason

The final chat message is not the deliverable. The inspectable bundle is.

HUMAN GATE

Scale to the full region and season?

APPROVE

REVISE

STOP

AGENT DOES

Assemble the data, QC, provenance, validation report, and run trace.

AUTOMATIC CHECK

Required files exist; schemas and sample checks pass; the stop reason is recorded.

HUMAN CHECKPOINT

Review the sample and provenance. Approve, revise, or stop before scale-up.

GAIA HazLab | Running case | Summer 2026

13/16

14 of 16

GAIA HazLab

PLACE THE GATE WHERE THE ACTION CHANGES RISK

Autonomy should track consequence and reversibility

LOW CONSEQUENCE

HIGH CONSEQUENCE

REVERSIBLE / SANDBOXED

HARD TO REVERSE / EXTERNAL

increasing need for explicit approval

Read documentation

Stage a 50 MB sample

Edit a feature branch

Start a large cloud job

Publish data or send a report

PRACTICAL RULE

Let read-only, local, cheap, and reversible actions proceed inside clear limits.

Require explicit approval before actions that are costly, external, hard to reverse, or outward-facing.

GAIA HazLab | Human checkpoints | Summer 2026

14/16

15 of 16

GAIA HazLab

AGENT DOES -> SYSTEM CHECKS -> HUMAN DECIDES

The same pattern across GAIA

GAIA JOB

AGENT DOES

SYSTEM CHECKS

HUMAN DECIDES

RESEARCH SOFTWARE

issue -> tested pull request

tests, CI, diff, minimal change

merge or close

DATA INTERROGATION

request -> staged cube + QC

schema, coverage, units, size

scale or publish

MODELING

experiment -> comparison + model card

held-out metrics, config, runtime

promote or interpret

GAIA TRANSLATOR

question -> cited workflow

retrieval, links, cited constraints

choose the method

COORDINATION

notes -> draft digest + actions

source-linked owners and dates

send or create

GAIA HazLab | Across GAIA | Summer 2026

15/16

16 of 16

GAIA HazLab

FIVE-MINUTE EXERCISE

Design one bounded agent task from your work

APPLY IT

YOUR GAIA AGENT TASK CARD

1 GOAL

What artifact or changed state should exist?

2 CONTEXT

Which files, data, documentation, or policies are authoritative?

3 TOOLS

What may the agent read, run, query, or modify?

4 CHECKS

How will each step expose whether it worked?

5 HUMAN GATES

Where must a scientist review or approve?

6 LIMITS

Time, cost, data volume, turns, and stopping condition.

5 MINUTES

Choose one recurring task.

1

Name the artifact.

2

Name one automatic check.

3

Name one human gate.

4

Name the limit that keeps the first run small.

PAIR SHARE: what remains a human decision?

BOUND IT. CHECK IT LOCALLY. CHOOSE THE GATES. KEEP THE ARTIFACTS.

GAIA HazLab | Apply it | Summer 2026

16/16