GAIA HazLab
LECTURE 1 | SUMMER 2026
Agents for GAIA
From research question to checked action
Agent anatomy | scientific workflows | human checkpoints
Delegate the work.
Keep the scientific decisions.
GAIA HazLab | Agents for GAIA | Summer 2026
1/16
GAIA HazLab
START WITH THE WORK, NOT THE PRODUCT
What work do we actually want to delegate?
Prepare data
RESEARCH QUESTION
analysis-ready sample + QC
Build software
ISSUE OR SPECIFICATION
tested pull request
Run models
EXPERIMENT BRIEF
comparison + model card
Translate
CROSS-DOMAIN QUESTION
cited workflow
Coordinate
NOTES AND STATUS
draft actions + digest
The design question is not only what the agent can do. It is which decisions remain ours.
GAIA HazLab | What to delegate | Summer 2026
2/16
GAIA HazLab
PART 1 OF 4
Agent anatomy: a model embedded in a system
HARNESS / CONTROLLER
LANGUAGE MODEL
proposes the
next step
GOAL +
CONTEXT
HUMAN
SCIENTIST
TOOLS
ENVIRONMENT
STATE + HISTORY
PERMISSIONS
FIVE PIECES
1 goal + context
2 model
3 tools
4 controller + permissions
5 environment + state
The model proposes.
The system executes.
GAIA HazLab | Agent anatomy | Summer 2026
3/16
GAIA HazLab
PART 2 OF 4
Agent anatomy: one cycle
OBSERVE
read context
and results
DECIDE
select the
next step
ACT
call a tool
INSPECT
read what
happened
ONE
CYCLE
EXAMPLE: DATA PREPARATION
1
Observe
Read the request and approved catalog metadata.
2
Decide
Query coverage before downloading any data.
3
Act
Call the catalog tool with region and dates.
4
Inspect
Coverage is incomplete; revise the source choice.
The result of one action changes the next action.
GAIA HazLab | Agent anatomy | Summer 2026
4/16
GAIA HazLab
PART 3 OF 4
Correctness is checked at every transition
Different steps require different kinds of checking.
GOAL -> PLAN
All requested variables, dates, limits, and outputs are represented.
Does the plan reflect the scientific intent?
PLAN -> ACTION
Tool exists; arguments are valid; action is permitted.
Is this the right scientific or operational step?
ACTION -> RESULT
Command completed; expected files or state changes exist.
Does the output mean what the agent says it means?
RESULT -> CONTINUE / STOP
Named completion criteria and thresholds are satisfied.
Is this enough to proceed, scale, publish, or stop?
Automatic checks establish state. Human checks establish meaning and acceptability.
GAIA HazLab | Agent anatomy | Summer 2026
5/16
GAIA HazLab
PART 4 OF 4
Scientists choose where to enter the loop
MORE HUMAN ATTENTION
MORE BOUNDED AUTONOMY
Human closes
every loop
CHAT / MANUAL
Approve the
plan
PLAN FIRST
Approve consequential
actions
PERMISSION GATES
Review milestones
and final product
BOUNDED RUN
Unattended inside
strict limits
HEADLESS / SCHEDULED
BEFORE THE RUN
Goal, sources of truth, allowed tools, limits, and stop condition.
AFTER THE PLAN
Interpretation, assumptions, sequence of work, and proposed checks.
BEFORE CONSEQUENCE
Expensive, external, irreversible, or outward-facing actions.
AT MILESTONES
Plots, diffs, sampled data, validation report, and final artifact.
Human-in-the-loop is not one setting. It is a separate design choice for each action.
GAIA HazLab | Agent anatomy | Summer 2026
6/16
GAIA HazLab
CHOOSE THE SIMPLEST FORM THAT FITS
Chat, script, workflow, or agent?
CHAT
Human chooses every next step
BEST FOR
Questions, synthesis, drafting, reflection
SCRIPT
Every branch is written in advance
BEST FOR
Stable, repeated, deterministic procedures
FIXED WORKFLOW
Stages are predefined; a model works inside them
BEST FOR
Repeatable pipelines with bounded language tasks
AGENT
The model chooses the next step from observations
BEST FOR
Open-ended, multi-step, checkable work
USE AN
AGENT WHEN
the task is multi-step
the next step depends on what it finds
tools can expose the state
success can be checked
AUDIENCE PAUSE
summarize one paper
convert 500 files
debug a failing pipeline
select data sources
GAIA HazLab | Choosing the mode | Summer 2026
7/16
GAIA HazLab
SCIENTISTS FRAME THE WORK AND SIGN OFF
Agents sit inside a scientific workflow
Frame the science
Interpret and release
01
Research
question
SCIENTIST
02
Assumptions
SCIENTIST
03
Goals +
constraints
SCIENTIST
04
Specification
CO-DESIGN
05
Experiment
CO-DESIGN
06
Implement
AGENT + RSE
07
Validate
CO-DESIGN
08
Handoff +
reuse
SCIENTIST
question is answerable
assumptions named
success defined
inputs + outputs
comparison is meaningful
tests + artifacts
automated + manual checks
ready for community use
CHECK AT EACH PHASE
Scientists own the framing, the meaning of the checks, the interpretation, and the decision to release.
GAIA HazLab | Scientific workflow | Summer 2026
8/16
GAIA HazLab
ILLUSTRATIVE GAIA CASE
Running case: a small analysis-ready hazard dataset
RUNNING CASE
7-day sample cube
precipitation | soil moisture | terrain
Stop before scaling beyond the sample.
GAIA AGENT TASK CARD
GOAL
Prepare a small, inspectable dataset for a western Washington landslide-nowcast pilot.
AUTHORITATIVE CONTEXT
Approved catalog records, variable definitions, and GAIA data recipes.
ALLOWED TOOLS
Catalog search, metadata reads, Python/Xarray, local plots, and file writes.
CONSTRAINTS
Seven days only. Sample under 100 MB. No full cloud job. No publishing.
DONE WHEN
The sample opens, variables and units are documented, and a QC figure is produced.
HUMAN GATE
A scientist approves data choice and sample quality before any scale-up.
This task is bounded by an artifact, a size limit, named checks, and an approval gate.
GAIA HazLab | Running case | Summer 2026
9/16
GAIA HazLab
RUNNING CASE
Step 1: interpret the scientific request
1
INTERPRET
2
PLAN
3
SAMPLE
4
DELIVER
SCIENTIST REQUEST
"Prepare a seven-day sample for a western Washington landslide-nowcast pilot. I need precipitation, soil moisture, terrain, a quick-look plot, and provenance. Do not scale yet."
Natural language contains both intent and ambiguity.
STRUCTURED INTERPRETATION
REGION
western Washington
PERIOD
7 days
VARIABLES
precipitation; soil moisture; terrain
OUTPUT
sample cube + QC + provenance
SCALE LIMIT
< 100 MB; no full run
TARGET RESOLUTION
not specified
TIME RESOLUTION
hourly or daily?
Pause: clarify resolution before planning.
AGENT DOES
Extract the region, period, variables, output, constraints, and unresolved questions.
AUTOMATIC CHECK
Required fields exist; coordinates and dates parse; scale limit is explicit.
HUMAN CHECKPOINT
Confirm variable definitions, target resolution, time resolution, and intended use.
GAIA HazLab | Running case | Summer 2026
10/16
GAIA HazLab
RUNNING CASE
Step 2: choose data and propose a plan
1
INTERPRET
2
PLAN
3
SAMPLE
4
DELIVER
ILLUSTRATIVE CATALOG ENTRIES
precip-grid
VARIABLE
precipitation
RESOLUTION
1 km | hourly
UNITS
mm h-1
ACCESS
coverage: complete
soil-moisture
VARIABLE
volumetric water content
RESOLUTION
9 km | daily
UNITS
m3 m-3
ACCESS
coverage: complete
terrain-dem
SELECTED
VARIABLE
elevation + slope
RESOLUTION
30 m | static
UNITS
m; degrees
ACCESS
coverage: complete
PROPOSED PLAN
1 fetch metadata only
2 align the sample region
3 stage seven days under 100 MB
4 harmonize coordinates and units
5 write quick-look plot and provenance
6 stop for approval
AGENT DOES
Read catalog metadata, select one source per variable, and produce a stepwise plan.
AUTOMATIC CHECK
Coverage intersects the request; variables and units exist; access and license are known.
HUMAN CHECKPOINT
Approve scientific suitability, scale assumptions, and the plan before data movement.
GAIA HazLab | Running case | Summer 2026
11/16
SELECTED
SELECTED
GAIA HazLab
RUNNING CASE
Step 3: stage a small sample and inspect it
1
INTERPRET
2
PLAN
3
SAMPLE
4
DELIVER
QUICK-LOOK MAP
sample only | synthetic preview
XARRAY SUMMARY
DIMENSIONS
time: 168 | y: 240 | x: 320
SIZE
61.8 MB
COORDINATES
time UTC | latitude | longitude
PRECIPITATION
float32 | mm h-1
SOIL_MOISTURE
float32 | m3 m-3
SLOPE
float32 | degrees
MISSING
1.8% | concentrated at boundary
Decision: inspect the pattern before any scale-up.
AGENT DOES
Stage only the bounded sample; harmonize coordinates and units; create a quick-look.
AUTOMATIC CHECK
Files open; dimensions, coordinates, units, missingness, and size are reported.
HUMAN CHECKPOINT
Inspect spatial and temporal patterns. Decide whether the sample is scientifically plausible.
GAIA HazLab | Running case | Summer 2026
12/16
GAIA HazLab
RUNNING CASE
Step 4: deliver the artifact and request approval to scale
1
INTERPRET
2
PLAN
3
SAMPLE
4
DELIVER
DELIVERABLE BUNDLE
plan.md
interpretation, assumptions, selected sources
data/sample.zarr
bounded, analysis-ready sample
qc/quicklook.png
visual inspection of the sample
provenance.yaml
source records, versions, transformations
validation.md
automatic checks and required manual checks
trace.json
tool calls, results, limits, and stop reason
The final chat message is not the deliverable. The inspectable bundle is.
HUMAN GATE
Scale to the full region and season?
APPROVE
REVISE
STOP
AGENT DOES
Assemble the data, QC, provenance, validation report, and run trace.
AUTOMATIC CHECK
Required files exist; schemas and sample checks pass; the stop reason is recorded.
HUMAN CHECKPOINT
Review the sample and provenance. Approve, revise, or stop before scale-up.
GAIA HazLab | Running case | Summer 2026
13/16
GAIA HazLab
PLACE THE GATE WHERE THE ACTION CHANGES RISK
Autonomy should track consequence and reversibility
LOW CONSEQUENCE
HIGH CONSEQUENCE
REVERSIBLE / SANDBOXED
HARD TO REVERSE / EXTERNAL
increasing need for explicit approval
Read documentation
Stage a 50 MB sample
Edit a feature branch
Start a large cloud job
Publish data or send a report
PRACTICAL RULE
Let read-only, local, cheap, and reversible actions proceed inside clear limits.
Require explicit approval before actions that are costly, external, hard to reverse, or outward-facing.
GAIA HazLab | Human checkpoints | Summer 2026
14/16
GAIA HazLab
AGENT DOES -> SYSTEM CHECKS -> HUMAN DECIDES
The same pattern across GAIA
GAIA JOB
AGENT DOES
SYSTEM CHECKS
HUMAN DECIDES
RESEARCH SOFTWARE
issue -> tested pull request
tests, CI, diff, minimal change
merge or close
DATA INTERROGATION
request -> staged cube + QC
schema, coverage, units, size
scale or publish
MODELING
experiment -> comparison + model card
held-out metrics, config, runtime
promote or interpret
GAIA TRANSLATOR
question -> cited workflow
retrieval, links, cited constraints
choose the method
COORDINATION
notes -> draft digest + actions
source-linked owners and dates
send or create
GAIA HazLab | Across GAIA | Summer 2026
15/16
GAIA HazLab
FIVE-MINUTE EXERCISE
Design one bounded agent task from your work
APPLY IT
YOUR GAIA AGENT TASK CARD
1 GOAL
What artifact or changed state should exist?
2 CONTEXT
Which files, data, documentation, or policies are authoritative?
3 TOOLS
What may the agent read, run, query, or modify?
4 CHECKS
How will each step expose whether it worked?
5 HUMAN GATES
Where must a scientist review or approve?
6 LIMITS
Time, cost, data volume, turns, and stopping condition.
5 MINUTES
Choose one recurring task.
1
Name the artifact.
2
Name one automatic check.
3
Name one human gate.
4
Name the limit that keeps the first run small.
PAIR SHARE: what remains a human decision?
BOUND IT. CHECK IT LOCALLY. CHOOSE THE GATES. KEEP THE ARTIFACTS.
GAIA HazLab | Apply it | Summer 2026
16/16