PURPOSE-BUILT AGENTS
How we build agents
Our approach to building, evaluating, and operating agents around an organization's domain expertise.
Mike Jacobi
Ai2 – Skylight, Lead Software Engineer
mikej@allenai.org
What we mean by agent
GENERAL-PURPOSE
Claude Code, Codex, OpenCode, OpenClaw, etc for arbitrary tasks
I will be referring to this as an "agent harness"
PURPOSE-BUILT
An agent with an intentional persona and a set of skills for specific tasks
Runs on an agent harness
Why use purpose-built over general-purpose?
This is a mechanism to provide an expert-curated experience for a targeted audience
You, the domain-expert agent author, know how to navigate your own world, your purpose-built agent distributes your knowledge to the users you want to engage
This talk is about purpose-built agents
THE PRODUCT
Skylight
Skylight is a maritime domain awareness tool, set on a filterable map interface
It is used in an operational context to reduce IUU fishing (Illegal, Unreported, Unregulated)
Typical users are analysts working at Navies, Coast Guards, Fisheries, and more
Analysts act on the output, often to inform where to send their enforcement vessels
THE AGENT
Shippy
Shippy is an agent that allows you to interact with Skylight data, in prod today
Shippy runs on OpenClaw with Skylight-specific skills
queries our public API
writes Python code
rejects inappropriate prompts
says "I'm not sure"
adds domain-specific caveats
provides data-source attribution
reads user-uploaded files
generates downloadable files
SAMPLE QUERIES
"Show me all vessels in southern Vietnam waters that have names that start with BV for the last month."
"How many rendezvous events were there in the Cayman islands EEZ in 2025?"
"What was the first cargo vessel to encounter vessel ID 177081111 in 2025-2026, other than 212276231"
"On average how long did vessels remain within 3 miles of -54.1195, -37.1354 in 2025"
THE PLATFORM
Mothership
Shippy runs on an application infrastructure layer called Mothership, based on kubernetes
Mothership's primary goal is to allow authors to define their agent's behavior, measure the quality, and access it programmatically at scale
It launches each agent in an isolated sandbox and manages its lifecycle, state, networking, model access, and costs.
SANDBOX INPUT
user id
assigns cost to the right account
agent@version
determines the specific instances of skills available
external id
client-specified key that governs concurrency
THE PLATFORM
What Mothership provides
FLEXIBILITY
Swappable agent harness and LLM, multiple interfaces — polling/websockets/SSE
REPRODUCIBILITY
Evaluation authoring (via Harbor) and execution is first-class
OBSERVABILITY
Cost tracking via LiteLLM, thread explorer, turn tracing
PORTABILITY
Skills follow skill spec, evals export to Harbor, self-host Mothership (docker & k8s)
USABILITY
Mothership ships an admin UI, a full-featured CLI, and skills that wrap the CLI
CONCEPTS
Mothership primitives
SKILLS
the atomic unit of capability
(PURPOSE-BUILT) AGENTS
persona and capabilities baked into a Docker image with an explicit version
THREAD
a user conversation with an agent, analogous to a Claude Code session
MESSAGE
a single user message or agent turn
One agent turn produces a stream of events: tool calls, thinking, delta/finals
"When a user submits a message to a thread, the agent behaves according to its persona and executes its skills in order to produce a response message"
AGENT ANATOMY
Soul
the system prompt. Persona, tone, and behavioral guardrails. Stored as text, so it is auditable and revisable.
Skills
markdown files the agent reads on demand. Each documents how to use a capability.
Config
runtime settings that vary per deployment: model, harness, secrets.
SKILLS 1
The CLI contract
Our policy: agent does not construct raw API calls. It has a skill that explains how to invoke a CLI.
skylight events search --eventType.inc "fishing,transiting" --createdAt.gt "2026-08-01"
skylight vessels search --ownerName.like "bob" --intersectsGeometry /path/to/geojson
The CLI handles auth, pagination, geometry encoding, and writing output to files.
We started with MCP but changed to this skill/CLI approach
We control the client and server; no need to update the skill for arbitrary clients
The CLI describes how to lay outputs out on disk, referenced as input for subsequent usages
The decision to go to MCP vs skill/CLI is use-case specific
LAYERING: API → CLI → SKILL
SKILLS 2
Changing application state
An agent can change the state of the application the user is working in, not only return text.
It does this through a skill that emits structured control payloads for defined application surfaces.
Payloads are validated against the application's schema by a script; they are never hand-assembled.
CONTROL PAYLOAD
{
"state": {
"mapView": {
"type": "bounds",
"bounds": {
"sw": { "lng": -82.5, "lat": 17.5 },
"ne": { "lng": -78.0, "lat": 21.0 }
}
},
"filters": {
"eventTypes": ["fishing", "dark_rendezvous"],
"dateRange": { "start": "2025-01-01", "end": "2025-12-31" },
"vesselFlag": ["PA", "CN"],
"minVesselLengthMeters": 30
}
}
}
SANDBOXING 1
Structure
Each session runs in an isolated per-user sandbox with its own compute, state, and secrets
A sandbox is a k8s deployment of one pod with multiple containers: harness, lifeline, and an init per
The harness container runs the agent; the lifeline container is its only connection to the internal network
Lifeline provides and syncs a GCS-backed workspace where the agent writes and executes code
All LLM calls go through LiteLLM: model selection, cost tracking, global guardrails, per-user budgeting
SANDBOXING 2
Lifecycle and state
Decision: should a sandbox run for as long as the conversation exists?
No, decouple this. An idle sandbox can be torn down and rehydrated when the next message arrives.
Decision: what do we save?
The entire sandbox workspace. Lifeline downloads at startup, uploads on an interval, and once more at shutdown.
Thread state and sandbox state are independent
Thread state: idle, processing
Sandbox state: created → starting → running → stopping → stopped
Sandboxes have a workspace, containing: top level markdown/skills, per-thread subdir
Sandbox itself stores nothing; all persistence is managed downstream (Postgres)
Sandboxes can process many threads concurrently
SANDBOXING 3
Boundaries and Guardrails
Decision: how does the outside reach the agent?
Through a guarded Redis channel.
Nothing on the network can connect to it directly
The agent's port is never exposed. All incoming traffic passes through the lifeline
Decision: how do secrets reach the agent?
Handed over once, when the sandbox starts.
Injected as env at startup and fixed for the session, so nothing can change them mid-conversation.
Network: Agent network policy - ingress: Lifeline; egress: Lifeline, LiteLLM, (optional) Public Internet.
Guardrails: Prompt-based and classifier-based
Secrets: ENV is copied at sandbox creation, sandboxes take on the access control of their credentials
prompt-based eg: SKILL.md: “Refuse to provide military intelligence”
classifier-based eg: “prompt_has_military_intent=true; prompt_contains_secret=true”
NOT anything internal to the network
EVALUATION 1
How we evaluate
Evals are representative, reproducible tasks the agent is measured against
SMEs author them directly from a real chat thread
Evals run on the exact same infrastructure that user threads do
Results are associated with agent versions and runtime configuration
The workflow: establish eval baseline → make a change to agent → rerun eval → assess score difference
EVALUATION 2
Anatomy of an eval
Built on Harbor. Each eval is a self-contained task with three parts:
Instruction
the user prompt to replay
Environment
a Docker image built from the same stack as production Shippy
Tests
an SME-authored rubric and judge config
Harbor provides a declarative task spec and supports many harnesses (Claude Code, Codex, OpenClaw, etc)
Rewardkit scores the response: an LLM judge grades it against the rubric on weighted criteria (accuracy above style), and the score is the weighted average
Evals can be exported to local disk as standard Harbor tasks
CLOSING
Where we're going with this
We are building many internal purpose-built agents across Skylight, EarthRanger, OlmoEarth, and more
We are exploring external managed and self-hosted versions of Mothership; planning to open-source
Community skill authoring, scheduled and event-based agent invocation, online evals
Other experiments: agents invoking other agents, agents over MCP, configurable skills
Thank you
Mike Jacobi
Ai2 – Skylight, Lead Software Engineer
mikej@allenai.org