1 of 17

PURPOSE-BUILT AGENTS

How we build agents

Our approach to building, evaluating, and operating agents around an organization's domain expertise.

Mike Jacobi

Ai2 – Skylight, Lead Software Engineer

mikej@allenai.org

2 of 17

What we mean by agent

GENERAL-PURPOSE

Claude Code, Codex, OpenCode, OpenClaw, etc for arbitrary tasks

I will be referring to this as an "agent harness"

PURPOSE-BUILT

An agent with an intentional persona and a set of skills for specific tasks

Runs on an agent harness

Why use purpose-built over general-purpose?

This is a mechanism to provide an expert-curated experience for a targeted audience

You, the domain-expert agent author, know how to navigate your own world, your purpose-built agent distributes your knowledge to the users you want to engage

This talk is about purpose-built agents

3 of 17

THE PRODUCT

Skylight

Skylight is a maritime domain awareness tool, set on a filterable map interface

It is used in an operational context to reduce IUU fishing (Illegal, Unreported, Unregulated)

Typical users are analysts working at Navies, Coast Guards, Fisheries, and more

Analysts act on the output, often to inform where to send their enforcement vessels

4 of 17

THE AGENT

Shippy

Shippy is an agent that allows you to interact with Skylight data, in prod today

Shippy runs on OpenClaw with Skylight-specific skills

queries our public API

writes Python code

rejects inappropriate prompts

says "I'm not sure"

adds domain-specific caveats

provides data-source attribution

reads user-uploaded files

generates downloadable files

SAMPLE QUERIES

"Show me all vessels in southern Vietnam waters that have names that start with BV for the last month."

"How many rendezvous events were there in the Cayman islands EEZ in 2025?"

"What was the first cargo vessel to encounter vessel ID 177081111 in 2025-2026, other than 212276231"

"On average how long did vessels remain within 3 miles of -54.1195, -37.1354 in 2025"

5 of 17

THE PLATFORM

Mothership

Shippy runs on an application infrastructure layer called Mothership, based on kubernetes

Mothership's primary goal is to allow authors to define their agent's behavior, measure the quality, and access it programmatically at scale

It launches each agent in an isolated sandbox and manages its lifecycle, state, networking, model access, and costs.

SANDBOX INPUT

user id

assigns cost to the right account

agent@version

determines the specific instances of skills available

external id

client-specified key that governs concurrency

6 of 17

THE PLATFORM

What Mothership provides

FLEXIBILITY

Swappable agent harness and LLM, multiple interfaces — polling/websockets/SSE

REPRODUCIBILITY

Evaluation authoring (via Harbor) and execution is first-class

OBSERVABILITY

Cost tracking via LiteLLM, thread explorer, turn tracing

PORTABILITY

Skills follow skill spec, evals export to Harbor, self-host Mothership (docker & k8s)

USABILITY

Mothership ships an admin UI, a full-featured CLI, and skills that wrap the CLI

7 of 17

CONCEPTS

Mothership primitives

SKILLS

the atomic unit of capability

(PURPOSE-BUILT) AGENTS

persona and capabilities baked into a Docker image with an explicit version

THREAD

a user conversation with an agent, analogous to a Claude Code session

MESSAGE

a single user message or agent turn

One agent turn produces a stream of events: tool calls, thinking, delta/finals

"When a user submits a message to a thread, the agent behaves according to its persona and executes its skills in order to produce a response message"

AGENT ANATOMY

Soul

the system prompt. Persona, tone, and behavioral guardrails. Stored as text, so it is auditable and revisable.

Skills

markdown files the agent reads on demand. Each documents how to use a capability.

Config

runtime settings that vary per deployment: model, harness, secrets.

8 of 17

SKILLS 1

The CLI contract

Our policy: agent does not construct raw API calls. It has a skill that explains how to invoke a CLI.

skylight events search --eventType.inc "fishing,transiting" --createdAt.gt "2026-08-01"

skylight vessels search --ownerName.like "bob" --intersectsGeometry /path/to/geojson

The CLI handles auth, pagination, geometry encoding, and writing output to files.

We started with MCP but changed to this skill/CLI approach

We control the client and server; no need to update the skill for arbitrary clients

The CLI describes how to lay outputs out on disk, referenced as input for subsequent usages

The decision to go to MCP vs skill/CLI is use-case specific

LAYERING: API → CLI → SKILL

9 of 17

SKILLS 2

Changing application state

An agent can change the state of the application the user is working in, not only return text.

It does this through a skill that emits structured control payloads for defined application surfaces.

Payloads are validated against the application's schema by a script; they are never hand-assembled.

CONTROL PAYLOAD

{

"state": {

"mapView": {

"type": "bounds",

"bounds": {

"sw": { "lng": -82.5, "lat": 17.5 },

"ne": { "lng": -78.0, "lat": 21.0 }

}

},

"filters": {

"eventTypes": ["fishing", "dark_rendezvous"],

"dateRange": { "start": "2025-01-01", "end": "2025-12-31" },

"vesselFlag": ["PA", "CN"],

"minVesselLengthMeters": 30

}

}

}

10 of 17

SANDBOXING 1

Structure

Each session runs in an isolated per-user sandbox with its own compute, state, and secrets

A sandbox is a k8s deployment of one pod with multiple containers: harness, lifeline, and an init per

The harness container runs the agent; the lifeline container is its only connection to the internal network

Lifeline provides and syncs a GCS-backed workspace where the agent writes and executes code

All LLM calls go through LiteLLM: model selection, cost tracking, global guardrails, per-user budgeting

11 of 17

SANDBOXING 2

Lifecycle and state

Decision: should a sandbox run for as long as the conversation exists?

No, decouple this. An idle sandbox can be torn down and rehydrated when the next message arrives.

Decision: what do we save?

The entire sandbox workspace. Lifeline downloads at startup, uploads on an interval, and once more at shutdown.

Thread state and sandbox state are independent

Thread state: idle, processing

Sandbox state: created → starting → running → stopping → stopped

Sandboxes have a workspace, containing: top level markdown/skills, per-thread subdir

Sandbox itself stores nothing; all persistence is managed downstream (Postgres)

Sandboxes can process many threads concurrently

12 of 17

SANDBOXING 3

Boundaries and Guardrails

Decision: how does the outside reach the agent?

Through a guarded Redis channel.

Nothing on the network can connect to it directly

The agent's port is never exposed. All incoming traffic passes through the lifeline

Decision: how do secrets reach the agent?

Handed over once, when the sandbox starts.

Injected as env at startup and fixed for the session, so nothing can change them mid-conversation.

Network: Agent network policy - ingress: Lifeline; egress: Lifeline, LiteLLM, (optional) Public Internet.

Guardrails: Prompt-based and classifier-based

Secrets: ENV is copied at sandbox creation, sandboxes take on the access control of their credentials

prompt-based eg: SKILL.md: “Refuse to provide military intelligence”

classifier-based eg: “prompt_has_military_intent=true; prompt_contains_secret=true”

NOT anything internal to the network

​

​

13 of 17

EVALUATION 1

How we evaluate

Evals are representative, reproducible tasks the agent is measured against

SMEs author them directly from a real chat thread

Evals run on the exact same infrastructure that user threads do

Results are associated with agent versions and runtime configuration

The workflow: establish eval baseline → make a change to agent → rerun eval → assess score difference

14 of 17

15 of 17

EVALUATION 2

Anatomy of an eval

Built on Harbor. Each eval is a self-contained task with three parts:

Instruction

the user prompt to replay

Environment

a Docker image built from the same stack as production Shippy

Tests

an SME-authored rubric and judge config

Harbor provides a declarative task spec and supports many harnesses (Claude Code, Codex, OpenClaw, etc)

Rewardkit scores the response: an LLM judge grades it against the rubric on weighted criteria (accuracy above style), and the score is the weighted average

Evals can be exported to local disk as standard Harbor tasks

16 of 17

CLOSING

Where we're going with this

We are building many internal purpose-built agents across Skylight, EarthRanger, OlmoEarth, and more

We are exploring external managed and self-hosted versions of Mothership; planning to open-source

Community skill authoring, scheduled and event-based agent invocation, online evals

Other experiments: agents invoking other agents, agents over MCP, configurable skills

17 of 17

Thank you

Mike Jacobi

Ai2 – Skylight, Lead Software Engineer

mikej@allenai.org