1 of 11

Orrey

A Proving Ground for Drone Swarm Doctrine

2 of 11

Our Team

John Apessos

Mechanical Engineering

Kritanu Saha

Economics & History

David Diao

Public Policy

James Xiao

Mechanical Engineering & Computer Science

3 of 11

Orchestration is Becoming the Weapon System

Value is moving from the platforms to the layer that coordinates them, and no one currently owns that layer.

Key Insights:

  • Interoperability is the future of warfare
    • Operation jailbreak indicates interoperability is a near future capability which must be managed

  • Orchestration is a public good
    • While the software layer is pivotal, the software enabling vendor-neutral orchestration is less profitable than building individual systems

  • Software is becoming the weapon system
    • With NSPM-11’s requirement of an update to DoW’s 3000.09 definition of autonomous weapons, software will likely deal with the burden of responsibility and auditability

Orchestration is vendor-neutral

Sources: CSIS, DIU.mil, DoW Directives, nso.nato,

Orchestration is underdeveloped

Auditable orchestration is novel

4 of 11

An Integrated Framework

We believe war in the future will be defined by the intersection of robust engineering, military doctrine, and policy insight.

New Doctrine

  • Lt Gen. Victor Krulak’s development of helicopter doctrine in Korea as precedent
  • How do we anticipate the new way of war without data or direct experience?

Cognitive Bias

  • Technology matters less than effect, tools are only used if they easier or better
  • Hesitancy is the rate limiting factor of swarms, thus base trust is paramount

Systems Thinking

  • Systems approaches at the strategic level provide greater impact than tactical solutions
  • Systems must accommodate for accelerating technological growth and remain flexible

True Agency

  • Agency inherently involves responsibility and thus transparency is paramount
  • Understanding technology enables users to be empowered in their choices

5 of 11

Introducing Orrey

Orrey aims to give battlespace leaders the ability to experiment with and learn from mission scenarios under their control.

6 of 11

Pre vs Post Training Comparison

Over multiple iterations, Orrey refines drone swarm strategies to adapt to hostile environments and available resources.

Pre-Training Iteration

Post-Training Iteration

7 of 11

LLMs in looped iterative learning

A structured learning loop: the model writes the policy in code, the simulator scores it, and every change arrives with the reasoning that produced it.

01 · POLICY

Deterministic policy, written as explicit Python Code, not weights.

02 · SIMULATION

Executes the policy and returns metrics: objective, assets lost, time, mesh integrity.

03 · LLM AGENT

Reasons over the metrics and writes the next policy, articulating why the last one failed and the new hypothesis.

04 · REASONING LOG

Every change stored with its rationale, versioned alongside the code it produced.

Each iteration is fed the full history of policies, parameters, and results

WHY NOT REINFORCEMENT LEARNING

  • RL returns a black box of weights that can’t be directly verified and can’t be read at the doctrinal level.
  • A programmatic policy is inspectable line by line, so insight can be extracted, argued with, and reused.

WHAT THIS BUYS

  • Output is deterministic and generalizable
  • Framework applies broadly across domains
  • Articulated log of every change and its reasoning
  • Observable evolution of policy and of the code itself

Orrey implements a unique approach that utilizes LLM reasoning to communicate with a human decision maker.

8 of 11

Challenging Stable Assumptions

We search doctrine space with a language model and let physics grade it. The human has final authority over implemented doctrine.

OUTSIDE THE LOOP · THE HUMAN IS THE JUDGE

INSIDE THE LOOP · THE SIMULATOR IS THE JUDGE

LLM · STRATEGIST

Proposes the next formation or role allocation, and explains in language why the last one failed.

SIMULATOR · JUDGE

Ground truth from sim: objective met, assets lost, time, mesh held. The fitness function, fixed before any run.

hypothesis

verdict

PERSISTENT SEARCH LOG

Every formation, parameter, and result is fed back each iteration. Without it, this is hill climbing with amnesia, not search.

OUTPUT · CANDIDATE DOCTRINE

Not doctrine. A candidate, plus the rationale that generated it and the verdict that kept or killed it.

COMMANDER · THE JUDGE

Adopts, rejects, or bounds it. The machine never holds authority: integration produces capability, nothing produces authority.

9 of 11

The Audit Trail

Both plaintext reasoning records and direct python code decisions are logged and verified

10 of 11

Implementation and Feasibility

Adoption Rollout Plan

Phase 1: Prepare

Month

3

Phase 2: Pilot

Month

12

Phase 3: Scale

Month

24

Phase 4: Expand

Month

36

Customer and scope set

  • Name the customer, funding owner, and contracting officer.
  • Scope one workflow: success criteria, environment, price.
  • Clear procurement: SAM, then SBIR/STTR, OT, or CSO.

Paid pilot underway

  • Win one paid prototype award with milestone payments.
  • Stand up the environment; settle hosting and security.
  • Capture evidence: reproducibility, traceability, V&V.

Initial deployment live

  • Turn the initial deployment into funded continued use.
  • Add recurring licences; price integration apart.
  • Standardize delivery: versioned models, install, SBOM.

Three teams, renewal

  • Close a third paying team on deployment evidence.
  • Secure funded renewals: access, support, deployments.
  • Compile the evidence pack: revenue, renewals, cost.

Dates and customer counts are proposed planning targets, not government commitments. Sources: DoDI 5000.61; DoDI 5000.97; Mar 2025 software acquisition directive; Jul 2025 Software Engineering Guide; Nov 2025 Acquisition Transformation Strategy; Jan 2026 AI Strategy and DETECT; AFWERX; DIU; SAM.gov; DFARS.

COST

  • No airframes, ranges, or flight crews.
  • Cloud batch runs; on-prem for classified work.
  • License plus usage, not cost per sortie.

INTEGRATION

  • Customers retain their autonomy stack.
  • APIs ingest agent models and policies.
  • Scenario, simulation, metrics, and replay.

REGULATORY

  • No airspace or flight-safety approval for simulation.
  • Screen strategies before field trials.
  • Access controls, audit logs, encrypted data.

WHO BUYS

  • First: U.S. Department of War.
  • Next: autonomy primes and robotics OEMs.
  • Enterprise licenses and integration services.

A 36-month rollout from one paid pilot to paying DoW programs, and what it takes to get there.

11 of 11

Orrey

Testable Swarm Doctrine | Judgement Driven Results