Agent environments

Environments built from production systems.

Working replicas of the software agents will operate in deployment: helpdesks, order desks, EHRs, repositories with CI. Each is seeded with de-identified real state, given tasks written from real tickets, and every run is graded by a practitioner against a published rubric.

episode runner · env: helpdesk-replica-02 · task: restore_accessRunning
$ origin run --env helpdesk-replica-02 --task restore_access --seed 2210[00.00] env ready · directory snapshot 2026-06-02 · 4,180 users · 62 groups · identities pseudonymised[00.01] task "An analyst lost access to the ledger app after moving teams. Restore what they need, nothing more."[00.03] agent → lookup_user("u-4471") · role: analyst · team: FP&A · moved 2 days ago[00.05] agent → diff_groups(previous, current) · missing: ledger-read, ledger-export[00.08] agent → check_policy("ledger-export") · requires manager approval[00.10] agent → grant("ledger-read") · ✓ applied[00.11] agent → request_approval("ledger-export", approver="mgr-0917") · ✓ ticket opened[00.12] agent → reply(user) · draft: what was restored, what is pending, why[00.13] episode complete · 6 actions · 13 s · grading…[grade] outcome ✓ access restored · least privilege ✓ · approval path ✓ · clarity 0.8 → score 0.91[grade] reviewer DO-EXP-0412 · rubric v2 · accepted into ORIGIN·DESK-1 / flow 087

01 · Anatomy

Three components, versioned together

State the agent observes, actions it can take, and a rubric written by someone who does the job. All three ship as one release, so a score is comparable across time and across models.

state · actions
ObserveActGrade

State with consequences

The agent sees what a human operator sees, and every action changes state. Grant the wrong permission and the directory records it.

rubric v2 · written by an ops lead
✓Outcome achieved0.40
✓Least privilege, no side-effects0.25
✓Approval path followed0.20
~Clear to the user0.15

Rubrics written by operators

Rubrics are written and applied by practitioners from the domain, not by string matching. Weights are published with every evaluation.

scoreboard · v2
cand-A
64%
cand-B
58%
cand-C
41%
baseline
23%

Versioned for comparability

Environment, tasks and rubric ship as one version. Re-run it in six months, or on a different model, and the score means the same thing.

02 · Evaluations

Public benchmarks overstate readiness

The same models, scored on a public suite and then on our evaluations. The gap is the distance between passing a test and doing the job. Figures are illustrative pre-launch targets.

ORIGIN·CODE-1
Software engineering · 220 tasks · licensed monorepos with live CI
Same models · public benchmark86–92%
Same models · this evaluation53–61%
220 tasks · gap to public benchmark−31 pts
ORIGIN·DESK-1
Operations · 160 flows · helpdesk, claims, order desk
Same models · public benchmark84–90%
Same models · this evaluation46–56%
160 flows · gap to public benchmark−34 pts
ORIGIN·CHART-1
Clinical reasoning · 130 cases · de-identified longitudinal charts
Same models · public benchmark88–94%
Same models · this evaluation57–66%
130 cases · gap to public benchmark−28 pts

03 · Environments

Six production systems, replicated

Each is a working replica of software a partner runs daily, with identities removed and behaviour intact. We build replicas of your own systems the same way.

E-01

Order desk & payments

The replica includes

Storefront, gateway with retries and webhooks, ledger, refunds, 12k seeded orders across 3 currencies.

Example tasks

Reverse a duplicated payment · match a settlement file · replay a dropped gateway callback

E-02

CRM & helpdesk

The replica includes

40 object types, ticket queues, email sync, calendar, 2 years of de-identified activity history.

Example tasks

Restore access after a team move · dedupe a merged contact · escalate a stalled deal

E-03

Clinic EHR

The replica includes

Order entry, results, medication reconciliation, discharge workflow, de-identified with clinical safeguards intact.

Example tasks

Reconcile a medication list · flag a contraindication · draft a discharge summary

E-04

Repository & CI

The replica includes

Editor, repository, test runner, CI with flaky-test history, code review with policy checks.

Example tasks

Fix a failing build · implement a scoped ticket · review a PR against policy

E-05

Warehouse floor

The replica includes

WMS with inventory, pick paths, exceptions and egocentric-video task references.

Example tasks

Resolve a stock discrepancy · re-route a delayed pick · escalate a damaged item

E-06

Finance back office

The replica includes

GL, AP/AR, bank feeds, approval chains, audit trail, month-end close checklist.

Example tasks

Match unreconciled payments · investigate a variance · prepare a close checklist

04 · Method

How an environment is built

01

Snapshot

A de-identified snapshot of the partner's production state, taken under the same licence terms as a dataset.

02

Stub

Surrounding services are stubbed to production behaviour: latency, retries and failure modes included.

03

Author

Practitioners from the domain write tasks from real tickets and incidents, each with its own rubric and weights.

04

Grade

Every run is graded against the rubric by a practitioner; environment, tasks and rubric ship as one version.

25%50%75%100%02.5k5k7.5k10kPublic task sets onlyWith DataOrigin environmentsSuccess on held-out production tasks vs. training episodes · illustrative

Illustrative curve. Public task sets stop producing training signal early; production replicas keep producing it.

05 · Delivery

Three ways to use them

D-01

Evaluate

Score a model · one-off or continuous

Submit a model or an agent harness. We run the full task set, grade it, and return per-task traces and a versioned score.

D-02

Train

RL · online or from graded trajectories

Use the environments directly as a reward source, or license graded trajectory sets for offline training.

D-03

Replicate your system

Your software · your tasks · your rubric

We replicate your own system and write tasks from your real backlog, so the agent you ship has been tested where it will run.