Agent environments
Environments built from production systems.
Working replicas of the software agents will operate in deployment: helpdesks, order desks, EHRs, repositories with CI. Each is seeded with de-identified real state, given tasks written from real tickets, and every run is graded by a practitioner against a published rubric.
01 · Anatomy
Three components, versioned together
State the agent observes, actions it can take, and a rubric written by someone who does the job. All three ship as one release, so a score is comparable across time and across models.
State with consequences
The agent sees what a human operator sees, and every action changes state. Grant the wrong permission and the directory records it.
Rubrics written by operators
Rubrics are written and applied by practitioners from the domain, not by string matching. Weights are published with every evaluation.
Versioned for comparability
Environment, tasks and rubric ship as one version. Re-run it in six months, or on a different model, and the score means the same thing.
02 · Evaluations
Public benchmarks overstate readiness
The same models, scored on a public suite and then on our evaluations. The gap is the distance between passing a test and doing the job. Figures are illustrative pre-launch targets.
03 · Environments
Six production systems, replicated
Each is a working replica of software a partner runs daily, with identities removed and behaviour intact. We build replicas of your own systems the same way.
Order desk & payments
The replica includes
Storefront, gateway with retries and webhooks, ledger, refunds, 12k seeded orders across 3 currencies.
Example tasks
Reverse a duplicated payment · match a settlement file · replay a dropped gateway callback
CRM & helpdesk
The replica includes
40 object types, ticket queues, email sync, calendar, 2 years of de-identified activity history.
Example tasks
Restore access after a team move · dedupe a merged contact · escalate a stalled deal
Clinic EHR
The replica includes
Order entry, results, medication reconciliation, discharge workflow, de-identified with clinical safeguards intact.
Example tasks
Reconcile a medication list · flag a contraindication · draft a discharge summary
Repository & CI
The replica includes
Editor, repository, test runner, CI with flaky-test history, code review with policy checks.
Example tasks
Fix a failing build · implement a scoped ticket · review a PR against policy
Warehouse floor
The replica includes
WMS with inventory, pick paths, exceptions and egocentric-video task references.
Example tasks
Resolve a stock discrepancy · re-route a delayed pick · escalate a damaged item
Finance back office
The replica includes
GL, AP/AR, bank feeds, approval chains, audit trail, month-end close checklist.
Example tasks
Match unreconciled payments · investigate a variance · prepare a close checklist
04 · Method
How an environment is built
Snapshot
A de-identified snapshot of the partner's production state, taken under the same licence terms as a dataset.
Stub
Surrounding services are stubbed to production behaviour: latency, retries and failure modes included.
Author
Practitioners from the domain write tasks from real tickets and incidents, each with its own rubric and weights.
Grade
Every run is graded against the rubric by a practitioner; environment, tasks and rubric ship as one version.
Illustrative curve. Public task sets stop producing training signal early; production replicas keep producing it.
05 · Delivery
Three ways to use them
Evaluate
Score a model · one-off or continuous
Submit a model or an agent harness. We run the full task set, grade it, and return per-task traces and a versioned score.
Train
RL · online or from graded trajectories
Use the environments directly as a reward source, or license graded trajectory sets for offline training.
Replicate your system
Your software · your tasks · your rubric
We replicate your own system and write tasks from your real backlog, so the agent you ship has been tested where it will run.