Production repositories with review threads, CI history and incident records attached
Licensed datasets · Agent environments · Expert network
Building training and data funnels for frontier AI.
Every funnel starts with a signed licence from the institution that produced the data and ends with a dataset ready for training, an agent environment or a private evaluation. In between: de-identification, review by practising specialists, and a provenance record for every item delivered.
Licensed sources → de-identification and practitioner review → delivered assets
Licensed · De-identified · Practitioner-reviewed
Coverage
Six verticals, licensed at the source
Ticket lifecycles, CRM activity and procurement threads with outcomes recorded
De-identified notes, imaging and orders with the clinical reasoning intact
Multi-speaker recordings of real work, transcribed, diarised and quality-checked
First-person footage of skilled manual work with step labels and ambient audio
Working papers, contracts and filings reviewed by chartered accountants and counsel
// Built for frontier labs · enterprise AI teams · sovereign AI programmes
01 · Why licensed data
Public web data no longer separates one model from another. The records that would, clinical decisions, code review, claims handling, financial work, were never published. They sit inside institutions, and reaching them takes a licence, not a crawler. That is the business DataOrigin is in.
Institutional records cover the tasks models fail at today: diagnosis, code review, claims handling, reconciliation.
Public text keeps the conclusion and drops the reasoning. Records authored by practitioners keep both.
A benchmark score is a claim. A graded run inside a replica of the production system is a measurement.
02 · Deliverables
What a delivery contains
Three product lines leave the pipeline: licensed datasets, agent environments and expert output. Each ships with the documentation a lab's legal, security and data teams will ask for.
01 · Licensed datasets
{ "id": "do-8f21-0412", "licence": "DO-L-2026-018", "pii_residual": 0, "reviewed_by": "DO-EXP-1142" }
JSONL · Parquet · signed manifest
02 · Agent environments
- production state snapshot, identities removed
- service stubs with production latency and failure modes
- seeded accounts, roles and approvers
- ✓ Reverse a duplicated payment
- ✓ Match an unpaid invoice
- Escalate a flagged claim
03 · Expert network · applied at every stage
03 · How we work
Three product lines, one standard
Engage on any one of the three or all together. Each ends with the same provenance record and the same practitioner sign-off.
01 · Licensed datasets
License, de-identify, review.
Data licensed directly from its owner, de-identified to an auditable standard, and reviewed by practitioners in the domain before delivery.
License
A licence or revenue-share agreement with the institution that owns the records, with scope, term and permitted use written in.
De-identify
Direct and quasi-identifiers removed, formats normalised, and every transformation logged against the record.
Review
Practitioners review the corpus before delivery. On a typical run, 42% of ingested records are rejected.
02 · Agent environments
Replicate the system, run the task.
A working copy of the software a job runs on, seeded with de-identified production state, with tasks written from real tickets and every run scored by a practitioner.
Replicate
A de-identified snapshot of production state, with the surrounding services stubbed to production behaviour.
Run
Agents attempt real tasks with the edge cases, approval steps and failure modes left in.
Score
Practitioners grade each run against a published rubric. Environment, tasks and rubric ship as one version.
03 · Expert network
Admit under 2%. Re-test weekly.
Practising physicians, engineers, accountants and counsel author training data and grade model output.
Verify
Credentials confirmed with the issuing body, a timed exam set by the network, then a paid trial graded blind.
Calibrate
Every expert grades a seeded gold set each week. Drift is retrained; persistent drift is removed.
Deploy
Authoring, grading, preference ranking and red-teaming, in your tooling or ours.
Provenance & compliance
Built to pass a data owner's legal review.
The most valuable data is the most regulated. DataOrigin is structured so that a hospital, a bank or an engineering organisation can license to us without lowering its standards, and so that a lab can audit exactly what it received.
De-identification before transfer
Direct and quasi-identifiers are removed inside the owner's environment. Each redaction is logged and residual risk is reviewed by a person before a record is packaged.
A licence behind every record
Each record traces to a licence ID, a consent basis and a permitted use, all written into the manifest. Nothing scraped, nothing of uncertain origin.
Owner-defined terms, enforced after delivery
Scope, exclusivity, term and revocation are set by the data owner and enforced contractually. We do not resell outside the agreed scope.
Auditable end to end
Encrypted pipelines, per-access logs and audit rights in the contract, so both the owner and the licensee can see who did what to the data, and when.
For AI labs and enterprise teams
Training or evaluating a model?
Tell us the capability gap. If the catalog covers it, you will have samples under NDA within days.
Request samplesFor data owners
Holding data that could train models?
License it on your terms. We handle de-identification, compliance and delivery; you keep control of scope, term and exclusivity.
License your dataNetwork
Sourced worldwide. Reviewed by practitioners worldwide.
Contact
Talk to us.
AI teams: tell us the capability gap and we will send samples under NDA. Data owners: tell us what you hold and we will walk you through licensing, de-identification and revenue share.