About

Independent data infrastructure for AI.

Public web data has been consumed. The next gains come from records that exist only inside institutions: chart notes, code reviews, audit trails, recordings of people doing their jobs. DataOrigin exists to license that data lawfully and make it usable for training and evaluation.

20182021202420272030crossoverTraining demandUsable public textLicensed from institutionsUsable public text vs. training demand · illustrative, indexed

Illustrative, indexed. Training demand keeps compounding; the usable public corpus does not.

01 · Why we exist

Three premises

Premise 01 · Supply

The next differentiator is the distribution, not the compute.

Every lab can buy the same chips and read the same web. What separates models now is the distribution they were trained on, and the most valuable distributions, how medicine, engineering, finance and operations are actually done, sit in institutional archives that have never been licensed.

Premise 02 · Provenance

A dataset without provenance is a liability.

We license from the owner, log every transformation, and write scope, term and revocation into the contract, so that for any record a lab can answer two questions: where did this come from, and are we permitted to use it?

Premise 03 · Judgement

Quality is a person, not a process.

The right answer in a chart note or a code review is not a majority vote. We admit under 2% of applicants, confirm credentials with the issuer, and re-calibrate every expert weekly, because a rubric is only as good as the person who wrote it.

02 · Operating principles

Four rules, no exceptions

R-01

Nothing scraped

Every record enters the catalog under a licence signed by the institution that owns it. If we cannot name the source, it is not delivered.

R-02

Owners set the terms

Scope, exclusivity, term and revocation are decided by the data owner and enforced after delivery. We do not resell outside the agreed scope.

R-03

De-identify, log, then audit

Identifiers are removed and every redaction is logged. Residual risk is reviewed by a person, and 42% of ingested records are rejected.

R-04

Practitioners are paid

Trials are paid, work is paid on time, and re-calibration is transparent. People at the top of their field do not work for exposure.

03 · Where we sit

Between the institution and the lab

Institutions hold the data and cannot train on it. Labs can train and cannot reach the data. DataOrigin is the layer between them, and stays independent of both.

Institutions

Hospitals, enterprises, studios and engineering orgs that generate valuable data as a by-product of their work.

DataOrigin · licensed datasets

Licensing, de-identification, normalisation, practitioner review, provenance and packaging.

DataOrigin · environments & evaluations

Working replicas of production systems, real tasks and practitioner grading, versioned together.

Frontier labs & AI teams

Labs, enterprise AI teams and sovereign programmes that train, evaluate and deploy models.

04 · The company

Independent by design

DataOrigin is an independent data infrastructure company. We do not train models and we do not compete with the institutions we license from. That is what lets both sides trust us with the middle: institutions license records they have never released, and labs receive datasets with provenance they can audit. The expert network is global and remote, staffed by practitioners who are verified, examined and paid, wherever they work.

ModelIndependent · no in-house model training
FocusLicensed data · Agent environments · Expert network
CustomersFrontier labs, enterprise AI teams, sovereign AI programmes
dataorigin · operating principlesStanding
01License before scalesigned
02Log every transformationprovenance
03Reject freely42% rejected
04Practitioners define qualityunder 2% admitted
05Owners keep controlrevocable
06Label what is illustrativepre-launch

We are pre-launch. Figures across this site are targets, not audited results, and no customer, partner or certification is implied.

For AI labs and enterprise teams

Training or evaluating a model?

Tell us the capability gap. If the catalog covers it, you will have samples under NDA within days.

Request samples

For data owners

Holding data that could train models?

License it on your terms. We handle de-identification, compliance and delivery; you keep control of scope, term and exclusivity.

License your data