Industry · Frontier AI labs

Frontier models need data nobody has published yet.

The public web has been read. What moves a frontier model now is new, human-made work from people qualified to judge it, licensed clearly enough that your counsel can sign off.

Photo needed · 21:9Researchers reviewing model evaluation resultsWide, daylight-filled research office with two people discussing results on monitors. Natural, unposed and calm, with clean desks and soft window light.Wide image under the opening
Overview

Frontier AI labs, with proof.

Photo needed · 4:5Mathematician writing proof on whiteboardPortrait close-up of a hand mid-stroke on a whiteboard covered in clear handwritten mathematics. Bright, sharp, editorial.Beside the overview

Frontier labs are rarely short of tokens. They are short of the right ones: graduate-level reasoning traces, preference judgments from people who can tell a correct proof from a confident one, and evaluation sets that have never touched the open internet. HUMXN sources that work from a credential-checked expert network and a registered catalogue of consenting creators, then verifies every item before it reaches your pipeline. You describe the capability gap. We write a dataset spec against it and deliver a pilot batch first.

Contamination is the quiet failure of modern benchmarks. If a test item exists anywhere public, you cannot be sure your model is reasoning rather than remembering. Held-out sets commissioned through HUMXN are written for you, fingerprinted at intake and checked for duplicates against everything already registered. Exclusive licences keep them yours. Each item carries an Ed25519-signed provenance record and sits in a hash-chained audit trail, so you can show later exactly what was delivered and when.

Rights matter more each quarter. Every work in the catalogue records its AI-training permission, licence terms and any model or property releases, and records whether it was human-created, AI-assisted or AI-generated. Generator signatures in EXIF and XMP, IPTC digital-source-type fields and C2PA records are detected at intake. That gives your data governance and legal teams a documented chain from creator to checkpoint, designed to support the disclosure and documentation work that rules such as the EU AI Act ask of model developers.

Data types
Reasoning tracesPreference pairsEvaluation itemsLong-form textCodeImagesAudio and speechVideo
Experts involved
MathematiciansPhysiciansAttorneysSoftware engineersChemistsFinancial analystsLinguists
What we deliver

Built to your specification.

Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for frontier ai labs.

01

Supervised fine-tuning demonstrations

Expert-written answers and step-by-step reasoning traces in mathematics, science, law, medicine and code, built to your rubric and reviewed by a second specialist.

02

Preference and RLHF comparisons

Pairwise and rubric-graded judgments from people qualified to assess correctness, not only fluency, with written rationales when your reward modeling needs them.

03

Held-out evaluation sets

Never-published test items commissioned for you, fingerprinted at intake and offered under exclusive licence so they stay out of future training corpora.

04

Red-teaming and safety prompts

Adversarial prompts and graded responses written by specialists in the domains where a wrong answer carries real risk, screened before delivery.

05

Licensed multimodal corpora

Human-made images, audio, video and documents with training permission and releases on file, for pre-training or mid-training where your model is thin.

06

Standing orders for ongoing supply

As capabilities shift, standing orders keep new, matching work arriving from the same network at the same review bar, with volume pricing.

Why it matters

Where unverified data falls short.

The nine layers of verification

01Benchmark contamination

Public test sets leak into training data. Commissioned, fingerprinted and exclusively licensed items give you evaluations you can trust to measure capability.

02Shallow preference signal

Raters who cannot verify an answer reward the most confident one. Credential-checked experts judge substance, and a second expert reviews the hard calls.

03Synthetic data drift

Training on model output compounds its errors. Human-origin checks and generator-signature detection keep machine-made work labeled or out of the set.

04Documented rights

Training permission, licence terms and releases are bound to each item, with a signed receipt listing every item by ID and hash.

Questions

Frontier AI labs: asked often.

How do you keep evaluation sets out of public training data?

Items are written for you rather than collected from the web. Each one is fingerprinted at intake and checked against everything already registered, so duplicates and near-duplicates surface before delivery. Exclusive licences mean the items are not offered to anyone else. The signed receipt lists every item by ID and hash, which lets you later check whether any of them has appeared somewhere it should not.

Who writes and grades the data?

Domain experts in our network, such as mathematicians, physicians, attorneys and software engineers. Every expert is credential-checked before taking work, and every label records its source: creator, AI-suggested or expert-reviewed. Difficult or high-stakes items go through a second-expert review pass. You set the rubric and acceptance criteria in the dataset spec, and we hold production to the bar agreed after the pilot.

What drives the price of a frontier dataset?

Mainly the expertise required, the time each item takes to write or grade, the review depth and the licence. A doctoral-level reasoning trace with a second-expert review costs more than a short preference judgment. Exclusive licences cost more than non-exclusive ones. Volume and standing orders bring the unit price down. We quote after the scoping call, once the spec is written.

Can we use the data for commercial model training?

Yes, within the licence you agree. Each item carries its AI-training permission and licence terms, plus any model or property releases. Works showing identifiable people are only offered for training when consent is on file. Your counsel should review the licence against your intended use; we provide the documentation, not legal advice.

What formats do you deliver in?

JSON Lines, CSV, Parquet, WebDataset, COCO captions and Croissant 1.0 metadata. Most frontier teams take JSON Lines or Parquet for text and WebDataset for multimodal work. Approved buyers can also pull deliveries through the API. Each delivery comes with a signed receipt and per-item provenance records.

How quickly can a project start?

After a short scoping call we write a dataset spec, usually within days. A pilot batch follows so you can test the data in your own pipeline before committing to volume. Timelines for production depend on how specialized the experts are and how long each item takes, and we describe them in the spec rather than promising a fixed number upfront.

Related

Tell us what your model needs to learn.

Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.