Solution · Custom datasets

Custom datasets, built to your spec.

Bring us the capability you are trying to build and the gaps you are seeing. We write the spec with you, source from consenting creators and vetted experts, and verify every item before it ships.

Photo needed · 21:9Data lead reviewing dataset specTwo people at a large daylight-lit table reviewing a printed specification and a laptop showing sample records. Calm, focused, editorial.Wide image under the opening
Overview

Custom datasets, with proof.

Photo needed · 4:5Annotated sample sheet close-upClose portrait-format shot of a pen annotating a printed page of sample items beside a laptop. Natural light, shallow depth of field.Beside the overview

Off-the-shelf corpora rarely match the distribution your model actually needs. A custom dataset starts from your failure cases: the domains, formats, edge conditions and label definitions that matter for the next training run. We turn that into a written dataset specification covering modality, volume, schema, acceptance criteria and licence terms, so both sides agree on what a correct item looks like before any sourcing begins. That document then governs the pilot, production and any standing order that follows.

Sourcing draws on two channels. The registered catalogue holds existing human-made work whose creators have granted AI-training permission and set licence terms. The expert network creates new material on commission: demonstrations, recordings, documents, annotations and reviews from people who are credential-checked in the relevant field. Most custom datasets combine both, using catalogue material for breadth and commissioned work for the long tail of cases that no existing archive covers well.

Every item passes the same intake checks regardless of channel. We fingerprint it against everything registered to catch duplicates and conflicting claims, record whether it is human-created, AI-assisted or AI-generated, check for generator signatures and C2PA records, screen for unsafe content and personal data, and bind the consent and licence terms to the item. You receive a signed receipt listing each item by ID and hash, so the dataset you trained on can be identified exactly later.

Data types
TextImagesAudio & speechVideo3D & CADDocumentsTabular dataSensor recordings
Experts involved
Domain specialistsProfessional writersAnnotatorsPhotographersVoice talentEngineersSecond-pass reviewers
What we deliver

Built to your specification.

Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for custom datasets.

01

Written dataset specification

Modality, schema, label definitions, volume, acceptance criteria and licence scope, agreed before sourcing starts.

02

Catalogue sourcing

Existing human-made work from registered creators who have granted AI-training permission on stated terms.

03

Commissioned expert work

New items created to your spec by credential-checked specialists where no existing material fits.

04

Pilot batch

A small, representative batch you test in your own pipeline before anything scales.

05

Verified production

Full-volume delivery held to the pilot's acceptance criteria, with a second-expert review pass on labels.

06

Standing orders

Ongoing supply of matching items at the same bar as your model's needs shift.

Why it matters

Where unverified data falls short.

The nine layers of verification

01Distribution, not volume

The value of a custom set is in covering the cases your model gets wrong. That requires sourcing against a precise spec rather than scraping more of what is already well represented.

02Rights you can document

Each item carries its AI-training permission, licence terms and any model or property releases, so counsel can review what was licensed and on what basis.

03Known origin

Recording whether each work is human-created, AI-assisted or AI-generated lets you control how much synthetic content enters your training mix.

04Exact reproducibility

A signed receipt with every item's ID and hash means the dataset behind a model version can be reconstructed and audited.

Questions

Custom datasets: asked often.

How is a custom dataset priced?

Pricing depends on modality, volume, how much of the set can come from the existing catalogue versus new commissioned work, the seniority of the experts involved, the depth of annotation and review, and the licence scope. Exclusive licences cost more than non-exclusive ones. We quote against the written specification after the scoping call, and volume pricing applies to larger orders and standing orders.

How long does a custom dataset take?

Timing depends mostly on how much new material must be commissioned and how specialized the experts are. Catalogue-heavy sets move faster than sets that need rare domain expertise or new recordings. We agree a schedule in the specification, deliver a pilot batch first, and report progress against the spec during production so you can plan training runs around real delivery dates.

What formats do you deliver in?

We deliver JSON Lines, CSV, Parquet, WebDataset, COCO captions and Croissant 1.0 metadata. Each record carries its labels, technical metadata, origin classification, licence reference and file hashes. Approved buyers can also pull deliveries through the API. If your pipeline expects a particular schema, we specify it in the dataset specification and validate against it before delivery.

Can we license the dataset exclusively?

Yes. Exclusive licences are available for commissioned work, and for catalogue items where the creator has agreed to exclusive terms. Exclusivity is recorded on each item, so the same work is not offered to another buyer for the licensed scope. Exclusivity affects price and sometimes sourcing time, so we settle it during scoping rather than after the pilot.

How do you measure quality on a custom set?

Quality is measured against acceptance criteria written into the specification and tested on the pilot batch. Labels are marked by source, whether creator, AI-suggested or expert-reviewed, and a second expert reviews a share of items agreed in the spec. Items that fail criteria are replaced rather than delivered, and you approve the pilot item by item before production starts.

What happens to items that fail verification?

They are not delivered. Items that duplicate or conflict with registered work, lack required consent or releases, fail safety screening, or carry undisclosed AI-generation signals are held back and replaced with matching items. The signed receipt only lists items that passed every check, so what you receive is exactly what was verified.

Related

Tell us what your model needs to learn.

Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.