Custom datasets, built to your spec.
Bring us the capability you are trying to build and the gaps you are seeing. We write the spec with you, source from consenting creators and vetted experts, and verify every item before it ships.
Custom datasets, with proof.
Off-the-shelf corpora rarely match the distribution your model actually needs. A custom dataset starts from your failure cases: the domains, formats, edge conditions and label definitions that matter for the next training run. We turn that into a written dataset specification covering modality, volume, schema, acceptance criteria and licence terms, so both sides agree on what a correct item looks like before any sourcing begins. That document then governs the pilot, production and any standing order that follows.
Sourcing draws on two channels. The registered catalogue holds existing human-made work whose creators have granted AI-training permission and set licence terms. The expert network creates new material on commission: demonstrations, recordings, documents, annotations and reviews from people who are credential-checked in the relevant field. Most custom datasets combine both, using catalogue material for breadth and commissioned work for the long tail of cases that no existing archive covers well.
Every item passes the same intake checks regardless of channel. We fingerprint it against everything registered to catch duplicates and conflicting claims, record whether it is human-created, AI-assisted or AI-generated, check for generator signatures and C2PA records, screen for unsafe content and personal data, and bind the consent and licence terms to the item. You receive a signed receipt listing each item by ID and hash, so the dataset you trained on can be identified exactly later.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for custom datasets.
Written dataset specification
Modality, schema, label definitions, volume, acceptance criteria and licence scope, agreed before sourcing starts.
Catalogue sourcing
Existing human-made work from registered creators who have granted AI-training permission on stated terms.
Commissioned expert work
New items created to your spec by credential-checked specialists where no existing material fits.
Pilot batch
A small, representative batch you test in your own pipeline before anything scales.
Verified production
Full-volume delivery held to the pilot's acceptance criteria, with a second-expert review pass on labels.
Standing orders
Ongoing supply of matching items at the same bar as your model's needs shift.
01Distribution, not volume
The value of a custom set is in covering the cases your model gets wrong. That requires sourcing against a precise spec rather than scraping more of what is already well represented.
02Rights you can document
Each item carries its AI-training permission, licence terms and any model or property releases, so counsel can review what was licensed and on what basis.
03Known origin
Recording whether each work is human-created, AI-assisted or AI-generated lets you control how much synthetic content enters your training mix.
04Exact reproducibility
A signed receipt with every item's ID and hash means the dataset behind a model version can be reconstructed and audited.
Custom datasets: asked often.
How is a custom dataset priced?
Pricing depends on modality, volume, how much of the set can come from the existing catalogue versus new commissioned work, the seniority of the experts involved, the depth of annotation and review, and the licence scope. Exclusive licences cost more than non-exclusive ones. We quote against the written specification after the scoping call, and volume pricing applies to larger orders and standing orders.
How long does a custom dataset take?
Timing depends mostly on how much new material must be commissioned and how specialized the experts are. Catalogue-heavy sets move faster than sets that need rare domain expertise or new recordings. We agree a schedule in the specification, deliver a pilot batch first, and report progress against the spec during production so you can plan training runs around real delivery dates.
What formats do you deliver in?
We deliver JSON Lines, CSV, Parquet, WebDataset, COCO captions and Croissant 1.0 metadata. Each record carries its labels, technical metadata, origin classification, licence reference and file hashes. Approved buyers can also pull deliveries through the API. If your pipeline expects a particular schema, we specify it in the dataset specification and validate against it before delivery.
Can we license the dataset exclusively?
Yes. Exclusive licences are available for commissioned work, and for catalogue items where the creator has agreed to exclusive terms. Exclusivity is recorded on each item, so the same work is not offered to another buyer for the licensed scope. Exclusivity affects price and sometimes sourcing time, so we settle it during scoping rather than after the pilot.
How do you measure quality on a custom set?
Quality is measured against acceptance criteria written into the specification and tested on the pilot batch. Labels are marked by source, whether creator, AI-suggested or expert-reviewed, and a second expert reviews a share of items agreed in the spec. Items that fail criteria are replaced rather than delivered, and you approve the pilot item by item before production starts.
What happens to items that fail verification?
They are not delivered. Items that duplicate or conflict with registered work, lack required consent or releases, fail safety screening, or carry undisclosed AI-generation signals are held back and replaced with matching items. The signed receipt only lists items that passed every check, so what you receive is exactly what was verified.
Tell us what your model needs to learn.
Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.

