Benchmarks your model has never seen.
Public benchmarks leak into pre-training corpora. We commission evaluation items that have never been online, and keep them that way.
Evaluation benchmarks, with proof.
Once a benchmark is public, it is a matter of time before its items, or close paraphrases, appear in web-scale training data. Scores then measure recall as much as capability, and gains stop transferring to real use. A private evaluation set avoids this by construction: items are written new by domain experts for your engagement, are never published, and are delivered only to you under licence. Text fingerprints are checked against registered material at intake to catch near-duplicates.
Good evaluation items need more care than training items. Each one has an unambiguous reference answer or a grading rubric that two experts would apply the same way, and a written justification. We design rubrics with you, pilot them on a sample of items, and revise wording where graders disagree. Difficulty is calibrated deliberately: items are tagged by skill and difficulty tier, and a set can be built so that current models do not saturate it.
Every item is independently solved or graded by a second expert before acceptance, which catches ambiguous prompts, wrong keys and items with more than one defensible answer. You receive the items, reference answers, rubrics, metadata such as skill and difficulty tags, and a signed receipt listing each item by ID and hash. Exclusive licences are standard for evaluation sets, so the same items are never offered to another buyer.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for evaluation benchmarks.
Held-out test sets
New items written for your engagement, never published and delivered only to you.
Reference answers and keys
Each item has a verified answer or a rubric, plus a written justification from its author.
Rubric design
Grading rubrics drafted with you and piloted until independent graders apply them consistently.
Difficulty calibration
Items tagged by skill and difficulty tier so the set discriminates between strong models.
Independent verification
A second expert solves or grades every item to catch ambiguity and wrong keys.
Refresh sets
Standing orders for new items over time, so the benchmark stays ahead of model progress.
01Contamination
Items that have been online can be memorized. Writing new items and never publishing them keeps scores tied to capability.
02Ambiguous items add noise
An item with two defensible answers penalizes good models at random. Independent solving by a second expert removes most of them.
03Saturation
A benchmark every model aces says nothing. Deliberate difficulty tiers keep headroom at the top of the scale.
04Provenance for reported results
Signed receipts and hashes let you show exactly which item set a reported score was measured on.
Evaluation benchmarks: asked often.
How do you keep evaluation items out of training data?
Items are written new for your engagement, are never published or listed in the catalogue, and are delivered only to you, usually under an exclusive licence. Writers are bound by confidentiality terms. At intake we compare text fingerprints against everything registered to catch reuse. What you do after delivery, such as access control in your own systems, remains in your hands.
How is difficulty calibrated?
Authors tag each item by skill and intended difficulty tier, and a second expert confirms the tag while solving it. If you share model outputs on a pilot set, we use them to check that tiers discriminate as intended and adjust the mix. The goal is a set where current strong models still have room to improve.
Can you build rubric-graded evaluations, not only exact-match?
Yes. For open-ended tasks such as clinical reasoning, legal analysis or long-form writing, we design a rubric with specific, checkable criteria and pilot it with multiple graders. Where graders disagree, the wording is revised. You can then use the rubric with expert graders from the network, with your own team, or as guidance for a model-based grader.
What does a private benchmark cost?
Evaluation items usually cost more per item than training data because each needs a verified key or rubric, a justification and an independent second solve. Price depends on domain, item length, difficulty, multimodality and whether you need grading services afterward. Sets are typically smaller than training sets, and exclusive licensing is included in the quote.
Do you also grade our model's outputs?
We can. Experts from the same network apply the rubric to your model's responses and record scores, rationales and grader IDs, with overlapping grading so agreement can be reported. This is useful for open-ended tasks where automatic scoring is unreliable. Model outputs you send are handled under your confidentiality terms.
Tell us what your model needs to learn.
Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.

