Preference data from people qualified to judge.
Reward models learn whatever the raters reward. We match comparisons to raters who can judge whether an answer is correct, not only whether it reads well.
RLHF & preference data, with proof.
Preference data for RLHF, DPO and reward-model training is only as good as the judgment behind each label. On technical prompts, a generalist rater tends to prefer the confident, well-formatted answer over the correct but terse one, and the reward model learns that bias. We route comparisons to raters qualified in the domain of the prompt: clinicians for medical answers, practicing lawyers for legal reasoning, mathematicians for proofs, engineers for code and system design.
We support the common collection formats. Pairwise comparisons with a preference strength, ranked lists over several completions, and absolute rubric grades on dimensions such as correctness, completeness, safety and instruction following. Raters write a short rationale for each judgment, which makes disagreements diagnosable and gives you material for critique models. Rubrics are written with you before collection starts and tested on a calibration set so raters apply them consistently.
Quality is measured, not asserted. A share of items is labeled by more than one rater so inter-annotator agreement can be computed and reported per batch and per rubric dimension. Low-agreement items are surfaced for adjudication by a senior reviewer instead of being averaged away. Raters qualify on a gold set before working and are monitored against it during production, and each label records which rater, which rubric version and which review pass produced it.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for rlhf & preference data.
Pairwise comparisons
A versus B judgments with preference strength, from raters matched to the prompt's domain.
Ranked completions
Orderings over several model outputs for listwise and reward-model training.
Rubric grading
Absolute scores on correctness, completeness, safety, style and instruction following, against a rubric you approve.
Written rationales
A short explanation with every judgment, captured as structured text alongside the label.
Agreement reporting
Inter-annotator agreement measured on overlapping items and reported per batch and rubric dimension.
Adjudication
Low-agreement items reviewed by a senior expert and resolved with a recorded decision.
01Style bias in reward models
Raters without domain knowledge reward fluency. Qualified raters catch the confident wrong answer, which is the failure that matters most for capable models.
02Disagreement is signal
Measured agreement and captured rationales show where the rubric is ambiguous or the task is genuinely contested, so you can fix the spec rather than train on noise.
03Rater accountability
Each label records its rater, rubric version and review pass, so you can audit or re-weight specific raters after the fact.
04No model-written labels in disguise
Labels are marked by source, so AI-suggested judgments are never presented as human preference.
RLHF & preference data: asked often.
How do you measure inter-annotator agreement?
We assign a share of items, agreed in the specification, to two or more raters and compute agreement on those overlaps per batch and per rubric dimension. The statistic used depends on the label type, for example a chance-corrected measure for categorical preferences. We report the figures with each delivery, and items below the agreed threshold go to a senior reviewer for adjudication.
How are raters qualified?
Raters are credential-checked in their field when they join the network. For each project they complete a qualification set of items with known answers built from your rubric, and only raters who meet the bar label production data. Gold items are mixed into production work so drift is caught, and raters who fall below the bar are removed from the project.
Do you supply the prompts and completions, or do we?
Either works. Most teams send prompts and model completions from their own policy so preferences reflect current behavior. We can also commission prompts from domain experts, which is useful when you need coverage of hard or rare cases. Completions you provide are handled under the confidentiality terms in your agreement and are not added to the catalogue.
What does preference data cost?
Cost is driven by rater expertise, the length and difficulty of the items, the number of rubric dimensions, whether rationales are required, and how much overlap you want for agreement measurement. Physician or attorney comparisons on long answers cost more than general-domain pairwise labels. We quote per item after scoping, with volume pricing for larger and standing orders.
Which formats do you deliver preference data in?
Typically JSON Lines with one record per comparison: prompt, completions, preference or scores, rationale, rater ID, rubric version and review status. Parquet is available for large volumes. We can match the schema your DPO or reward-model trainer expects, and we validate deliveries against it.
Can you collect multi-turn or agentic preferences?
Yes. Raters can compare full conversations or agent trajectories rather than single turns, grading each step or the final outcome according to the rubric. These tasks take longer per item and usually need raters with hands-on experience of the tools or workflows involved, which we account for in rater selection and pricing.
Tell us what your model needs to learn.
Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.

