Solution · RLHF & preference data

Preference data from people qualified to judge.

Reward models learn whatever the raters reward. We match comparisons to raters who can judge whether an answer is correct, not only whether it reads well.

Photo needed · 21:9Expert rater comparing two answersA focused professional at a bright desk with a wide monitor showing two text panels side by side, notes beside the keyboard. Editorial, natural light.Wide image under the opening
Overview

RLHF & preference data, with proof.

Photo needed · 4:5Rubric notes beside keyboardPortrait close-up of a printed scoring rubric with handwritten marks next to a keyboard. Soft daylight, calm tones.Beside the overview

Preference data for RLHF, DPO and reward-model training is only as good as the judgment behind each label. On technical prompts, a generalist rater tends to prefer the confident, well-formatted answer over the correct but terse one, and the reward model learns that bias. We route comparisons to raters qualified in the domain of the prompt: clinicians for medical answers, practicing lawyers for legal reasoning, mathematicians for proofs, engineers for code and system design.

We support the common collection formats. Pairwise comparisons with a preference strength, ranked lists over several completions, and absolute rubric grades on dimensions such as correctness, completeness, safety and instruction following. Raters write a short rationale for each judgment, which makes disagreements diagnosable and gives you material for critique models. Rubrics are written with you before collection starts and tested on a calibration set so raters apply them consistently.

Quality is measured, not asserted. A share of items is labeled by more than one rater so inter-annotator agreement can be computed and reported per batch and per rubric dimension. Low-agreement items are surfaced for adjudication by a senior reviewer instead of being averaged away. Raters qualify on a gold set before working and are monitored against it during production, and each label records which rater, which rubric version and which review pass produced it.

Data types
Pairwise comparisonsRanked listsRubric scoresWritten rationalesMulti-turn dialoguesCode outputs
Experts involved
PhysiciansAttorneysMathematiciansSoftware engineersScientistsFinancial analystsProfessional writersSenior adjudicators
What we deliver

Built to your specification.

Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for rlhf & preference data.

01

Pairwise comparisons

A versus B judgments with preference strength, from raters matched to the prompt's domain.

02

Ranked completions

Orderings over several model outputs for listwise and reward-model training.

03

Rubric grading

Absolute scores on correctness, completeness, safety, style and instruction following, against a rubric you approve.

04

Written rationales

A short explanation with every judgment, captured as structured text alongside the label.

05

Agreement reporting

Inter-annotator agreement measured on overlapping items and reported per batch and rubric dimension.

06

Adjudication

Low-agreement items reviewed by a senior expert and resolved with a recorded decision.

Why it matters

Where unverified data falls short.

The nine layers of verification

01Style bias in reward models

Raters without domain knowledge reward fluency. Qualified raters catch the confident wrong answer, which is the failure that matters most for capable models.

02Disagreement is signal

Measured agreement and captured rationales show where the rubric is ambiguous or the task is genuinely contested, so you can fix the spec rather than train on noise.

03Rater accountability

Each label records its rater, rubric version and review pass, so you can audit or re-weight specific raters after the fact.

04No model-written labels in disguise

Labels are marked by source, so AI-suggested judgments are never presented as human preference.

Questions

RLHF & preference data: asked often.

How do you measure inter-annotator agreement?

We assign a share of items, agreed in the specification, to two or more raters and compute agreement on those overlaps per batch and per rubric dimension. The statistic used depends on the label type, for example a chance-corrected measure for categorical preferences. We report the figures with each delivery, and items below the agreed threshold go to a senior reviewer for adjudication.

How are raters qualified?

Raters are credential-checked in their field when they join the network. For each project they complete a qualification set of items with known answers built from your rubric, and only raters who meet the bar label production data. Gold items are mixed into production work so drift is caught, and raters who fall below the bar are removed from the project.

Do you supply the prompts and completions, or do we?

Either works. Most teams send prompts and model completions from their own policy so preferences reflect current behavior. We can also commission prompts from domain experts, which is useful when you need coverage of hard or rare cases. Completions you provide are handled under the confidentiality terms in your agreement and are not added to the catalogue.

What does preference data cost?

Cost is driven by rater expertise, the length and difficulty of the items, the number of rubric dimensions, whether rationales are required, and how much overlap you want for agreement measurement. Physician or attorney comparisons on long answers cost more than general-domain pairwise labels. We quote per item after scoping, with volume pricing for larger and standing orders.

Which formats do you deliver preference data in?

Typically JSON Lines with one record per comparison: prompt, completions, preference or scores, rationale, rater ID, rubric version and review status. Parquet is available for large volumes. We can match the schema your DPO or reward-model trainer expects, and we validate deliveries against it.

Can you collect multi-turn or agentic preferences?

Yes. Raters can compare full conversations or agent trajectories rather than single turns, grading each step or the final outcome according to the rubric. These tasks take longer per item and usually need raters with hands-on experience of the tools or workflows involved, which we account for in rater selection and pricing.

Related

Tell us what your model needs to learn.

Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.