Demonstrations worth imitating.
In fine-tuning, the model copies what it is shown. We commission demonstrations from specialists who would be trusted to give that answer in their own work.
SFT data, with proof.
Supervised fine-tuning data sets both what a model knows how to do and how it behaves while doing it. Small, high-quality sets often move a model further than large noisy ones, which puts the weight on who writes each demonstration. We commission prompt and response pairs from credential-checked experts in the target domain, so the content is correct and the response shows the judgment, caveats and structure a practitioner would actually use.
Many teams need more than final answers. We collect step-by-step reasoning traces for math, science, code and analytical tasks, with each step checkable, and multi-turn dialogues that cover clarification, correction and refusal. Writers work from a style guide agreed with you: tone, length, formatting, citation conventions, when to ask a question rather than assume, and how to handle uncertainty. The guide is versioned, and every item records which version it was written against.
Each demonstration is reviewed by a second expert who checks correctness, completeness and adherence to the style guide before it is accepted. Reviewers can return items with comments, and the review outcome is stored with the item. Writers attest that the work is their own, intake checks look for generator signatures and text fingerprints that match registered or known material, and AI-assisted drafting is recorded rather than hidden, so you control what enters your mix.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for sft data.
Expert demonstrations
Prompt and response pairs written by specialists in the domain, to your coverage plan.
Reasoning traces
Step-by-step worked solutions in math, science, code and analysis, with each step checkable.
Multi-turn dialogues
Conversations that include clarifying questions, corrections, follow-ups and appropriate refusals.
Style guide authoring
A versioned guide for tone, structure, formatting and uncertainty handling, drafted with your team.
Second-expert review
Every item checked by another specialist for correctness and guide adherence before acceptance.
Prompt coverage design
A taxonomy of task types and difficulty levels so the set covers what your model is missing.
01The model copies errors too
A subtly wrong demonstration teaches the wrong answer with full confidence. Domain experts and a second review pass are the main defense.
02Consistency across writers
A shared, versioned style guide keeps many experts producing responses that read as one assistant rather than a mix of voices.
03Synthetic contamination
Origin is recorded per item and generator signatures are checked, so model-written text does not quietly enter a set sold as human-written.
SFT data: asked often.
How many SFT examples do we need?
It depends on the behavior you are teaching and your base model. Many teams start with a focused pilot set covering their highest-priority task types, measure the effect on their own evaluations, then scale the categories that moved. We help design a coverage plan so volume goes where it changes the model, rather than prescribing a fixed number up front.
Do you write reasoning traces or only final answers?
Both. For math, science, code and analytical tasks we commission full worked solutions where each step can be checked, and the reviewer verifies the steps as well as the final answer. Trace format, such as how much intermediate notation to show or where to state assumptions, is set in the style guide so traces are consistent across writers.
How do you stop writers using AI to draft responses?
Writers attest to how each item was produced, and intake records it as human-created, AI-assisted or AI-generated. We check for generator signatures in metadata and compare text fingerprints against registered work. Second-expert review also catches the characteristic patterns of unedited model output. If your spec requires fully human-written items, AI-assisted items are excluded.
What drives the cost of SFT data?
The main factors are the expertise required, the length and difficulty of each demonstration, whether reasoning traces are needed, multi-turn depth, and the review depth you specify. A long clinical or legal answer reviewed by a second specialist costs more than a short general-domain response. We quote per item against the agreed spec, with volume pricing for larger orders.
Can we use our own style guide?
Yes. If you have one, writers are onboarded to it and tested on a calibration set before production. If not, we draft one with you during scoping. Either way the guide is versioned, every item records the version it was written against, and changes mid-project are tracked so you can filter or revise older items.
Tell us what your model needs to learn.
Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.

