Clinical AI data, written and checked by clinicians.
In medicine, a fluent wrong answer is worse than no answer. Healthcare models need data written and judged by people who practice, documented well enough for a clinical safety review.
Healthcare, with proof.
Healthcare models are asked to summarize encounters, answer clinical questions, draft documentation and triage messages. Each of those tasks fails in its own way: a missed contraindication, a hallucinated lab value, a confident answer outside the evidence. HUMXN commissions credential-checked physicians, nurses, pharmacists and other clinicians to write cases, demonstrations and reference answers to your specification, and to grade model output against a rubric you agree. You bring the use case. We write the dataset spec and deliver a pilot batch for your clinical team to assess.
Clinical judgment is not a single opinion. Our rubric-based grading records which expert made each call, labels are marked by source, and contested items go to a second clinician for review. That gives your evaluation team disagreement data as well as consensus, which is often where the real risk sits. Clinical vignettes can be written from scratch by clinicians, so a dataset can teach realistic reasoning without starting from patient records at all.
Where documents or tables are part of a dataset, they are scanned for personal data and masked before delivery, and location and personal metadata never reach buyers. Works showing identifiable people are offered for training only with consent on file. Every item carries an Ed25519-signed provenance record and appears on a signed receipt. This is designed to support your own compliance work under frameworks such as HIPAA, GDPR or FDA guidance; it is not a substitute for your privacy and regulatory review.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for healthcare.
Clinician-written vignettes
Realistic cases written from scratch by practicing clinicians across specialties, with the history, findings and reasoning your model should learn.
Graded model answers
Rubric-based grading of model responses for accuracy, safety, completeness and appropriate escalation, with written rationales.
Documentation demonstrations
Expert-written notes, summaries and patient-facing explanations built from fictional encounters, for fine-tuning documentation tools.
Medical image annotation
Labels and region annotations from credential-checked specialists, with a second-expert review pass on difficult or ambiguous findings.
Clinical safety red-teaming
Adversarial prompts on dosing, contraindications and self-harm, written by clinicians, with graded responses showing the safe answer.
Held-out clinical evaluations
Never-published question sets in the specialties you target, exclusively licensed so your benchmark measures reasoning, not recall.
01Confident errors
Models trained on general web text sound clinical without being correct. Clinician-graded data teaches the difference and measures it.
02Patient privacy
Documents and tables are scanned for personal data and masked before delivery, and personal metadata never reaches you.
03Expert disagreement
Medicine has gray areas. Second-clinician review and source-marked labels preserve disagreement instead of hiding it.
04Traceable evidence
Clinical safety reviews need to know where data came from. Signed provenance and per-item receipts provide that record.
Healthcare: asked often.
Do you provide real patient records?
Our healthcare work centers on data created and reviewed by clinicians: vignettes written from scratch, graded model answers, documentation demonstrations and expert annotation. Where documents or tables are part of a dataset, they are scanned for personal data and masked before delivery. Your team should still review any dataset against its own privacy obligations; we provide documentation, not legal advice.
Does this support our HIPAA and GDPR obligations?
We do not claim certifications we do not hold. Our controls, including personal-data scanning and masking, metadata stripping, consent records and signed provenance, are designed to support your compliance work under HIPAA, GDPR and similar frameworks. Whether a dataset fits your obligations is a question for your privacy and legal teams, and we will share whatever documentation they need.
How are clinical experts vetted?
Every clinician is credential-checked before taking work, and identity and tax verification are completed before payment. Projects specify the specialty and seniority needed, so a cardiology rubric is graded by people who practice cardiology. Hard items get a second-clinician review, and each label records whether it came from a creator, an AI suggestion or an expert review.
What does clinical grading cost?
Cost follows specialty, seniority, time per item and review depth. Grading a short patient message is quicker than reviewing a complex differential with written rationale. Second-clinician review and exclusive licences add cost, while volume and standing orders reduce it. We quote after scoping, once the rubric and acceptance criteria are written.
How long does a healthcare dataset take?
Scoping produces a written spec, usually within days. A pilot batch follows so your clinical team can assess quality before volume. Production time depends on the specialties involved and how long each item takes; narrow subspecialties take longer to staff. We describe the timeline in the spec rather than quoting a fixed figure upfront.
Which formats do you deliver?
Text and graded responses usually arrive as JSON Lines or Parquet, tables as CSV or Parquet, and annotated images with COCO-format captions or labels. Croissant 1.0 metadata can describe the full dataset. Every delivery includes a signed receipt listing each item by ID and hash, and approved buyers can use the API.
Tell us what your model needs to learn.
Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.

