Legal AI data, drafted and checked by attorneys.
Legal models are judged on citations, clauses and consequences. The data that trains and tests them should come from people who practice law, with every item accounted for.
Legal, with proof.
Legal AI products summarize contracts, flag risky clauses, draft correspondence, answer research questions and prepare for litigation. Each task has a characteristic failure: an invented citation, a missed indemnity carve-out, an answer that ignores jurisdiction. HUMXN commissions credential-checked attorneys and legal specialists to write demonstrations, annotate documents and grade model answers to a rubric you agree. You describe the practice area and the tasks. We scope a dataset spec, then deliver a pilot batch for your legal and product teams to assess.
Good legal annotation is more than highlighting. A clause label is only useful if the reviewer understands how it interacts with definitions, governing law and the rest of the agreement. Our contract and litigation experts label to your taxonomy, write rationales where your model needs them, and pass disputed items to a second attorney. Labels record their source, whether creator, AI-suggested or expert-reviewed, so your team can weight them accordingly. Held-out evaluation sets are written for you and exclusively licensed.
Legal documents are full of names, addresses and account details. Documents and tables are scanned for personal data and masked before delivery, and personal metadata never reaches buyers. Catalogue works carry their AI-training permission and licence terms, so you are not building on material you cannot use. Every item carries an Ed25519-signed provenance record and appears on a signed receipt by ID and hash, which gives your general counsel a clear record of what went into the model.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for legal.
Contract clause annotation
Clause identification, risk flags and deviation-from-standard labels across agreement types, drawn to your taxonomy by contract attorneys.
Drafting demonstrations
Attorney-written clauses, memos, letters and redlines from realistic fictional matters, for fine-tuning drafting and editing tools.
Legal reasoning traces
Step-by-step analysis of issues under stated jurisdictions and facts, written by practicing lawyers to show how a conclusion is reached.
Graded research answers
Rubric-based grading of model answers for accuracy, citation support, jurisdiction and hedging, with written rationales.
Litigation document review
Relevance, privilege-style and issue tags applied by experienced reviewers to document sets prepared for training.
Held-out legal evaluations
Never-published question and drafting tasks, exclusively licensed, for measuring legal capability without contamination.
01Hallucinated authority
Models invent citations that look real. Attorney grading that checks support for every claim teaches and measures that failure.
02Jurisdiction blindness
A correct answer in one state can be wrong in another. Experts write and grade with the jurisdiction stated.
03Confidential details
Documents are scanned for personal data and masked before delivery, so names and account details do not reach your training set.
04Usable rights
Training permission and licence terms are bound to each item and listed on a signed receipt your counsel can review.
Legal: asked often.
Who annotates and grades legal data?
Credential-checked attorneys and legal specialists, matched to the practice area in your spec. Contract work goes to contract lawyers; litigation review goes to people with litigation experience. Each label records whether it came from a creator, an AI suggestion or expert review, and disputed items get a second attorney. You set the rubric and the bar the pilot has to clear.
Is this legal advice?
No. The data our experts create is training and evaluation material for your models, not advice to you or anyone else. Drafting demonstrations use realistic fictional matters. Whether a dataset and its licence fit your intended use is a question for your own counsel, and we provide the documentation they need to answer it.
How do you handle confidential information in documents?
Documents and tables are scanned for personal data and masked before delivery, and personal metadata never reaches buyers. Where a dataset is built from fictional matters, there is no client information to begin with. Confidentiality terms for your project are agreed during scoping. Every delivered item is listed on a signed receipt by ID and hash.
Can you cover multiple jurisdictions and languages?
We scope jurisdictions explicitly, because the same question can have different answers in different places. Where we have qualified experts for a jurisdiction, they write and grade in it; where coverage is thin, we say so before you commit. Legal translators can work on multilingual material, with a second reviewer checking terminology.
How is legal data priced?
Price reflects practice area, seniority, time per item, review depth and licence. A full redline with rationale costs more than a clause tag. Second-attorney review and exclusive licences add cost, and volume or standing orders reduce it. We quote after the spec and acceptance criteria are agreed.
What formats do you deliver?
Annotations, answers and traces arrive as JSON Lines or Parquet, tabular labels as CSV, and documents in their original formats such as PDF or Word where your pipeline needs them. Croissant 1.0 metadata can describe the dataset. Approved buyers can also take deliveries through the API.
Tell us what your model needs to learn.
Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.

