Industry · Legal

Legal AI data, drafted and checked by attorneys.

Legal models are judged on citations, clauses and consequences. The data that trains and tests them should come from people who practice law, with every item accounted for.

Photo needed · 21:9Attorneys reviewing contract at tableWide editorial scene of a daylight-filled law office with two attorneys working through a marked-up contract. Composed, quiet and professional.Wide image under the opening
Overview

Legal, with proof.

Photo needed · 4:5Pen marking contract clausePortrait close-up of a hand annotating a contract margin with a fine pen. Bright paper, shallow depth of field.Beside the overview

Legal AI products summarize contracts, flag risky clauses, draft correspondence, answer research questions and prepare for litigation. Each task has a characteristic failure: an invented citation, a missed indemnity carve-out, an answer that ignores jurisdiction. HUMXN commissions credential-checked attorneys and legal specialists to write demonstrations, annotate documents and grade model answers to a rubric you agree. You describe the practice area and the tasks. We scope a dataset spec, then deliver a pilot batch for your legal and product teams to assess.

Good legal annotation is more than highlighting. A clause label is only useful if the reviewer understands how it interacts with definitions, governing law and the rest of the agreement. Our contract and litigation experts label to your taxonomy, write rationales where your model needs them, and pass disputed items to a second attorney. Labels record their source, whether creator, AI-suggested or expert-reviewed, so your team can weight them accordingly. Held-out evaluation sets are written for you and exclusively licensed.

Legal documents are full of names, addresses and account details. Documents and tables are scanned for personal data and masked before delivery, and personal metadata never reaches buyers. Catalogue works carry their AI-training permission and licence terms, so you are not building on material you cannot use. Every item carries an Ed25519-signed provenance record and appears on a signed receipt by ID and hash, which gives your general counsel a clear record of what went into the model.

Data types
ContractsLegal memosClause annotationsGraded answersReasoning tracesCorrespondence
Experts involved
Contract attorneysLitigatorsParalegalsCompliance specialistsLegal translators
What we deliver

Built to your specification.

Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for legal.

01

Contract clause annotation

Clause identification, risk flags and deviation-from-standard labels across agreement types, drawn to your taxonomy by contract attorneys.

02

Drafting demonstrations

Attorney-written clauses, memos, letters and redlines from realistic fictional matters, for fine-tuning drafting and editing tools.

03

Legal reasoning traces

Step-by-step analysis of issues under stated jurisdictions and facts, written by practicing lawyers to show how a conclusion is reached.

04

Graded research answers

Rubric-based grading of model answers for accuracy, citation support, jurisdiction and hedging, with written rationales.

05

Litigation document review

Relevance, privilege-style and issue tags applied by experienced reviewers to document sets prepared for training.

06

Held-out legal evaluations

Never-published question and drafting tasks, exclusively licensed, for measuring legal capability without contamination.

Why it matters

Where unverified data falls short.

The nine layers of verification

01Hallucinated authority

Models invent citations that look real. Attorney grading that checks support for every claim teaches and measures that failure.

02Jurisdiction blindness

A correct answer in one state can be wrong in another. Experts write and grade with the jurisdiction stated.

03Confidential details

Documents are scanned for personal data and masked before delivery, so names and account details do not reach your training set.

04Usable rights

Training permission and licence terms are bound to each item and listed on a signed receipt your counsel can review.

Questions

Legal: asked often.

Who annotates and grades legal data?

Credential-checked attorneys and legal specialists, matched to the practice area in your spec. Contract work goes to contract lawyers; litigation review goes to people with litigation experience. Each label records whether it came from a creator, an AI suggestion or expert review, and disputed items get a second attorney. You set the rubric and the bar the pilot has to clear.

Is this legal advice?

No. The data our experts create is training and evaluation material for your models, not advice to you or anyone else. Drafting demonstrations use realistic fictional matters. Whether a dataset and its licence fit your intended use is a question for your own counsel, and we provide the documentation they need to answer it.

How do you handle confidential information in documents?

Documents and tables are scanned for personal data and masked before delivery, and personal metadata never reaches buyers. Where a dataset is built from fictional matters, there is no client information to begin with. Confidentiality terms for your project are agreed during scoping. Every delivered item is listed on a signed receipt by ID and hash.

Can you cover multiple jurisdictions and languages?

We scope jurisdictions explicitly, because the same question can have different answers in different places. Where we have qualified experts for a jurisdiction, they write and grade in it; where coverage is thin, we say so before you commit. Legal translators can work on multilingual material, with a second reviewer checking terminology.

How is legal data priced?

Price reflects practice area, seniority, time per item, review depth and licence. A full redline with rationale costs more than a clause tag. Second-attorney review and exclusive licences add cost, and volume or standing orders reduce it. We quote after the spec and acceptance criteria are agreed.

What formats do you deliver?

Annotations, answers and traces arrive as JSON Lines or Parquet, tabular labels as CSV, and documents in their original formats such as PDF or Word where your pipeline needs them. Croissant 1.0 metadata can describe the dataset. Approved buyers can also take deliveries through the API.

Tell us what your model needs to learn.

Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.