Documents and tables, with personal data masked.
Real documents are full of personal information. We license document and tabular data with the rights to train, and mask personal data before it reaches you.
Document & structured data, with proof.
Document understanding, extraction and agentic workflows need training material that looks like real work: contracts, filings, invoices, reports, forms, spreadsheets and technical manuals, with their layouts, tables and inconsistencies intact. That kind of material is hard to obtain with clear rights and without exposing personal data. HUMXN sources documents and structured data from professionals and organizations who have granted AI-training permission, and commissions realistic documents from experts where originals cannot be shared.
Labels follow your schema: document classification, key-value extraction, table structure, layout regions, entity and relation tags, question-answer pairs over documents, and summaries. Lawyers, accountants, clinicians and analysts label material in their field, because what counts as a material clause or a reconciling item is not obvious from layout. A second-expert pass reviews an agreed share, and each label records its source. PDF, Word and common spreadsheet formats are supported.
Every document and table is scanned for personal data at intake, and detected names, identifiers, contact details and similar fields are masked before delivery. Text fingerprints catch duplicates and near-duplicates against everything registered. Embedded metadata such as author and location is stripped. Each item carries a signed provenance record, and your receipt lists it by ID and hash, so the training set can be reconstructed and reviewed later.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for document & structured data.
Licensed documents
Contracts, reports, forms, filings and manuals from contributors who granted AI-training permission.
Commissioned documents
Realistic documents written by experts where originals cannot be shared.
Tabular and structured data
Spreadsheets and tables with schema documentation and column-level descriptions.
Extraction labels
Key-value pairs, entities, relations and table structure to your schema.
Document QA pairs
Questions and grounded answers over long documents, written by domain experts.
Personal data masking
Detected personal data masked before delivery, with embedded metadata removed.
01Personal data is everywhere
Real documents contain names, addresses and identifiers. Scanning and masking before delivery limits what reaches your environment.
02Layouts matter
Clean synthetic text does not teach a model to read a scanned form or a merged-cell table. Real layouts do.
03Domain judgment in labels
Which clause is material or which figure reconciles requires lawyers and accountants, not generalists.
04Near-duplicate templates
Text fingerprints catch the same template submitted many times with small edits.
Document & structured data: asked often.
How do you handle personal data in documents?
Every document and table is scanned at intake, and detected personal data such as names, identifiers and contact details is masked before delivery. Embedded metadata such as author and location is removed. Masking reduces exposure but is not a substitute for your own data-protection assessment, and your counsel should consider the obligations that apply to your use.
Can you provide real-world documents rather than synthetic ones?
Where contributors have granted AI-training permission, yes, with personal data masked. Where originals cannot be shared, experts write realistic documents that follow real conventions and layouts. Each item records its origin, so you know which documents are licensed originals and which were commissioned.
What labeling do you support for documents?
Classification, key-value extraction, entity and relation tagging, table structure, layout regions, summaries and question-answer pairs grounded in the document. Domain experts label material in their field, and a second-expert pass reviews an agreed share. Each label records its source, so AI-suggested labels are distinguishable from expert-reviewed ones.
What delivery formats are available?
Source files are delivered in their original format, such as PDF or Word, with labels and metadata in JSON Lines, CSV or Parquet. Tabular data can be delivered as Parquet or CSV with a schema description. Croissant 1.0 metadata describes the set, and approved buyers can retrieve deliveries through the API.
What affects pricing for document data?
Price depends on document type and length, licensed versus commissioned material, the depth of labeling, the expertise required, review depth and exclusivity. Long contracts labeled by attorneys cost more than invoice key-value extraction. We quote per document or per page after scoping.
Tell us what your model needs to learn.
Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.

