Public sector AI data, with a record for every item.
Public agencies are accountable for what their systems learned from. Data for government AI should arrive documented: who made it, under what terms, and what was removed.
Public sector, with proof.
Government agencies and their vendors use models to answer citizen questions, process forms, summarize case files, translate services and support caseworkers. Accountability is the defining constraint. Procurement teams, auditors and the public may ask what a model was trained and tested on, and an answer of web scrape is increasingly hard to defend. HUMXN sources data from consenting creators and commissions credential-checked experts to create, label and grade it, then verifies every item. You define the services and tasks. We scope a dataset spec and deliver a pilot batch first.
Public-facing language models must be accurate, plain and fair across the people they serve. Our experts write demonstrations in plain language, grade model answers against your rubric, and build held-out evaluation sets that test performance on the questions residents actually ask. Linguists and translators extend that work across the languages your services are offered in. Labels are marked by source, and contested items go to a second expert. Documents and forms are scanned for personal data and masked before delivery, and personal metadata never reaches buyers.
Every delivered item carries an Ed25519-signed provenance record, sits in a hash-chained signed audit trail and is anchored with OpenTimestamps in Bitcoin, so its existence at a point in time can be checked independently. A signed receipt lists each item by ID and hash. That documentation is designed to support the transparency, records and risk-assessment work public bodies carry out under frameworks such as the EU AI Act and national guidance. It is not a guarantee of compliance, and your counsel should make that call.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for public sector.
Plain-language demonstrations
Expert-written answers to resident questions about services, eligibility and processes, clear and accurate, for fine-tuning assistants.
Form and document extraction
Labeled fields and structures from realistic forms and case documents, scanned for personal data and masked before delivery.
Multilingual service content
Translations and localized content by professional translators and linguists, with second review for terminology.
Graded model answers
Rubric grading for accuracy, clarity, tone and appropriate referral, with written rationales your team can audit.
Held-out evaluation sets
Never-published test sets covering the questions and documents your systems will meet, exclusively licensed for your agency.
Safety and fairness test prompts
Adversarial and edge-case prompts written by specialists, for red-teaming public-facing systems before launch.
01Accountability
Agencies must explain what systems learned from. Signed provenance, an audit trail and timestamp anchoring make that answerable.
02Personal information
Forms and case files are full of it. Documents are scanned and masked, and personal metadata never reaches you.
03Language access
Services must work across languages. Professional translators and linguists build and review multilingual data.
04Unclear rights
Scraped data carries unknown terms. Every item we deliver has documented training permission and licence terms.
Public sector: asked often.
What documentation do we receive with a dataset?
Each item has an Ed25519-signed provenance record and sits in a hash-chained signed audit trail, anchored with OpenTimestamps in Bitcoin. Delivery includes a signed receipt listing every item by ID and hash, plus licence terms and consent records. Croissant 1.0 metadata can describe the dataset for your catalog or records.
Does this satisfy the EU AI Act or our national rules?
We do not offer compliance guarantees or legal advice. Our provenance, consent and privacy controls are designed to support the documentation and transparency work these frameworks ask of public bodies and their suppliers. Your legal and compliance teams should decide whether a dataset meets your obligations.
How is personal information handled?
Documents and tables are scanned for personal data and masked before delivery, and location and personal metadata never reach buyers. Many public sector datasets use realistic fictional forms and cases written by specialists, which avoids citizen data entirely. Works showing identifiable people require consent on file.
Can you work through our procurement process?
We start with a scoping call and a written dataset spec, which gives your procurement team a concrete document to work from. A pilot batch lets you evaluate quality before a larger commitment. Contracting terms, vendor requirements and security questionnaires are discussed during scoping.
Which languages can you cover?
It depends on the languages in your service plan and on qualified linguists and translators in our network. We confirm coverage during scoping and say plainly where it is thin. Translations go through a second review for terminology and plain-language standards.
Tell us what your model needs to learn.
Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.

