Industry · Public sector

Public sector AI data, with a record for every item.

Public agencies are accountable for what their systems learned from. Data for government AI should arrive documented: who made it, under what terms, and what was removed.

Photo needed · 21:9Caseworker helping resident at counterWide editorial scene of a modern, daylight-filled public office with a caseworker and resident reviewing a form. Respectful, calm and real.Wide image under the opening
Overview

Public sector, with proof.

Photo needed · 4:5Hands completing printed service formPortrait close-up of a pen completing a form on a clean desk. Bright and simple.Beside the overview

Government agencies and their vendors use models to answer citizen questions, process forms, summarize case files, translate services and support caseworkers. Accountability is the defining constraint. Procurement teams, auditors and the public may ask what a model was trained and tested on, and an answer of web scrape is increasingly hard to defend. HUMXN sources data from consenting creators and commissions credential-checked experts to create, label and grade it, then verifies every item. You define the services and tasks. We scope a dataset spec and deliver a pilot batch first.

Public-facing language models must be accurate, plain and fair across the people they serve. Our experts write demonstrations in plain language, grade model answers against your rubric, and build held-out evaluation sets that test performance on the questions residents actually ask. Linguists and translators extend that work across the languages your services are offered in. Labels are marked by source, and contested items go to a second expert. Documents and forms are scanned for personal data and masked before delivery, and personal metadata never reaches buyers.

Every delivered item carries an Ed25519-signed provenance record, sits in a hash-chained signed audit trail and is anchored with OpenTimestamps in Bitcoin, so its existence at a point in time can be checked independently. A signed receipt lists each item by ID and hash. That documentation is designed to support the transparency, records and risk-assessment work public bodies carry out under frameworks such as the EU AI Act and national guidance. It is not a guarantee of compliance, and your counsel should make that call.

Data types
Forms and documentsPlain-language textGraded answersTranslationsSpeechStructured tables
Experts involved
Policy specialistsAttorneysLinguistsTranslatorsPlain-language writersAccessibility specialists
What we deliver

Built to your specification.

Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for public sector.

01

Plain-language demonstrations

Expert-written answers to resident questions about services, eligibility and processes, clear and accurate, for fine-tuning assistants.

02

Form and document extraction

Labeled fields and structures from realistic forms and case documents, scanned for personal data and masked before delivery.

03

Multilingual service content

Translations and localized content by professional translators and linguists, with second review for terminology.

04

Graded model answers

Rubric grading for accuracy, clarity, tone and appropriate referral, with written rationales your team can audit.

05

Held-out evaluation sets

Never-published test sets covering the questions and documents your systems will meet, exclusively licensed for your agency.

06

Safety and fairness test prompts

Adversarial and edge-case prompts written by specialists, for red-teaming public-facing systems before launch.

Why it matters

Where unverified data falls short.

The nine layers of verification

01Accountability

Agencies must explain what systems learned from. Signed provenance, an audit trail and timestamp anchoring make that answerable.

02Personal information

Forms and case files are full of it. Documents are scanned and masked, and personal metadata never reaches you.

03Language access

Services must work across languages. Professional translators and linguists build and review multilingual data.

04Unclear rights

Scraped data carries unknown terms. Every item we deliver has documented training permission and licence terms.

Questions

Public sector: asked often.

What documentation do we receive with a dataset?

Each item has an Ed25519-signed provenance record and sits in a hash-chained signed audit trail, anchored with OpenTimestamps in Bitcoin. Delivery includes a signed receipt listing every item by ID and hash, plus licence terms and consent records. Croissant 1.0 metadata can describe the dataset for your catalog or records.

Does this satisfy the EU AI Act or our national rules?

We do not offer compliance guarantees or legal advice. Our provenance, consent and privacy controls are designed to support the documentation and transparency work these frameworks ask of public bodies and their suppliers. Your legal and compliance teams should decide whether a dataset meets your obligations.

How is personal information handled?

Documents and tables are scanned for personal data and masked before delivery, and location and personal metadata never reach buyers. Many public sector datasets use realistic fictional forms and cases written by specialists, which avoids citizen data entirely. Works showing identifiable people require consent on file.

Can you work through our procurement process?

We start with a scoping call and a written dataset spec, which gives your procurement team a concrete document to work from. A pilot batch lets you evaluate quality before a larger commitment. Contracting terms, vendor requirements and security questionnaires are discussed during scoping.

Which languages can you cover?

It depends on the languages in your service plan and on qualified linguists and translators in our network. We confirm coverage during scoping and say plainly where it is thin. Translations go through a second review for terminology and plain-language standards.

Tell us what your model needs to learn.

Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.