Language data labeled by trained linguists.
Language models handle the languages and dialects they saw most. We commission trained linguists to label the structure, sound and meaning they missed.
Linguistics, with proof.
Speech and language models struggle where structure is rich and data is thin: morphologically complex words split badly by tokenizers, tone and stress ignored in transcription, dialect features treated as errors, pragmatic meaning taken literally. HUMXN commissions credential-checked linguists, phoneticians and fieldworkers to produce phonetic transcription, morphological glossing, syntactic and semantic annotation and dialect labels, to a written scheme that makes every decision consistent.
Speech and text are sourced from consenting speakers and creators, with consent bound to each recording or document. Audio is fingerprinted at intake with acoustic fingerprints and checked for duplicates, documents are scanned for personal information and masked, and label provenance records whether each label came from the author, a model suggestion or an expert reviewer. A second linguist reviews every item, and each delivery carries a signed receipt.
For linguists, this is paid, remote work that uses the analytical skills of your training: transcribing in IPA, glossing morphology, parsing sentences and describing variation. Rates are shown before you accept each job, and you choose work that fits around research or teaching. Native or near-native command of less-resourced languages is especially valuable. Your identity is not published, and approved work is paid automatically through Stripe.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for linguistics.
Phonetic transcription
Narrow and broad IPA transcription with stress, tone and prosody marked, aligned to consented speech recordings.
Morphological glossing
Interlinear glosses and segmentation for morphologically rich languages, following Leipzig-style conventions or your scheme.
Syntactic annotation
Dependency and constituency parses, including for languages with free word order or limited existing treebanks.
Semantic and pragmatic labels
Annotation of reference, implicature, politeness and speech acts, where literal readings mislead models.
Dialect and variety labels
Speech and text labeled by region, variety and register, so models stop treating variation as error.
Low-resource language data
Consented recordings and texts in less-resourced languages, transcribed and glossed by linguists with field experience.
01Inconsistent schemes
Linguistic labels are only useful if applied consistently. Written guidelines and second-linguist review keep them so.
02Dialect erasure
Models trained on standard varieties misread others. Variety-labeled data from consenting speakers corrects that.
03Speaker consent
Voices are personal. Every recording carries a consent record bound to the item and an acoustic fingerprint.
Paid work in your field.
Remote and flexible, with the rate stated before you accept. Every job is reviewed by a second expert, and approved work is paid automatically.
- A graduate degree in linguistics or a closely related field
- Training in phonetic transcription with the IPA
- Experience in a subfield such as syntax, morphology, phonology or semantics
- Fieldwork or documentation experience with less-resourced languages
- Native or near-native command of one or more languages beyond English
Transcribe speech
Produce IPA transcriptions with stress, tone and prosody, aligned to consented recordings.
Gloss and parse
Segment and gloss words and annotate sentence structure following a written scheme.
Label variation and meaning
Tag dialect, register and pragmatic features, and grade model interpretations.
Review another linguist's labels
Act as second reviewer on a colleague's transcriptions and annotations before delivery.
Linguistics: asked often.
Which languages can you cover?
Requests are scoped by language, variety and annotation type, then matched to linguists verified in that language and subfield. We are particularly suited to work that needs trained analysts rather than general transcribers, including less-resourced languages. A pilot batch confirms coverage, the scheme and the standard before the request scales.
Where does the speech data come from?
Recordings come from consenting speakers, with consent and any releases bound to each item before it enters a dataset. Audio is fingerprinted at intake and checked for duplicates, and every item carries signed provenance. You receive a signed receipt listing each recording and its labels by ID and hash.
What qualifications do I need?
Most jobs ask for a graduate degree in linguistics and demonstrated skill in the relevant area, such as IPA transcription or syntactic annotation. Native command of a less-resourced language combined with linguistic training can qualify you for specialist work. We verify identity and credentials and may set a short, paid assessment.
How do I get paid?
Rates are set by field and task and are shown before you accept. Once a second linguist approves your work, it is credited to your balance and paid automatically through Stripe after your identity and tax details are verified. Every job appears on an itemized ledger.
How much time does the work take?
There is no minimum commitment. Opportunities are offered when they match your verified languages and subfield, each with a stated scope and deadline, and you choose which to accept. The work is remote and can fit around research, teaching or fieldwork.
Can we license the data exclusively?
Yes. Commissioned linguistic datasets can be licensed exclusively, with scope, term and exclusivity agreed in writing. Speech data can be delivered as WebDataset or Parquet and annotations as JSONL or Croissant, each with a signed receipt listing every item by ID and hash. Standing orders can extend a dataset to new languages or varieties.
Linguistics, done by experts.
Need expert-made data in this field, or have the expertise to make it? Start here.

