Native speakers, not machine translation.
Translated English data teaches a model translated English. We commission data written natively in each language, by the people who speak it.
Multilingual data, with proof.
Models trained mostly on English and then extended with machine-translated data tend to sound translated: literal idioms, wrong register, cultural references that do not land, and weak handling of dialects and code-switching. HUMXN commissions multilingual data from native speakers and professional translators in the network, writing natively in the target language rather than translating, so prompts and responses reflect how people actually ask and answer in that language.
The same data types we offer in English are available across languages: SFT demonstrations, preference comparisons, evaluation items, red-teaming prompts, speech recordings and annotation. Where you need parallel data, professional translators produce translations that are reviewed for adequacy and fluency, and localization adapts content rather than translating it word for word. A style guide per language sets register, formality conventions, script and orthography choices, and dialect coverage.
Every item is reviewed by a second native speaker before acceptance. Origin is recorded, so machine-translated or AI-assisted text is labeled and can be excluded. Text fingerprints catch duplicates against everything registered. Language and dialect metadata is attached to each item. For less widely spoken languages we confirm network coverage during scoping before quoting, rather than promising coverage we cannot staff.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for multilingual data.
Native-written SFT data
Demonstrations written directly in each target language, to a per-language style guide.
Multilingual preference data
Comparisons judged by native speakers on correctness, fluency and cultural fit.
Professional translation
Parallel data from professional translators, reviewed for adequacy and fluency.
Localization
Content adapted to local conventions, references and register rather than translated literally.
Multilingual evaluation sets
Held-out items written natively, never published.
Dialect and code-switching data
Regional variants and mixed-language text where your users write that way.
01Translationese
Machine-translated data carries source-language structure. Native writing teaches the language as it is used.
02Register and politeness
Formality systems differ by language. Per-language style guides and native reviewers get them right.
03Hidden machine translation
Recorded origin keeps machine-translated text from entering a set sold as native-written.
Multilingual data: asked often.
Which languages do you cover?
Coverage depends on the native speakers and translators in the network, which varies by language and domain. Widely spoken languages are generally easier to staff, especially for specialist domains. We confirm coverage for each language and domain combination during scoping and report actual distributions against the spec during production.
Is your multilingual data translated from English?
Not unless you ask for parallel data. By default, native speakers write directly in the target language, which avoids translated phrasing and cultural mismatch. When you need translations, professional translators produce them and a second native speaker reviews them. Origin is recorded for each item, including any machine-translation assistance.
How do you check quality across languages you don't read?
Every item is reviewed by a second native speaker, and reviewers follow the same per-language style guide as writers. Where relevant, overlapping items are labeled by multiple raters so agreement can be measured and reported per language. Reviewer comments and outcomes are stored with each item, so you can audit decisions.
What affects multilingual data pricing?
Price depends on the language and how many qualified native speakers are available, domain expertise, task type, whether translation or native writing is required, review depth and exclusivity. Specialist content in less widely spoken languages costs more. We quote per item and per language after scoping.
Can you handle dialects and code-switching?
Yes, where we can staff native speakers of the variant. Dialect and code-switching requirements go into the style guide and the target distribution, and each item is tagged with language and dialect metadata. If a variant cannot be staffed reliably, we say so during scoping.
Tell us what your model needs to learn.
Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.

