Solution · Speech & audio data

Speech data with the speaker's consent.

A voice is personal data. We source speech and audio from people who agreed to AI training, and bind that consent to every recording.

Photo needed · 21:9Voice actor in recording boothA person speaking into a studio microphone in a bright, acoustically treated booth, script on a stand, engineer visible through glass. Warm, editorial.Wide image under the opening
Overview

Speech & audio data, with proof.

Photo needed · 4:5Microphone and printed scriptPortrait close-up of a condenser microphone and pop filter with a marked-up script page behind it. Soft daylight.Beside the overview

Speech recognition, text-to-speech and audio-language models need recordings that match real conditions: accents, ages, microphones, rooms, background noise and speaking styles. They also need a clear basis for using someone's voice. HUMXN sources speech from consenting speakers and voice professionals, and audio from creators who have granted AI-training permission. Each recording carries the consent and licence terms that apply to it, and speaker metadata is limited to what the speaker agreed to share.

Collection follows a written specification covering scripted or spontaneous speech, read prompts, conversational pairs, recording conditions, sample rate and bit depth, and target distributions across languages, dialects and speaker attributes. Transcripts are produced or corrected by trained transcribers and reviewed by a second listener, with conventions for disfluencies, overlaps, non-speech events and timestamps agreed up front. Non-speech audio such as music, ambience and sound events is tagged to the label set you specify.

Every recording is fingerprinted at intake with a SHA-256 hash and a Chromaprint acoustic fingerprint, then checked against everything registered, which catches re-uploads and re-encoded duplicates. Origin is recorded as human-created, AI-assisted or AI-generated, so synthetic voices do not enter a set sold as real speech. Deliveries include audio, transcripts and metadata in your chosen format, with a signed receipt listing each recording by ID and hash.

Data types
Scripted speechConversational speechTranscriptsTimestamps & alignmentsMusicSound eventsSpeaker metadata
Experts involved
Voice actorsTranscribersLinguistsPhoneticiansAudio engineersMusicians
What we deliver

Built to your specification.

Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for speech & audio data.

01

Scripted speech

Read prompts recorded to your script, conditions and speaker distribution.

02

Conversational speech

Spontaneous two-party and multi-party conversations with speaker turns marked.

03

Verified transcripts

Transcription with timestamps and agreed conventions, reviewed by a second listener.

04

Voice talent recordings

Studio-quality sessions from professional voice actors for TTS and expressive speech.

05

Non-speech audio

Music, ambience and sound events from consenting creators, tagged to your label set.

06

Speaker metadata

Language, dialect, age band and recording conditions, limited to what speakers agreed to share.

Why it matters

Where unverified data falls short.

The nine layers of verification

01Voices need consent

A recording identifies a person. Consent bound to each file is the basis for training on it, and for answering questions about it later.

02Duplicates hide in re-encodes

The same clip re-encoded at a different bitrate defeats file hashes. Acoustic fingerprints catch it.

03Synthetic voices in real-speech sets

Recorded origin and generator checks keep TTS output out of data sold as human speech.

04Transcript conventions matter

Inconsistent handling of disfluencies and overlaps adds label noise. Agreed conventions and a second listener reduce it.

Questions

Speech & audio data: asked often.

How do you obtain consent from speakers?

Speakers and voice professionals agree to AI-training use and the applicable licence terms before recording, and that record is bound to each file. Speaker metadata is limited to what they agreed to share. Recordings without consent on file are not offered for training. Consent terms are visible to you as part of the licence for each recording.

Which languages and accents can you cover?

Coverage depends on the speakers in the network and the catalogue for each language, and can be expanded through commissioned collection. We set target distributions for languages, dialects and speaker attributes in the specification and report actual coverage against them during production. For less widely spoken languages we confirm feasibility during scoping before quoting.

What audio formats and transcript formats do you deliver?

Audio is delivered in lossless formats at the sample rate and bit depth set in your specification. Transcripts and metadata come as JSON Lines, CSV or Parquet, and large sets can be packaged as WebDataset shards with Croissant 1.0 metadata. Timestamps, speaker turns and non-speech tags follow the conventions agreed during scoping.

How do you check that recordings are real human speech?

Each recording's origin is recorded at intake, and metadata is checked for generator signatures and C2PA records. Acoustic fingerprints are compared against everything registered to catch duplicates. Commissioned recordings are captured in sessions tied to verified speakers. If your spec excludes AI-generated or AI-processed audio, those items are held back.

What drives the price of speech data?

Price depends on the language and how easy speakers are to find, recording conditions, scripted versus conversational speech, transcription depth, review requirements, voice-talent rates for studio work, and exclusivity. Exclusive voice recordings for TTS usually cost the most. We quote per recorded hour or per item after scoping.

Related

Tell us what your model needs to learn.

Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.