Industry · Media & entertainment

Creative data, licensed from the people who made it.

Generative media models are only as defensible as the work they learn from. We license it from creators who said yes, and keep the record that proves it.

Photo needed · 21:9Musicians recording in bright studioWide editorial shot of a daylight recording studio with musicians mid-take and microphones in frame. Warm, natural and unposed.Wide image under the opening
Overview

Media & entertainment, with proof.

Photo needed · 4:5Editor's hands on mixing consolePortrait close-up of hands on faders with soft studio light. Tactile and precise.Beside the overview

Studios, streaming platforms, music companies and generative media teams build models for editing, captioning, dubbing, music generation, image and video synthesis, and content understanding. The common question is no longer whether the model works but whether the training data was licensed. HUMXN sources music, audio, film, photography, illustration and writing from a registered catalogue of consenting creators, and commissions new work to your specification. Every work records its AI-training permission and licence terms. You define the content you need. We deliver a pilot batch first.

Creative datasets are only as clean as their intake. Audio is identified with Chromaprint acoustic fingerprints, video with frame hashes, images with crop-tolerant perceptual hashes and text with text fingerprints, and every work is checked for duplicates and conflicts against everything registered. That catches the same track uploaded twice or a film still claimed by two people. Each work records whether it is human-created, AI-assisted or AI-generated, and generator signatures and C2PA records are detected, so synthetic material does not pass as human craft.

Creative work also needs expert description. Musicians, editors, writers and translators label structure, mood, technique and shot type, and write captions dense enough to ground a multimodal model. Labels record their source, and contested calls get a second review. Licensed image deliveries carry forensic delivery marks, items carry Ed25519-signed provenance and a C2PA manifest where the format allows, and the signed receipt lists every work by ID and hash. Creators are paid for their work, and you get a record your counsel can read.

Data types
Music and stemsSpeech and voiceFilm footagePhotographyIllustrationScripts and proseCaptions
Experts involved
MusiciansMusic theoristsFilm editorsPhotographersScreenwritersTranslatorsVoice performers
What we deliver

Built to your specification.

Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for media & entertainment.

01

Licensed music and audio

Tracks, stems, performances and sound design from consenting creators, with training permission and licence terms bound to each work.

02

Film and video footage

Licensed footage across genres and shot types, with releases on file for identifiable people and frame-level fingerprints at intake.

03

Photography and illustration

Human-made images with human-origin records, forensic delivery marks on licensed deliveries and C2PA manifests where supported.

04

Expert creative annotation

Captions, structure, mood, instrumentation, technique and shot-type labels from musicians, editors and visual artists.

05

Writing and scripts

Licensed fiction, nonfiction and screenwriting, plus commissioned writing in styles and genres your model is thin on.

06

Dubbing and localization data

Translated scripts and voice recordings from professional translators and performers, with consent for voice use on file.

Why it matters

Where unverified data falls short.

The nine layers of verification

01Licensing exposure

Training on unlicensed creative work invites disputes. Every work we deliver carries explicit AI-training permission from its creator.

02Duplicate and disputed works

Acoustic, frame and perceptual fingerprints catch duplicates and conflicting claims before they enter your dataset.

03Synthetic contamination

AI-generated media mislabeled as human work degrades quality. Human-origin records and signature detection keep it identified.

04Voice and likeness

Recordings and images of identifiable people are offered for training only with consent and releases on file.

Questions

Media & entertainment: asked often.

Do creators actually consent to AI training?

Yes. Every work in the registered catalogue records the creator's AI-training permission and the licence terms they accepted. Works showing identifiable people are offered for training only with consent and releases on file. Creators register their own work and are paid when it is licensed, and the signed receipt lists each work you receive by ID and hash.

How do you stop the same work being sold twice or claimed by the wrong person?

Every work is fingerprinted at intake: Chromaprint for audio, frame hashes for video, perceptual hashes for images, text fingerprints for writing. Each new upload is checked for duplicates and conflicts against everything registered, and conflicts are held for review rather than delivered.

Can we get an exclusive licence?

Yes, for commissioned work and for catalogue works where the creator offers it. Exclusive licences cost more than non-exclusive ones and are set out in the licence terms attached to each item. Your counsel should review the terms against your intended use; we provide the documentation, not legal advice.

What formats do you deliver media datasets in?

Audio, video and images are typically delivered in WebDataset shards, image captions in COCO format, and metadata and labels in JSON Lines or Parquet, with Croissant 1.0 describing the dataset. Licensed image deliveries carry forensic delivery marks. Approved buyers can use the API.

How is creative data priced?

By medium, licence scope, exclusivity, annotation depth and volume. Multitrack stems with expert annotation cost more than captioned stills, and exclusive licences cost more than shared ones. Standing orders and volume reduce the unit price. We quote after scoping, once the spec is written.

Related

Tell us what your model needs to learn.

Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.