Creative data, licensed from the people who made it.
Generative media models are only as defensible as the work they learn from. We license it from creators who said yes, and keep the record that proves it.
Media & entertainment, with proof.
Studios, streaming platforms, music companies and generative media teams build models for editing, captioning, dubbing, music generation, image and video synthesis, and content understanding. The common question is no longer whether the model works but whether the training data was licensed. HUMXN sources music, audio, film, photography, illustration and writing from a registered catalogue of consenting creators, and commissions new work to your specification. Every work records its AI-training permission and licence terms. You define the content you need. We deliver a pilot batch first.
Creative datasets are only as clean as their intake. Audio is identified with Chromaprint acoustic fingerprints, video with frame hashes, images with crop-tolerant perceptual hashes and text with text fingerprints, and every work is checked for duplicates and conflicts against everything registered. That catches the same track uploaded twice or a film still claimed by two people. Each work records whether it is human-created, AI-assisted or AI-generated, and generator signatures and C2PA records are detected, so synthetic material does not pass as human craft.
Creative work also needs expert description. Musicians, editors, writers and translators label structure, mood, technique and shot type, and write captions dense enough to ground a multimodal model. Labels record their source, and contested calls get a second review. Licensed image deliveries carry forensic delivery marks, items carry Ed25519-signed provenance and a C2PA manifest where the format allows, and the signed receipt lists every work by ID and hash. Creators are paid for their work, and you get a record your counsel can read.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for media & entertainment.
Licensed music and audio
Tracks, stems, performances and sound design from consenting creators, with training permission and licence terms bound to each work.
Film and video footage
Licensed footage across genres and shot types, with releases on file for identifiable people and frame-level fingerprints at intake.
Photography and illustration
Human-made images with human-origin records, forensic delivery marks on licensed deliveries and C2PA manifests where supported.
Expert creative annotation
Captions, structure, mood, instrumentation, technique and shot-type labels from musicians, editors and visual artists.
Writing and scripts
Licensed fiction, nonfiction and screenwriting, plus commissioned writing in styles and genres your model is thin on.
Dubbing and localization data
Translated scripts and voice recordings from professional translators and performers, with consent for voice use on file.
01Licensing exposure
Training on unlicensed creative work invites disputes. Every work we deliver carries explicit AI-training permission from its creator.
02Duplicate and disputed works
Acoustic, frame and perceptual fingerprints catch duplicates and conflicting claims before they enter your dataset.
03Synthetic contamination
AI-generated media mislabeled as human work degrades quality. Human-origin records and signature detection keep it identified.
04Voice and likeness
Recordings and images of identifiable people are offered for training only with consent and releases on file.
Media & entertainment: asked often.
Do creators actually consent to AI training?
Yes. Every work in the registered catalogue records the creator's AI-training permission and the licence terms they accepted. Works showing identifiable people are offered for training only with consent and releases on file. Creators register their own work and are paid when it is licensed, and the signed receipt lists each work you receive by ID and hash.
How do you stop the same work being sold twice or claimed by the wrong person?
Every work is fingerprinted at intake: Chromaprint for audio, frame hashes for video, perceptual hashes for images, text fingerprints for writing. Each new upload is checked for duplicates and conflicts against everything registered, and conflicts are held for review rather than delivered.
Can we get an exclusive licence?
Yes, for commissioned work and for catalogue works where the creator offers it. Exclusive licences cost more than non-exclusive ones and are set out in the licence terms attached to each item. Your counsel should review the terms against your intended use; we provide the documentation, not legal advice.
What formats do you deliver media datasets in?
Audio, video and images are typically delivered in WebDataset shards, image captions in COCO format, and metadata and labels in JSON Lines or Parquet, with Croissant 1.0 describing the dataset. Licensed image deliveries carry forensic delivery marks. Approved buyers can use the API.
How is creative data priced?
By medium, licence scope, exclusivity, annotation depth and volume. Multitrack stems with expert annotation cost more than captioned stills, and exclusive licences cost more than shared ones. Standing orders and volume reduce the unit price. We quote after scoping, once the spec is written.
Tell us what your model needs to learn.
Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.

