Video with the rights to train on it.
Video models need footage with motion, cause and effect, and clear rights. We supply clips from consenting creators, captioned over time.
Video datasets, with proof.
Video generation, video understanding and world models all depend on footage that is diverse in subject, motion and camera behavior, and on captions that describe what happens over time rather than a single frame. HUMXN sources video from filmmakers, videographers and other creators who have granted AI-training permission, and commissions new footage where a distribution is missing, such as specific actions, environments, camera moves or egocentric perspectives.
Annotation follows your schema: clip-level captions, dense temporal captions with timestamps, action and event segments, camera motion tags, shot boundaries and object tracks. Experts handle domain footage such as surgical video, industrial processes or sports technique. Each label records its source, and a second-expert pass reviews an agreed share. Technical metadata such as resolution, frame rate, codec and duration is recorded for every clip so you can filter before training.
Clips are fingerprinted at intake with file hashes and frame hashes, so re-encoded or trimmed copies of registered footage are caught. Origin is classified, and generator signatures and C2PA records are checked. Footage showing identifiable people is offered only with consent on file, releases are bound to the clip, and location metadata is stripped. Each clip carries a signed provenance record and appears on your receipt by ID and hash.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for video datasets.
Licensed footage
Clips from consenting filmmakers and videographers across subjects, settings and styles.
Commissioned capture
New footage to a shot list for actions, environments or camera moves the catalogue lacks.
Temporal captions
Timestamped descriptions of what happens across a clip, written to your style guide.
Action and event segments
Start and end times for actions and events against your label taxonomy.
Camera and shot metadata
Camera motion, shot boundaries and technical properties for every clip.
Egocentric video
First-person recordings of tasks, with participant consent on file.
01Time is the hard part
Single-frame captions miss motion and causality. Temporal captions describe what changes, which is what video models must learn.
02Trimmed and re-encoded copies
Frame hashes catch footage that file hashes miss after trimming or transcoding.
03People on screen
Consent and releases bound to each clip make footage with identifiable people usable.
Video datasets: asked often.
What kinds of video can you source?
Catalogue footage covers the subjects our registered creators work in, and commissioned capture fills gaps: specific actions, environments, camera moves, egocentric perspectives or domain footage such as industrial processes. We confirm what can come from the catalogue and what must be commissioned during scoping, which drives both price and schedule.
How are temporal captions written and checked?
Caption writers follow a style guide agreed with you that defines granularity, timestamp precision and vocabulary. They describe actions, changes and camera behavior over time. A second reviewer checks an agreed share for accuracy and timing, and each caption records its source, including whether an AI-suggested draft was corrected by a person.
How do you catch duplicate footage?
Every clip gets a file hash and frame hashes at intake, compared against everything registered. That catches clips that were trimmed, re-encoded or resized from registered footage, which ordinary file hashes would miss. Conflicting ownership claims are flagged and held back until they are resolved, so disputed footage never reaches your delivery.
Which delivery formats do you support for video?
Clips are delivered as files with metadata and labels in JSON Lines, CSV or Parquet, or as WebDataset shards for streaming training, with Croissant 1.0 metadata describing the set. Resolution, frame rate and codec requirements are set in the specification. Approved buyers can retrieve deliveries through the API.
What affects the price of video data?
Main factors are catalogue versus commissioned footage, resolution and duration, releases required, caption density, segment or tracking annotation, domain expertise and exclusivity. Densely captioned, commissioned footage with people on screen costs the most. We quote per clip or per hour of footage after scoping.
Tell us what your model needs to learn.
Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.

