Solution · Video datasets

Video with the rights to train on it.

Video models need footage with motion, cause and effect, and clear rights. We supply clips from consenting creators, captioned over time.

Photo needed · 21:9Videographer filming a street sceneA videographer walking with a gimbal-mounted camera through a bright, clean urban scene, a few people in soft focus. Wide editorial frame.Wide image under the opening
Overview

Video datasets, with proof.

Photo needed · 4:5Editor marking a video timelinePortrait close-up of a hand on a trackpad beside a monitor showing a timeline with colored segments. Soft daylight.Beside the overview

Video generation, video understanding and world models all depend on footage that is diverse in subject, motion and camera behavior, and on captions that describe what happens over time rather than a single frame. HUMXN sources video from filmmakers, videographers and other creators who have granted AI-training permission, and commissions new footage where a distribution is missing, such as specific actions, environments, camera moves or egocentric perspectives.

Annotation follows your schema: clip-level captions, dense temporal captions with timestamps, action and event segments, camera motion tags, shot boundaries and object tracks. Experts handle domain footage such as surgical video, industrial processes or sports technique. Each label records its source, and a second-expert pass reviews an agreed share. Technical metadata such as resolution, frame rate, codec and duration is recorded for every clip so you can filter before training.

Clips are fingerprinted at intake with file hashes and frame hashes, so re-encoded or trimmed copies of registered footage are caught. Origin is classified, and generator signatures and C2PA records are checked. Footage showing identifiable people is offered only with consent on file, releases are bound to the clip, and location metadata is stripped. Each clip carries a signed provenance record and appears on your receipt by ID and hash.

Data types
Stock-style footageEgocentric videoTemporal captionsAction segmentsObject tracksCamera motion tagsTechnical metadata
Experts involved
FilmmakersVideographersCaption writersSurgeonsSports coachesIndustrial engineersAnnotation reviewers
What we deliver

Built to your specification.

Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for video datasets.

01

Licensed footage

Clips from consenting filmmakers and videographers across subjects, settings and styles.

02

Commissioned capture

New footage to a shot list for actions, environments or camera moves the catalogue lacks.

03

Temporal captions

Timestamped descriptions of what happens across a clip, written to your style guide.

04

Action and event segments

Start and end times for actions and events against your label taxonomy.

05

Camera and shot metadata

Camera motion, shot boundaries and technical properties for every clip.

06

Egocentric video

First-person recordings of tasks, with participant consent on file.

Why it matters

Where unverified data falls short.

The nine layers of verification

01Time is the hard part

Single-frame captions miss motion and causality. Temporal captions describe what changes, which is what video models must learn.

02Trimmed and re-encoded copies

Frame hashes catch footage that file hashes miss after trimming or transcoding.

03People on screen

Consent and releases bound to each clip make footage with identifiable people usable.

Questions

Video datasets: asked often.

What kinds of video can you source?

Catalogue footage covers the subjects our registered creators work in, and commissioned capture fills gaps: specific actions, environments, camera moves, egocentric perspectives or domain footage such as industrial processes. We confirm what can come from the catalogue and what must be commissioned during scoping, which drives both price and schedule.

How are temporal captions written and checked?

Caption writers follow a style guide agreed with you that defines granularity, timestamp precision and vocabulary. They describe actions, changes and camera behavior over time. A second reviewer checks an agreed share for accuracy and timing, and each caption records its source, including whether an AI-suggested draft was corrected by a person.

How do you catch duplicate footage?

Every clip gets a file hash and frame hashes at intake, compared against everything registered. That catches clips that were trimmed, re-encoded or resized from registered footage, which ordinary file hashes would miss. Conflicting ownership claims are flagged and held back until they are resolved, so disputed footage never reaches your delivery.

Which delivery formats do you support for video?

Clips are delivered as files with metadata and labels in JSON Lines, CSV or Parquet, or as WebDataset shards for streaming training, with Croissant 1.0 metadata describing the set. Resolution, frame rate and codec requirements are set in the specification. Approved buyers can retrieve deliveries through the API.

What affects the price of video data?

Main factors are catalogue versus commissioned footage, resolution and duration, releases required, caption density, segment or tracking annotation, domain expertise and exclusivity. Densely captioned, commissioned footage with people on screen costs the most. We quote per clip or per hour of footage after scoping.

Related

Tell us what your model needs to learn.

Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.