Driving data for the scenes that break models.
Common miles are easy to collect. The scenes that matter are rare, varied and hard to label well, and they need documentation your safety team can stand behind.
Autonomous vehicles, with proof.
Autonomous driving stacks are trained on enormous volumes of ordinary driving, then fail on the unusual: a cyclist half-hidden by a van, a lane closure marked only by hand-placed cones, glare on wet asphalt at dusk. HUMXN sources footage and sensor data against a written edge-case specification, from consenting creators in our registered catalogue, and commissions experts to label and review it. You tell us which scenarios your models miss. We scope them into a dataset spec and deliver a pilot batch you can run through your own evaluation.
Labels in driving data are only useful when they are consistent. Bounding boxes, segmentation, occlusion flags, intent tags for pedestrians and cyclists all depend on clear guidelines and reviewers who follow them. Our annotators work to the rubric in your spec, labels record whether they came from a creator, an AI suggestion or expert review, and a second expert reviews disputed frames. Video is fingerprinted frame by frame at intake, so duplicate clips and near-duplicates surface before they inflate your training set.
Road footage captures bystanders, license plates and private property. Releases are bound to each item, works showing identifiable people are offered for training only with consent on file, and location and personal metadata never reach buyers. Each item ships with an Ed25519-signed provenance record and a signed receipt by ID and hash. That documentation is designed to support the data traceability work your safety case and regulators may ask for, without standing in for your own legal review.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for autonomous vehicles.
Edge-case driving video
Footage sourced against a written scenario list, such as night rain, construction zones, unusual road users and atypical intersections.
Scene and object annotation
Bounding boxes, segmentation masks, occlusion and truncation flags, and lane and sign labels drawn to your annotation guideline.
Road-user intent labels
Pedestrian, cyclist and driver intent tags applied by trained reviewers, for prediction models that must anticipate what happens next.
Point cloud and 3D data
Point clouds and 3D scene assets in PLY and related formats, labeled for objects and free space where your perception stack needs them.
Scenario descriptions for planning
Natural-language descriptions and expert critiques of driving scenes, for vision-language models and planning evaluation.
Held-out evaluation sets
Exclusive, never-published scenario sets for regression testing, fingerprinted so you can confirm they stay out of training.
01The long tail
Rare scenes are rare in every fleet log. Sourcing to a scenario list concentrates effort on the cases that actually cause disengagements.
02Label inconsistency
Ambiguous occlusion and intent calls vary by annotator. A written rubric and second-expert review on disputed frames keep labels consistent.
03Bystander privacy
Street footage captures people who never agreed to it. Releases, consent checks and metadata stripping are bound to every delivered item.
04Traceable training data
Safety reviews ask where data came from. Signed provenance and per-item receipts give you a record to point to.
Autonomous vehicles: asked often.
Can you source specific driving scenarios?
Yes. During scoping we turn the failure modes you are seeing into a written scenario list, with conditions such as weather, time of day, road type and road users. Sourcing then targets that list rather than general driving. The pilot batch shows how closely we can match it before you commit to volume, and standing orders can extend coverage as your models improve.
How are license plates and faces handled?
Works showing identifiable people are offered for training only with consent on file, and releases are bound to each item. Location and personal metadata are stripped before delivery and never reach buyers. If your privacy review requires additional masking of faces or plates in imagery, we agree that in the spec so it is applied before delivery.
Which annotation types do you support for driving data?
Bounding boxes, segmentation masks, occlusion and truncation flags, lane and sign labels, intent tags and free-text scene descriptions, all drawn to your guideline. Each label records its source, and disputed frames go through a second-expert review. If your taxonomy is unusual, we pilot it on a small batch first so ambiguities surface early.
What delivery formats are available?
Images and captions in COCO format, video in WebDataset shards, labels and metadata in JSON Lines or Parquet, and point clouds in PLY. Croissant 1.0 metadata can describe the full dataset. Approved buyers can pull deliveries through the API. Every delivery includes a signed receipt listing each item by ID and hash.
How does pricing work for AV datasets?
Price depends on scenario rarity, annotation density, the number of label types, review depth and licence scope. Dense segmentation of a rare night scene costs more than boxes on daytime highway frames. Exclusive licences cost more than shared ones, and volume or standing orders bring the unit price down. We quote after the spec is agreed.
Tell us what your model needs to learn.
Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.

