Robot learning data, recorded by trained operators.
Robot policies learn from demonstrations, and demonstrations are only as good as the people and records behind them. We source both, then verify every episode before it ships.
Robotics, with proof.
Robot foundation models and task-specific policies share one constraint: there is no internet-scale archive of physical interaction to scrape. Every useful episode has to be recorded by someone, on real hardware, in a real scene. HUMXN commissions trained teleoperators and sources recordings from consenting creators to your specification, covering the objects, tasks, environments and embodiments your policy needs to generalize across. You define the task distribution. We scope it into a dataset spec and deliver a pilot batch you can load into your training stack before anything scales.
The details that matter in robot data are easy to lose. Multi-view camera streams must stay synchronized with joint states, gripper commands and force readings. Each episode needs an honest outcome label, and failure-and-recovery episodes are often worth more than clean successes, because they teach a policy what to do after a slip. Our reviewers check episodes for completeness and label accuracy, and a second expert reviews contested outcomes. Labels record their source, so you can tell an operator's call from an expert-reviewed one.
Recording in homes, labs and workplaces means recording people and places. Model and property releases are bound to each item, and recordings showing identifiable people are only offered for training with consent on file. Location and personal metadata never reach buyers. Every delivered episode carries an Ed25519-signed provenance record and appears on a signed receipt by ID and hash. For teams preparing safety cases or customer audits, that is a clear record of where each demonstration came from and under what terms.
Built to your specification.
Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for robotics.
Teleoperation demonstrations
Manipulation episodes recorded by trained operators on the tasks you specify, from pick-and-place to multi-step assembly, with consistent operator protocols.
Multi-view synchronized video
Wrist, head and third-person camera streams aligned with robot state, so policies can learn from the viewpoints they will see at deployment.
Proprioception and force data
Joint positions, velocities, gripper state and force-torque readings recorded alongside video, delivered in structured formats your loaders expect.
Success, failure and recovery labels
Episode outcomes and grasp success labels reviewed by experts, including deliberate failure and recovery episodes that teach correction.
Egocentric human video
First-person recordings of people doing real tasks with their hands, with releases on file, for pre-training visual representations.
3D object and scene assets
Scanned objects, meshes and point clouds in STL, OBJ, PLY and GLB formats for simulation, grasp planning and narrowing the sim-to-real gap.
01The sim-to-real gap
Policies trained only in simulation stumble on real friction, lighting and clutter. Real demonstrations in varied scenes give your model the variation it lacks.
02Embodiment diversity
A policy that has seen one arm in one lab rarely transfers. Sourcing across grippers, platforms and environments broadens what it can handle.
03Unreliable outcome labels
A mislabeled success teaches the wrong behavior. Expert review and a second-expert pass on contested episodes keep outcome labels trustworthy.
04People in the frame
Real-world recordings capture faces and homes. Releases and consent are bound to each item, and personal metadata never reaches you.
Robotics: asked often.
What does a teleoperation dataset include?
Typically synchronized multi-view video, joint states, gripper commands and, where the hardware supports it, force-torque readings, plus a task description and an outcome label per episode. The exact streams, rates and camera placements are set in the dataset spec. We can include failure and recovery episodes in a ratio you choose, because they often teach a policy more than clean successes do.
Can you record on our robot platform?
Recording protocols, task lists and hardware are agreed during scoping. Where your project needs a particular embodiment, we discuss how operators get access to it and what that means for timeline and cost. Where embodiment diversity matters more than one platform, we source across several. Either way the pilot batch shows you real episodes before you commit to volume.
How do you verify the quality of demonstrations?
Each episode is checked for complete, synchronized streams and a correct outcome label. Reviewers flag hesitations, collisions and protocol breaks against the criteria in your spec. Contested outcomes go to a second expert. Every label records whether it came from the operator, an AI suggestion or expert review, so you can filter by confidence in training.
What formats do you deliver robotics data in?
Episode data can be delivered as Parquet or JSON Lines with video in WebDataset shards, with Croissant 1.0 metadata describing the fields. 3D assets come as STL, OBJ, PLY, GLB or point clouds. If your stack expects a particular layout, we agree it in the spec. Approved buyers can also take deliveries through the API.
How is privacy handled in home and workplace recordings?
Model and property releases are bound to each recording, and any recording showing identifiable people is only offered for training when consent is on file. Location and personal metadata are stripped and never reach buyers. Every delivered item is listed on a signed receipt by ID and hash, so your team has a clear record for its own privacy review.
How is robotics data priced?
Cost follows operator time, hardware requirements, the number of synchronized streams, review depth and the licence. Long-horizon, multi-step tasks cost more per episode than short pick-and-place. Exclusive licences cost more than shared ones. Standing orders and volume reduce the unit price. We quote once the spec is agreed.
Tell us what your model needs to learn.
Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.

