Teleoperation, manipulation and sensor data
Demonstrations from trained operators, synchronized video, proprioception and force, with task success reviewed by robotics specialists.
Labs, robotics companies and ML teams bring us a capability they're trying to build. We source the data from consenting creators and vetted experts, verify every item, and deliver it licensed and ready to train on.

Licensed, human-made corpora in the domains your model is thin on, with clear rights for commercial use.
Expert-written demonstrations and reasoning traces, built to your rubric and reviewed by a second specialist.
Pairwise and rubric-graded comparisons from people qualified to judge the answer, not just its tone.
Held-out, never-published test sets written by domain experts, so your benchmarks measure capability, not memorization.
Teleoperation and manipulation demonstrations, egocentric video and sensor streams, with success labels.
Images, audio and video with dense, expert captions and the releases to use people and places in them.
You see real samples before anything scales, and every batch after that is held to the acceptance criteria you approved.
Thirty minutes with a data lead. We pin down the capability, the failure modes you're seeing, the volume and the licence you need.
We source a small, representative batch from the catalogue and the expert network so you can test it in your pipeline before committing.
Sourcing scales to the full volume with the acceptance criteria from the pilot. Every item is verified before it reaches you.
As your model improves its needs shift. Standing orders keep new, matching work flowing from the same network at the same bar.
Demonstrations from trained operators, synchronized video, proprioception and force, with task success reviewed by robotics specialists.
Consented recordings with speaker releases, transcripts and expert phonetic labels.
Reasoning traces, rubrics, preference pairs and evaluations written by people who practice the field.
Camera originals with releases, human-made origin checks and dense captions.
Meshes, point clouds and drawings with units and measured-versus-modelled labels.
Egocentric, instructional and cinematic video with frame-level fingerprints and scene annotations.
Contracts, filings and tables with schema, personal-data scans and masking before anything ships.
Delivered in the format you already train on, with labels, technical metadata, origin and the hashes to check every file.

One record per item with labels, provenance and keyed download links.
Columnar exports for large tabular and caption sets.
Sharded tar archives of files plus metadata, for streaming training.
Captions in COCO format and dataset cards in Croissant 1.0 JSON-LD.
By modality, the expertise involved, volume, exclusivity and licence term. Catalogue works carry the creator's training price; commissioned work is quoted against your spec. Larger orders get volume discounts.
Yes. An exclusive licence closes the work to other buyers for its term. Exclusivity is priced into the quote and written into the licence.
Documents and structured data are scanned for personal data and masked before delivery. Works showing identifiable people are only offered when consent for AI training is on file.
Approved buyers get an API key for searching curated selections, placing orders, posting requests and pulling exports, plus standing orders that deliver new matching work automatically.
Five minutes to send a brief. A data lead replies within one business day with questions and a first sourcing plan.