Solution · Image datasets

Images with releases, captions and rights.

Every image arrives with the creator's AI-training permission, any releases it needs, and a caption written by someone who knows what is in the frame.

Photo needed · 21:9Photographer on a bright location shootA photographer framing a still-life scene on a large sunlit set, assistant adjusting a reflector. Wide, airy editorial composition.Wide image under the opening
Overview

Image datasets, with proof.

Photo needed · 4:5Contact sheet with caption notesPortrait close-up of a printed photo contact sheet with neat pencil captions beside several frames. Natural light.Beside the overview

Vision and multimodal models are increasingly limited by how well their images are described and by whether they can be used at all. HUMXN sources photography, illustration and design work from creators who have granted AI-training permission on stated terms, and commissions new shoots where the catalogue has gaps. Images showing identifiable people are only offered for training with consent on file, and model and property releases are bound to the item, so rights questions have documented answers.

Captions and labels are written to your schema: dense descriptive captions, short alt-style captions, object and attribute tags, bounding boxes or segmentation, and domain-specific labels from experts such as radiologists, botanists or engineers. Each label records its source, whether creator, AI-suggested or expert-reviewed, and a second-expert pass checks a share agreed in the spec. Captions can be delivered in COCO captions format alongside JSON Lines, Parquet or WebDataset shards.

Intake fingerprints every image with SHA-256 and a perceptual hash that tolerates crops, resizes and re-compression, then checks it against everything registered. EXIF and XMP generator signatures, IPTC digital-source-type and C2PA records are read to classify origin. Location and personal metadata are stripped before delivery. Licensed deliveries carry forensic delivery marks, and each image has an Ed25519-signed provenance record and appears by ID and hash on your receipt.

Data types
PhotographyIllustrationProduct imageryDense captionsBounding boxesSegmentation masksScientific imagery
Experts involved
PhotographersIllustratorsRadiologistsBotanistsEngineersCaption writersAnnotation reviewers
What we deliver

Built to your specification.

Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for image datasets.

01

Licensed photography

Work from consenting photographers across subjects, styles and conditions, with terms bound to each image.

02

Commissioned shoots

New photography to your shot list where the catalogue does not cover the distribution you need.

03

Dense captions

Detailed descriptions written to your style guide, with short captions alongside where useful.

04

Detection and segmentation

Boxes, polygons and masks to your class list, reviewed for consistency.

05

Expert domain labels

Labels from specialists for medical, scientific, industrial and botanical imagery.

06

Releases on file

Model and property releases bound to images that show identifiable people or private property.

Why it matters

Where unverified data falls short.

The nine layers of verification

01Captions limit multimodal models

Short, generic captions teach shallow grounding. Dense captions from people who understand the subject teach more per image.

02Near-duplicates inflate sets

Crops and re-compressions of the same photo pass file hashes. Crop-tolerant perceptual hashes catch them.

03AI images in human sets

Generator signatures, IPTC source type and C2PA records are checked so synthetic images are labeled, not mixed in silently.

04People and property

Consent and releases on file are what make images with identifiable people usable for training.

Questions

Image datasets: asked often.

Are the images licensed for commercial model training?

Each image carries the licence terms its creator agreed to, including AI-training permission, and those terms are listed with the delivery. Commercial training use is the typical scope, and exclusive terms are available for commissioned work and where creators agree. Your counsel can review the terms for every image before you train.

How do you detect duplicates and near-duplicates?

Every image gets a SHA-256 hash and a perceptual hash that tolerates crops, resizes and re-compression. Both are checked against everything registered at intake, so the same photo submitted twice, or a cropped copy of a registered work, is caught. Conflicting ownership claims are flagged and held back until resolved.

Do you include AI-generated images?

Only if your spec allows them, and always labeled. Intake records each image as human-created, AI-assisted or AI-generated, and reads EXIF and XMP generator signatures, IPTC digital-source-type and C2PA records. If you want human-made images only, anything classified otherwise is excluded from your delivery.

What formats do image deliveries use?

Images are delivered as files with metadata in JSON Lines, CSV or Parquet, or packaged as WebDataset shards for streaming training. Captions are available in COCO captions format, and Croissant 1.0 metadata describes the set. Each record includes labels, label source, licence reference, origin and file hashes.

What determines image dataset pricing?

Pricing depends on how much comes from the catalogue versus commissioned shoots, subject difficulty, releases required, caption depth, annotation type such as boxes or masks, expert labeling, and exclusivity. Dense captions and segmentation cost more than tags. We quote per image against the agreed spec, with volume pricing for larger orders.

Related

Tell us what your model needs to learn.

Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.