Solution · Red teaming & safety data

Red teaming by people who know the domain.

Generic jailbreak lists find generic failures. Domain experts find the ones specific to medicine, law, chemistry, finance and code.

Photo needed · 21:9Security specialist testing a modelA professional at a clean, bright workstation with two monitors showing a chat interface and a notes document. Serious but calm.Wide image under the opening
Overview

Red teaming & safety data, with proof.

Photo needed · 4:5Taxonomy chart on deskPortrait close-up of a printed category and severity grid with a few handwritten marks, pen nearby. Neutral daylight.Beside the overview

The most consequential model failures are often domain-specific: a dosage answer that is plausible and dangerous, legal advice that sounds authoritative and is wrong, a code completion with an exploitable flaw. Finding these needs people who know where the line is in their field. We commission adversarial prompts from credential-checked experts who probe your model against a defined scope, including multi-turn escalation, role-play framings and indirect requests, and record each attempt with its outcome.

Work is organized by a harm taxonomy agreed with you before collection starts, either your existing policy categories or one we draft together. Each prompt and response is labeled by category, severity and whether the model's behavior was acceptable, over-refusing or harmful, with a rationale. Alongside attacks, experts write preferred responses: safe completions that are still genuinely helpful, and refusals phrased the way your policy specifies. That gives you material for training and evaluation, not only a list of failures.

Sensitive outputs are handled deliberately. Work runs in scoped tasks with access limited to the experts assigned, content passes block-list hash screening and a safety classifier, and material that falls outside the agreed scope is escalated rather than collected. Personal data is masked before delivery. Expert participation is voluntary per task with clear content warnings, and the most sensitive categories are restricted to experts with the relevant professional background.

Data types
Adversarial promptsMulti-turn attack transcriptsHarm labelsSeverity ratingsSafe completionsRefusal exemplars
Experts involved
PhysiciansPharmacologistsChemistsAttorneysSecurity engineersFinancial compliance specialistsLinguists
What we deliver

Built to your specification.

Every engagement starts from a written spec and a pilot batch. These are the most common requests we source for red teaming & safety data.

01

Expert adversarial prompts

Single and multi-turn attacks written by specialists in the domains your policy covers.

02

Harm taxonomy design

Categories, severity levels and decision rules agreed with your policy team before collection.

03

Response labeling

Each model output labeled as acceptable, over-refusal or harmful, with category, severity and rationale.

04

Safe completions

Expert-written preferred responses that stay helpful while respecting your policy.

05

Over-refusal sets

Benign prompts near the boundary that a well-calibrated model should answer.

06

Controlled handling

Scoped access, safety screening and escalation rules for sensitive material.

Why it matters

Where unverified data falls short.

The nine layers of verification

01Domain-specific harm

Many dangerous answers look reasonable to a non-expert. Specialists recognize the failure and can explain why it matters.

02Over-refusal costs too

A model that refuses legitimate professional questions is less useful. Boundary sets measure both directions.

03Handling sensitive content

Red-team data includes harmful material by design. Scoped access, screening and escalation rules keep collection inside agreed limits.

Questions

Red teaming & safety data: asked often.

How is a red-teaming engagement scoped?

We start from your usage policy and the risks you most need to measure, then agree a harm taxonomy, the domains in scope, attack styles such as multi-turn or role-play, and what falls outside scope entirely. That document sets which experts are assigned, what they may attempt, and escalation rules. Nothing outside the agreed scope is collected.

How do you handle dangerous or illegal content?

Collection is limited to the agreed scope and run in tasks visible only to assigned experts. Content passes block-list hash screening and a safety classifier, and anything that crosses a hard line is escalated and excluded rather than delivered. Personal data is masked before delivery. You should still apply your own handling controls once data reaches your systems.

Do you only find failures, or also provide training data?

Both. Each attack comes with labels for category, severity and outcome, and experts write preferred responses: safe, helpful completions and well-phrased refusals that follow your policy. Paired with the failing responses, that material can be used for preference training, fine-tuning or as a held-out safety evaluation.

Can red-teaming results support our regulatory work?

Structured findings with a taxonomy, severity labels and a signed record of each item can be useful evidence when you document testing for frameworks your organization must consider. We do not provide legal advice or certify compliance. Your counsel should decide how this material fits your obligations.

What determines the cost of red teaming?

Cost depends on the domains and seniority of experts, whether attacks are single or multi-turn, labeling depth, whether safe completions are written, and the sensitivity of the categories involved. Specialist chemistry or clinical red teaming costs more than general-policy probing. We quote against the scoped plan, with pilot rounds before larger campaigns.

Related

Tell us what your model needs to learn.

Send a brief in five minutes. A data lead replies within one business day with questions and a first sourcing plan.