Skip to main content

Red Teaming

Red teaming probes a system with adversarial cases to find unsafe, insecure or policy-violating behaviour before it ships. It is not the same activity as measuring quality, and it does not replace it.

AsksAnswers
EvaluationDoes this behave as intended on expected input?A rate.
Red teamingHow does it behave under abusive, adversarial or unexpected input?A list of failures.

A mature programme runs both. Evaluation tells you the policy works; red teaming tells you how it breaks.

Scope it before you start​

Agree in writing what is in scope, who is authorised to run the exercise, and what the testers may do. Only probe systems you own or are expressly authorised to test — an integration usually sits in front of a model, a data store and a set of downstream actions that may belong to someone else.

Two boundaries worth stating explicitly:

  • Production versus a copy. Adversarial traffic against production pollutes your own metrics and may trip your own abuse controls.
  • Downstream effects. If a pass verdict causes an action — a message sent, a record written, a robot moving — decide whether the exercise stops at the verdict or continues into the action.

What to attack​

A policy-backed integration has four surfaces, and they fail differently.

The policy text​

Rules are written in plain English, which is what makes them auditable and also what makes them contestable. Attack the wording:

  • The carve-out. Nearly every useful rule has an unless. Craft input that asserts the exemption without satisfying it.
  • Scope drift. A rule about "customers" — does it cover a prospect, a former customer, a named third party?
  • Unstated assumptions. A rule that assumes English, or a particular format, or that the subject is a person.

The integration around the check​

Most real failures are here rather than in the verdict.

  • Is every path checked, or only the obvious one? Retries, edited messages, attachments, batch imports and admin tools are the usual gaps.
  • Is the content checked the same content that is used? A check on the input before normalisation, truncation or template expansion is a check on something else.
  • What happens on error? If an unreachable service, a timeout or an out-of-credit error is caught and treated as a pass, the safety layer is optional under load — which is exactly when an attacker will put it under load.

The verdict handling​

  • Does anything treat a falsy check on the verdict as safe? unknown is a string, and a truthy one.
  • Is block enforced, or logged? A verdict that only writes a log line is a metric, not a control.
  • Can a user see the reason text? It is written to be legible, which also makes it a map of the policy if you expose it verbatim.

The operational surface​

  • Key handling — see Revoke compromised API keys.
  • Who can call Policies.save? Overwriting a policy set silently changes what every check means. save() is an upsert by default; { ifNotExists: true } is how you stop a second caller from taking a name.
  • Who can call Policies.remove? A deleted set makes checks against it fail rather than pass, but they do fail.

Run it as a loop​

  1. Define what would count as a finding, before testing. "The model said something odd" is not actionable; "content matching rule 2 received pass" is.
  2. Generate cases — by hand for the ones that need judgement, and programmatically for volume. Paraphrase, translate, obfuscate, and restate every case you already have.
  3. Run them against a copy of the integration, capturing the policy set's text with every result (Policies.get returns it). Without it you cannot later tell a real regression from a policy that was edited underneath you.
  4. Triage. Separate a policy that is wrong from an integration that is wrong. They have different fixes: one is a wording change, the other is code.
  5. Fix, then keep the case. Every confirmed finding becomes a permanent test. This is the part that is usually skipped, and it is the part that makes the exercise compound.
Draft

Red teaming dataset and testing code​

Tracked as CR-5 in the feature backlog. Nothing described here ships today; the section is kept in draft so an integration knows what will be provided and what it still has to write itself in the meantime.

Chidakashi will publish an adversarial dataset and the code that runs it, so that the loop above starts from a corpus rather than from a blank file. The step that stops most red team exercises is step 2 — generating enough cases to be worth triaging.

What it is intended to contain:

PartWhat it gives you
The corpusAdversarial inputs grouped by the surface they attack: carve-out abuse, paraphrase and indirection, injection through content, boundary cases.
Expected verdictsEach case labelled against the policy set it is written for, so a run produces a pass/fail list rather than a pile of verdicts to read.
The runnerCode that executes the corpus against your own deployment and policy sets, recording the policy text each case ran against.
A regression formatThe same file shape for cases you add yourself, so your findings and the shipped corpus run in one pass.

Two limits it will not remove. A published corpus is a floor, not a certificate — every case in it is a case an attacker can read too. And the labels hold against the policy sets the corpus ships with; run it against your own wording and some expectations will be wrong for good reasons.

Until it exists, the loop under Run it as a loop is hand-rolled: your own cases, your own runner, your own expected verdicts.

Draft

Red teaming platform for adversarial testing​

Tracked as CR-6 in the feature backlog, and not available today.

Where the dataset above is something you run, the platform is somewhere the exercise lives: a hosted surface for generating, running and tracking adversarial testing against your policy sets over time.

What it is intended to do:

  • Generate cases against your policy text, rather than against a generic taxonomy. A rule with an unless clause implies the attacks on that clause; that derivation is mechanical and worth automating.
  • Run campaigns on a schedule, so a policy edit is followed by a re-run rather than by an assumption.
  • Track findings across runs, keyed on the exact policy text. This is the part a local script does badly — telling a real regression from a policy that was edited underneath the suite needs history, not one run.
  • Separate a policy failure from an integration failure in the report, because they have different owners and different fixes.

It does not replace the manual exercise. Judgement cases, downstream effects and the scoping conversation above stay human.

After the exercise​

Write down what was tested and what was not. A red team report that lists only findings reads as though everything else is safe, which is rarely what the exercise established. The untested surface is a finding of its own.