Red Teaming
Red teaming probes a system with adversarial cases to find unsafe, insecure or policy-violating behaviour before it ships. It is not the same activity as measuring quality, and it does not replace it.
| Asks | Answers | |
|---|---|---|
| Evaluation | Does this behave as intended on expected input? | A rate. |
| Red teaming | How does it behave under abusive, adversarial or unexpected input? | A list of failures. |
A mature programme runs both. Evaluation tells you the policy works; red teaming tells you how it breaks.
Scope it before you start
Agree in writing what is in scope, who is authorised to run the exercise, and what the testers may do. Only probe systems you own or are expressly authorised to test — an integration usually sits in front of a model, a data store and a set of downstream actions that may belong to someone else.
Two boundaries worth stating explicitly:
- Production versus a copy. Adversarial traffic against production pollutes your own metrics and may trip your own abuse controls.
- Downstream effects. If a
passverdict causes an action — a message sent, a record written, a robot moving — decide whether the exercise stops at the verdict or continues into the action.
What to attack
A policy-backed integration has four surfaces, and they fail differently.
The policy text
Rules are written in plain English, which is what makes them auditable and also what makes them contestable. Attack the wording:
- The carve-out. Nearly every useful rule has an
unless. Craft input that asserts the exemption without satisfying it. - Scope drift. A rule about "customers" — does it cover a prospect, a former customer, a named third party?
- Unstated assumptions. A rule that assumes English, or a particular format, or that the subject is a person.
The integration around the check
Most real failures are here rather than in the verdict.
- Is every path checked, or only the obvious one? Retries, edited messages, attachments, batch imports and admin tools are the usual gaps.
- Is the content checked the same content that is used? A check on the input before normalisation, truncation or template expansion is a check on something else.
- What happens on error? If an unreachable service, a timeout or an out-of-credit error is caught and treated as a pass, the safety layer is optional under load — which is exactly when an attacker will put it under load.
The verdict handling
- Does anything treat a falsy check on the verdict as safe?
unknownis a string, and a truthy one. - Is
blockenforced, or logged? A verdict that only writes a log line is a metric, not a control. - Can a user see the
reasontext? It is written to be legible, which also makes it a map of the policy if you expose it verbatim.
The operational surface
- Key handling — see Revoke compromised API keys.
- Who can call
Policies.save? Overwriting a policy set silently changes what every check means.save()is an upsert by default;{ ifNotExists: true }is how you stop a second caller from taking a name. - Who can call
Policies.remove? A deleted set makes checks against it fail rather than pass, but they do fail.
Run it as a loop
- Define what would count as a finding, before testing. "The model said
something odd" is not actionable; "content matching rule 2 received
pass" is. - Generate cases — by hand for the ones that need judgement, and programmatically for volume. Paraphrase, translate, obfuscate, and restate every case you already have.
- Run them against a copy of the integration, capturing the policy set's
text with every result (
Policies.getreturns it). Without it you cannot later tell a real regression from a policy that was edited underneath you. - Triage. Separate a policy that is wrong from an integration that is wrong. They have different fixes: one is a wording change, the other is code.
- Fix, then keep the case. Every confirmed finding becomes a permanent test. This is the part that is usually skipped, and it is the part that makes the exercise compound.
Red teaming dataset and testing code
Tracked as CR-5 in the feature backlog. Nothing described here ships today; the section is kept in draft so an integration knows what will be provided and what it still has to write itself in the meantime.
Chidakashi will publish an adversarial dataset and the code that runs it, so that the loop above starts from a corpus rather than from a blank file. The step that stops most red team exercises is step 2 — generating enough cases to be worth triaging.
What it is intended to contain:
| Part | What it gives you |
|---|---|
| The corpus | Adversarial inputs grouped by the surface they attack: carve-out abuse, paraphrase and indirection, injection through content, boundary cases. |
| Expected verdicts | Each case labelled against the policy set it is written for, so a run produces a pass/fail list rather than a pile of verdicts to read. |
| The runner | Code that executes the corpus against your own deployment and policy sets, recording the policy text each case ran against. |
| A regression format | The same file shape for cases you add yourself, so your findings and the shipped corpus run in one pass. |
Two limits it will not remove. A published corpus is a floor, not a certificate — every case in it is a case an attacker can read too. And the labels hold against the policy sets the corpus ships with; run it against your own wording and some expectations will be wrong for good reasons.
Until it exists, the loop under Run it as a loop is hand-rolled: your own cases, your own runner, your own expected verdicts.
Red teaming platform for adversarial testing
Tracked as CR-6 in the feature backlog, and not available today.
Where the dataset above is something you run, the platform is somewhere the exercise lives: a hosted surface for generating, running and tracking adversarial testing against your policy sets over time.
What it is intended to do:
- Generate cases against your policy text, rather than against a generic
taxonomy. A rule with an
unlessclause implies the attacks on that clause; that derivation is mechanical and worth automating. - Run campaigns on a schedule, so a policy edit is followed by a re-run rather than by an assumption.
- Track findings across runs, keyed on the exact policy text. This is the part a local script does badly — telling a real regression from a policy that was edited underneath the suite needs history, not one run.
- Separate a policy failure from an integration failure in the report, because they have different owners and different fixes.
It does not replace the manual exercise. Judgement cases, downstream effects and the scoping conversation above stay human.
After the exercise
Write down what was tested and what was not. A red team report that lists only findings reads as though everything else is safe, which is rarely what the exercise established. The untested surface is a finding of its own.