Safety Best Practices
A policy check is one control in a system, not the whole of it. The practices below are what an integration needs around the check for the result to mean something in production.
Adversarial testing
Test your integration against inputs written by someone trying to get past it, not only against the inputs you expect. A policy set that holds on representative traffic can still be trivially steerable.
What to probe:
- Paraphrase and indirection. The same prohibited content stated obliquely, in another language, or as a hypothetical.
- Carve-out abuse. Most useful policies have an
unlessclause. Write inputs that falsely claim the carve-out applies, and confirm the verdict does not turn on the claim alone. - Injection through content. If the text you check was authored by a user, assume it contains instructions aimed at whatever reads it next.
- Boundary content. Inputs that sit deliberately close to the line, in both directions. False positives cost as much as false negatives once a policy is enforcing.
The dashboard's Playground is the right place to do this before the policy reaches production: write the policy, throw adversarial inputs at it, and revise the wording rather than the code.
Keep the cases you find. A set of adversarial inputs with expected verdicts is a regression suite, and it is the thing that tells you whether a policy edit made the policy better or merely different.
Human in the loop
Route to a human wherever an automated verdict is not sufficient on its own — high-stakes decisions, irreversible actions, and anything a customer can appeal.
The platform gives you one explicit hand-off point. Image checks return
unknown when the image is too blurry, dark, occluded, cropped or otherwise
unreadable for the policy. unknown is not a pass.
import { isSafe } from '@chidakashi-ai/safety';
const policySet = 'generated-media-safety';
const result = await client.checkImage(file, policySet, { showReason: true });
if (!isSafe(result)) {
await reviewQueue.create({
file,
policySet,
verdict: result.verdict,
reason: result.reason,
});
}
isSafe() returns true only for pass, so unknown — and any verdict a
future deployment adds — lands in the queue rather than slipping through a
truthiness check.
Give the reviewer what they need to decide. Ask for the reason with
{ showReason: true } and keep it with the original content and the policy set
name. A queue entry without the content and the reason is a decision a human
cannot actually make.
Safety identifiers
Tracked as CR-1 in the feature backlog. The practice below is what an integration should do today; the parameter that would make it first-class does not exist yet.
Attach a stable, pseudonymous identifier to every check you make, so that abuse can be traced back to one end user rather than to your whole integration.
Hash the username or email rather than sending it; for pre-login traffic, a session identifier works. The goal is a value that is stable per user and meaningless outside your systems.
The platform has no safetyIdentifier parameter today. checkText and
checkImage take content, a policy set name and showReason, and nothing
else. Until one exists, keep the mapping on your side: record your identifier
with each check you make, so a verdict can be joined back to an end user without
the identifier ever leaving your infrastructure.
const result = await client.checkText(message, 'threats-and-weapons', { showReason: true });
await auditLog.write({
subject: hashedUserId, // yours, never sent to the platform
checkId: crypto.randomUUID(), // your own handle on this check
policySet: 'threats-and-weapons',
verdict: result.verdict,
reason: result.reason,
});
This is worth doing before you need it. The first time you have to answer "who was doing this, and did it happen more than once", the join has to already exist.
Revoke compromised API keys
An API key is the whole credential. Treat exposure as compromise and revoke it immediately — logged in plaintext, committed to a repository, pasted into a ticket, or shared with a system that should not have had it.
Practices that keep a rotation cheap:
- Keep keys in a secret manager, never in source and never in browser code. Anyone holding a key can spend your organization's credit.
- Use a separate project, and so separate keys, per environment, so revoking a development key does not take production down with it. See Projects.
- Revoke a key on the dashboard's API keys page. It stops working at once.
- Make policy-management calls server-side.
Policies.saveandPolicies.removechange what your checks mean.
Rotation is only fast if nothing has the key hardcoded. Read it from the
environment — SAFETY_API_KEY — and a revocation is a config change rather
than a deploy.
Let users report issues
Tracked as CR-2 in the feature backlog. Reporting is something you build on your side today — there is no platform endpoint that accepts a report or looks a check back up for you.
Give the people affected by a verdict a visible way to say it was wrong, and have a human read what comes in. Both directions matter: content wrongly blocked, and content that got through and should not have.
Put your own identifier for the check in the report, and keep the content, the policy set name, the verdict and the reason against it. Without that a complaint is an anecdote. A report that cannot be traced to a check cannot be acted on.
Reports are also your best source of adversarial test cases. Every false negative a user reports is an input that belongs in the regression suite described above.
Know what the check does not cover
A verdict answers one question: does this content violate this policy set.
- Text returns
blockorpass. There is nounknownfor text. - Image returns
block,passorunknown. - A check names one policy set and is judged against that set alone. Two sets never interact, so content passing a loose set says nothing about a strict one.
- Image policies decide from visible evidence in one frame. Intent, consent, and what happens next are not visible. See Moderating uploads for what a still can and cannot settle.
Calibrate what you tell your own users accordingly.