Skip to main content
Coming soon

Safety Checks

Everything else in this section is a control you apply. Safety checks are the reverse: safeguards the platform applies to traffic, independent of the policy sets a project holds.

Neither of the two below is in place today. No traffic is currently monitored or limited by either mechanism, and no error code described here can be returned by the current release. The page documents the intended shape so an integration can be built to survive them arriving.

Cybersecurity checks​

Traffic would be monitored for signals of cybersecurity misuse — the platform being used as a step in an attack rather than as a safety control. Where signals passed a threshold, access could be limited while the activity was reviewed.

Two things follow for an integration, and both are worth building now:

Expect a distinct error, and keep it apart from a verdict. Today every SDK error is a plain Error carrying the service's message (see Errors). A safeguard rejection is not a policy block — the content was never judged — and an integration that treats every failure as a block will silently start blocking legitimate content.

try {
const result = await client.checkText(text, 'threats-and-weapons');
if (!isSafe(result)) block();
} catch (err) {
escalate(err.message); // not the same as a block verdict
}

Per-user identifiers limit the blast radius. Where activity can be attributed to one end user, a restriction can be scoped to that user. Where it cannot, the only available scope is the project or organization — meaning one user's misuse takes the whole integration down. The platform has no identifier parameter yet, which is the gap described in Safety identifiers; keeping the mapping on your side now is what makes scoping possible later.

Calibration is the hard part of any such system. Legitimate security research and defensive work look, in traffic, much like the thing being watched for.

Misalignment monitoring​

Where a Kavach verdict gates an action rather than a message — the robotics case, where a check sits in front of an actuator — the question stops being whether content is permitted and becomes whether the sequence of actions is consistent with what was asked.

Three properties this would need, and they are what an integration should assume:

  • A flag is a request for review, not a finding. It does not establish that a policy was violated or that anything acted contrary to instruction. Monitoring can miss real problems and flag legitimate work.
  • It is asynchronous. An action may complete before monitoring reaches a concern about it. A stop does not undo what already happened, which is why pre-execution checking on the device is the control that matters and this is the one that catches what it missed.
  • A stopped sequence does not resume. Preserve your own records of the checks and the actions already taken, show them to whoever owns the task, and have them compare what happened against what was intended before anything restarts.

None of this replaces application-level safeguards. Human approval for consequential actions, bounded tool scope and audit logging remain yours to implement — see Human in the loop.

What exists today​

MechanismStatus
Policy-backed checks on text and imageLive.
Cybersecurity checksNot in place.
Misalignment monitoringNot in place.
Per-end-user safety identifiersNot available.

Changes to this table will appear in the Changelog.