step.ai.check, so a policy you write once can guard a tool call, a prompt, a completion, or any
call of your own.
Turning it on
schema adapter loads one at the moment it is first
used. @cfworker/json-schema is the one to install: it does no code generation, which is what lets
it run under a strict CSP and on edge runtimes where new Function is unavailable.
@duraton/agent-kit it is the same option on agent():
inputSchema is allowed through - there is nothing to check it against.
The adapters that ship
schema and pii are deterministic and call nothing, so they are safe to leave on. moderation
calls out, which is why the classifier is yours:
getGuardrail("moderation") therefore returns an
adapter that faults rather than allowing - a name lookup that could not find a classifier must
never read as a clean bill of health.
pii masks by default rather than refusing, because the request minus the identifier is usually
still the request. Set action to deny or halt for data that must not travel at all, and
observe to measure a new rule against real traffic before it refuses anything.
What a block looks like to the model
A refused call does not throw. It comes back as that tool’s result, so the model reads why it was refused and can correct itself on the next turn:issue-credit with { "reason": "duplicate" } against the schema above.
Every failing keyword is listed, located by JSON Pointer, so the model can fix them all in one
turn rather than one per turn - and each extra turn would be another model call and another
durable step.
This is the same shape a human denial produces from an approval gate, and for the same reason: MCP
classes invalid input data as a tool execution error reported in the result, not a protocol
error. A thrown error would end the run and teach the model nothing.
When a tripwire halts the run
halt is the strongest verdict, and it is not an error. Inside step.ai.loop the loop ends with
its own stop reason and names the policy that decided:
stopReason: "guardrail" sits alongside "max-iterations" and "approval-budget" in the same
closed set, so a halted agent reads as a ceiling that fired rather than as a run that broke, and
result.haltedBy is the guardrail’s name.
That distinction is load-bearing rather than cosmetic. A tool handler runs inside a durable step, so
a thrown error commits as that step’s failure - a control that worked would be recorded as an
agent that broke, and every replay of that run would reproduce the failure forever. The halt travels
out of the turn as a value instead, for the same reason bail() does.
Two consequences follow:
- Every sibling tool call in the halting turn still completes and still lands on the record. A tool whose side effect already happened is never dropped just because another call was refused.
- A halt outranks a
bail()in the same turn.bailis the agent deciding it is done, and a policy that refused the turn is not something the agent gets to overrule.
step.ai.check has no loop to end, so a halt there fails the run non-retriably
instead: the check’s own step commits the failure, so the run stops at its last checkpoint rather
than re-taking a decision that is already committed until its attempt budget runs out.
Checking a value on its own
step.ai.check runs the same policies over anything - a prompt before it reaches a provider, a
completion before you use it, a payload before it leaves the runner - and, optionally, guards one
call with them.
Guardrail[]
required
The policies to ask, in order; the first verdict that is not allow wins. A guardrail that does not declare the subject’s placement is skipped.
GuardrailInput
What to check before the guarded call - the placement, the value, and optionally a schema and a tool name.
(result: T) => GuardrailInput
Derives what to check from the call’s result. Omitted, nothing is checked afterwards.
(signal: AbortSignal) => T | Promise<T>
The call the checks guard. Omitted, check is a plain policy evaluation over a value you already have.
boolean
Race the input check against the call instead of gating the call on it.
GuardrailVerdict
required
The verdict that decided: the output check’s when one ran and the input check cleared, else the input check’s.
boolean
required
Whether the check refused to clear what it was given. deny and approve trip; mask and rewrite do not, because they hand back a replacement.
unknown
What to use in place of the value you offered - the replacement under mask or rewrite, else the value as given. Absent when the check tripped.
T
The guarded call’s own return. Absent when no call was given, when a check tripped before or during it, and when the output check refused it.
Running the check beside the call
Withparallel: true the check and the call start together, so you pay the check’s latency
concurrently instead of in front of the call. The moment the policy trips, the call’s AbortSignal
is aborted and its result never reaches you:
A
deny from step.ai.check does not throw. There is no model to hand a refusal back to, so the
trip comes back as tripped for your own code to act on. Only halt ends the run.The verdict is a durable step
Each guardrail pass in a loop writes its own step - a sibling of the tool step and the approval step, never a suffix of either. Astep.ai.check writes three: its input check, the call it guards,
and its output check. Two consequences follow from that and from nothing else:
- A replay reads the recorded verdict. The detector does not run again, so a re-run of the same run cannot disagree with the original, and a verdict is an auditable fact rather than something re-derived on every pass.
- A refused call writes no tool step at all. The verdict step exists, the tool step does not, which is what makes “the handler never ran” checkable from the run record.
issue-credit denied by schema. The row is drawn from the run’s event stream, which carries the
action and the guardrail but never the verdict’s reason, so it is safe to read for someone who
should not see the value that was refused.
When a detector is down
A guardrail that throws does not read as “clean”. It produces a verdict whoseaction is deny
and whose outcome is partial, meaning the check could not complete. Nothing in the loop treats
partial as an allow.
Writing a guardrail
A guardrail declares which placements it understands and returns a verdict. It never mutates the caller’s state - it says what should happen and the loop applies it.allow wins, so a later guardrail can never overturn an earlier refusal.
Placements
Where a guardrail runs.tool-args is the placement the loop evaluates today; the rest are part of
the contract so an adapter written now stays valid as they are wired.
Actions
What a verdict asks for.observe is how a new rule earns its place: turn it on, let it record what it would have refused,
and only then promote it to deny.
Which dialect a schema is read in
JSON Schema leaves the dialect of a schema with no$schema up to the implementation
(2020-12 core §8.1.1).
Duraton reads the dialect the schema declares when it declares one, and otherwise uses 2020-12.
Override the fallback if your tool schemas are written against an older draft:
4, 7, 2019-09, 2020-12.