> ## Documentation Index
> Fetch the complete documentation index at: https://docs.duraton.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Guardrails

> Check a model's tool arguments before the tool runs. The verdict is a durable step, so a replay reads what was decided instead of deciding again.

A model authors the arguments its tools are called with. Left unchecked, those arguments reach your
handler - and whatever it talks to - exactly as the model wrote them. MCP puts this on the server
side without qualification: *"Servers **MUST**: Validate all tool inputs"*
([Tools, Security Considerations](https://modelcontextprotocol.io/specification/2025-06-18/server/tools)).

A **guardrail** is that check, as a port. Duraton ships three adapters and the port is what lets
another drop in without touching the loop or the adapters already there. The same port also backs
`step.ai.check`, so a policy you write once can guard a tool call, a prompt, a completion, or any
call of your own.

## Turning it on

```sh theme={null}
npm install @cfworker/json-schema
```

The SDK bundles no JSON Schema engine, so the `schema` adapter loads one at the moment it is first
used. `@cfworker/json-schema` is the one to install: it does no code generation, which is what lets
it run under a strict CSP and on edge runtimes where `new Function` is unavailable.

```ts theme={null}
import { createSchemaGuardrail } from "@duraton/sdk";

const result = await ctx.step.ai.loop("agent", {
  prompt: `Resolve this ticket: ${ticket}`,
  maxIterations: 6,
  guardrails: [createSchemaGuardrail()],
  tools: {
    "issue-credit": {
      inputSchema: {
        type: "object",
        properties: { amount: { type: "number" } },
        required: ["amount"],
        additionalProperties: false,
      },
      handler: (input) => credit(input),
    },
  },
  turn,
});
```

With `@duraton/agent-kit` it is the same option on `agent()`:

```ts theme={null}
const result = await agent(ctx, "agent", {
  model: "claude-opus-4-8",
  prompt: `Resolve this ticket: ${ticket}`,
  tools: [issueCredit],
  maxIterations: 6,
  guardrails: [createSchemaGuardrail()],
});
```

A tool that declares no `inputSchema` is allowed through - there is nothing to check it against.

## The adapters that ship

| Adapter      | What it checks                                                                            | Default action | Needs                   |
| ------------ | ----------------------------------------------------------------------------------------- | -------------- | ----------------------- |
| `schema`     | A tool call against that tool's own `inputSchema`                                         | `deny`         | `@cfworker/json-schema` |
| `pii`        | Text for email addresses, E.164 phone numbers, Luhn-valid card numbers and IPv4 addresses | `mask`         | nothing                 |
| `moderation` | Text, through a classifier you supply                                                     | `halt`         | your classifier         |

`schema` and `pii` are deterministic and call nothing, so they are safe to leave on. `moderation`
calls out, which is why the classifier is yours:

```ts theme={null}
import { createModerationGuardrail, createPiiGuardrail } from "@duraton/sdk";

const policies = [
  createPiiGuardrail({ rules: ["credit-card", "email"] }),
  createModerationGuardrail({
    classify: async (text) => {
      const verdict = await myClassifier(text);
      return { flagged: verdict.blocked, categories: verdict.categories };
    },
    threshold: 0.8,
  }),
];
```

Duraton holds no moderation model and ships no default one: what counts as acceptable belongs to the
people running the agent, not to the runtime. `getGuardrail("moderation")` therefore returns an
adapter that **faults** rather than allowing - a name lookup that could not find a classifier must
never read as a clean bill of health.

`pii` masks by default rather than refusing, because the request minus the identifier is usually
still the request. Set `action` to `deny` or `halt` for data that must not travel at all, and
`observe` to measure a new rule against real traffic before it refuses anything.

## What a block looks like to the model

A refused call does **not** throw. It comes back as that tool's result, so the model reads why it
was refused and can correct itself on the next turn:

```json theme={null}
{
  "valid": false,
  "tool": "issue-credit",
  "error": "arguments for tool \"issue-credit\" do not match its inputSchema - #: Instance does not have required property \"amount\".; #: Property \"reason\" does not match additional properties schema.; #/reason: False boolean schema."
}
```

That is the model calling `issue-credit` with `{ "reason": "duplicate" }` against the schema above.
Every failing keyword is listed, located by JSON Pointer, so the model can fix them all in one
turn rather than one per turn - and each extra turn would be another model call and another
durable step.

This is the same shape a human denial produces from an approval gate, and for the same reason: MCP
classes invalid input data as a *tool execution error reported in the result*, not a protocol
error. A thrown error would end the run and teach the model nothing.

<Warning>
  The refusal carries no decider. Only a person's decision on an approval names a person; a policy
  refusing a call is not a person saying no, and the two never share a shape.
</Warning>

## When a tripwire halts the run

`halt` is the strongest verdict, and it is **not** an error. Inside `step.ai.loop` the loop ends with
its own stop reason and names the policy that decided:

```ts theme={null}
const result = await ctx.step.ai.loop("agent", { guardrails: policies, /* ... */ });

if (result.stopReason === "guardrail") {
  await notifyCompliance({ run: ctx.runId, policy: result.haltedBy });
}
```

`stopReason: "guardrail"` sits alongside `"max-iterations"` and `"approval-budget"` in the same
closed set, so a halted agent reads as a ceiling that fired rather than as a run that broke, and
`result.haltedBy` is the guardrail's name.

That distinction is load-bearing rather than cosmetic. A tool handler runs inside a durable step, so
a thrown error commits as that **step's failure** - a control that worked would be recorded as an
agent that broke, and every replay of that run would reproduce the failure forever. The halt travels
out of the turn as a value instead, for the same reason [`bail()`](/agent-kit/agents-and-tools) does.

Two consequences follow:

* **Every sibling tool call in the halting turn still completes and still lands on the record.** A
  tool whose side effect already happened is never dropped just because another call was refused.
* **A halt outranks a `bail()` in the same turn.** `bail` is the agent deciding it is done, and a
  policy that refused the turn is not something the agent gets to overrule.

Outside a loop, `step.ai.check` has no loop to end, so a `halt` there fails the run **non-retriably**
instead: the check's own step commits the failure, so the run stops at its last checkpoint rather
than re-taking a decision that is already committed until its attempt budget runs out.

## Checking a value on its own

`step.ai.check` runs the same policies over anything - a prompt before it reaches a provider, a
completion before you use it, a payload before it leaves the runner - and, optionally, guards one
call with them.

```ts theme={null}
import { createModerationGuardrail } from "@duraton/sdk";

const gate = await ctx.step.ai.check<string>("moderate", {
  guardrails: [createModerationGuardrail({ classify })],
  input: { placement: "pre-prompt", value: question },
  output: (answer) => ({ placement: "post-model", value: answer }),
  call: () => askTheModel(question),
});

if (gate.tripped) return { refused: gate.verdict.by };
return { answer: gate.result };
```

<ResponseField name="guardrails" type="Guardrail[]" required>
  The policies to ask, in order; the first verdict that is not allow wins. A guardrail that does not declare the subject's placement is skipped.
</ResponseField>

<ResponseField name="input" type="GuardrailInput">
  What to check before the guarded call - the placement, the value, and optionally a schema and a tool name.
</ResponseField>

<ResponseField name="output" type="(result: T) => GuardrailInput">
  Derives what to check from the call's result. Omitted, nothing is checked afterwards.
</ResponseField>

<ResponseField name="call" type="(signal: AbortSignal) => T | Promise<T>">
  The call the checks guard. Omitted, check is a plain policy evaluation over a value you already have.
</ResponseField>

<ResponseField name="parallel" type="boolean">
  Race the input check against the call instead of gating the call on it.
</ResponseField>

It returns:

<ResponseField name="verdict" type="GuardrailVerdict" required>
  The verdict that decided: the output check's when one ran and the input check cleared, else the input check's.
</ResponseField>

<ResponseField name="tripped" type="boolean" required>
  Whether the check refused to clear what it was given. deny and approve trip; mask and rewrite do not, because they hand back a replacement.
</ResponseField>

<ResponseField name="value" type="unknown">
  What to use in place of the value you offered - the replacement under mask or rewrite, else the value as given. Absent when the check tripped.
</ResponseField>

<ResponseField name="result" type="T">
  The guarded call's own return. Absent when no call was given, when a check tripped before or during it, and when the output check refused it.
</ResponseField>

### Running the check beside the call

With `parallel: true` the check and the call start together, so you pay the check's latency
concurrently instead of in front of the call. The moment the policy trips, the call's `AbortSignal`
is aborted and its result never reaches you:

```ts theme={null}
const gate = await ctx.step.ai.check<string>("moderate", {
  guardrails: [createModerationGuardrail({ classify })],
  input: { placement: "pre-prompt", value: question },
  parallel: true,
  call: (signal) => askTheModel(question, { signal }),
});
```

The trade is a call you may throw away, which is the right trade when the classifier is fast and the
model call is the slow part. A cancelled call still commits its step, recorded as cancelled - the run
says a policy stopped it rather than leaving a hole where a call should be.

<Note>
  A `deny` from `step.ai.check` does not throw. There is no model to hand a refusal back to, so the
  trip comes back as `tripped` for your own code to act on. Only `halt` ends the run.
</Note>

## The verdict is a durable step

Each guardrail pass in a loop writes its own step - a sibling of the tool step and the approval
step, never a suffix of either. A `step.ai.check` writes three: its input check, the call it guards,
and its output check. Two consequences follow from that and from nothing else:

* **A replay reads the recorded verdict.** The detector does not run again, so a re-run of the same
  run cannot disagree with the original, and a verdict is an auditable fact rather than something
  re-derived on every pass.
* **A refused call writes no tool step at all.** The verdict step exists, the tool step does not,
  which is what makes "the handler never ran" checkable from the run record.

A loop with no guardrails writes no guardrail steps, so turning this on costs nothing until you do.

In the console, a turn that ran guardrails says how many it checked, and any verdict that did not
allow the call gets its own row naming the tool, what happened to it, and which guardrail decided -
`issue-credit denied by schema`. The row is drawn from the run's event stream, which carries the
action and the guardrail but never the verdict's `reason`, so it is safe to read for someone who
should not see the value that was refused.

## When a detector is down

A guardrail that throws does not read as "clean". It produces a verdict whose `action` is `deny`
and whose `outcome` is `partial`, meaning the check could not complete. Nothing in the loop treats
`partial` as an allow.

| `outcome`  | Meaning                                                            |
| ---------- | ------------------------------------------------------------------ |
| `complete` | Every guardrail ran                                                |
| `partial`  | At least one did not run; the result is not a clean bill of health |
| `failed`   | The check itself failed outright                                   |

## Writing a guardrail

A guardrail declares which placements it understands and returns a verdict. It never mutates the
caller's state - it says what should happen and the loop applies it.

```ts theme={null}
import type { Guardrail } from "@duraton/sdk";

const noExternalRecipients: Guardrail = {
  name: "recipient-policy",
  placements: ["tool-args"],
  check: async ({ value, tool }) => {
    const to = (value as { to?: string }).to ?? "";
    if (to.endsWith("@example.com")) {
      return { action: "allow", outcome: "complete", detections: [] };
    }
    return {
      action: "deny",
      outcome: "complete",
      detections: [{ rule: "recipient.external", detected: true }],
      reason: `${tool} may only send to internal recipients`,
    };
  },
};
```

Pass it alongside the others. They run in the order you list them and the first verdict that is not
`allow` wins, so a later guardrail can never overturn an earlier refusal.

### Placements

Where a guardrail runs. `tool-args` is the placement the loop evaluates today; the rest are part of
the contract so an adapter written now stays valid as they are wired.

| Placement     | What it inspects                                            |
| ------------- | ----------------------------------------------------------- |
| `pre-prompt`  | The prompt and system text about to reach a provider        |
| `post-model`  | A completed model response, before anything is done with it |
| `tool-args`   | The arguments a model proposed, before the tool step runs   |
| `tool-result` | A tool's output, before it re-enters the transcript         |
| `egress`      | A URL, host or payload about to leave the runner            |

### Actions

What a verdict asks for.

| Action    | Effect at `tool-args`                                                                           |
| --------- | ----------------------------------------------------------------------------------------------- |
| `allow`   | The tool runs with the model's own arguments                                                    |
| `observe` | Detected but deliberately not enforced - the run is unchanged, the detection is recorded        |
| `mask`    | The tool runs with `payload` instead; the original never reaches the handler or the step record |
| `rewrite` | As `mask`, for a repaired rather than a redacted value                                          |
| `deny`    | The tool does not run; the reason goes back to the model as that tool's result                  |
| `approve` | The call is parked on a human, even if the tool itself was not marked as needing approval       |
| `halt`    | The loop ends with `stopReason: "guardrail"`, naming the rule; a `step.ai.check` fails the run  |

`observe` is how a new rule earns its place: turn it on, let it record what it *would* have refused,
and only then promote it to `deny`.

## Which dialect a schema is read in

JSON Schema leaves the dialect of a schema with no `$schema` up to the implementation
([2020-12 core §8.1.1](https://json-schema.org/draft/2020-12/json-schema-core#name-the-schema-keyword)).
Duraton reads the dialect the schema declares when it declares one, and otherwise uses **2020-12**.
Override the fallback if your tool schemas are written against an older draft:

```ts theme={null}
createSchemaGuardrail({ draft: "7" });
```

Supported drafts: `4`, `7`, `2019-09`, `2020-12`.

## Edits a person makes are checked too

When a tool is approval-gated and the reviewer edits the arguments before approving, those edits are
hand-typed JSON that the model never proposed - so they are re-checked, under a durable step of
their own. A refusal there fails the run rather than returning a
refusal to the model: telling the model its own arguments were wrong would be false when a person
is the one who broke them.
