> ## Documentation Index
> Fetch the complete documentation index at: https://docs.duraton.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# AI steps

> Make every model call run once: the step.ai reference for generate, wrap, embed, and the durable agent loop, with providers, cost, and cache options.

`step.ai` makes a model call a durable step. Like [`step.run`](/reference/sdk/steps#step), each call
takes a stable `id`, records its result under that id, and returns the saved result on replay instead
of calling the model again. So a retry after a crash never re-spends on work that already completed.

New to AI steps? The [AI quickstart](/start/ai-quickstart) walks you from a first
`generate` call to spend landing in the console.

Duraton stores the AI metadata (model, token counts, latency) as an opaque journal block. It never
parses it and never stores your prompt, the response text, or your API key - those stay in your runner.

```ts theme={null}
const { output } = await ctx.step.ai.generate<Triage>("classify", {
  model: "claude-opus-4-8",
  prompt: `Classify this ticket: ${subject}`,
  output: triageSchema,
});
```

## The step.ai API

<ResponseField name="generate" type="<T>(id, opts) => Promise<StructuredResult<T>>" required>
  One model call as a durable step, with optional structured output + durable re-ask.
</ResponseField>

<ResponseField name="wrap" type="<T>(id, fn) => Promise<T>" required>
  Make any caller-supplied AI call durable - bring your own client.
</ResponseField>

<ResponseField name="embed" type="(id, opts) => Promise<EmbedResult>" required>
  Batch embeddings with per-batch checkpointing.
</ResponseField>

<ResponseField name="loop" type="<T>(id, opts) => Promise<LoopResult<T>>" required>
  A durable agent loop: one durable step per turn and per tool call.
</ResponseField>

<ResponseField name="check" type="<T>(id, opts) => Promise<CheckResult<T>>" required>
  One guardrail policy pass as a durable step, optionally guarding a call of your own.
</ResponseField>

## `step.ai.generate`

Make one model call as a durable step. The built-in provider is Anthropic; the request is validated,
sent, and the result memoized under `id`.

```ts theme={null}
const result = await ctx.step.ai.generate("draft-reply", {
  model: "claude-opus-4-8",
  prompt: `Write a one-line apology for ticket ${ticketId}.`,
});
// result.text, result.model, result.usage.inputTokens, result.usage.outputTokens
```

<ResponseField name="model" type="string" required>
  The model id to call, e.g. "claude-opus-4-8".
</ResponseField>

<ResponseField name="prompt" type="string" required>
  The user prompt.
</ResponseField>

<ResponseField name="system" type="string">
  An optional system prompt.
</ResponseField>

<ResponseField name="maxTokens" type="number">
  Max output tokens.
</ResponseField>

<ResponseField name="temperature" type="number">
  Sampling temperature.
</ResponseField>

<ResponseField name="output" type="Record<string, unknown>">
  A JSON Schema the response must satisfy. Set it to get a parsed, validated result on output, with durable re-ask on failure.
</ResponseField>

<ResponseField name="reask" type="number">
  Max durable re-asks when output validation fails (default 1). Each re-ask is its own memoized step; 0 disables re-asking.
</ResponseField>

<ResponseField name="validate" type="(value: unknown) => string | undefined">
  A deeper check beyond "valid JSON": return an error string to reject the value (triggers a re-ask), or undefined to accept.
</ResponseField>

<ResponseField name="provider" type="ProviderName">
  The provider adapter to use. Defaults to "anthropic".
</ResponseField>

<ResponseField name="stream" type="boolean">
  Stream tokens as they arrive: each delta is journaled as an ai\_chunk timeline frame for a live, replayable view. The durable result is still the complete text. Falls back to a plain generate when the transport has no live channel. See Streaming below.
</ResponseField>

<ResponseField name="fallback" type="FallbackCandidate[]">
  Backup models tried in order after model when a call fails with a retryable error (429/5xx/timeout). The first to return wins. Each candidate is \{ model, provider? }. See Fallback chains below.
</ResponseField>

<ResponseField name="cache" type="boolean | CacheOptions">
  Opt this call into the inference cache: an identical prior call is served without a provider call (zero spend). Engages only when temperature is explicitly \<= 0.2. true uses a 24h TTL; CacheOptions is \{ ttlMs?, seed? }. See Inference cache below.
</ResponseField>

<ResponseField name="promptCache" type="PromptCacheScope">
  Ask the provider to bill this call's repeated prefix at its own cached rate - a different thing from cache above, which skips the provider call entirely. "prefix" caches the stable head, "conversation" also follows a transcript that grows call over call. The provider must declare the prompt-cache capability, or the call is refused rather than quietly billed in full. See Prompt caching below.
</ResponseField>

<ResponseField name="apiKey" type="string">
  Passed through per call and never stored; omit to fall back to the provider SDK env var (e.g. ANTHROPIC\_API\_KEY).
</ResponseField>

`generate` returns a `StructuredResult<T>`:

<ResponseField name="text" type="string" required>
  The response text.
</ResponseField>

<ResponseField name="model" type="string" required>
  The model that actually answered - may differ from the requested one (e.g. a server-side fallback). This is what the journal records.
</ResponseField>

<ResponseField name="provider" type="ProviderName" required>
  The provider that served the call.
</ResponseField>

<ResponseField name="usage" type="TokenUsage" required>
  Four disjoint token axes: inputTokens counts the uncached input only, so inputTokens + cacheReadTokens + cacheCreationTokens is everything the provider processed. The two cache axes are filled in only when the provider reports them. See Prompt caching below.
</ResponseField>

<ResponseField name="stopReason" type="string">
  Why generation stopped.
</ResponseField>

<ResponseField name="output" type="T">
  The parsed, validated value - present only when you passed output.
</ResponseField>

### Structured output and durable re-ask

Pass `output` (a JSON Schema) to constrain the model and get a typed, validated value back on
`result.output`. If the response fails to parse or validate, `generate` re-prompts with the validation
error - each re-ask is its own memoized step, so the retry survives a crash and never repeats a
committed attempt. Add `validate` for semantic rules the schema can't express.

```ts theme={null}
const { output } = await ctx.step.ai.generate<{ category: string; priority: string }>("triage", {
  model: "claude-opus-4-8",
  prompt: `Triage: ${subject}`,
  output: {
    type: "object",
    properties: { category: { type: "string" }, priority: { type: "string" } },
    required: ["category", "priority"],
  },
  reask: 2,
  validate: (v) => (["low", "normal", "high"].includes((v as { priority: string }).priority) ? undefined : "priority out of range"),
});
```

<Note>
  The `apiKey` you pass is used for that one call and never written to the journal or the run store.
  Omit it to let the provider SDK read its conventional env var.
</Note>

### Streaming

Pass `stream: true` to feed the model's tokens to a live viewer as they arrive. Each delta is appended to
the run's durable timeline as an `ai_chunk` frame, so a viewer sees the text build in real time and a late
or reconnecting viewer replays it from token 0. The return value is unchanged - `result.text` is still the
complete response, memoized on replay - so streaming affects only what a viewer sees while the step runs.

```ts theme={null}
const result = await ctx.step.ai.generate("summarize-thread", {
  model: "claude-opus-4-8",
  prompt: `Summarize this thread:\n\n${thread}`,
  stream: true,
});
// result.text is the full summary; deltas streamed live on the way there.
```

Live deltas travel over the [connect](/reference/sdk/connect) runner's socket. See the [Streaming](/ai/streaming) concept for the timeline frames,
resumability, and the `useStream` React hook.

### Fallback chains

Pass `fallback` - an ordered list of backup models - to keep a call resilient when a model is rate-limited
or down. The primary `model` is tried first; if it fails with a **retryable** error (429, a 5xx, or a
timeout), the call **advances** to the next candidate, and the first one to return wins. Its result is the
step's durable output, so a caller never sees the failover.

```ts theme={null}
const answer = await ctx.step.ai.generate("answer", {
  model: "claude-opus-4-8",
  prompt: question,
  fallback: [{ model: "claude-sonnet-4-6" }, { model: "claude-haiku-4-5" }],
});
```

Each candidate is `{ model, provider? }`; `provider` defaults to the call's provider, so a chain can span
providers once you have more than one adapter configured. The step's journal records the outcome:

<ResponseField name="chain" type="string[]" required>
  The models tried, in order (the primary plus each fallback).
</ResponseField>

<ResponseField name="used" type="string" required>
  The model that actually served the call.
</ResponseField>

<ResponseField name="reason" type="string">
  Why the chain advanced - the classified failures of the skipped models, e.g. "claude-opus-4-8: 429". Absent when the primary served.
</ResponseField>

The console renders this as a chain pill on the AI step, and it rides the opaque journal so an agent
reading the run over MCP sees the same `chain` / `used` / `reason`.

<Note>
  Only 429, 5xx, and timeout advance the chain. A terminal 4xx (a bad request, an auth failure) fails the
  step immediately - another model won't fix a malformed request. An exhausted chain also fails the step,
  re-throwing the last error, so the workflow's own durable retry policy still applies. Fallback is
  per-call resilience, distinct from the [flow-control](/core/flow-control) spend controls
  (`cap` / `tokenThrottle`).
</Note>

### Inference cache

Set `cache` to reuse the result of an identical earlier call instead of paying for it again. On a **hit**
the provider is never called, so the step commits with **zero spend** - the cache is the one control that
*reduces* spend rather than capping it, and a cached call counts nothing against `cap` /
`tokenThrottle`. Where step memoization already makes a **replay** free, the cache makes an identical call
in a **different run** free too.

```ts theme={null}
const answer = await ctx.step.ai.generate("answer", {
  model: "claude-opus-4-8",
  prompt: question,
  temperature: 0,   // required: caching engages only for a deterministic call
  cache: true,      // or { ttlMs, seed }
});
```

The key is an exact match over the runner's app, the model, the prompt, and every output-affecting
parameter, so no entry ever crosses a project boundary and a different config never returns a stale answer.
The step's journal records the outcome:

<ResponseField name="hit" type="boolean" required>
  Whether the call was served from the cache (true = the provider was not called).
</ResponseField>

<ResponseField name="key" type="string" required>
  The entry key (metadata only - the cached completion is held runner-side and never reaches Duraton).
</ResponseField>

<ResponseField name="ageMs" type="number">
  How long ago the entry was stored, on a hit.
</ResponseField>

The console renders this as a cache pill on the AI step, and it rides the opaque journal so an agent
reading the run over MCP sees the same `hit` / `key` / `ageMs`.

<Note>
  Caching is **exact-match** and engages only when `temperature` is explicitly set to `0.2` or lower -
  caching a sampled (high-temperature) answer would freeze one draw, and an *unset* temperature is treated
  as non-deterministic (a provider default is often 1.0). The default TTL is 24h, overridable per call with
  `{ ttlMs }`; `{ seed }` overrides the default project seed (the runner's app) to scope entries further.
</Note>

## `step.ai.wrap`

Makes a model call you already write yourself - through the OpenAI SDK, the Anthropic SDK, the Vercel
AI SDK, or anything else - a durable step, with no other change to the call site. `wrap` returns your
function's value unchanged; when it recognizes the response shape it enriches the journal with the
model and token counts and records which library it wrapped.

```ts theme={null}
import OpenAI from "openai";
const openai = new OpenAI();

const completion = await ctx.step.ai.wrap("classify", () =>
  openai.chat.completions.create({
    model: "gpt-4o",
    messages: [{ role: "user", content: subject }],
  }),
);
```

Recognized shapes: the Anthropic SDK, the OpenAI SDK, and the Vercel AI SDK. An unrecognized value
still becomes a durable `wrap` step - you just get less metadata on the journal.

## `step.ai.embed`

Turn a list of inputs into vectors, one durable batch at a time. Anthropic has no embeddings API, so
you supply the embedding call (`embed`); Duraton owns the batching and per-batch checkpointing. If a
batch fails, only that batch re-runs on retry - committed batches are not re-embedded.

```ts theme={null}
const { vectors } = await ctx.step.ai.embed("embed-kb", {
  model: "voyage-3",
  inputs: ["duplicate charge policy", "annual plan refunds", "refund SLA"],
  embed: (batch) => voyage.embed(batch),
  batchSize: 2,
});
```

<ResponseField name="model" type="string" required>
  Names the embedding model - recorded on the journal.
</ResponseField>

<ResponseField name="inputs" type="string[]" required>
  The inputs to embed.
</ResponseField>

<ResponseField name="embed" type="(batch: string[]) => number[][] | Promise<number[][]>" required>
  Your embedding call for one batch: inputs in, one vector per input out.
</ResponseField>

<ResponseField name="batchSize" type="number">
  Inputs per durable batch (default 100). Each batch checkpoints independently.
</ResponseField>

`embed` returns `{ vectors }` - one vector per input, in input order.

## `step.ai.loop`

A durable agent loop. Each turn is your own model call (bring-your-own, normalized to tool calls or a
final answer); the loop executes the tools the turn requested and feeds the results into the next turn,
until the model returns a final answer, `stop` fires, or `maxIterations` is reached.

Every turn is a durable step, and so is every tool call it makes, so an agent that crashes mid-run
resumes at the last committed turn.

Writing `turn` yourself is the low-level path. To declare a model, instructions and tools and have
the turn composed for you, use the [agent kit](/agent-kit) - it composes this loop rather than
replacing it, so everything below still applies.

```ts theme={null}
const agent = await ctx.step.ai.loop<{ resolution: string }>("agent", {
  prompt: `Resolve the ticket about: ${subject}`,
  maxIterations: 6,
  tools: {
    "search-kb": { handler: (input) => searchKb(input) },
    "lookup-order": { workflow: "orders.lookup", app: "orders" },
  },
  turn: (ctx, iteration) => callModel(ctx.prompt, ctx.history, iteration),
});
// agent.final, agent.iterations, agent.stopReason
```

<ResponseField name="prompt" type="string" required>
  The task the agent is working on; surfaced to turn via ctx.prompt.
</ResponseField>

<ResponseField name="turn" type="(ctx, iteration) => LoopTurn | Promise<LoopTurn>" required>
  One model turn, a pure function of ctx - keep it deterministic for replay.
</ResponseField>

<ResponseField name="maxIterations" type="number" required>
  Hard cap on turns; the loop halts before exceeding it.
</ResponseField>

<ResponseField name="tools" type="Record<string, LoopTool>">
  The tools the model may call, keyed by name.
</ResponseField>

<ResponseField name="approval" type="ApprovalRule">
  The gate for every tool that has not answered for itself. Resolved against each tool's own approval - see the agent kit for the order.
</ResponseField>

<ResponseField name="maxApprovals" type="number">
  Ceiling on the human decisions this loop may ask for. Past it the loop halts with stopReason "approval-budget". Omitted, there is none.
</ResponseField>

<ResponseField name="stop" type="(ctx) => boolean">
  Optional early stop after a completed turn; must be pure for replay.
</ResponseField>

Your `turn` returns a `LoopTurn` - either tool calls to run, or a final answer:

<ResponseField name="toolCalls" type="Array<{ id, name, input }>">
  Tools to run this turn; each name must be a key in tools.
</ResponseField>

<ResponseField name="final" type="unknown">
  The final answer. Returning this ends the loop with stopReason "final".
</ResponseField>

<ResponseField name="model" type="string">
  The model that produced the turn - recorded on the journal.
</ResponseField>

<ResponseField name="tokensIn" type="number">
  Input tokens for this turn.
</ResponseField>

<ResponseField name="tokensOut" type="number">
  Output tokens for this turn.
</ResponseField>

### Which loop composed a turn

Every turn's journal carries a `loopVersion` alongside the model and the token counts. The model
says what answered; this says what asked:

```json theme={null}
{ "kind": "loop", "iteration": 0, "loopVersion": "typescript.1", "model": "claude-sonnet-5" }
```

It is written when the turn is composed and never rewritten, so replaying a year-old run shows the
version that produced each turn rather than the version replaying it. Without it a replay quietly
mixes old journal entries with new loop behaviour: the run replays, the numbers look reasonable, and
the conclusion drawn from them is wrong.

The value is prefixed with the language that implemented the loop, as in `typescript.1`, so an entry
read later needs nothing but itself to be understood. Only turns carry it: a `generate` or an `embed`
was not composed by the loop.

<Note>
  Read it, do not pin on it. A bump means turns composed after it may differ from turns composed
  before, which is a reason to compare two runs carefully - not a reason to refuse the older one.
</Note>

A tool is either a **handler** (a local function) or a **workflow** (another Duraton workflow, called
as a linked child run):

<ResponseField name="handler" type="(input: unknown) => unknown | Promise<unknown>">
  A local function tool.
</ResponseField>

<ResponseField name="workflow" type="string">
  The name of a workflow to run as this tool; its call becomes a linked child run - the tool's step carries childRunId and the spawned run carries parentRunId, parentStep and parentAttempt, so the call is followable both ways.
</ResponseField>

<ResponseField name="app" type="string">
  The workflow tool's app; addressed like step.runWorkflow.
</ResponseField>

<ResponseField name="runner" type="string">
  Pin the workflow tool to a specific runner.
</ResponseField>

<ResponseField name="requiresApproval" type="boolean">
  Park the run on a human before this tool runs. See Gating a tool on a human below.
</ResponseField>

<ResponseField name="approval" type="ApprovalRule">
  Annotates that gate and decides whether it is raised at all. See Gating a tool on a human below.
</ResponseField>

### Gating a tool on a human

A tool marked `requiresApproval` does not run until someone decides on it. The run parks in
`needs_attention` holding no worker, exactly as [`step.approval`](/reference/sdk/steps#approval) does -
it is the same gate, raised for you:

```ts theme={null}
tools: {
  "issue-refund": { requiresApproval: true, handler: (input) => issueRefund(input) },
}
```

Three things follow, and each is a deliberate choice:

|                                |                                                                                                                                                                           |
| ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **The decider's edits win**    | The reviewer sees the input the model proposed and may change it. The tool runs with the decided input, not the proposed one.                                             |
| **A refusal is a tool result** | A denial comes back to the model as that tool's result (`{ approved: false, ... }`) rather than as a thrown error, so the agent can react to a no. The run does not fail. |
| **The gate is durable**        | Each approval is its own durable step, so resuming replays the loop without re-asking anyone. A denied call runs no tool and writes no tool step.                         |

A tool's `approval` says more about that gate and can decide whether it is raised at all, and the
loop's own `approval` is the default for every tool that has not answered for itself. Both are the
same `ApprovalRule`: the environment a condition reads is on
[Approvals](/ai/approvals#deciding-whether-to-gate-at-all), and the order the two levels resolve in
is on the [agent kit](/agent-kit/approvals#which-rule-applies-to-a-tool).

`maxIterations` bounds how many turns a loop may run; a per-run [spend
cap](/ai/cost-controls) bounds how much it may spend. When a run reaches its cap the loop
halts before its next turn's model call and the run fails with a `BudgetError` - the committed turns
stay, and the halted turn is the loop's last (failed) iteration.

`ctx.history` gives each turn the prior turns' `toolCalls` and `toolResults`, so your model call can
see what it has already tried. `loop` returns:

<ResponseField name="final" type="T">
  The final answer, if the loop reached one.
</ResponseField>

<ResponseField name="iterations" type="number" required>
  How many turns ran.
</ResponseField>

<ResponseField name="stopReason" type="&#x22;final&#x22; | &#x22;max-iterations&#x22; | &#x22;stopped&#x22; | &#x22;approval-budget&#x22; | &#x22;bail&#x22; | &#x22;guardrail&#x22;" required>
  Why the loop ended. approval-budget means the loop reached maxApprovals; bail means a tool returned bail(); guardrail means a policy returned halt.
</ResponseField>

<ResponseField name="stoppedBy" type="string">
  Which named stop condition ended the loop, on stopReason "stopped". Absent when the opts.stop closure ended it, because a closure has no name to report.
</ResponseField>

<ResponseField name="haltedBy" type="string">
  Which guardrail halted the loop, on stopReason "guardrail". The name only, never the verdict's reason.
</ResponseField>

## `step.ai.check`

Ask a guardrail policy about a value, as a durable step. The verdict memoizes, so a replay reads what
was decided instead of asking again. Supply `call` to guard one call with the policy: the check gates
it, or with `parallel: true` races it and aborts its signal the moment the policy trips.

```ts theme={null}
const gate = await ctx.step.ai.check<string>("moderate", {
  guardrails: [createModerationGuardrail({ classify })],
  input: { placement: "pre-prompt", value: question },
  output: (answer) => ({ placement: "post-model", value: answer }),
  call: () => askTheModel(question),
});
```

<ResponseField name="guardrails" type="Guardrail[]" required>
  The policies to ask, in order; the first verdict that is not allow wins. A guardrail that does not declare the subject's placement is skipped.
</ResponseField>

<ResponseField name="input" type="GuardrailInput">
  What to check before the guarded call: placement, value, and optionally schema and tool.
</ResponseField>

<ResponseField name="output" type="(result: T) => GuardrailInput">
  Derives what to check from the call's result. Omitted, nothing is checked afterwards.
</ResponseField>

<ResponseField name="call" type="(signal: AbortSignal) => T | Promise<T>">
  The call the checks guard. Omitted, check is a plain policy evaluation.
</ResponseField>

<ResponseField name="parallel" type="boolean">
  Race the input check against the call instead of gating the call on it.
</ResponseField>

`check` returns `{ verdict, tripped, value?, result? }` - `value` is the replacement under a mask or
rewrite verdict, and `result` is the guarded call's own return, withheld whenever a check refused it.
A `halt` verdict fails the check's step non-retriably rather than returning; every other refusal
comes back as `tripped`. See [Guardrails](/ai/guardrails) for the placements, the actions, and the
adapters that ship.

It records the input check, the guarded call and the output check as three separate durable steps.
The verdict step journals `kind: "check"` with the placement, the action and the deciding
guardrail - never the verdict's reason.

## Providers

`step.ai.generate` resolves its `provider` name to an **`AIProvider`** adapter through a port, so you
can supply your own instead of the built-in registry. Pass `resolveProvider` to
[`connect`](/reference/sdk/connect) and every `generate` call in that runner
goes through it - the call sites are unchanged. The same resolver is also reachable directly as
`ctx.resolveProvider`, which is what `step.ai.loop` (and the [agent kit](/agent-kit)'s
`agent()`) resolves its own model call through, since a loop's `turn` only ever sees
`(ctx, iteration)` and has no `opts.provider`-shaped call site of its own to inject into.

```ts theme={null}
import { connect } from "@duraton/sdk";
import { type AIProvider, createAnthropicProvider, getProvider } from "@duraton/sdk/ai";

const recording: AIProvider = {
  name: "anthropic",
  generate: (req) => fixtures[req.prompt] ?? getProvider("anthropic").generate(req),
};

connect({
  app: "support-app",
  workflows,
  resolveProvider: (name) => (name === "anthropic" ? recording : getProvider(name)),
});
```

| Export                           | Type                                                           | Description                                                                                                                                                                                                      |
| -------------------------------- | -------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `PROVIDERS`                      | `readonly ["anthropic", "aisdk"]`                              | The closed set of provider names the SDK ships an adapter for.                                                                                                                                                   |
| `ProviderName`                   | `"anthropic" \| "aisdk"`                                       | The type derived from `PROVIDERS`; what `GenerateOptions.provider` accepts.                                                                                                                                      |
| `getProvider(name)`              | `(name: ProviderName) => AIProvider`                           | The built-in registry: one adapter per name. The default `resolveProvider`.                                                                                                                                      |
| `createAnthropicProvider(opts?)` | `(opts?: { fetch?, baseURL? }) => AIProvider`                  | The Anthropic adapter. It loads `@anthropic-ai/sdk` lazily, so that package is an optional peer dependency you install only if you call Anthropic.                                                               |
| `createAisdkProvider(opts?)`     | `(opts?: { resolveModel? }) => AIProvider`                     | The Vercel AI SDK adapter - one adapter for every provider the AI SDK supports. It loads `ai` lazily, so that package is an optional peer dependency. See [Using any model provider](#using-any-model-provider). |
| `ProviderResolver`               | `(name: ProviderName) => AIProvider`                           | The `resolveProvider` option's type.                                                                                                                                                                             |
| `ToolDeclaration`                | `{ name, description?, inputSchema }`                          | A tool the model may call, in Duraton's own vocabulary - no provider SDK type.                                                                                                                                   |
| `ToolCall`                       | `{ id, name, input }`                                          | One tool the model asked to run. `input` is raw model output and is not validated against the declaration.                                                                                                       |
| `ToolResult`                     | `{ id, output }`                                               | What one tool call returned, under the `id` of the call it answers.                                                                                                                                              |
| `TranscriptTurn`                 | `{ toolCalls, toolResults }`                                   | One completed exchange: the tools the model asked to run, and what they returned.                                                                                                                                |
| `PROVIDER_CAPABILITIES`          | `readonly ["transcript", "structured-output", "prompt-cache"]` | The closed set of request fields an adapter opts into.                                                                                                                                                           |
| `providerSupports(p, c)`         | `(p: AIProvider, c: ProviderCapability) => boolean`            | Whether an adapter honours a capability. A caller asks before sending the field.                                                                                                                                 |
| `PROMPT_CACHE_SCOPES`            | `readonly ["prefix", "conversation"]`                          | The closed set of prompt-cache scopes. See [Prompt caching](#prompt-caching).                                                                                                                                    |
| `PromptCacheScope`               | `"prefix" \| "conversation"`                                   | The type derived from `PROMPT_CACHE_SCOPES`; what `GenerateOptions.promptCache` accepts.                                                                                                                         |

An `AIProvider` implements `generate(req)` and, optionally, `stream(req, onDelta)` (an adapter without
it falls back to `generate`, so `stream: true` still returns the right text) and `classifyError(err)`
(which decides whether a failure is retryable, and so whether a [fallback chain](#fallback-chains)
advances - an unclassified error is treated as terminal). It may also declare `capabilities` -
see [Conversations](#conversations).

### Using any model provider

The `aisdk` provider delegates to the [Vercel AI SDK](https://ai-sdk.dev), so one adapter reaches
every provider the AI SDK supports - OpenAI, Google, Mistral, Bedrock, Groq and the rest - without
Duraton shipping an adapter per vendor. Install `ai` alongside the provider package you want:

```bash theme={null}
npm i ai @ai-sdk/openai
```

Point `resolveModel` at that package and wire it through `resolveProvider`:

```ts theme={null}
import { openai } from "@ai-sdk/openai";
import { connect, createAisdkProvider, getProvider } from "@duraton/sdk";

await connect({
  workflows,
  resolveProvider: (name) =>
    name === "aisdk"
      ? createAisdkProvider({ resolveModel: (model) => openai(model) })
      : getProvider(name),
});
```

Then ask for it per call:

```ts theme={null}
const answer = await ctx.step.ai.generate("draft", {
  provider: "aisdk",
  model: "gpt-5.1",
  prompt: "Summarise this ticket.",
});
```

Without `resolveModel` the model string is passed to the AI SDK as-is, which resolves it through its
global provider - the Vercel AI Gateway, requiring `AI_GATEWAY_API_KEY`. Supply `resolveModel`
whenever you want to call a provider directly rather than route through the gateway.

Two behaviours are worth knowing:

* **`apiKey` on the call is ignored by this adapter.** The AI SDK carries credentials on the model,
  so the key belongs to whatever `resolveModel` returns (`openai({ apiKey })`).
* **Duraton still owns the loop and the retries.** The adapter makes exactly one model call per
  step and disables the AI SDK's own retries, so a rate limit checkpoints and reschedules durably
  instead of blocking a worker. Tools are declared to the AI SDK without an executor, so every tool
  call comes back to [`step.ai.loop`](#step-ai-loop) and stays a durable, approvable step.

The `anthropic` adapter is not deprecated by this. It imports no framework, which is what keeps
`GenerateRequest` from drifting into any one vendor's types.

### Tool calling

A `GenerateRequest` may carry `tools: ToolDeclaration[]`, and a `GenerateResult` may answer with
`toolCalls: ToolCall[]`. If you write your own adapter, honour both halves:

| Rule                                                                                 | Why                                                                                    |
| ------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------- |
| No `tools`, or an empty array, means send **no** tools field to your provider at all | A call that declares none must be indistinguishable from one made before tools existed |
| Leave `toolCalls` **unset** when the model requested none - never an empty array     | `toolCalls?.length` is the only check a caller makes                                   |
| A turn may carry text **and** tool calls; return both                                | Neither displaces the other                                                            |
| Hand `input` back as the provider produced it                                        | Validation is the caller's job, not the adapter's                                      |

Tell a tool turn from a final answer by whether `toolCalls` is present - never by reading
`stopReason`, which stays your provider's own raw string.

### Conversations

A multi-turn agent has a conversation, and `GenerateRequest.transcript` carries it: the exchanges
that have already happened, oldest first, with `prompt` as the task that opened them. A provider's
native tool-use protocol has a shape for this - Anthropic answers an assistant `tool_use` block
with a user `tool_result` block - and the built-in adapter maps the transcript onto it.

Ignoring the field would silently lose the history rather than lose a nicety, so an adapter has to
say it reads it:

```ts theme={null}
const recording: AIProvider = {
  name: "anthropic",
  capabilities: ["transcript"],
  generate: (req) => callMyProvider(req.prompt, req.transcript ?? []),
};
```

An adapter that declares nothing keeps working exactly as it did. The
[agent kit](/agent-kit) checks with `providerSupports` and renders the turns into the prompt
as text for anything that has not opted in, so no adapter is broken by the field existing.

<Note>
  A turn carries the tool calls and their results, not the assistant's prose. `step.ai.loop`
  records exactly that much per turn, and a transcript built from anything else would stop being
  identical on replay.
</Note>

<Note>
  The API key rides each `GenerateRequest` and is never stored by the SDK, never journaled, and never
  sent to Duraton. Omit it and the adapter falls back to its provider SDK's conventional env var (for
  Anthropic, `ANTHROPIC_API_KEY`). Your model keys stay in your runner.
</Note>

### Prompt caching

`promptCache` opts a call into **the provider's own prompt cache**: the provider still runs the call,
and the only thing that changes is what it bills for the part of the request it has already seen. It
is a different control from the [inference cache](#inference-cache), which skips the provider call
altogether, and the two can be set on the same call.

It is a capability like `transcript` above - an adapter that can place a cache breakpoint declares
`"prompt-cache"` - and the scope travels from the call site to the adapter unchanged: `GenerateOptions`
(or the agent kit's `AgentOptions`) to `GenerateRequest.promptCache`. **The adapter alone decides where
the breakpoints go.** `step.ai.loop` places none of its own; a turn is cached because the option reached
the provider, not because the loop did anything to the request.

The built-in `anthropic` adapter places them like this:

| Scope            | What the request carries                                                                                                                                                                                    |
| ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `"prefix"`       | One breakpoint closing the static head: the last `system` block, or the last tool declaration when the call has no system prompt.                                                                           |
| `"conversation"` | That breakpoint, plus a top-level `cache_control` that re-places itself on the last cacheable block as the request grows, so turn N+1 reads turn N's transcript back instead of paying to process it again. |
| unset            | Nothing at all - the request is byte-identical to one made before the option existed.                                                                                                                       |

No TTL is sent. `{ type: "ephemeral" }` is the only cache type the Messages API defines, and Duraton
does not override the API's own default lifetime (five minutes, today), because a read refreshes the
entry and an agent's next turn starts well inside that window.

The `aisdk` adapter does **not** declare the capability. The AI SDK carries cache control on
`providerOptions` keyed by the provider's own name, and the adapter never learns which provider
`resolveModel` returned - so it cannot place a breakpoint the answering model would honour.

Asking an adapter that has not declared it **fails the call** instead of sending it uncached. Both
`step.ai.generate` and `agent()` throw a plain `Error`; there is no error class and no code, and the
one stable thing to match on is the shared substring `cannot place a cache breakpoint` in the message.
The refusal is deliberate: an adapter that dropped the field would still answer correctly, just at the
full input price forever, and both cache axes would read zero - exactly what a cache that was asked for
and missed looks like.

`promptCache` is a `step.ai.generate` and `agent()` option only. It is not on `step.ai.loop`, where the `turn` you write owns the request it sends.

## Cost

Duraton holds no model price list, so a call's `cost` is absent unless your runner supplies it. Pass
`resolveCost` - the `CostSource` port - and each `step.ai` call is priced from the axes the journal
already holds. Supplying it is what makes `cap: { maxCost }` bite; `maxTokens`
needs nothing, because tokens are metered from every call.

```ts theme={null}
import type { CostSource } from "@duraton/sdk/ai";

const PRICES: Record<string, { in: number; out: number }> = {
  "claude-opus-4-8": { in: 5 / 1_000_000, out: 25 / 1_000_000 },
};

const resolveCost: CostSource = ({ model, tokensIn = 0, tokensOut = 0 }) => {
  const p = model ? PRICES[model] : undefined;
  return p ? tokensIn * p.in + tokensOut * p.out : undefined;   // undefined leaves cost absent
};

connect({ app: "support-app", workflows, resolveCost });
```

<ResponseField name="kind" type="AIStepKind" required>
  Which step.ai call this was: "generate", "wrap", "embed", or "loop" (one agent turn).
</ResponseField>

<ResponseField name="model" type="string">
  The model that served the call.
</ResponseField>

<ResponseField name="tokensIn" type="number">
  Input tokens.
</ResponseField>

<ResponseField name="tokensOut" type="number">
  Output tokens.
</ResponseField>

<ResponseField name="cacheReadTokens" type="number">
  Tokens read from the provider's own prompt cache, when it reports them.
</ResponseField>

<ResponseField name="cacheCreationTokens" type="number">
  Tokens written to the provider's own prompt cache, when it reports them.
</ResponseField>

Returning `undefined` leaves the cost absent - Duraton never fabricates a zero - and a call that
already carries an explicit cost is left untouched.

## Cache store

The [inference cache](#inference-cache) is backed by the **`AICache`** port, so the store is swappable.
The default is `createMemoryCache()`: a process-local `Map` with per-entry TTL and LRU eviction, bounded
at **1000** entries. Pass `cache` to `connect` to swap it - for a store shared across runner
processes, say.

```ts theme={null}
import { createMemoryCache } from "@duraton/sdk/ai";

connect({
  app: "support-app",
  workflows,
  cache: createMemoryCache({ maxEntries: 10_000 }),
});
```

<ResponseField name="get" type="(key: string) => CacheEntry | undefined" required>
  A live entry for this key, or undefined on a miss (absent or expired). Best-effort: a miss costs one real call, never correctness, so it must not throw.
</ResponseField>

<ResponseField name="set" type="(key: string, product: GenerateStepProduct, ttlMs: number) => void" required>
  Store a call's product under key for ttlMs, overwriting any existing entry.
</ResponseField>

The store is only ever consulted for a call that opted in with `cache` - its mere presence changes
nothing. The cached completion is held **runner-side**: Duraton's journal records only the cache
metadata (`hit`, `key`, `ageMs`), never the payload. A store you share across processes must seed its
keys deliberately, since the default seed (the runner's app) assumes the process boundary isolates it.

## Related

* [AI agents](/ai/ai-steps) - when to reach for `generate` vs `loop`, workflow tools, the agent patterns, and the replay rules.
* [Cost controls](/ai/cost-controls) - `cap`, `tokenThrottle`, the inference cache, and fallback chains.
* [Agent kit](/agent-kit) - `agent()` and `tool()` on top of `step.ai.loop`.
* [Guardrails](/ai/guardrails) - the policy port, the adapters that ship, and how a tripwire halts a run.
