step.ai makes a model call a durable step. Like step.run, each call
takes a stable id, records its result under that id, and returns the saved result on replay instead
of calling the model again. So a retry after a crash never re-spends on work that already completed.
New to AI steps? The AI quickstart walks you from a first
generate call to spend landing in the console.
Duraton stores the AI metadata (model, token counts, latency) as an opaque journal block. It never
parses it and never stores your prompt, the response text, or your API key - those stay in your runner.
The step.ai API
<T>(id, opts) => Promise<StructuredResult<T>>
required
One model call as a durable step, with optional structured output + durable re-ask.
<T>(id, fn) => Promise<T>
required
Make any caller-supplied AI call durable - bring your own client.
(id, opts) => Promise<EmbedResult>
required
Batch embeddings with per-batch checkpointing.
<T>(id, opts) => Promise<LoopResult<T>>
required
A durable agent loop: one durable step per turn and per tool call.
<T>(id, opts) => Promise<CheckResult<T>>
required
One guardrail policy pass as a durable step, optionally guarding a call of your own.
step.ai.generate
Make one model call as a durable step. The built-in provider is Anthropic; the request is validated,
sent, and the result memoized under id.
string
required
The model id to call, e.g. “claude-opus-4-8”.
string
required
The user prompt.
string
An optional system prompt.
number
Max output tokens.
number
Sampling temperature.
Record<string, unknown>
A JSON Schema the response must satisfy. Set it to get a parsed, validated result on output, with durable re-ask on failure.
number
Max durable re-asks when output validation fails (default 1). Each re-ask is its own memoized step; 0 disables re-asking.
(value: unknown) => string | undefined
A deeper check beyond “valid JSON”: return an error string to reject the value (triggers a re-ask), or undefined to accept.
ProviderName
The provider adapter to use. Defaults to “anthropic”.
boolean
Stream tokens as they arrive: each delta is journaled as an ai_chunk timeline frame for a live, replayable view. The durable result is still the complete text. Falls back to a plain generate when the transport has no live channel. See Streaming below.
FallbackCandidate[]
Backup models tried in order after model when a call fails with a retryable error (429/5xx/timeout). The first to return wins. Each candidate is { model, provider? }. See Fallback chains below.
boolean | CacheOptions
Opt this call into the inference cache: an identical prior call is served without a provider call (zero spend). Engages only when temperature is explicitly <= 0.2. true uses a 24h TTL; CacheOptions is { ttlMs?, seed? }. See Inference cache below.
PromptCacheScope
Ask the provider to bill this call’s repeated prefix at its own cached rate - a different thing from cache above, which skips the provider call entirely. “prefix” caches the stable head, “conversation” also follows a transcript that grows call over call. The provider must declare the prompt-cache capability, or the call is refused rather than quietly billed in full. See Prompt caching below.
string
Passed through per call and never stored; omit to fall back to the provider SDK env var (e.g. ANTHROPIC_API_KEY).
generate returns a StructuredResult<T>:
string
required
The response text.
string
required
The model that actually answered - may differ from the requested one (e.g. a server-side fallback). This is what the journal records.
ProviderName
required
The provider that served the call.
TokenUsage
required
Four disjoint token axes: inputTokens counts the uncached input only, so inputTokens + cacheReadTokens + cacheCreationTokens is everything the provider processed. The two cache axes are filled in only when the provider reports them. See Prompt caching below.
string
Why generation stopped.
T
The parsed, validated value - present only when you passed output.
Structured output and durable re-ask
Passoutput (a JSON Schema) to constrain the model and get a typed, validated value back on
result.output. If the response fails to parse or validate, generate re-prompts with the validation
error - each re-ask is its own memoized step, so the retry survives a crash and never repeats a
committed attempt. Add validate for semantic rules the schema can’t express.
The
apiKey you pass is used for that one call and never written to the journal or the run store.
Omit it to let the provider SDK read its conventional env var.Streaming
Passstream: true to feed the model’s tokens to a live viewer as they arrive. Each delta is appended to
the run’s durable timeline as an ai_chunk frame, so a viewer sees the text build in real time and a late
or reconnecting viewer replays it from token 0. The return value is unchanged - result.text is still the
complete response, memoized on replay - so streaming affects only what a viewer sees while the step runs.
useStream React hook.
Fallback chains
Passfallback - an ordered list of backup models - to keep a call resilient when a model is rate-limited
or down. The primary model is tried first; if it fails with a retryable error (429, a 5xx, or a
timeout), the call advances to the next candidate, and the first one to return wins. Its result is the
step’s durable output, so a caller never sees the failover.
{ model, provider? }; provider defaults to the call’s provider, so a chain can span
providers once you have more than one adapter configured. The step’s journal records the outcome:
string[]
required
The models tried, in order (the primary plus each fallback).
string
required
The model that actually served the call.
string
Why the chain advanced - the classified failures of the skipped models, e.g. “claude-opus-4-8: 429”. Absent when the primary served.
chain / used / reason.
Only 429, 5xx, and timeout advance the chain. A terminal 4xx (a bad request, an auth failure) fails the
step immediately - another model won’t fix a malformed request. An exhausted chain also fails the step,
re-throwing the last error, so the workflow’s own durable retry policy still applies. Fallback is
per-call resilience, distinct from the flow-control spend controls
(
cap / tokenThrottle).Inference cache
Setcache to reuse the result of an identical earlier call instead of paying for it again. On a hit
the provider is never called, so the step commits with zero spend - the cache is the one control that
reduces spend rather than capping it, and a cached call counts nothing against cap /
tokenThrottle. Where step memoization already makes a replay free, the cache makes an identical call
in a different run free too.
boolean
required
Whether the call was served from the cache (true = the provider was not called).
string
required
The entry key (metadata only - the cached completion is held runner-side and never reaches Duraton).
number
How long ago the entry was stored, on a hit.
hit / key / ageMs.
Caching is exact-match and engages only when
temperature is explicitly set to 0.2 or lower -
caching a sampled (high-temperature) answer would freeze one draw, and an unset temperature is treated
as non-deterministic (a provider default is often 1.0). The default TTL is 24h, overridable per call with
{ ttlMs }; { seed } overrides the default project seed (the runner’s app) to scope entries further.step.ai.wrap
Makes a model call you already write yourself - through the OpenAI SDK, the Anthropic SDK, the Vercel
AI SDK, or anything else - a durable step, with no other change to the call site. wrap returns your
function’s value unchanged; when it recognizes the response shape it enriches the journal with the
model and token counts and records which library it wrapped.
wrap step - you just get less metadata on the journal.
step.ai.embed
Turn a list of inputs into vectors, one durable batch at a time. Anthropic has no embeddings API, so
you supply the embedding call (embed); Duraton owns the batching and per-batch checkpointing. If a
batch fails, only that batch re-runs on retry - committed batches are not re-embedded.
string
required
Names the embedding model - recorded on the journal.
string[]
required
The inputs to embed.
(batch: string[]) => number[][] | Promise<number[][]>
required
Your embedding call for one batch: inputs in, one vector per input out.
number
Inputs per durable batch (default 100). Each batch checkpoints independently.
embed returns { vectors } - one vector per input, in input order.
step.ai.loop
A durable agent loop. Each turn is your own model call (bring-your-own, normalized to tool calls or a
final answer); the loop executes the tools the turn requested and feeds the results into the next turn,
until the model returns a final answer, stop fires, or maxIterations is reached.
Every turn is a durable step, and so is every tool call it makes, so an agent that crashes mid-run
resumes at the last committed turn.
Writing turn yourself is the low-level path. To declare a model, instructions and tools and have
the turn composed for you, use the agent kit - it composes this loop rather than
replacing it, so everything below still applies.
string
required
The task the agent is working on; surfaced to turn via ctx.prompt.
(ctx, iteration) => LoopTurn | Promise<LoopTurn>
required
One model turn, a pure function of ctx - keep it deterministic for replay.
number
required
Hard cap on turns; the loop halts before exceeding it.
Record<string, LoopTool>
The tools the model may call, keyed by name.
ApprovalRule
The gate for every tool that has not answered for itself. Resolved against each tool’s own approval - see the agent kit for the order.
number
Ceiling on the human decisions this loop may ask for. Past it the loop halts with stopReason “approval-budget”. Omitted, there is none.
(ctx) => boolean
Optional early stop after a completed turn; must be pure for replay.
turn returns a LoopTurn - either tool calls to run, or a final answer:
Array<{ id, name, input }>
Tools to run this turn; each name must be a key in tools.
unknown
The final answer. Returning this ends the loop with stopReason “final”.
string
The model that produced the turn - recorded on the journal.
number
Input tokens for this turn.
number
Output tokens for this turn.
Which loop composed a turn
Every turn’s journal carries aloopVersion alongside the model and the token counts. The model
says what answered; this says what asked:
typescript.1, so an entry
read later needs nothing but itself to be understood. Only turns carry it: a generate or an embed
was not composed by the loop.
Read it, do not pin on it. A bump means turns composed after it may differ from turns composed
before, which is a reason to compare two runs carefully - not a reason to refuse the older one.
(input: unknown) => unknown | Promise<unknown>
A local function tool.
string
The name of a workflow to run as this tool; its call becomes a linked child run - the tool’s step carries childRunId and the spawned run carries parentRunId, parentStep and parentAttempt, so the call is followable both ways.
string
The workflow tool’s app; addressed like step.runWorkflow.
string
Pin the workflow tool to a specific runner.
boolean
Park the run on a human before this tool runs. See Gating a tool on a human below.
ApprovalRule
Annotates that gate and decides whether it is raised at all. See Gating a tool on a human below.
Gating a tool on a human
A tool markedrequiresApproval does not run until someone decides on it. The run parks in
needs_attention holding no worker, exactly as step.approval does -
it is the same gate, raised for you:
A tool’s
approval says more about that gate and can decide whether it is raised at all, and the
loop’s own approval is the default for every tool that has not answered for itself. Both are the
same ApprovalRule: the environment a condition reads is on
Approvals, and the order the two levels resolve in
is on the agent kit.
maxIterations bounds how many turns a loop may run; a per-run spend
cap bounds how much it may spend. When a run reaches its cap the loop
halts before its next turn’s model call and the run fails with a BudgetError - the committed turns
stay, and the halted turn is the loop’s last (failed) iteration.
ctx.history gives each turn the prior turns’ toolCalls and toolResults, so your model call can
see what it has already tried. loop returns:
T
The final answer, if the loop reached one.
number
required
How many turns ran.
"final" | "max-iterations" | "stopped" | "approval-budget" | "bail" | "guardrail"
required
Why the loop ended. approval-budget means the loop reached maxApprovals; bail means a tool returned bail(); guardrail means a policy returned halt.
string
Which named stop condition ended the loop, on stopReason “stopped”. Absent when the opts.stop closure ended it, because a closure has no name to report.
string
Which guardrail halted the loop, on stopReason “guardrail”. The name only, never the verdict’s reason.
step.ai.check
Ask a guardrail policy about a value, as a durable step. The verdict memoizes, so a replay reads what
was decided instead of asking again. Supply call to guard one call with the policy: the check gates
it, or with parallel: true races it and aborts its signal the moment the policy trips.
Guardrail[]
required
The policies to ask, in order; the first verdict that is not allow wins. A guardrail that does not declare the subject’s placement is skipped.
GuardrailInput
What to check before the guarded call: placement, value, and optionally schema and tool.
(result: T) => GuardrailInput
Derives what to check from the call’s result. Omitted, nothing is checked afterwards.
(signal: AbortSignal) => T | Promise<T>
The call the checks guard. Omitted, check is a plain policy evaluation.
boolean
Race the input check against the call instead of gating the call on it.
check returns { verdict, tripped, value?, result? } - value is the replacement under a mask or
rewrite verdict, and result is the guarded call’s own return, withheld whenever a check refused it.
A halt verdict fails the check’s step non-retriably rather than returning; every other refusal
comes back as tripped. See Guardrails for the placements, the actions, and the
adapters that ship.
It records the input check, the guarded call and the output check as three separate durable steps.
The verdict step journals kind: "check" with the placement, the action and the deciding
guardrail - never the verdict’s reason.
Providers
step.ai.generate resolves its provider name to an AIProvider adapter through a port, so you
can supply your own instead of the built-in registry. Pass resolveProvider to
connect and every generate call in that runner
goes through it - the call sites are unchanged. The same resolver is also reachable directly as
ctx.resolveProvider, which is what step.ai.loop (and the agent kit’s
agent()) resolves its own model call through, since a loop’s turn only ever sees
(ctx, iteration) and has no opts.provider-shaped call site of its own to inject into.
An
AIProvider implements generate(req) and, optionally, stream(req, onDelta) (an adapter without
it falls back to generate, so stream: true still returns the right text) and classifyError(err)
(which decides whether a failure is retryable, and so whether a fallback chain
advances - an unclassified error is treated as terminal). It may also declare capabilities -
see Conversations.
Using any model provider
Theaisdk provider delegates to the Vercel AI SDK, so one adapter reaches
every provider the AI SDK supports - OpenAI, Google, Mistral, Bedrock, Groq and the rest - without
Duraton shipping an adapter per vendor. Install ai alongside the provider package you want:
resolveModel at that package and wire it through resolveProvider:
resolveModel the model string is passed to the AI SDK as-is, which resolves it through its
global provider - the Vercel AI Gateway, requiring AI_GATEWAY_API_KEY. Supply resolveModel
whenever you want to call a provider directly rather than route through the gateway.
Two behaviours are worth knowing:
apiKeyon the call is ignored by this adapter. The AI SDK carries credentials on the model, so the key belongs to whateverresolveModelreturns (openai({ apiKey })).- Duraton still owns the loop and the retries. The adapter makes exactly one model call per
step and disables the AI SDK’s own retries, so a rate limit checkpoints and reschedules durably
instead of blocking a worker. Tools are declared to the AI SDK without an executor, so every tool
call comes back to
step.ai.loopand stays a durable, approvable step.
anthropic adapter is not deprecated by this. It imports no framework, which is what keeps
GenerateRequest from drifting into any one vendor’s types.
Tool calling
AGenerateRequest may carry tools: ToolDeclaration[], and a GenerateResult may answer with
toolCalls: ToolCall[]. If you write your own adapter, honour both halves:
Tell a tool turn from a final answer by whether
toolCalls is present - never by reading
stopReason, which stays your provider’s own raw string.
Conversations
A multi-turn agent has a conversation, andGenerateRequest.transcript carries it: the exchanges
that have already happened, oldest first, with prompt as the task that opened them. A provider’s
native tool-use protocol has a shape for this - Anthropic answers an assistant tool_use block
with a user tool_result block - and the built-in adapter maps the transcript onto it.
Ignoring the field would silently lose the history rather than lose a nicety, so an adapter has to
say it reads it:
providerSupports and renders the turns into the prompt
as text for anything that has not opted in, so no adapter is broken by the field existing.
A turn carries the tool calls and their results, not the assistant’s prose.
step.ai.loop
records exactly that much per turn, and a transcript built from anything else would stop being
identical on replay.The API key rides each
GenerateRequest and is never stored by the SDK, never journaled, and never
sent to Duraton. Omit it and the adapter falls back to its provider SDK’s conventional env var (for
Anthropic, ANTHROPIC_API_KEY). Your model keys stay in your runner.Prompt caching
promptCache opts a call into the provider’s own prompt cache: the provider still runs the call,
and the only thing that changes is what it bills for the part of the request it has already seen. It
is a different control from the inference cache, which skips the provider call
altogether, and the two can be set on the same call.
It is a capability like transcript above - an adapter that can place a cache breakpoint declares
"prompt-cache" - and the scope travels from the call site to the adapter unchanged: GenerateOptions
(or the agent kit’s AgentOptions) to GenerateRequest.promptCache. The adapter alone decides where
the breakpoints go. step.ai.loop places none of its own; a turn is cached because the option reached
the provider, not because the loop did anything to the request.
The built-in anthropic adapter places them like this:
No TTL is sent.
{ type: "ephemeral" } is the only cache type the Messages API defines, and Duraton
does not override the API’s own default lifetime (five minutes, today), because a read refreshes the
entry and an agent’s next turn starts well inside that window.
The aisdk adapter does not declare the capability. The AI SDK carries cache control on
providerOptions keyed by the provider’s own name, and the adapter never learns which provider
resolveModel returned - so it cannot place a breakpoint the answering model would honour.
Asking an adapter that has not declared it fails the call instead of sending it uncached. Both
step.ai.generate and agent() throw a plain Error; there is no error class and no code, and the
one stable thing to match on is the shared substring cannot place a cache breakpoint in the message.
The refusal is deliberate: an adapter that dropped the field would still answer correctly, just at the
full input price forever, and both cache axes would read zero - exactly what a cache that was asked for
and missed looks like.
promptCache is a step.ai.generate and agent() option only. It is not on step.ai.loop, where the turn you write owns the request it sends.
Cost
Duraton holds no model price list, so a call’scost is absent unless your runner supplies it. Pass
resolveCost - the CostSource port - and each step.ai call is priced from the axes the journal
already holds. Supplying it is what makes cap: { maxCost } bite; maxTokens
needs nothing, because tokens are metered from every call.
AIStepKind
required
Which step.ai call this was: “generate”, “wrap”, “embed”, or “loop” (one agent turn).
string
The model that served the call.
number
Input tokens.
number
Output tokens.
number
Tokens read from the provider’s own prompt cache, when it reports them.
number
Tokens written to the provider’s own prompt cache, when it reports them.
undefined leaves the cost absent - Duraton never fabricates a zero - and a call that
already carries an explicit cost is left untouched.
Cache store
The inference cache is backed by theAICache port, so the store is swappable.
The default is createMemoryCache(): a process-local Map with per-entry TTL and LRU eviction, bounded
at 1000 entries. Pass cache to connect to swap it - for a store shared across runner
processes, say.
(key: string) => CacheEntry | undefined
required
A live entry for this key, or undefined on a miss (absent or expired). Best-effort: a miss costs one real call, never correctness, so it must not throw.
(key: string, product: GenerateStepProduct, ttlMs: number) => void
required
Store a call’s product under key for ttlMs, overwriting any existing entry.
cache - its mere presence changes
nothing. The cached completion is held runner-side: Duraton’s journal records only the cache
metadata (hit, key, ageMs), never the payload. A store you share across processes must seed its
keys deliberately, since the default seed (the runner’s app) assumes the process boundary isolates it.
Related
- AI agents - when to reach for
generatevsloop, workflow tools, the agent patterns, and the replay rules. - Cost controls -
cap,tokenThrottle, the inference cache, and fallback chains. - Agent kit -
agent()andtool()on top ofstep.ai.loop. - Guardrails - the policy port, the adapters that ship, and how a tripwire halts a run.