workflow next
to retry, plus three per-call options that cut spend and absorb provider failures.
The reference tables for the two workflow fields live in
flow control; the per-call options are in
step.ai.generate. This page is how to choose between them.
cap - bound one run
cap is a hard ceiling on a single run’s AI spend. The run halts before the step.ai call
that would cross a set axis - that call never runs - and fails with a BudgetError. Everything
committed before the halt stays committed.
At least one axis is set. The ceiling is crossed by at most the one call that reaches it: a call’s
cost is unknown until it returns, so the call that pushes spend to the limit completes and the
next one halts.
What you see. A capped run is an ordinary
failed run in the console, carrying BudgetError as
its terminal error - there is no separate “capped” state. In an agent loop, the halted turn is the
loop’s last (failed) iteration. Raise the cap and replay the run and it starts
fresh with spend back at zero.
maxTokens always bites - tokens are metered from every model call. maxCost bites only when your
runner prices its calls through a resolveCost map; Duraton holds no price list, so without one the
cost axis stays inert.tokenThrottle - bound the rate
AtokenThrottle protects your own provider quota: at most tokens spent per perMs across the
runs sharing a key. Duraton debits each AI step’s actual token usage after the step commits, and
when the bucket is drained it delays new run starts for that key - the run waits in the queue
holding no runner, then runs normally.
Because a step’s tokens are known only after it runs, the throttle gates a run’s start on the key’s
recent usage: a key that has recently spent heavily has its next runs spread out, a fresh key
starts immediately. It shapes the rate of starts, not any single run - pair it with a
cap to also bound
one run.
What you see. A throttled run sits in queued with a future start time, then runs normally.
There is no new state to handle.
cache - don’t pay twice for the same call
The inference cache is the one control that reduces spend rather than bounding it. On a hit the provider is never called, so the step commits with zero spend and counts nothing againstcap or tokenThrottle. Where step memoization makes
a replay free, the cache makes an identical call in a different run free too.
The key is an exact match over the seed, the provider, the model, the prompt, and every
output-affecting parameter (
system, temperature, maxTokens, output), so a changed prompt or
model never returns a stale answer. apiKey is never part of the key.
Caching engages only when
temperature is explicitly set to 0.2 or lower. An unset
temperature is treated as non-deterministic (a provider default is often 1.0), so cache: true
with no temperature is a documented no-op - the call runs and is charged.hit, key,
and ageMs, with zero tokens recorded. The console’s AI view has a cache hit rate for the
window.
promptCache - pay the cached rate for a repeated prefix
promptCache opts a call into the provider’s own prompt cache: the provider still runs the call,
and the only thing that changes is what it charges for the part of the request it has already seen.
That is the opposite trade to cache above - the inference
cache skips the provider call entirely on an exact repeat,
while this one keeps calling and re-prices the repeated prefix. The two are independent and can be
set together.
- step.ai.generate
- agent()
"prefix" caches the static head - the tool declarations and the system prompt - which is what
many calls sharing one long instruction block but ending differently need. "conversation" caches
that head and follows the transcript as it grows, so turn N+1 reads turn N’s exchanges back instead
of paying to process them again; that is the scope an agent wants. Duraton sends no TTL of its own, so
an entry lives for whatever the provider defaults to (five minutes, on Anthropic).
It refuses rather than quietly ignoring you. The option only reaches a provider that declares the
prompt-cache capability; asking any other adapter fails the call with an error saying the provider
cannot place a cache breakpoint. Of the two built-in adapters only anthropic declares it - see
Prompt caching. Silence would be worse than a failure here:
an adapter that dropped the field would still answer correctly, just at the full input price forever,
and both cache axes would read zero - indistinguishable from a cache that was asked for and missed.
What you see. The step’s AI journal gains cacheReadTokens and cacheCreationTokens once the
provider reports them, and the run inspector’s AI pane shows them as cache read and cache
write. They are separate axes from tokensIn, not a slice of it - see token and cost
spend. The console’s cache hit rate is the inference
cache’s: a call that used only promptCache never enters that denominator.
A cache write costs more than a plain call and a read costs a fraction of one, so a prefix cached
and never read back is a loss. It pays from the second call on - a long system prompt many calls
share, or an agent transcript re-sent every turn.
fallback - survive a rate-limited model
A fallback chain keeps one call alive when a model is rate-limited or down. The primarymodel is tried first; a retryable failure (429, a 5xx, or a
timeout) advances to the next candidate, and the first to return wins. Its result is the step’s
durable output, so the caller never sees the failover.
Only 429, 5xx, and timeout advance the chain. A terminal 4xx (a malformed request, an auth failure)
fails the step immediately - another model will not fix a bad request - and an exhausted chain fails
the step too, re-throwing the last error so the workflow’s own retry
policy still applies. Falling back to a cheaper model changes what
the call costs, so a chain interacts with
cap through whichever model actually served.
What you see. The step shows a chain pill carrying chain (the models tried, in order), used
(the one that served), and reason (why the chain advanced, e.g. "claude-opus-4-8: 429").
Watching spend
For totals rather than ceilings, read the window’s spend, tokens, average latency, and cache-hit rate, broken down by hour, by model, and by workflow, from AI observability. An agent reads the same numbers through theai_spend MCP tool.