Skip to main content
AI spend has two shapes of failure: one run runs away, or a burst saturates your provider quota. Duraton gives each its own control, declared on workflow next to retry, plus three per-call options that cut spend and absorb provider failures.
The reference tables for the two workflow fields live in flow control; the per-call options are in step.ai.generate. This page is how to choose between them.

cap - bound one run

cap is a hard ceiling on a single run’s AI spend. The run halts before the step.ai call that would cross a set axis - that call never runs - and fails with a BudgetError. Everything committed before the halt stays committed.
At least one axis is set. The ceiling is crossed by at most the one call that reaches it: a call’s cost is unknown until it returns, so the call that pushes spend to the limit completes and the next one halts. What you see. A capped run is an ordinary failed run in the console, carrying BudgetError as its terminal error - there is no separate “capped” state. In an agent loop, the halted turn is the loop’s last (failed) iteration. Raise the cap and replay the run and it starts fresh with spend back at zero.
maxTokens always bites - tokens are metered from every model call. maxCost bites only when your runner prices its calls through a resolveCost map; Duraton holds no price list, so without one the cost axis stays inert.

tokenThrottle - bound the rate

A tokenThrottle protects your own provider quota: at most tokens spent per perMs across the runs sharing a key. Duraton debits each AI step’s actual token usage after the step commits, and when the bucket is drained it delays new run starts for that key - the run waits in the queue holding no runner, then runs normally.
Because a step’s tokens are known only after it runs, the throttle gates a run’s start on the key’s recent usage: a key that has recently spent heavily has its next runs spread out, a fresh key starts immediately. It shapes the rate of starts, not any single run - pair it with a cap to also bound one run. What you see. A throttled run sits in queued with a future start time, then runs normally. There is no new state to handle.

cache - don’t pay twice for the same call

The inference cache is the one control that reduces spend rather than bounding it. On a hit the provider is never called, so the step commits with zero spend and counts nothing against cap or tokenThrottle. Where step memoization makes a replay free, the cache makes an identical call in a different run free too.
The key is an exact match over the seed, the provider, the model, the prompt, and every output-affecting parameter (system, temperature, maxTokens, output), so a changed prompt or model never returns a stale answer. apiKey is never part of the key.
Caching engages only when temperature is explicitly set to 0.2 or lower. An unset temperature is treated as non-deterministic (a provider default is often 1.0), so cache: true with no temperature is a documented no-op - the call runs and is charged.
What you see. A cache-served step shows a cache pill in the run inspector carrying hit, key, and ageMs, with zero tokens recorded. The console’s AI view has a cache hit rate for the window.

promptCache - pay the cached rate for a repeated prefix

promptCache opts a call into the provider’s own prompt cache: the provider still runs the call, and the only thing that changes is what it charges for the part of the request it has already seen. That is the opposite trade to cache above - the inference cache skips the provider call entirely on an exact repeat, while this one keeps calling and re-prices the repeated prefix. The two are independent and can be set together.
"prefix" caches the static head - the tool declarations and the system prompt - which is what many calls sharing one long instruction block but ending differently need. "conversation" caches that head and follows the transcript as it grows, so turn N+1 reads turn N’s exchanges back instead of paying to process them again; that is the scope an agent wants. Duraton sends no TTL of its own, so an entry lives for whatever the provider defaults to (five minutes, on Anthropic). It refuses rather than quietly ignoring you. The option only reaches a provider that declares the prompt-cache capability; asking any other adapter fails the call with an error saying the provider cannot place a cache breakpoint. Of the two built-in adapters only anthropic declares it - see Prompt caching. Silence would be worse than a failure here: an adapter that dropped the field would still answer correctly, just at the full input price forever, and both cache axes would read zero - indistinguishable from a cache that was asked for and missed. What you see. The step’s AI journal gains cacheReadTokens and cacheCreationTokens once the provider reports them, and the run inspector’s AI pane shows them as cache read and cache write. They are separate axes from tokensIn, not a slice of it - see token and cost spend. The console’s cache hit rate is the inference cache’s: a call that used only promptCache never enters that denominator.
A cache write costs more than a plain call and a read costs a fraction of one, so a prefix cached and never read back is a loss. It pays from the second call on - a long system prompt many calls share, or an agent transcript re-sent every turn.

fallback - survive a rate-limited model

A fallback chain keeps one call alive when a model is rate-limited or down. The primary model is tried first; a retryable failure (429, a 5xx, or a timeout) advances to the next candidate, and the first to return wins. Its result is the step’s durable output, so the caller never sees the failover.
Only 429, 5xx, and timeout advance the chain. A terminal 4xx (a malformed request, an auth failure) fails the step immediately - another model will not fix a bad request - and an exhausted chain fails the step too, re-throwing the last error so the workflow’s own retry policy still applies. Falling back to a cheaper model changes what the call costs, so a chain interacts with cap through whichever model actually served. What you see. The step shows a chain pill carrying chain (the models tried, in order), used (the one that served), and reason (why the chain advanced, e.g. "claude-opus-4-8: 429").

Watching spend

For totals rather than ceilings, read the window’s spend, tokens, average latency, and cache-hit rate, broken down by hour, by model, and by workflow, from AI observability. An agent reads the same numbers through the ai_spend MCP tool.