> ## Documentation Index
> Fetch the complete documentation index at: https://docs.duraton.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Cost controls

> Stop an agent before it overspends: cap halts before the call, tokenThrottle spaces out runs, the cache replays identical calls free, fallback survives an outage.

AI spend has two shapes of failure: one run runs away, or a burst saturates your provider quota.
Duraton gives each its own control, declared on [`workflow`](/reference/sdk/defining-workflows) next
to `retry`, plus three per-call options that cut spend and absorb provider failures.

```ts theme={null}
workflow({
  name: "summarize",
  cap: { maxCost: 0.25 },                                  // ceiling on ONE run's spend -> fail
  tokenThrottle: { tokens: 100_000, perMs: 60_000 },       // token rate -> delay the start
  handler,
});
```

| Control | Scope | Acts | When it fires |
| - | - | - | - |
| `cap` | one run | before a `step.ai` call | The run **fails** with a `BudgetError`. |
| `tokenThrottle` | runs sharing a key | at run start, debited after each AI step | The next runs **start later**. |
| `cache` | one `generate` call | before the provider call | An identical prior call is served with **zero spend**. |
| `promptCache` | one `generate` call, or every turn of an `agent` | on the provider's side | A repeated prefix is billed at the provider's **cached rate**. |
| `fallback` | one `generate` call | after a retryable failure | The call **advances** to the next model. |

The reference tables for the two workflow fields live in
[flow control](/ai/cost-controls); the per-call options are in
[`step.ai.generate`](/reference/sdk/ai-steps#step-ai-generate). This page is how to choose between them.

## cap - bound one run

`cap` is a hard ceiling on a **single run's** AI spend. The run halts **before** the `step.ai` call
that would cross a set axis - that call never runs - and fails with a `BudgetError`. Everything
committed before the halt stays committed.

```ts theme={null}
workflow({
  name: "research.agent",
  cap: { maxCost: 0.25, maxTokens: 40_000 },
  handler: async (ctx) => {
    await ctx.step.ai.loop("agent", {
      prompt: ctx.event.data.brief,
      maxIterations: 20,
      tools: { search: { handler: (q) => search(q) } },
      turn: (c, i) => callModel(c.prompt, c.history, i),
    });
  },
});
```

| Property | Type | Default | Description |
| - | - | - | - |
| `maxCost` | `number` (USD) | off | Halt before a `step.ai` call once the run's summed cost reaches this. |
| `maxTokens` | `number` | off | Halt before a `step.ai` call once the run's summed tokens (in + out) reach this. |

At least one axis is set. The ceiling is crossed by at most the one call that reaches it: a call's
cost is unknown until it returns, so the call that pushes spend to the limit completes and the
**next** one halts.

**What you see.** A capped run is an ordinary `failed` run in the console, carrying `BudgetError` as
its terminal error - there is no separate "capped" state. In an agent loop, the halted turn is the
loop's last (failed) iteration. Raise the cap and [replay](/reference/api/runs) the run and it starts
fresh with spend back at zero.

<Note>
  `maxTokens` always bites - tokens are metered from every model call. `maxCost` bites only when your
  runner prices its calls through a `resolveCost` map; Duraton holds no price list, so without one the
  cost axis stays inert.
</Note>

## tokenThrottle - bound the rate

A `tokenThrottle` protects **your own provider quota**: at most `tokens` spent per `perMs` across the
runs sharing a key. Duraton debits each AI step's **actual** token usage after the step commits, and
when the bucket is drained it delays new run **starts** for that key - the run waits in the queue
holding no runner, then runs normally.

```ts theme={null}
workflow({
  name: "enrich.contact",
  tokenThrottle: { tokens: 100_000, perMs: 60_000, key: "customerId" }, // 100k tokens/min per customer
  handler,
});
```

| Property | Type | Default | Description |
| - | - | - | - |
| `tokens` | `number` | required | The token budget per window (in + out) across the runs sharing the key. Positive. |
| `perMs` | `number` (ms) | required | The window the token budget refills over. Positive. |
| `key` | `string` | whole workflow | An event-data path (e.g. `"customerId"`, `"user.id"`); each value gets its own independent rate. |

Because a step's tokens are known only after it runs, the throttle gates a run's start on the key's
**recent** usage: a key that has recently spent heavily has its next runs spread out, a fresh key
starts immediately. It shapes the rate of starts, not any single run - pair it with a `cap` to also bound
one run.

**What you see.** A throttled run sits in `queued` with a future start time, then runs normally.
There is no new state to handle.

## cache - don't pay twice for the same call

The [inference cache](/reference/sdk/ai-steps#inference-cache) is the one control that *reduces* spend
rather than bounding it. On a hit the provider is never called, so the step commits with **zero
spend** and counts nothing against `cap` or `tokenThrottle`. Where step memoization makes
a **replay** free, the cache makes an identical call in a **different run** free too.

```ts theme={null}
const answer = await ctx.step.ai.generate("answer", {
  model: "claude-opus-4-8",
  prompt: `Answer from this policy doc:\n${doc}\n\nQ: ${question}`,
  temperature: 0, // required - caching engages only for a deterministic call
  cache: { ttlMs: 3_600_000 },
});
```

| Property | Type | Default | Description |
| - | - | - | - |
| `cache` | `boolean` \| `CacheOptions` | off | `true` opts the call in with the defaults; an object overrides them. |
| `cache.ttlMs` | `number` (ms) | `86_400_000` (24h) | How long an entry stays servable. |
| `cache.seed` | `string` | your app name | Scopes entries further; entries never cross a project boundary. |

The key is an exact match over the seed, the provider, the model, the prompt, and every
output-affecting parameter (`system`, `temperature`, `maxTokens`, `output`), so a changed prompt or
model never returns a stale answer. `apiKey` is never part of the key.

<Note>
  Caching engages only when `temperature` is **explicitly** set to `0.2` or lower. An unset
  `temperature` is treated as non-deterministic (a provider default is often 1.0), so `cache: true`
  with no `temperature` is a documented no-op - the call runs and is charged.
</Note>

**What you see.** A cache-served step shows a cache pill in the run inspector carrying `hit`, `key`,
and `ageMs`, with zero tokens recorded. The console's **AI** view has a **cache hit** rate for the
window.

## promptCache - pay the cached rate for a repeated prefix

`promptCache` opts a call into **the provider's own prompt cache**: the provider still runs the call,
and the only thing that changes is what it charges for the part of the request it has already seen.
That is the opposite trade to `cache` above - the [inference
cache](/reference/sdk/ai-steps#inference-cache) skips the provider call entirely on an exact repeat,
while this one keeps calling and re-prices the repeated prefix. The two are independent and can be
set together.

<Tabs>
  <Tab title="step.ai.generate">
    ```ts theme={null}
    const answer = await ctx.step.ai.generate("answer", {
      model: "claude-opus-4-8",
      system: policyDoc, // the stable head every question re-sends
      prompt: question,
      promptCache: "prefix",
    });
    ```
  </Tab>

  <Tab title="agent()">
    ```ts theme={null}
    const result = await agent(ctx, "triage", {
      model: "claude-opus-4-8",
      prompt: ticket.body,
      tools: [searchKb],
      maxIterations: 6,
      promptCache: "conversation",
    });
    ```
  </Tab>
</Tabs>

| Property | Type | Default | Description |
| - | - | - | - |
| `promptCache` | `"prefix"` \| `"conversation"` | off | Which part of the request the provider is asked to hold. Unset, the request is byte-identical to one made before the option existed. |

`"prefix"` caches the **static head** - the tool declarations and the system prompt - which is what
many calls sharing one long instruction block but ending differently need. `"conversation"` caches
that head *and* follows the transcript as it grows, so turn N+1 reads turn N's exchanges back instead
of paying to process them again; that is the scope an agent wants. Duraton sends no TTL of its own, so
an entry lives for whatever the provider defaults to (five minutes, on Anthropic).

**It refuses rather than quietly ignoring you.** The option only reaches a provider that declares the
`prompt-cache` capability; asking any other adapter fails the call with an error saying the provider
`cannot place a cache breakpoint`. Of the two built-in adapters only `anthropic` declares it - see
[Prompt caching](/reference/sdk/ai-steps#prompt-caching). Silence would be worse than a failure here:
an adapter that dropped the field would still answer correctly, just at the full input price forever,
and both cache axes would read zero - indistinguishable from a cache that was asked for and missed.

**What you see.** The step's AI journal gains `cacheReadTokens` and `cacheCreationTokens` once the
provider reports them, and the run inspector's AI pane shows them as **cache read** and **cache
write**. They are separate axes from `tokensIn`, not a slice of it - see [token and cost
spend](/ai/observability#token-and-cost-spend). The console's **cache hit** rate is the inference
cache's: a call that used only `promptCache` never enters that denominator.

<Note>
  A cache write costs more than a plain call and a read costs a fraction of one, so a prefix cached
  and never read back is a loss. It pays from the second call on - a long system prompt many calls
  share, or an agent transcript re-sent every turn.
</Note>

## fallback - survive a rate-limited model

A [fallback chain](/reference/sdk/ai-steps#fallback-chains) keeps one call alive when a model is
rate-limited or down. The primary `model` is tried first; a **retryable** failure (429, a 5xx, or a
timeout) advances to the next candidate, and the first to return wins. Its result is the step's
durable output, so the caller never sees the failover.

```ts theme={null}
const answer = await ctx.step.ai.generate("answer", {
  model: "claude-opus-4-8",
  prompt: question,
  fallback: [{ model: "claude-sonnet-4-6" }, { model: "claude-haiku-4-5" }],
});
```

| Property | Type | Default | Description |
| - | - | - | - |
| `fallback` | `FallbackCandidate[]` | off | Backup models tried in order after `model`. |
| `fallback[].model` | `string` | required | The candidate model id. |
| `fallback[].provider` | `ProviderName` | the call's provider | The candidate's provider, so a chain can span providers. |

Only 429, 5xx, and timeout advance the chain. A terminal 4xx (a malformed request, an auth failure)
fails the step immediately - another model will not fix a bad request - and an exhausted chain fails
the step too, re-throwing the last error so the workflow's own [retry
policy](/core/retries) still applies. Falling back to a cheaper model changes what
the call costs, so a chain interacts with `cap` through whichever model actually served.

**What you see.** The step shows a chain pill carrying `chain` (the models tried, in order), `used`
(the one that served), and `reason` (why the chain advanced, e.g. `"claude-opus-4-8: 429"`).

## Watching spend

For totals rather than ceilings, read the window's spend, tokens, average latency, and cache-hit rate,
broken down by hour, by model, and by workflow, from [AI observability](/ai/observability). An agent
reads the same numbers through the [`ai_spend`](/integrations/mcp-server) MCP tool.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.