> ## Documentation Index
> Fetch the complete documentation index at: https://docs.duraton.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Streaming

> Show a viewer tokens as the model produces them and still get one durable result - the stream replays from token 0, the memoized value is the full text.

A [`step.ai.generate`](/reference/sdk/ai-steps#step-ai-generate) call can **stream** the model's tokens as
they arrive: each delta is appended to the run's durable
[timeline](/core/realtime) as an `ai_chunk` frame. The step's durable result is
still the complete text, reused on replay, so streaming changes what a viewer sees while the step
runs, never what the step produces. An agent turn streams the same way - see
[Streaming agent turns](#streaming-agent-turns).

## Turning on streaming

Pass `stream: true` to `generate`:

```ts theme={null}
const result = await ctx.step.ai.generate("summarize-thread", {
  model: "claude-opus-4-8",
  prompt: `Summarize this support thread:\n\n${thread}`,
  stream: true,
});
// result.text is the complete summary - identical to a non-streaming call.
```

The return value is unchanged: `result.text` is the full response and the usage counts are the same.
`stream: true` only adds the live delta feed alongside the durable result.

<Note>
  Live deltas travel over the [connect](/reference/sdk/connect) runner's socket back to Duraton.
</Note>

## Streaming agent turns

An agent's turns stream on the same terms. `stream: true` on [`agent()`](/agent-kit/agents-and-tools)
opts every turn in; on [`step.ai.loop`](/reference/sdk/ai-steps#step-ai-loop) it does the same for a
loop whose `turn` you write yourself, where `ctx.onDelta` is where that turn sends its deltas:

<Tabs>
  <Tab title="agent()">
    ```ts theme={null}
    const result = await agent(ctx, "triage", {
      model: "claude-opus-4-8",
      prompt: ticket.body,
      tools: [searchKb],
      maxIterations: 6,
      stream: true,
    });
    ```
  </Tab>

  <Tab title="step.ai.loop">
    ```ts theme={null}
    const result = await ctx.step.ai.loop("triage", {
      prompt: ticket.body,
      maxIterations: 6,
      stream: true,
      tools: { "search-kb": { handler: searchKb } },
      turn: (loop, iteration) => callModel(loop.prompt, loop.history, iteration, loop.onDelta),
    });
    ```
  </Tab>
</Tabs>

A turn that ignores `onDelta` behaves exactly as it did before the option existed.

| | On a streamed turn |
| - | - |
| The turn's result | Unchanged - the complete result, and the value a replay reads |
| Tool calls | Unchanged - the visible text streams while the turn's tool calls arrive intact |
| The timeline | One `ai_chunk` frame per delta, keyed to the turn's own durable step: turn 0 of `triage` streams under `triage:iter:0` |
| The turn's journal | Gains `stream: true` and `ttftMs`, so per-turn time to first token is on the record; the completion text stays out of the journal, as on any AI step |
| A replayed turn | Streams nothing - the memoized turn is returned without calling the model again |

Both built-in providers stream (`anthropic` and `aisdk`), and an adapter that implements no `stream`
falls back to a plain call, so `stream: true` degrades rather than failing.

<Note>
  It is opt-in for what it costs: a timeline row per delta, and the model's completion text on the
  timeline. An agent that only needs the durable record should not pay for either.
</Note>

## On the timeline

Streamed deltas ride the same per-run timeline as status transitions and logs, as an `ai_chunk` frame
[kind](/core/realtime#watching-a-run):

| `kind` | Fields beyond `seq` / `ts` / `runId` |
| - | - |
| `ai_chunk` | `step`, `attempt`, `index`, `delta`, `ttftMs?` |

Tail them with `runs.watch` and reconstruct the text by concatenating deltas in `index` order:

```ts theme={null}
import { createClient } from "@duraton/sdk/client";

const duraton = createClient({ url: process.env.DURATON_URL! });

let text = "";
for await (const frame of duraton.runs.watch(runId)) {
  if (frame.kind === "ai_chunk" && frame.step === "summarize-thread") {
    text += frame.delta;
    if (frame.ttftMs !== undefined) console.log("time to first token:", frame.ttftMs, "ms");
  }
}
```

`ttftMs` (time to first token) rides only the **first** delta of a stream, so you can surface latency the
moment generation begins. `index` is a per-stream counter; frames are appended in order and each carries
the run's monotonic `seq`, so the [lossless reconnect](/core/realtime#resuming-after-a-drop)
rules apply unchanged - resume past the last `seq` you saw and you never miss or double-count a delta.

## Resumability

Because every delta is a durable row, the stream is **replayable**, not ephemeral:

* A viewer that opens the run *after* generation started replays every delta from token 0, rebuilds the
  full text, then tails the rest live.
* A refresh mid-stream loses nothing: `runs.watch` replays the history, so the text reconstructs exactly.
* Deltas are keyed by `(step, attempt)`. If a crash re-runs the step on a later
  [attempt](/core/retries), its stream carries a new `attempt`, so a resumed stream never
  mixes with the abandoned one - render only the latest attempt's deltas.

On [replay](/core/durable-execution), the generate step is memoized from its recorded result and
returns the complete text without calling the model again - so a replay does not re-stream, and there is
no re-spend on a recorded call.

## In React

`@duraton/react` is the headless-hooks package behind the Duraton console, not something you
install - it exposes a `useStream` hook that does the reconstruction above for you, riding the same
durable timeline described in [the client reference](/reference/sdk/client), so it is replay-safe and
reconnecting by construction. The shape below is illustrative of the pattern, not a copy-paste
install target; build the equivalent against the durable timeline directly if you need it in your
own app:

```tsx theme={null}
import { useStream } from "@duraton/react";

function SummaryStream({ runId }: { runId: string }) {
  const { text, ttftMs, tokenCount, streaming } = useStream(runId, "summarize-thread");
  return (
    <div>
      <p>{text}{streaming && <span className="caret" />}</p>
      <small>
        {tokenCount} tokens{ttftMs !== undefined && ` · ttft ${ttftMs}ms`}
      </small>
    </div>
  );
}
```

`useStream(runId, step)` returns the reconstructed `text`, the `ttftMs` once the first delta lands, a live
`tokenCount`, and a `streaming` flag that stays `true` until the step reaches a terminal status. It picks
the latest attempt automatically, so a crash-retried stream renders cleanly. The hook must be used under a
`DuratonProvider` holding a [client](/reference/sdk/client).

## In the console

A streaming generate step gets a **Stream** tab in the run's step detail. It shows the text building live
with a caret, the running token count, and the time to first token; after the step finishes it keeps the
reconstructed text (no caret) - the same replay-from-token-0 view the API exposes.

<Note>
  See streaming running end to end in the
  [examples](/start/recipes#make-a-model-call-durable).
</Note>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.