Skip to main content
A step.ai.generate call can stream the model’s tokens as they arrive: each delta is appended to the run’s durable timeline as an ai_chunk frame. The step’s durable result is still the complete text, reused on replay, so streaming changes what a viewer sees while the step runs, never what the step produces. An agent turn streams the same way - see Streaming agent turns.

Turning on streaming

Pass stream: true to generate:
The return value is unchanged: result.text is the full response and the usage counts are the same. stream: true only adds the live delta feed alongside the durable result.
Live deltas travel over the connect runner’s socket back to Duraton.

Streaming agent turns

An agent’s turns stream on the same terms. stream: true on agent() opts every turn in; on step.ai.loop it does the same for a loop whose turn you write yourself, where ctx.onDelta is where that turn sends its deltas:
A turn that ignores onDelta behaves exactly as it did before the option existed. Both built-in providers stream (anthropic and aisdk), and an adapter that implements no stream falls back to a plain call, so stream: true degrades rather than failing.
It is opt-in for what it costs: a timeline row per delta, and the model’s completion text on the timeline. An agent that only needs the durable record should not pay for either.

On the timeline

Streamed deltas ride the same per-run timeline as status transitions and logs, as an ai_chunk frame kind: Tail them with runs.watch and reconstruct the text by concatenating deltas in index order:
ttftMs (time to first token) rides only the first delta of a stream, so you can surface latency the moment generation begins. index is a per-stream counter; frames are appended in order and each carries the run’s monotonic seq, so the lossless reconnect rules apply unchanged - resume past the last seq you saw and you never miss or double-count a delta.

Resumability

Because every delta is a durable row, the stream is replayable, not ephemeral:
  • A viewer that opens the run after generation started replays every delta from token 0, rebuilds the full text, then tails the rest live.
  • A refresh mid-stream loses nothing: runs.watch replays the history, so the text reconstructs exactly.
  • Deltas are keyed by (step, attempt). If a crash re-runs the step on a later attempt, its stream carries a new attempt, so a resumed stream never mixes with the abandoned one - render only the latest attempt’s deltas.
On replay, the generate step is memoized from its recorded result and returns the complete text without calling the model again - so a replay does not re-stream, and there is no re-spend on a recorded call.

In React

@duraton/react is the headless-hooks package behind the Duraton console, not something you install - it exposes a useStream hook that does the reconstruction above for you, riding the same durable timeline described in the client reference, so it is replay-safe and reconnecting by construction. The shape below is illustrative of the pattern, not a copy-paste install target; build the equivalent against the durable timeline directly if you need it in your own app:
useStream(runId, step) returns the reconstructed text, the ttftMs once the first delta lands, a live tokenCount, and a streaming flag that stays true until the step reaches a terminal status. It picks the latest attempt automatically, so a crash-retried stream renders cleanly. The hook must be used under a DuratonProvider holding a client.

In the console

A streaming generate step gets a Stream tab in the run’s step detail. It shows the text building live with a caret, the running token count, and the time to first token; after the step finishes it keeps the reconstructed text (no caret) - the same replay-from-token-0 view the API exposes.
See streaming running end to end in the examples.