step.ai.generate call can stream the model’s tokens as
they arrive: each delta is appended to the run’s durable
timeline as an ai_chunk frame. The step’s durable result is
still the complete text, reused on replay, so streaming changes what a viewer sees while the step
runs, never what the step produces. An agent turn streams the same way - see
Streaming agent turns.
Turning on streaming
Passstream: true to generate:
result.text is the full response and the usage counts are the same.
stream: true only adds the live delta feed alongside the durable result.
Live deltas travel over the connect runner’s socket back to Duraton.
Streaming agent turns
An agent’s turns stream on the same terms.stream: true on agent()
opts every turn in; on step.ai.loop it does the same for a
loop whose turn you write yourself, where ctx.onDelta is where that turn sends its deltas:
- agent()
- step.ai.loop
onDelta behaves exactly as it did before the option existed.
Both built-in providers stream (
anthropic and aisdk), and an adapter that implements no stream
falls back to a plain call, so stream: true degrades rather than failing.
It is opt-in for what it costs: a timeline row per delta, and the model’s completion text on the
timeline. An agent that only needs the durable record should not pay for either.
On the timeline
Streamed deltas ride the same per-run timeline as status transitions and logs, as anai_chunk frame
kind:
Tail them with
runs.watch and reconstruct the text by concatenating deltas in index order:
ttftMs (time to first token) rides only the first delta of a stream, so you can surface latency the
moment generation begins. index is a per-stream counter; frames are appended in order and each carries
the run’s monotonic seq, so the lossless reconnect
rules apply unchanged - resume past the last seq you saw and you never miss or double-count a delta.
Resumability
Because every delta is a durable row, the stream is replayable, not ephemeral:- A viewer that opens the run after generation started replays every delta from token 0, rebuilds the full text, then tails the rest live.
- A refresh mid-stream loses nothing:
runs.watchreplays the history, so the text reconstructs exactly. - Deltas are keyed by
(step, attempt). If a crash re-runs the step on a later attempt, its stream carries a newattempt, so a resumed stream never mixes with the abandoned one - render only the latest attempt’s deltas.
In React
@duraton/react is the headless-hooks package behind the Duraton console, not something you
install - it exposes a useStream hook that does the reconstruction above for you, riding the same
durable timeline described in the client reference, so it is replay-safe and
reconnecting by construction. The shape below is illustrative of the pattern, not a copy-paste
install target; build the equivalent against the durable timeline directly if you need it in your
own app:
useStream(runId, step) returns the reconstructed text, the ttftMs once the first delta lands, a live
tokenCount, and a streaming flag that stays true until the step reaches a terminal status. It picks
the latest attempt automatically, so a crash-retried stream renders cleanly. The hook must be used under a
DuratonProvider holding a client.
In the console
A streaming generate step gets a Stream tab in the run’s step detail. It shows the text building live with a caret, the running token count, and the time to first token; after the step finishes it keeps the reconstructed text (no caret) - the same replay-from-token-0 view the API exposes.See streaming running end to end in the
examples.