Skip to main content
step.ai.loop’s history array grows by one entry every turn and nothing trims it. An agent that runs long enough eventually hits the model’s context window and fails - the only failure mode of its kind, since every other bound the loop enforces (maxIterations, maxApprovals) is a ceiling you set on purpose. A context trimmer is that bound, as a port. Duraton ships three adapters - count, token-budget and summarize - and the port is what lets a later one drop in without touching the loop or the adapters already there.

Turning it on

With @duraton/agent-kit it is a compact descriptor on agent() instead - see Bounding context for that surface. Omitted, an agent behaves exactly as it did before this option existed: the whole history reaches every turn.

The three strategies

count and token-budget never remove anything the run itself remembers - each is a fresh view computed before every turn, so the loop’s own history (what a replay reads, what a later stop predicate sees) still has every iteration. Recomputing costs nothing, so there is nothing to gain by remembering the last result. summarize is different: it pays for its own output; the loop’s own history is replaced by the result, so a later turn starts from the already-reduced baseline instead of paying to re-summarize the same growing prefix on every turn.
@duraton/agent-kit’s agent() builds generate for you from the same provider/model/apiKey every turn already uses - the raw SDK option above takes it directly because step.ai.loop has no provider of its own to resolve one from.

How a summary is represented

A summary does not add a new kind of history entry - LoopIteration stays exactly { toolCalls, toolResults }, the same shape it has always been. Instead, the earliest surviving iteration’s own toolResults[].output is overwritten with the summary text, reusing that entry’s real id and name rather than fabricating one:
This mirrors Anthropic’s own clear_tool_uses context-editing behaviour: an existing entry’s content is cleared or replaced in place, never a new one invented. A second summarization folds the same way - the reused entry’s current content (raw or already a summary) feeds the next summarization call, so history never carries more than one summary-bearing entry at once.

The summary is a durable step

Every context-trim pass - count and token-budget included, not only summarize - writes its own step, a sibling of the turn step and the guardrail step, never a suffix of either. Two consequences follow from that and from nothing else:
  • A replay reads the recorded result. summarize’s model call does not run again, so a re-run of the same run cannot disagree with the original, and the summary is an auditable fact rather than something re-derived on every pass.
  • Every adapter gets the same guarantee, whether or not it happens to call a model. count and token-budget are pure functions that would be replay-safe either way; wrapping them identically means adding a future model-calling adapter never needs new plumbing.
A loop with no context option writes no context-trim steps, so turning this on costs nothing until you do.

Writing a trimmer

A trimmer takes the accumulated history and the prompt, and returns a (possibly unchanged) result. It never mutates the caller’s history array - it says what the new view should be and the loop applies it.
persist is what separates a cheap per-turn view (false - recomputed every time, the loop’s own record is untouched) from a result that should become the new baseline (true - the loop replaces its own accumulator, so a later turn’s trim pass sees the reduced history rather than the original). Only set it true when the work being saved by not recomputing is worth the loop’s history no longer holding what was replaced.