step.ai call records a journal block: the model that answered, its token
usage, and the call shape. Every surface on this page is a read over that journal.
The journal holds metering facts - model name, token counts, an optional supplied cost. Your prompt,
the response text, and your provider key stay in your runner: the provider key is passed per call to
the provider SDK and is never sent to Duraton.
Token and cost spend
duraton.ai.spend() (GET /ai/spend) rolls the journal up across
a project: window totals plus breakdowns by hour, by model, and by workflow.
number
required
Total input + output tokens in the window.
number
Total supplied cost - absent when no call in the window reported one.
number
required
Journaled model calls (steps that recorded a model).
number
required
The bucket width used for the hourly series.
AISpendPoint[]
required
Spend per time bucket: ts, tokens, cost?, calls.
AIModelSpend[]
required
Per-model rollup: model, tokens, cost?, calls, avgLatencyMs? (mean latency for that model, absent when none reported one).
AIWorkflowSpend[]
required
Per-workflow rollup: workflow, app, tokens, cost?, runs.
number
Mean call latency over calls that reported one - absent when none did, never 0-filled.
number
required
Calls in the window served from the inference cache.
number
required
Calls that used the cache at all; cacheHits / cacheEligible is the hit rate.
spend({ app, workflow, since, bucket }) scopes the rollup: app and workflow filter it, since
sets the window (an RFC3339 string or a Date), and bucket sets the hourly bucket width in seconds
(default 3600).
Tokens always, cost only when supplied
Duraton meters tokens and holds no price list, sotokens is always present while cost appears only
where a call supplied one. The distinction between “no cost reported” and “zero cost” is preserved: every
cost field stays absent rather than defaulting to 0. Supply prices with the resolveCost option on
connect() and every rollup above carries cost too.
Metering is read-only. To enforce a ceiling, declare a spend cap or a
token throttle on the workflow. maxTokens always applies;
maxCost bites only once resolveCost supplies a price.
The four token axes
A call meters four token counts, and they are disjoint.tokensIn is the uncached input
only, so tokensIn + cacheReadTokens + cacheCreationTokens is everything the provider processed, and
tokensOut is what it wrote back. Nothing is double-counted, so the axes can be summed without
knowing which provider answered: an adapter for a vendor that reports an inclusive prompt total
subtracts before filling them in.
The two cache axes are the provider’s own prompt cache - what a call opts into
with promptCache - and not the inference cache, whose hit rate is cacheHits / cacheEligible
above. They stay absent rather than 0 when the provider reported neither, the same rule cost
follows.
Their names differ by surface. The usage your generate call returns spells the first two
inputTokens and outputTokens; a step’s ai journal and the run usage on the wire spell them
tokensIn and tokensOut. The read axis is cacheReadTokens on all three. The write axis is the
one real rename: cacheCreationTokens in the SDK and on the journal, cacheWriteTokens on the
wire - so a run’s usage read back from the API names it differently from the result your own call
handed you.
Conversation sessions
Runs that belong to the same conversation form a session. Set an event’ssession to a stable
conversation id and every run it starts joins that session; omit it and each run is a session of one.
string
required
The conversation id (or a run’s own id when it started none).
number
required
Runs in the session.
Partial<Record<RunStatus, number>>
required
Runs per status; only statuses present appear.
number
Total AI tokens across the session - absent when no run made a model call.
number
Total supplied cost - absent when no run in the session reported one.
string
required
When the earliest run in the session started.
string
required
When the most recent run started.
list({ app, since, limit }) filters by app, bounds the window with since, and caps how many sessions
come back (default 100, max 500).
Run counts over time
runs.stats() returns the current per-status counts plus p50/p95 latency; runs.timeseries() returns
them bucketed over time, with a latency summary (average, longest, p50, p95) per bucket. These are the
two reads behind the console’s run charts.
Both accept
app, workflow, and since; timeseries also takes bucket (width in seconds, default
3600).
Latency percentiles
p50Ms is the median finished-run duration and p95Ms the 95th percentile, both in whole
milliseconds and over the same population as avgMs / maxMs (the scope’s, or bucket’s, finished
runs). They are continuous percentiles with linear interpolation between adjacent durations, so a p95
may fall between two observed values rather than on one - the standard reading of “95% of runs finished
at or below this.”
Percentiles use a single continuous definition (linear interpolation between adjacent durations),
so a percentile is never an average relabelled: p95 is the duration at or below which 95% of the
matched runs finished.
Logs and the run timeline
ctx.log lines and every status transition append to one durable
per-run timeline. Read it as history with runs.logs(id), or tail it live with
runs.watch(id).
Reading an agent loop as it runs
An agent’s durable steps are the record of what it did.agent_event frames are what make
that record readable as turns while the run is still going - they arrive on the same
per-run timeline, so one runs.watch sees them beside logs and status changes.
cost is present only when a cost source priced the call - Duraton holds no price list, so
an unpriced turn carries no cost rather than a zero. Likewise a turn served from the
inference cache reports a real tokensIn: 0, which is a different fact from carrying no
token count at all.
These frames carry no free text by construction: no prompt, no completion, no tool
arguments, no tool results, and no guardrail reason (a reason quotes the value it
refused). That is the same invariant the AI journal holds, and it is what
makes a trace safe to hand an operator who does not own the agent’s code.
stream flag and ttftMs, that turn’s time to first token, so latency is readable per turn
rather than per run. Its deltas arrive as ai_chunk frames rather than agent_event ones, and those
do carry the model’s completion text - the journal still holds none.
Traces in your own backend
The spans for your step bodies are emitted in your runner process by the SDK, and exported by whatever OpenTelemetry provider you register there (aNodeSDK, for example). Register none and every
step span is a no-op. Duraton runs the engine for you, so its internal telemetry is the platform’s to
operate; what reaches your own backend is what your runner emits.
step.ai span carries the OpenTelemetry GenAI semantic
conventions, so a vendor-neutral backend
reads it as a model call with no custom mapping. Only response-side facts the journal holds are emitted;
a field it does not hold (temperature, max tokens, the requested model before a fallback) is absent, never
guessed.
The span is named
{operation} {model} (e.g. chat claude-opus-4-8), or the bare operation when the
model is unknown - the convention’s own naming rule. Facts the convention has no attribute for live under
a duraton.* vendor prefix, so a standard backend ignores them:
From an AI assistant
The same rollups are MCP tools:ai_spend returns the spend rollup and list_sessions
returns the conversation list, both scoped to the caller’s project. Paired with the read tools for runs
and steps, an assistant can answer “which workflow is burning the most tokens” against your live project.