> ## Documentation Index
> Fetch the complete documentation index at: https://docs.duraton.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# AI observability

> Show someone what an agent did and what it cost: token and cost rollups, conversation sessions, run time-series, and GenAI spans read from the durable journal.

Every [`step.ai`](/ai/ai-steps) call records a journal block: the model that answered, its token
usage, and the call shape. Every surface on this page is a read over that journal.

```ts theme={null}
import { createClient } from "@duraton/sdk/client";

const duraton = createClient({
  url: process.env.DURATON_URL!,
  apiKey: process.env.DURATON_API_KEY,
});

const spend = await duraton.ai.spend({ since: "2026-07-01T00:00:00Z" });
console.log(spend.tokens, "tokens across", spend.calls, "calls");
```

<Note>
  The journal holds metering facts - model name, token counts, an optional supplied cost. Your prompt,
  the response text, and your provider key stay in your runner: the provider key is passed per call to
  the provider SDK and is never sent to Duraton.
</Note>

## Token and cost spend

`duraton.ai.spend()` ([`GET /ai/spend`](/reference/api/runs#ai-spend-&-sessions)) rolls the journal up across
a project: window totals plus breakdowns by hour, by model, and by workflow.

```ts theme={null}
const spend = await duraton.ai.spend({ app: "assistant", since: "2026-07-01T00:00:00Z" });
for (const m of spend.byModel) console.log(m.model, m.tokens, m.cost ?? "(no price)");
```

<ResponseField name="tokens" type="number" required>
  Total input + output tokens in the window.
</ResponseField>

<ResponseField name="cost" type="number">
  Total supplied cost - absent when no call in the window reported one.
</ResponseField>

<ResponseField name="calls" type="number" required>
  Journaled model calls (steps that recorded a model).
</ResponseField>

<ResponseField name="bucketSeconds" type="number" required>
  The bucket width used for the hourly series.
</ResponseField>

<ResponseField name="hourly" type="AISpendPoint[]" required>
  Spend per time bucket: ts, tokens, cost?, calls.
</ResponseField>

<ResponseField name="byModel" type="AIModelSpend[]" required>
  Per-model rollup: model, tokens, cost?, calls, avgLatencyMs? (mean latency for that model, absent when none reported one).
</ResponseField>

<ResponseField name="byWorkflow" type="AIWorkflowSpend[]" required>
  Per-workflow rollup: workflow, app, tokens, cost?, runs.
</ResponseField>

<ResponseField name="avgLatencyMs" type="number">
  Mean call latency over calls that reported one - absent when none did, never 0-filled.
</ResponseField>

<ResponseField name="cacheHits" type="number" required>
  Calls in the window served from the inference cache.
</ResponseField>

<ResponseField name="cacheEligible" type="number" required>
  Calls that used the cache at all; cacheHits / cacheEligible is the hit rate.
</ResponseField>

`spend({ app, workflow, since, bucket })` scopes the rollup: `app` and `workflow` filter it, `since`
sets the window (an RFC3339 string or a `Date`), and `bucket` sets the hourly bucket width in seconds
(default `3600`).

### Tokens always, cost only when supplied

Duraton meters tokens and holds no price list, so `tokens` is always present while `cost` appears only
where a call supplied one. The distinction between "no cost reported" and "zero cost" is preserved: every
`cost` field stays absent rather than defaulting to `0`. Supply prices with the `resolveCost` option on
`connect()` and every rollup above carries cost too.

Metering is read-only. To *enforce* a ceiling, declare a [spend cap](/ai/cost-controls) or a
[token throttle](/ai/cost-controls) on the workflow. `maxTokens` always applies;
`maxCost` bites only once `resolveCost` supplies a price.

### The four token axes

A call meters **four** token counts, and they are disjoint. `tokensIn` is the **uncached** input
only, so `tokensIn + cacheReadTokens + cacheCreationTokens` is everything the provider processed, and
`tokensOut` is what it wrote back. Nothing is double-counted, so the axes can be summed without
knowing which provider answered: an adapter for a vendor that reports an inclusive prompt total
subtracts before filling them in.

The two cache axes are [the provider's own prompt cache](/ai/cost-controls) - what a call opts into
with `promptCache` - and not the inference cache, whose hit rate is `cacheHits` / `cacheEligible`
above. They stay absent rather than `0` when the provider reported neither, the same rule `cost`
follows.

Their names differ by surface. The `usage` your `generate` call returns spells the first two
`inputTokens` and `outputTokens`; a step's `ai` journal and the run usage on the wire spell them
`tokensIn` and `tokensOut`. The read axis is `cacheReadTokens` on all three. The write axis is the
one real rename: `cacheCreationTokens` in the SDK and on the journal, **`cacheWriteTokens`** on the
wire - so a run's usage read back from the API names it differently from the result your own call
handed you.

## Conversation sessions

Runs that belong to the same conversation form a **session**. Set an event's `session` to a stable
conversation id and every run it starts joins that session; omit it and each run is a session of one.

```ts theme={null}
await duraton.events.send({
  name: "chat.message",
  app: "assistant",
  session: conversationId, // the OpenTelemetry gen_ai.conversation.id
  data: { text },
});

for (const s of await duraton.sessions.list({ app: "assistant" })) {
  console.log(s.session, s.runCount, "runs,", s.aiTokens ?? 0, "tokens");
}
```

<ResponseField name="session" type="string" required>
  The conversation id (or a run's own id when it started none).
</ResponseField>

<ResponseField name="runCount" type="number" required>
  Runs in the session.
</ResponseField>

<ResponseField name="statusCounts" type="Partial<Record<RunStatus, number>>" required>
  Runs per status; only statuses present appear.
</ResponseField>

<ResponseField name="aiTokens" type="number">
  Total AI tokens across the session - absent when no run made a model call.
</ResponseField>

<ResponseField name="aiCost" type="number">
  Total supplied cost - absent when no run in the session reported one.
</ResponseField>

<ResponseField name="firstStartedAt" type="string" required>
  When the earliest run in the session started.
</ResponseField>

<ResponseField name="lastStartedAt" type="string" required>
  When the most recent run started.
</ResponseField>

`list({ app, since, limit })` filters by app, bounds the window with `since`, and caps how many sessions
come back (default `100`, max `500`).

## Run counts over time

`runs.stats()` returns the current per-status counts plus p50/p95 latency; `runs.timeseries()` returns
them bucketed over time, with a latency summary (average, longest, p50, p95) per bucket. These are the
two reads behind the console's run charts.

```ts theme={null}
const stats = await duraton.runs.stats({ app: "assistant" });
console.log(stats.succeeded, "/", stats.total, "-", stats.successRate);

const series = await duraton.runs.timeseries({ since: "2026-07-01T00:00:00Z", bucket: 3600 });
for (const b of series.buckets) console.log(b.ts, b.total, b.avgMs ?? "(no terminal run)");
```

| Read                   | Returns                                                                                                                                                                     |
| ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `GET /runs/stats`      | `total`, `active`, `queued`, `running`, `succeeded`, `failed`, `successRate`, plus `p50Ms` / `p95Ms` over the filter's finished runs (both absent when none have finished). |
| `GET /runs/timeseries` | Buckets of `ts`, per-status `counts`, `total`, plus `avgMs` / `maxMs` / `p50Ms` / `p95Ms` over the bucket's terminal runs (all absent when a bucket has none).              |

Both accept `app`, `workflow`, and `since`; `timeseries` also takes `bucket` (width in seconds, default
`3600`).

### Latency percentiles

`p50Ms` is the median finished-run duration and `p95Ms` the 95th percentile, both in whole
milliseconds and over the same population as `avgMs` / `maxMs` (the scope's, or bucket's, finished
runs). They are continuous percentiles with linear interpolation between adjacent durations, so a p95
may fall between two observed values rather than on one - the standard reading of "95% of runs finished
at or below this."

Percentiles use a single continuous definition (linear interpolation between adjacent durations),
so a percentile is never an average relabelled: `p95` is the duration at or below which 95% of the
matched runs finished.

## Logs and the run timeline

[`ctx.log`](/core/logging) lines and every status transition append to one durable
per-run timeline. Read it as history with `runs.logs(id)`, or tail it live with
[`runs.watch(id)`](/core/realtime).

```ts theme={null}
for await (const frame of duraton.runs.watch(runId)) {
  if (frame.kind === "log") console.log(frame.level, frame.message);
}
```

## Reading an agent loop as it runs

An agent's durable steps are the record of what it did. `agent_event` frames are what make
that record readable as turns while the run is still going - they arrive on the same
per-run timeline, so one `runs.watch` sees them beside logs and status changes.

```ts theme={null}
for await (const frame of duraton.runs.watch(runId)) {
  if (frame.kind !== "agent_event") continue;
  if (frame.event === "turn.finished") console.log(frame.iteration, frame.cost);
  if (frame.event === "agent.finished") console.log(frame.stopReason, frame.stoppedBy);
}
```

| Event                                  | Raised when                   | Carries                                               |
| -------------------------------------- | ----------------------------- | ----------------------------------------------------- |
| `turn.started`                         | a turn begins                 | `agent`, `step`, `iteration`                          |
| `turn.finished`                        | its model call completes      | `model`, `tokensIn`, `tokensOut`, `cost`, `latencyMs` |
| `tool.started` / `tool.finished`       | a tool call runs              | `tool`, `callId`                                      |
| `guardrail.verdict`                    | a policy judges a call        | `guardrail`, `action`                                 |
| `approval.raised` / `approval.decided` | a gated call reaches a person | `decision`, `decidedBy`                               |
| `agent.finished`                       | the loop ends                 | `stopReason`, `stoppedBy`, `guardrail`                |

`cost` is present only when a cost source priced the call - Duraton holds no price list, so
an unpriced turn carries no cost rather than a zero. Likewise a turn served from the
inference cache reports a real `tokensIn: 0`, which is a different fact from carrying no
token count at all.

<Note>
  These frames carry no free text by construction: no prompt, no completion, no tool
  arguments, no tool results, and no guardrail reason (a reason quotes the value it
  refused). That is the same invariant the [AI journal](/ai/ai-steps) holds, and it is what
  makes a trace safe to hand an operator who does not own the agent's code.
</Note>

They stay on the per-run watch and never join the project-wide transition stream, which
carries run and step status only. A run makes many passes over the same loop - parking on
a human, resuming after a crash - so an event is deduped on the run, attempt, step and
event name where it is stored: a re-reported event is a no-op, never a duplicate row.

A turn that [streams](/ai/streaming#streaming-agent-turns) records two more facts on its own journal:
the `stream` flag and `ttftMs`, that turn's time to first token, so latency is readable per turn
rather than per run. Its deltas arrive as `ai_chunk` frames rather than `agent_event` ones, and those
do carry the model's completion text - the journal still holds none.

## Traces in your own backend

The spans for your **step bodies** are emitted **in your runner process** by the SDK, and exported by
whatever OpenTelemetry provider you register there (a `NodeSDK`, for example). Register none and every
step span is a no-op. Duraton runs the engine for you, so its internal telemetry is the platform's to
operate; what reaches your own backend is what your runner emits.

```ts theme={null}
import { NodeSDK } from "@opentelemetry/sdk-node";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http";

new NodeSDK({ traceExporter: new OTLPTraceExporter({ url: "https://otlp.example.com/v1/traces" }) }).start();
```

Duraton sends the run's W3C trace context with every invoke, so one run is one trace id: every pass span
and every step span of that run - across retries, and across the runners it touches - joins it.

A `step.ai` span carries the OpenTelemetry [GenAI semantic
conventions](https://github.com/open-telemetry/semantic-conventions-genai), so a vendor-neutral backend
reads it as a model call with no custom mapping. Only response-side facts the journal holds are emitted;
a field it does not hold (temperature, max tokens, the requested model before a fallback) is absent, never
guessed.

| Attribute                                  | On a `step.ai` span                                               |
| ------------------------------------------ | ----------------------------------------------------------------- |
| `gen_ai.operation.name`                    | `chat` for `generate` / `wrap` / `loop`, `embeddings` for `embed` |
| `gen_ai.provider.name`                     | The provider that served the call, when known                     |
| `gen_ai.response.model`                    | The model that actually answered (post-fallback)                  |
| `gen_ai.usage.input_tokens`                | Input token count                                                 |
| `gen_ai.usage.output_tokens`               | Output token count                                                |
| `gen_ai.usage.cache_read.input_tokens`     | Input tokens served from a provider cache, when reported          |
| `gen_ai.usage.cache_creation.input_tokens` | Input tokens written to a provider cache, when reported           |
| `gen_ai.response.finish_reasons`           | Why generation stopped, as a one-element array                    |

The span is named `{operation} {model}` (e.g. `chat claude-opus-4-8`), or the bare operation when the
model is unknown - the convention's own naming rule. Facts the convention has no attribute for live under
a `duraton.*` vendor prefix, so a standard backend ignores them:

| Attribute                                   | Meaning                                                                                             |
| ------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| `duraton.run.id` / `duraton.step.name`      | Correlate the span back to its run and step                                                         |
| `duraton.ai.kind`                           | The exact call kind (`generate` / `wrap` / `embed` / `loop`) that `gen_ai.operation.name` collapses |
| `duraton.ai.wraps`                          | The client library a `wrap` recognized (e.g. `openai`)                                              |
| `duraton.ai.batches` / `duraton.ai.dims`    | `embed` batch count and vector dimensions                                                           |
| `duraton.ai.iteration` / `duraton.ai.tools` | `loop` turn index and the tools available                                                           |
| `duraton.ai.reask`                          | The durable re-ask attempt index                                                                    |

## From an AI assistant

The same rollups are [MCP](/integrations/mcp-server) tools: `ai_spend` returns the spend rollup and `list_sessions`
returns the conversation list, both scoped to the caller's project. Paired with the read tools for runs
and steps, an assistant can answer "which workflow is burning the most tokens" against your live project.

## In the console

| View     | Shows                                                                                                                                      | Reads                    |
| -------- | ------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------ |
| **AI**   | Tiles for spend, tokens, avg latency, and cache-hit rate, then charts by hour, model, and workflow; a **Sessions** table of conversations. | `/ai/spend`, `/sessions` |
| **Runs** | Token and cost per run.                                                                                                                    | `/runs`                  |
