Skip to main content
When a step throws, Duraton retries that step - not the whole workflow. Steps that already succeeded keep their recorded results and are not re-executed. A workflow does not retry unless it opts in with retry.

What happens on failure

  1. A step throws.
  2. Duraton waits the backoff delay, then runs that one step again.
  3. This repeats until the step succeeds or reaches maxAttempts.
  4. If the step fails on its last attempt, the run fails.

Backoff

By default Duraton waits 1000 ms between step attempts, the same before every attempt. Set backoff to shape how the delay grows with each attempt, initialDelayMs to change the base delay, and maxDelayMs to cap it:
To override the delay for a single attempt at runtime (for example honoring an upstream 429’s Retry-After), throw RetryAfterError with the delay you want. Two other delays exist and are not this one:

Per-step retry

A workflow’s retry sets the default for every step. Pass a retry option to a single ctx.step.run to override it for that step alone - useful for a polling step that needs many short attempts without forcing that budget onto the rest of the run:
The step’s own policy governs its attempt budget and backoff; every other step keeps the workflow default. The per-step retry option and configurable backoff are available in the TypeScript SDK.

Controlling retries from a step

Two error types let a step override the default retry behavior. Import them from the Duraton SDK.

NonRetriableError

Fail the run immediately with NonRetriableError, skipping any remaining attempts. Use it for failures retrying cannot fix, such as a validation error or missing configuration.

RetryAfterError

Retry after a delay you choose with RetryAfterError instead of the policy’s configured backoff, for example honoring an upstream 429’s Retry-After.
RetryAfterError does not grant extra attempts - it only changes when the next attempt runs. Once the step reaches maxAttempts the run fails as usual.

Transport errors

If Duraton cannot reach your runner at all (the process is down, or it returns a server error), the invoke is retried on its own budget: 3 attempts, 1000 ms apart, independent of maxAttempts. A network blip does not consume a step’s retry budget. After the third failed invoke, the run fails.

onFailure

Declare an onFailure handler to run compensation or notification logic when a run fails. It fires once the run has exhausted its retries (or hit a non-retriable error) and been marked failed, and it receives the original event plus ctx.error.
onFailure is itself a durable execution: its steps are recorded and retried like any handler. It cannot un-fail the run - the failed run stays failed. Use it to compensate (release a hold, reverse a write) or to notify, not to retry the work.

Failed runs

A run is marked failed only once it is terminal and out of retries - a run still retrying stays non-terminal - so the set of failed runs is exactly the set of runs that permanently failed, each carrying its terminal error.
In the console, failed runs appear in the Runs view with their failure reason inline; the Failed stat tile filters the list to them in one click.
See the error handling example running end to end in Examples.