Skip to main content

When something fails

A step can fail. What happens next is your choice, and there are three mechanisms.

1. Retries​

Steps that support repetition can try again automatically.

SettingWhat it does
Number of attemptsFrom 1 to 10. The default is one — no repetition
Wait between attemptsIt grows with each attempt, with jitter, so they do not all repeat at the same instant

Not every failure is retried: only the transient ones — the engine unavailable, a timeout, an expired credential. Configuration failures are not retried, because trying again would give the same result.

Deterministic steps — Prepare data, Condition — offer no repetition for the same reason. And Generate with a model does not retry on its own because an ambiguous answer may already have been billed.

While waiting to retry, the run releases resources and comes back on its own.

2. The error path​

Steps that support it can have an if it fails exit, connected to another part of the flow.

With an error path:

  • the step stays marked as failed — it genuinely failed;
  • the run continues along the diversion;
  • the diversion receives the failure's reason: the code, the message, whether it was transient, and which attempt it happened on.

The error path is only taken after retries are exhausted. And it is a step like any other: if the handling fails, the run fails.

Without an error path, the run ends as a failure and the following steps are cancelled.

3. Declared degradation​

When an error is handled by a diversion, the agent's final result says so: it carries the list of steps that broke and the reason.

Handled is not "did not happen". A flow that handles errors and delivers a clean success hides exactly what someone would need to know in order to trust the number they received.

What the error message contains​

Every message goes through a cleanup before being stored: credentials, tokens, and database addresses are stripped. What remains is enough to diagnose, and nothing that should not circulate.

Timeouts​

There are four deadlines, and they add up over the run:

DeadlineWhere
The run'sDeclared in the agent. Once exceeded, the run ends as expired
The step'sOptional, per step. It has to fit inside the run's deadline
The approval'sOptional, on the approval step
The platform's internalRecovery of interrupted runs

Cancellation​

A run can be cancelled at any moment. Cancellation is cooperative: it is checked between steps, never in the middle of a call in flight — interrupting midway would leave the operation on the other side with no known outcome.

Child runs are cancelled along with it.

Useful patterns​

GoalHow to build it
Continue even if a system is downError path → Prepare data with a default value → continue
Notify someone when it failsError path → AI task or a notification automation
Require every source of a joinCombine results with "fail" when a source is missing
Tolerate a bad item in a listFor each item with "continue and record"

Common errors​

SymptomCauseWhat to do
The run fails without retryingThe default is one attemptConfigure the retries on the step
The step retries and fails the same wayThe failure is not transientFix the configuration; retrying will not help
The result says success, but something brokeAn error was handled by a diversionLook at the steps with handled failures in the result
The run ends as expiredThe total deadline ran outRaise the timeout or reduce the work

Next steps​