Direct answer
Make every state-changing action idempotent or explicitly non-retryable. Give the run and action stable identifiers, check whether the intended effect already happened, save checkpoints, and reconcile ambiguous outcomes before trying again. A better prompt cannot prevent duplicate side effects after a timeout.
A timeout does not mean the action failed
Suppose an agent sends an email, creates an invoice, updates a CRM record, or submits a payment. The external API completes the action, but the response times out before the agent receives it. A naive retry repeats the side effect.
The model sees “tool call failed.” The business sees two emails, two records, or two charges.
This is why tool reliability cannot live only in prompt instructions such as “do not perform the same action twice.” The application must enforce the rule at the boundary where state changes.
Separate intent from execution
Represent each state-changing step as an action record before calling the tool. Store the run ID, tenant, action type, target, normalized arguments, approval state, attempt count, and an idempotency key.
The execution layer should then:
- check whether the action already succeeded;
- call the external system with the same key when supported;
- record the provider response and durable result;
- reconcile uncertain outcomes before another attempt;
- refuse retries after the allowed limit.
The agent may decide what should happen. The transaction layer decides whether it is safe to happen now.
Not every operation deserves the same retry policy
Read-only retrieval can usually retry with backoff. Draft generation is cheap to repeat if only one result is accepted. Creating a customer, sending a message, changing a status, moving money, or deleting data requires stronger control.
Classify tools by consequence:
- Read: retry within latency and rate limits.
- Prepare: regenerate, but version the draft.
- Reversible write: retry only with idempotency and a verified rollback path.
- Irreversible write: require approval or a single controlled attempt followed by reconciliation.
This classification belongs in tool metadata and application logic, not tribal knowledge.
Checkpoints make long workflows recoverable
An agent with ten steps should not restart at step one because step eight failed. Persist completed steps, outputs, approvals, and external references. Resume from a safe checkpoint after validating that earlier state still holds.
Zenveus production automation work repeatedly showed the value of this surrounding system. One AI sales pipeline added input validation, retries, fallbacks, execution logs, and immediate failure alerts. It grew from 30–50 to more than 80 leads per day, while invalid records reaching AI enrichment fell to zero and monthly pipeline stalls stopped requiring manual recovery.
The result did not come from asking the model to “be more reliable.” It came from making failure cheap and visible.
Put a ceiling on recovery
Retries need a maximum attempt count, elapsed-time limit, cost budget, and escalation path. Repeating the same failing call is not resilience; it is a loop.
After the threshold, stop. Preserve the checkpoint, explain which action is uncertain, and route the case to a person or reconciliation job. That operator should see the intended action, attempts, provider responses, current external state, and safe next options.
The acceptance test
For every state-changing tool, simulate these cases: success with a lost response, timeout before success, duplicate delivery, stale credentials, rate limiting, and partial downstream completion.
Then ask one question: can the workflow run again without duplicating business state?
If the answer depends on the model remembering what it did, the agent is not safe to retry.
Related Zenveus service: Agentic AI Development