Ground Truth.
AI, checked against the source.

Learn · Intermediate

Idempotency and crash recovery: making agent retries safe

Idempotency means repeating the same logical operation has the same intended effect as performing it once. Durable execution records progress so work can recover after a crash, but safe recovery also requires deciding whether an external action already happened. Together, these ideas make long-running AI agents more reliable without pretending that retries automatically make every side effect occur exactly once.

The uncertainty that a crash creates

Suppose an agent submits a payment, the payment service accepts it, and the agent process stops before recording the reply. After restart, its local record says the operation has no confirmed result. That record cannot distinguish a payment that failed to leave the machine from one that succeeded but lost its acknowledgment.

Repeating the call might finish the intended task, or charge twice. Refusing to repeat it might avoid duplication, or leave the customer unpaid. This uncertainty comes from communication and persistence boundaries. It does not disappear when the model becomes better at reasoning.

A useful analogy is ordering a meal over a noisy phone line. If the restaurant receives the order and the confirmation is cut off, calling again without an order number can create two meals. An order number lets the restaurant recognize the same request. Crucially, the restaurant must keep that record; writing the number only in your own notebook is insufficient.

Idempotency is about effects

Setting a document’s status to approved can be idempotent: doing it twice leaves the same status. Appending an approval event twice is a different operation and can create duplicate records. Similarly, setting a balance to a specified value differs from adding an amount to the balance. Whether an operation is idempotent depends on the intended effect, not on whether its function name sounds harmless.

For operations such as payments, a service can implement an idempotency key. The caller supplies a stable identifier for one logical action. The service stores that identifier with the result and recognizes later attempts as repetitions. Reusing a key with different payment details should be rejected rather than silently interpreted as a new operation.

The record and the effect must be coordinated. If a service performs the payment and only later writes its duplicate-detection record, a crash between those steps recreates the problem. A transactional design can make recording the key and committing the effect one indivisible change within that service’s own boundary. Remote systems introduce additional boundaries that must be handled explicitly.

What a durable runtime contributes

A durable runtime records steps and checkpoints. After a process stops, another process reads the record, reuses completed results where appropriate, and resumes unfinished work. That is different from agent memory, which stores information useful to future reasoning. A memory saying that a payment was discussed is not a transaction record establishing whether it committed.

Earendil’s Pi Durable design illustrates the distinction. It records model and tool tasks, retries interrupted model requests, and repeats tool calls only when marked safe. Its request identifier deduplicates submissions to the runtime. Its payment example separately supplies an external idempotency key. Those are two layers with different responsibilities.

A runtime can therefore offer reliable recovery of its own work graph without guaranteeing that every connected service has identical semantics. Application builders must specify safe replay behavior. They may need to query an external operation’s status, use a stable key, stop for review, or perform a compensating operation when an effect cannot be undone directly.

Why exactly once needs a boundary

The phrase exactly once is meaningful only when it specifies what is counted and where. A queue may deliver a message more than once while the receiving service applies the intended database change once. A runtime may accept a submission once while a remote tool’s result remains uncertain. Those statements can all be true simultaneously.

Jerome Saltzer, David Reed, and David Clark’s End-to-End Arguments in System Design explains a broader principle: some correctness properties require cooperation at the endpoints, even when lower layers provide helpful mechanisms. An agent framework cannot establish a payment service’s duplicate-handling behavior merely by promising reliable local execution.

This is especially relevant to tool-using models. The model proposes actions; the surrounding software carries them out. Retry logic belongs in that software and in the receiving systems. Asking the model to remember not to repeat a payment is weaker than having the service enforce a durable rule.

Reliability and authority remain separate

An operation can be perfectly idempotent and still unauthorized. Repeatedly setting someone else’s account status can have one stable effect while violating access rules. Scoped credentials and sandboxing address that separate question. Durable execution makes work recoverable; it does not decide which work is allowed.

The practical design exercise is to classify each tool by its effect. Identify read-only calls, repeatable state-setting calls, and actions needing external duplicate protection. Then test interruption before submission, after external acceptance, and before local recording. A successful recovery must preserve the intended effect and an intelligible record, rather than merely produce a reassuring final answer.

Key papers
End-to-End Arguments in System Design — Saltzer, Reed, and Clark (1984)

Key questions

Why can a tool call be uncertain after a crash?

The external service may complete the action before the agent records its response. Local storage alone cannot tell whether the action failed or only its acknowledgment was lost.

Where must an idempotency key be enforced?

The system that commits the external effect must durably recognize the key. A key stored only in the caller’s memory cannot prevent a duplicate payment at the receiving service.

Does durable execution make every tool safe to retry?

Durable execution preserves progress but does not automatically protect external effects from duplication. Each tool needs an appropriate replay rule, idempotency mechanism, or recovery procedure.
Cite this

APA

Ground Truth. (2026, October 2). Idempotency and crash recovery: making agent retries safe. Ground Truth. https://groundtruth.day/learn/idempotency-and-agent-crash-recovery.html

BibTeX

@misc{groundtruth:idempotency-and-agent-crash-recovery,
  title  = {Idempotency and crash recovery: making agent retries safe},
  author = {{Ground Truth}},
  year   = {2026},
  month  = {oct},
  url    = {https://groundtruth.day/learn/idempotency-and-agent-crash-recovery.html}
}

Topics: agents · reliability · distributed-systems · durable-execution · fundamentals