Here is a composite example of something that happens often. A support agent is resolving a billing ticket. It calls create_refund against the ERP, and the call times out after 30 seconds. The framework retries, and the second call succeeds. The customer gets two refunds, because the first call never failed. The ERP committed the refund and the response just didn't arrive in time. Meanwhile a second agent in the same fleet, one that triages tickets, reads the ticket, adds an escalated tag, and writes the whole tag array back. The billing agent wrote refund-issued to the same array four seconds earlier, and that tag is now gone.
No error appeared in either case. Each tool call returned 200, the traces look clean, and the system of record is wrong.
These are two old distributed-systems bugs: duplicate writes and lost updates. Payment and database engineers solved both years ago. Agent fleets bring them back because most agent frameworks treat a tool call as a function call, while a write to a CRM, ERP, or helpdesk is a network request against shared state. Our position is that a write tool should not get production permissions until it carries an idempotency key and a version precondition. If it retries without them, it is corrupting data without telling anyone.
Two failure modes, two different guards
People mix these two bugs up, and the mix-up leads to the wrong fix. Idempotency stops one operation from being applied twice. Optimistic locking stops two different operations from silently overwriting each other. Each guard covers its own bug, so you need both.
| Failure | Trigger | What breaks | Guard |
|---|---|---|---|
| Duplicate write | Retry after a timeout, a model re-issuing a tool call, a replayed workflow step | Two refunds, two contacts, two tickets, two emails | Idempotency key |
| Lost update | Two agents (or an agent and a human) doing read-modify-write on the same record | Tags, notes, status, or field values silently reverted | Version precondition (If-Match / ETag) |
| Interleaved multi-step edit | An agent updates the order, then the invoice, while another agent edits the same pair | Records consistent individually, inconsistent together | Lease lock or a single owner per record |
Agents trigger all three more often than traditional integrations do. A deterministic sync job calls the API on a fixed schedule with fixed payloads. An agent loop can call the same tool twice because the model didn't trust the first response, because the harness resumed from a checkpoint, or because a retry policy fired at the HTTP layer while the model was also retrying at the reasoning layer. Anthropic's Building Effective Agents recommends investing in tool design with the same care you would put into a human-facing interface. For write tools, that care has to include what happens when the tool runs twice.
If you've already mapped which system owns each field, you know where these guards belong. They go on every write into that system.
Idempotency keys: derive them from intent, not from arguments
Stripe's idempotent requests documentation is the reference design most teams copy. The client sends an Idempotency-Key header. Stripe stores the status code and body of the first request made with that key, and any later request with the same key gets the saved result back without the operation running again. If the same key arrives with different parameters, Stripe returns an error rather than guessing. Keys expire after a retention window, so the guard covers retries and not the record's entire history.
The pattern carries over to agents with one important change: who generates the key and from what.
Two shortcuts fail in agent systems:
- A random UUID per tool call. Each retry creates a new UUID, so the server sees two different operations. The key gives you nothing.
- A hash of the full tool arguments. This looks deterministic, but when a model re-issues a call it often rewords free-text fields. "Refund for duplicate charge" becomes "Refunding duplicate charge on order 456." The hash changes and the duplicate gets through.
Build the key from fields that identify the business intent, and keep the model out of it entirely. The harness has everything it needs: the run ID, the step or plan-node ID, the tool name, the target record ID, and the operation's natural key (order ID plus refund reason code, contact email plus source, and so on).
import hashlib
def idempotency_key(run_id: str, step_id: str, tool: str,
target_id: str, intent_fields: dict) -> str:
# Only stable, structured fields. Never free text the model wrote.
canonical = "|".join([run_id, step_id, tool, target_id] +
[f"{k}={intent_fields[k]}" for k in sorted(intent_fields)])
return hashlib.sha256(canonical.encode()).hexdigest()[:64]
key = idempotency_key(
run_id="run_8f2c", step_id="refund-1", tool="create_refund",
target_id="order_456", intent_fields={"reason": "duplicate_charge", "amount_cents": 4900},
)
Deciding the key's scope is a business decision. Leave run_id out and the key covers every run, so a second run that tries to refund the same order for the same reason gets the original result back. For refunds, that's probably what you want. For "send a follow-up email," scoping to the run may be correct, because a legitimate second follow-up next week shouldn't be blocked. Write down the scope for each tool when you write the allowed-actions document.
Aleksei Aleinikov's walkthrough of idempotency keys in practice covers the server-side mechanics, including storing the key and response together and handling a second request that arrives while the first is still in flight. Your tool layer needs those same mechanics when the downstream API doesn't provide them, which brings us to the harder case.
When the system of record has no idempotency support
Stripe is unusual. Many helpdesk, CRM, and ERP APIs accept no idempotency header at all. You have three options, from most to least preferred.
Upsert by external ID. Many CRMs let you write a record keyed on an identifier you control. Writing the same external ID twice updates the existing record instead of creating a new one. Airbyte's guide to idempotency in data pipelines treats upserts and merge semantics as the main way to make repeated loads safe, and agent writes benefit from the same approach. If your agent creates contacts, leads, or cases, give each one an external ID derived from the intent and upsert on it.
An idempotency ledger in your tool layer. Before calling the downstream API, insert the key into a table with status pending. After the call, store the result and mark it done. A retry checks the ledger first. If it finds done, return the stored result. If it finds pending, the earlier call may or may not have committed, so query the system of record to reconcile (search for the refund or contact you meant to create) before you write again. This is the one case where you can't avoid the ambiguity, and the reconcile query is what resolves it.
Search-before-create. This is the weakest option, because two agents can both search, both find nothing, and both create. Use it only when the volume per record is low and a human reviews duplicates.
The Idempotency for Agents piece from Data Science Collective argues that this logic belongs in the tool wrapper, where the model can't skip it. We agree. A system prompt that says "don't create duplicates" is a request. A ledger check is a control.
Optimistic locking: make stale writes fail loudly
Lost updates come from the read-modify-write cycle. Agent A reads the ticket at version 7. Agent B reads it at version 7. A writes, and the ticket is now at version 8. B writes a payload computed from version 7 and overwrites A's change. Neither agent did anything wrong as far as its own logic goes.
HTTP has had a fix for this since the 1990s. RFC 9110's conditional requests let a client send If-Match: "<etag>" (or If-Unmodified-Since) with a write. If the resource has changed since the client read it, the server refuses with 412 Precondition Failed instead of applying the stale write. Salesforce's REST API documents conditional request headers for record operations, so this works on at least one of the large systems of record agents write to most.
For agents, the rule is: the version travels with the data. Whenever a read tool returns a record, it also returns the ETag or LastModifiedDate. The write tool requires it as a parameter and refuses to run without it. The model never invents a version. The harness passes along the one from the most recent read.
Here's what a write tool should return on a 412:
{
"status": "conflict",
"message": "Record changed since you read it.",
"current": { "tags": ["billing", "escalated"], "status": "open", "version": "\"v9\"" },
"your_change": { "tags_add": ["refund-issued"] }
}
Don't retry blindly on a conflict. Return the fresh state and let the agent decide again, because the other agent's change may have made this write unnecessary (someone else already closed the ticket) or wrong (the refund was escalated to a human). Allow two conflict cycles at most and then route the case to a human queue. If two agents keep fighting over the same record, the fix is to give that record a single owner, and more retries won't get you there.
Two design habits reduce how often conflicts happen in the first place:
- Send deltas, not whole records. A PATCH that adds one tag touches less than a PUT that replaces the whole tag array. Where the API supports add/remove operations on collections, use them.
- Keep the read-to-write window short. An agent that reads a record, spends 90 seconds reasoning, then writes has a 90-second conflict window. Re-read right before writing, or split the step so the slow reasoning happens before the read.
Leases for the edits that cannot interleave
Some operations span several records that have to move together: update the order, regenerate the invoice, post a note on the account. Per-record optimistic locking keeps each write correct but doesn't prevent another agent from slipping in between your writes.
For these cases, take a lease: a lock with an expiry, held in Redis, Postgres, or your workflow engine, keyed on the business entity (lock:order_456). Agents are bad lock holders. They stall on slow model calls, get killed by budget limits, or wait on a human approval for an hour. So:
| Lease rule | Why |
|---|---|
| Short expiry (seconds to a few minutes), renewed by heartbeat | A crashed or killed agent releases the record automatically |
| Fencing token (an incrementing number) passed with every write | A stalled agent whose lease expired can't write over the new holder's work |
| Never hold a lease across a human approval | Approvals take hours. Release the lease, then re-acquire it and re-validate after approval |
| Lock the business entity, not the agent | Two steps of the same agent can conflict too |
If a workflow needs long leases regularly, treat that as a sign the work should be split into durable, checkpointed steps with a single owner per entity. For most fleets, routing every write for a given record through one queue partition removes the need for locks entirely. It is the simplest concurrency control there is, and teams skip it because a fleet of parallel agents looks more impressive in a demo.
The write-tool admission checklist
Here is the gate we recommend before any write tool gets production credentials. It complements scoped agent identity: identity decides whether the agent may write, and these guards decide whether the write is safe to repeat and safe to race.
| Check | Pass condition |
|---|---|
| Idempotency key | Derived by the harness from run, step, target, and intent fields, never generated by the model |
| Key scope | Documented per tool (per run, per entity, or global) with a retention window |
| Downstream support | Native idempotency header, upsert by external ID, or a ledger in the tool layer |
| Ambiguous outcome handling | A timeout triggers a reconcile query before any retry |
| Version precondition | Every update sends If-Match / If-Unmodified-Since, or the API's equivalent version field |
| Version source | Taken from the most recent read tool response, passed through by the harness |
| Conflict response | 412 returns current state to the agent, with at most 2 conflict cycles before escalation |
| Partial updates | PATCH or collection add/remove instead of full-record PUT where the API allows |
| Multi-record edits | Lease with expiry and fencing token, or single-owner routing |
| Retry layering | Only one layer (HTTP client or agent loop) retries a given write |
The last row catches more teams than you'd expect. An HTTP client with three automatic retries, inside an agent loop that also retries failed tools, inside a workflow engine that replays failed steps, can send the same write up to 27 times (3 × 3 × 3). Pick one layer to own retries for writes and turn the others off for those tools.
Writeback patterns differ by system (a helpdesk note is append-only, while an ERP journal entry may be immutable once posted), but every checklist row applies to each one.
Prove it with the timeout that happens after commit
Unit tests won't find these bugs, because the bugs only appear under bad timing. The test that matters is a commit-then-timeout injection: let the downstream call succeed, then drop or delay the response past the client timeout. Run the agent and count the records in the system of record. One record means you pass. Two means you've found the bug in staging.
Add a race test too. Start two agents against the same record with conflicting goals and confirm one of them gets a 412 and handles it. If both report success and the record shows only one change, you have a lost update. Our fault injection guide covers how to wire these failures into a test harness. Keep both scenarios in the regression suite permanently, because a framework upgrade can quietly add a retry layer.
Watch for these in production:
- Duplicate-detection queries on created entities (same external ID or same natural key within a short window)
- 412 rates by tool and record type, since a rising rate means two agents are fighting over the same records
- Ledger entries stuck in
pending, which point to unresolved ambiguous outcomes - Lease expirations, which point to stalled agents
Getting your write tools fleet-ready
If you run one agent today and plan to run several against the same CRM or helpdesk, add these guards now, while the number of write tools is small enough to audit in an afternoon. Adding them after a duplicate-refund incident means auditing every record those tools have touched.
OpenNash does this as part of the design phase of an engagement: we list every write tool, decide each one's key scope and conflict policy, build the ledger or upsert path where the downstream API lacks one, and add the commit-then-timeout and race tests to CI before deployment. Teams that keep everything inside a single vendor platform should first check whether that platform already enforces idempotency and versioning on its native actions. Many do for their own objects and don't for custom HTTP actions. The OpenNash work fits best when agents write across several systems of record and the client wants to own the tool layer afterward.
Start by listing every tool your agents can call that changes state somewhere, and mark which ones send an idempotency key and which send a version precondition. If any row has neither, book a call and we'll go through the gaps and a remediation plan with you.