APIs, integration & security — in depth
AgenticLong read

Idempotency Keys in Agent-Driven Transactional Email

Stable keys prevent agents from flooding inboxes with duplicate emails during retries.

Senior Correspondent, Agentic Systems · · 11 min read
Cover illustration for “Idempotency Keys in Agent-Driven Transactional Email”
Agentic · October 9, 2026 · 11 min read · 2,580 words

Agent-driven email fails differently than human-triggered email because of retry behavior. When an HTTP call times out, or when an orchestration framework restarts a failed step, the agent re-invokes the same tool with the same arguments. A password-reset workflow that crashes after the email has already gone out, but before the workflow records that fact, will restart and send the password-reset email again. Without a deduplication mechanism sitting between the agent and the outbound send, nothing stops a third attempt, or a tenth.

This retry instinct is not a defect in the agent. Stateless retry is the correct behavior for model inference: an agent that gives up after one failed call is a worse agent than one that tries again. The danger appears when that same retry logic touches a side effect that carries consequences outside the model, a card charge, a ticket creation, an email send. A model can safely re-run a calculation. A send cannot safely be re-run without some guarantee that the second attempt will be recognized as the first one and replayed as the first one.

At scale, this produces what looks like an email storm: a recipient who should have received one message receives hundreds of identical copies before anyone notices. Because an agent works without a human reviewer at each step, the failure can run for minutes or hours unsupervised, compounding as it goes. Duplicate sends inflate complaint rates and erode recipient trust well before a human catches the pattern. The only viable strategy is prevention, built before the first message leaves the queue, not detection after the fact.

What an idempotency key does at the API layer

Prevention, in this context, has a name: the idempotency key. The agent submits a key alongside each send request. The API stores that key together with the outcome of the send. If a second request arrives carrying the same key, the API returns the original response. The retry still happens, as retries should, but the side effect does not repeat.

The contract has a clear shape on the API's side. On first receipt of a key, the API stores it, executes the send, and stores the result against that key. On a subsequent receipt of the same key, within whatever window the API has defined, the API returns the stored result. Outside that window, the API treats the request as new.

The key itself is a string generated by the client, not the API. The API uses it purely as a lookup, which puts weight on a detail easy to overlook: the API does not trust that string unconditionally. The IETF's draft specification for idempotency key headers states that servers should "always validate the key as per its published specification before processing," and that if a key is reused with a different payload, the server "SHOULD reply with a HTTP 422 status code." A well-built API checks format and rejects mismatches. A key that was poorly constructed in the first place defeats this validation. If the client generates a different key on every retry of the same logical send, the API has no way to recognize that the two requests are, in fact, the same request. Correctness here depends entirely on the client.

That dependency raises an immediate question the rest of this piece works through: how long should the API remember a key before treating a repeat as new. A twenty-four-hour window suits an agent that retries within seconds or minutes of a failure. It does not suit an agent that runs on a daily or weekly schedule and only discovers a gap in its own records days later. That mismatch, between the window the API offers and the retry interval the agent actually needs, becomes one of the sharper edges in key construction, and it is where bad key design starts to show its cracks.

How a badly chosen key defeats deduplication

Idempotency keys fail in a small number of recognizable ways, each with its own mechanism and its own fix.

The most common mistake is generating the key from the current time at the moment of the call. A timestamp feels like a reasonable candidate for a unique identifier, and that instinct is exactly backward: the key needs to be stable across retries of the same logical send, not unique across all time. A timestamp changes on every invocation, including the retry, so the API sees what looks like a brand-new send each time the agent tries again. The deduplication check never has a chance to fire.

A second version of the same mistake occurs when UUIDs are generated inside the tool call itself, a uuid4() invoked fresh at call time. The logic is identical to the timestamp case: a new random value on every invocation defeats the purpose of a stable key just as thoroughly as a new timestamp does.

A subtler variant of this problem appears when the language model itself is made responsible for key generation. Developers sometimes assume an instruction to the model, telling it to "reuse the same key for retries," is enough, but it is not. Even given identical inputs, a probabilistic model offers no guarantee that it reproduces the same UUID or the same string on a second pass, and even small formatting differences, a missing dash, a different case, an extra space, are enough for the API to treat what is logically one send as two unrelated ones.

A fourth failure mode has nothing to do with how the key is built and everything to do with the window discussed in the previous section. If the deduplication window has genuinely expired by the time a legitimate retry arrives, the API behaves exactly as its contract specifies: it treats the request as new, because by the terms of that contract, it is new. The resulting duplicate is indistinguishable from a legitimate fresh send, which makes this failure mode silent in a way the others are not. Nothing in the API's logs flags an error, because nothing has gone wrong at the API layer. The problem sits upstream, in a mismatch between the agent's retry interval and the window the API was configured to remember.

A fifth failure mode involves fan-out: a single workflow step that triggers more than one side effect, a charge and a confirmation email, for instance. A key that covers only one of those effects leaves the other unprotected, so a retry of the step duplicates whichever action the key didn't scope. The fix is either a distinct key per side effect, tracked through a clear state machine, or a saga pattern with compensating actions for partial failures. A single key wrapped around a multi-effect step will not hold.

The correct key construction strategy: deterministic inputs only

Every one of those failure modes traces back to the same root cause: the key was not a deterministic function of the send's identity. The fix is to make it one. A correctly built idempotency key is computed from stable inputs that describe the logical send itself, such that the same logical send always produces the same key, and any genuinely different send always produces a different one.

Four components, combined, do this reliably. The workflow run ID, assigned by the orchestration layer to this specific execution, stays stable across every retry of that run and changes only when a genuinely new run begins. The step index or step name identifies which step within the run triggered the send, so that two separate email steps inside the same run produce two separate keys. The recipient address scopes the key to a specific person, so a fan-out send to multiple recipients in a single step generates one distinct key per recipient. And a hash of the message content, subject and body combined, catches the case where the same step, for the same recipient, genuinely needs to send different content, a legitimate second message that should go through.

The construction looks roughly like this in a language-agnostic sketch:

key = run_id + ":" + step_name + ":" + recipient + ":" + hash(subject + body)

The delimiter convention barely matters; what matters is that every component is deterministic and that the hash is computed over the actual content being sent, not over something that varies independently of it.

The construction only holds if it lives in code, not in the model's behavior. The orchestration layer, or the wrapper around the email tool, generates the key before the API is ever invoked. The model receives that key as a parameter to pass through with its tool call. It does not construct the key itself. This single design decision removes the formatting-variation failure mode entirely, because the function generating the key is deterministic code, not a probabilistic process whose output can drift from one run to the next.

The same discipline applies on the inbound side, in reverse. Email infrastructure retries delivery, so a webhook handler that launches a fresh agent loop on every delivery attempt of the same inbound message will duplicate replies and corrupt the state of the conversation thread. The fix mirrors the outbound case: store the message_id of each processed inbound email in a durable store, and check that store before starting a new agent loop.

How the email API should deduplicate on receipt

A deterministic key only delivers its protection if the API on the receiving end implements deduplication correctly. A key that is ignored, stored inconsistently, or allowed to expire silently produces the same outcome as sending no key at all: a duplicate message in the recipient's inbox.

A correct implementation rests on a few specific guarantees. Storage of the key and its result has to be durable and atomic, written in the same transaction as the send itself, so that a crash between "key stored" and "result stored" cannot leave a gap that produces a phantom duplicate on the next retry. Key comparison has to be normalized, case folding and whitespace stripping applied before any match is checked, so that trivial formatting differences in the client's key don't slip past deduplication that should have caught them. The deduplication window itself needs to be documented explicitly and sized to the retry behavior of the systems calling it: a twenty-four-hour window works for an agent retrying within the hour, and fails for a scheduled agent that only discovers a missed send days later. Window policy belongs to the architecture of the system, not to an implementation footnote. And when a duplicate key does arrive, the API should return the original result in full, the original message ID, the original status code replayed exactly as first issued (a stored 201 stays a 201), optionally marked with a distinguishing header such as Idempotent-Replay: true. Downgrading that response to a generic 200 erases information the caller needs.

The implication for anyone building or buying an email API meant for autonomous agents is that idempotency has to be a first-class feature on every send endpoint, not an optional header available on a handful of endpoints the designer happened to anticipate. Agents retry against whichever endpoint they're calling, not just the ones that were built with retries in mind.

Deduplication and pre-send validation operate at different points in the lifecycle and serve different purposes. A pre-send linting layer, the kind of guardrail system found in AgentSend's architecture, intercepts malformed or non-compliant sends at the API boundary and returns a machine-readable error with a specific fix, before the deduplication check is ever reached. A retry that would have produced a broken duplicate gets caught at that earlier stage instead. The two controls are additive: one keeps bad sends from reaching the queue at all, the other keeps good sends from being repeated.

Even with both controls in place, idempotency keys protect individual tool calls, not the workflow surrounding them. Two edge cases sit outside what a key, however well constructed, can cover. The first is an orchestration restart: when a framework restarts a failed run and assigns it a new run ID, any key built from that run ID differs from the key used in the prior, incomplete run, and the send that actually completed before the crash goes out again. The defense here is checkpointing at the orchestration layer, recording which steps have finished and what they returned, in durable storage, before the workflow advances. On restart, completed steps are replayed from that checkpoint. Idempotency keys guard against retries within a single run. Checkpointing guards against re-execution across runs. Production systems need both.

The second edge case is window expiry on a schedule. An agent that runs daily, finds no record of having sent a particular message, and generates a key identical to the one it used the last time it ran, is behaving exactly as it should. If the API's deduplication window has already closed by the time that key arrives, the API has no way to recognize it as a repeat. The defense is an application-layer dedup table, a durable record of sent message IDs maintained independently of whatever window the API enforces, checked before the send tool is ever invoked. Circuit breakers that halt sends once bounce or complaint rates cross a threshold, and effect budgets that cap how many actions an agent can take without human approval, round out the set of orchestration-level controls that idempotency alone was never meant to provide.

The three-layer deduplication architecture for production agent email

Diagram: Three-Layer Deduplication Architecture. Visualizes: Show three stacked defensive layers that together prevent duplicate agent email sends, with each layer's scope and mechanism labeled.

Put together, the defenses described above form three layers, each covering ground the others cannot.

The first layer is the idempotency key on the API call itself: a deterministic key built from the workflow run ID, step name, recipient, and content hash, generated in code rather than by the model, submitted with every send, and deduplicated atomically by the API inside a documented window.

The second layer is an application-level dedup table, a durable store, whether a database or a cache with persistence, recording the message ID of every send the application has confirmed. The application checks this table before it even calls the send tool, so a send already known to have completed never reaches the API a second time. This layer catches window expiry, and it is the right place to track processed inbound message_id values for webhook handling.

The third layer is workflow checkpointing at the orchestration level: completed steps and their results recorded durably, so that a restart replays from the checkpoint. This keeps the run ID, and therefore the idempotency key built from it, stable across a restart rather than shifting to a new value with a new run assignment.

No single layer covers every failure mode on its own. A breach at one layer is caught by the next, and a production system that leans on only one of the three will eventually meet the scenario that layer was never designed to handle. An email API that treats idempotency as a built-in feature, with a documented window, machine-readable duplicate responses, and per-agent scoped credentials for auditing, takes the implementation burden off the application layer and lets the dedup table and the checkpointing logic focus on the narrower set of cases the API's own window genuinely cannot reach. That division of labor, guardrails at the point of send, idempotency enforcement built into the API rather than added on afterward, reflects a simple premise: agents are now first-class senders of email, and the real risk they introduce is not a visible crash but a silent, compounding duplicate that nobody is watching for.

Filed underAgentic

More in Agentic