Idempotency keys: revisit your retries before they charge twice

A customer taps “Pay”, the spinner keeps turning, the mobile network drops and the app tries again. Between that phone and your database, an SDK, a proxy or a queue consumer may also try again. Each retry is reasonable on its own. Together they can create two orders, two refunds or two charges for one intent. The retry is not the bug. The bug is a server that cannot tell a retry from a new request.

This post covers what “idempotent” means in HTTP, the IETF Idempotency-Key draft and its status, how Stripe, Adyen and Amazon EC2 document their keys, and a design for your own API in PostgreSQL: scope, fingerprint, stored response, in-progress duplicates, expiry and atomicity. It also covers queue consumers, backoff with jitter, common mistakes, a checklist and a sandbox drill. Facts come from the RFCs, the IETF datatracker, each vendor’s own docs, the AWS Builders’ Library and Architecture Blog, and the PostgreSQL and Kafka docs, all checked on 6 October 2026.

The short version

  • Duplicates are normal. Any system with timeouts and retries delivers some requests more than once.
  • In HTTP, GET, HEAD, OPTIONS, TRACE, PUT and DELETE are idempotent by definition (RFC 9110). POST is not, and PATCH is “neither safe nor idempotent” (RFC 5789).
  • An idempotency key is an ID the client picks once per intent. The server stores it with a fingerprint of the request and the final response, and replays that response when the same request comes again.
  • The IETF draft for the Idempotency-Key header reached version 07 on 15 October 2025, and the datatracker now lists it as expired. The pattern is widely used, but it is not an RFC.
  • Providers differ: Stripe allows keys up to 255 characters and may prune them after 24 hours, Adyen keeps keys for 7 to 14 days, and EC2 client tokens are up to 64 ASCII characters. Publish your own rules.
  • The key record and the business write must commit together. In PostgreSQL, a primary key on (account, key) plus INSERT ... ON CONFLICT DO NOTHING in the same transaction does most of the work.
  • Kafka’s idempotent producer, on by default since 3.0 (with a caveat), prevents duplicate writes to the log. It does not make your consumer’s side effects idempotent.

Why duplicates are normal, not rare

One request can be resent by the user (double-click, refresh), the mobile app, the client SDK, a gateway or service mesh, a queue that redelivers after a consumer crash, or a batch job that is re-run. Stripe’s docs say most of its client libraries can generate keys and retry automatically once configured, and AWS says some of its SDKs retry automatically on network errors, server faults and rate limiting.

The hard case is a timeout after the server has already done the work. To the client, “never arrived” and “done, but the reply was lost” look the same. The AWS Builders’ Library article on idempotent APIs uses this kind of example: a tool asks EC2 for exactly one instance, gets no response, and must choose between retrying (maybe two instances) and not retrying (maybe none).

Retries also multiply. As simple arithmetic, if the app, the SDK and a gateway each make up to three attempts, one tap can become up to 27 backend calls.

What idempotent means in HTTP

RFC 9110, section 9.2.2, defines a method as idempotent when the intended effect on the server of many identical requests is the same as the effect of one. Of the methods it defines, PUT, DELETE and the safe methods are idempotent, and section 9.2.1 lists the safe ones as GET, HEAD, OPTIONS and TRACE.

The definition is about the intended effect, not the response: the response to a repeat may differ, and the server may still log every request. The same section says a client should not automatically retry a non-idempotent request unless it knows the request is safe to repeat or was never applied, a proxy must not do so at all, and a client should not automatically retry a failed automatic retry.

RFC 5789 says PATCH is neither safe nor idempotent, though it can be made idempotent, for example with a conditional request. The operations we most fear repeating (create a payment, place an order, send money) are usually POSTs. An idempotency key is the client and server agreeing, per request, that this particular POST may be repeated safely.

The Idempotency-Key header draft: what it says and where it stands

The HTTPAPI working group draft “The Idempotency-Key HTTP Header Field” is by Jayadeba Jena and Sanjay Dalal. On 6 October 2026 the datatracker shows the latest revision as draft-ietf-httpapi-idempotency-key-header-07, dated 15 October 2025, with the state “Expired” (the draft says it expires on 18 April 2026). Treat it as useful work in progress, not a standard. Version 07 says, in summary:

  • Syntax and uniqueness. The value is a Structured Field String, such as Idempotency-Key: "8e03978e-40d5-43e8-bc93-6894a57f9324", ideally a UUID or similar random value. It must be unique and must not be reused with a different payload; the resource owner defines uniqueness.
  • Expiry. A resource may expire keys and should publish its expiry policy. It must publish its idempotency rules.
  • Fingerprint. A server may combine the key with a fingerprint: a checksum of all or selected fields, field matching, or a request digest or signature.
  • Enforcement. First request: process normally. Retry after the original finished: return that original result, success or error. Retry while it is still running: conflict.
  • Errors. 400 when a required key is missing, 422 when a key is reused with a different payload, 409 while the original is still being processed.
  • Security. Low-entropy keys let attackers guess keys and read other clients’ cached responses. Validate the key format and build the lookup key from the client’s key plus client attributes only the server knows.

Its implementation list includes Stripe, Adyen and Dwolla, and APIs that use other names for the same idea, such as PayPal’s PayPal-Request-Id.

How real APIs do it

Stripe

From Stripe’s API reference and its page on advanced error handling:

  • The Idempotency-Key header is accepted on all POST requests; on GET and DELETE it has no effect. Keys can be up to 255 characters: V4 UUIDs or other high-entropy strings, or a key derived from your own object such as a cart ID, but not sensitive data like email addresses.
  • Stripe saves the status code and body of the first request for a key, success or failure, and replays it, including 500 errors. Results are saved only once the endpoint starts executing, so validation failures and concurrent conflicts are not saved. Rate limiters run before the idempotency layer, so a 429 can differ on retry.
  • Keys can be pruned once they are at least 24 hours old, after which a reused key creates a new request. Different parameters with the same key return an error.
  • Replays carry Idempotent-Replayed: true. For a 500, Stripe advises treating the result as indeterminate and not retrying with a new key, because the original may have had side effects.

Adyen

Adyen’s API idempotency page says:

  • POST requests take an idempotency-key header of up to 64 characters; a UUID is recommended. Keys are stored at company account level, valid for 7 to 14 days, and not checked across regions.
  • A duplicate sent before the first completes gets 422 or 409 with error code 704. A transient-error: true header means retry later with the same key. If the idempotency store is down, the API returns 503 with code 703.

AWS EC2 client tokens

The EC2 developer guide says some actions, such as TerminateInstances, are idempotent by default. A longer list of actions accepts a ClientToken: a unique, case-sensitive string of up to 64 ASCII characters. A retry with the same token and the same parameters succeeds without doing anything more. A retry with the same token but different parameters (other than Region or Availability Zone) fails with IdempotentParameterMismatch. For RunInstances the token is scoped either to the Region or to an Availability Zone, depending on how the zone is specified.

What they agree on

The shape is the same: a client-generated random key, scoped to an account, stored with the original parameters, with a limited lifetime. Limits, retention and error codes differ, so “we support idempotency keys” is not a complete contract. Document the header, format, scope, retention, mismatch error and in-progress behaviour.

Ideas worth keeping from the AWS Builders’ Library

Malcolm Featonby’s article “Making retries safe with idempotent APIs” is the best architecture reference here. The ideas worth keeping, in my own words:

  • Ask for intent, do not guess it. Treating identical parameters as a duplicate fails when the caller really wants two identical things, such as two instances. A caller-provided ID makes intent explicit and visible in logs.
  • One atomic unit. Recording the token and the mutation must succeed or fail together.
  • Same meaning, not “already exists”. An “already exists” error leaves the caller unsure whether its request did the work. A semantically equivalent response lets SDKs retry without the calling code noticing.
  • Late retries and bounded memory. For RunInstances, AWS honours the original contract even if the instance was since terminated, and keeps the token for the resource’s life plus a period after which late retries are not expected.
  • Changed parameters are a different request. Same token, different parameters gets a validation error, so the original parameters must be stored.

Designing your own idempotency layer

Scope the key

Never look up a key globally. Use (account_id, idempotency_key), with the account taken from authentication, not from the body. This follows the draft’s security advice and stops two customers’ keys colliding. If one key could reach several operations, include the route in the scope or fingerprint.

Store enough to replay

Store the fingerprint, a status (in_progress or completed), the response code and body, and created and expiry times. The stored response is what makes a retry get the same payment ID, not a new one and not an error.

Fingerprint the request

Hash the method, route and a canonical body, such as JSON with sorted keys. Same key with a different fingerprint: reject it (the draft suggests 422). Never silently run the new request, and never replay the old response for a different one.

Decide what happens to concurrent duplicates

  • One database transaction. Insert the key row, do the business write and save the response in one transaction. The PostgreSQL docs on index uniqueness checks explain that when a conflicting row was inserted by a transaction that has not committed, the new inserter waits for it to commit or roll back. On commit, the duplicate finds the completed row and replays it; on rollback, it runs fresh. Simple, and good when all the work is in one database.
  • Commit an in_progress row first. When the work calls another service, such as a payment gateway, do not hold a transaction open across it. Commit the key as in_progress, make the call, then save the result as completed. A duplicate that finds in_progress gets 409 with Retry-After. Pass the same key, or one derived from it, to the downstream API, and add a lease and a recovery job for rows stuck after a crash.

Decide which errors to store

  • Rejected before any work (validation, auth): do not store. Stripe does not save validation failures either.
  • A real business outcome such as “card declined”: store and replay it.
  • A 5xx: in the one-transaction design, the rollback removes the key row too, so the retry correctly starts fresh. If a side effect may have happened outside your database, keep the key in_progress and reconcile with the downstream system before any new attempt.

Pick a retention period

Choose the period from your clients’ longest retry window (queue redelivery, mobile apps coming back online, re-run jobs), not from storage cost alone, and publish it. Delete expired rows in small batches.

A minimal design in PostgreSQL

The tables below are a sandbox sketch. The second unique constraint on payments is a safety net in case a code path ever skips the key table.

CREATE TABLE idempotency_keys (
  account_id     bigint      NOT NULL,
  idem_key       text        NOT NULL CHECK (length(idem_key) BETWEEN 8 AND 255),
  request_hash   text        NOT NULL,
  status         text        NOT NULL CHECK (status IN ('in_progress', 'completed')),
  response_code  int,
  response_body  jsonb,
  created_at     timestamptz NOT NULL DEFAULT now(),
  expires_at     timestamptz NOT NULL,
  PRIMARY KEY (account_id, idem_key)
);

CREATE TABLE payments (
  id          bigserial   PRIMARY KEY,
  account_id  bigint      NOT NULL,
  idem_key    text        NOT NULL,
  amount      bigint      NOT NULL,
  currency    text        NOT NULL,
  created_at  timestamptz NOT NULL DEFAULT now(),
  UNIQUE (account_id, idem_key)
);

The request flow for the one-transaction design:

handle POST /payments:
  account = from auth token
  key     = Idempotency-Key header      -> 400 if missing or badly formed
  hash    = sha256(method + route + canonical JSON body)

  BEGIN
    INSERT INTO idempotency_keys (account_id, idem_key, request_hash, status, expires_at)
    VALUES (account, key, hash, 'in_progress', now() + interval '24 hours')
    ON CONFLICT (account_id, idem_key) DO NOTHING
    RETURNING idem_key

    if no row returned:                 -- key seen before (we may have waited here)
      row = SELECT request_hash, status, response_code, response_body
              FROM idempotency_keys WHERE account_id = account AND idem_key = key
      if row.request_hash != hash: ROLLBACK, return 422
      if row.status != 'completed': ROLLBACK, return 409 with Retry-After
      ROLLBACK, return row.response_code, row.response_body (mark as replayed)

    INSERT INTO payments (...) RETURNING id
    UPDATE idempotency_keys
       SET status = 'completed', response_code = 201, response_body = ...
     WHERE account_id = account AND idem_key = key
  COMMIT
  return 201 with the payment

The PostgreSQL 18 INSERT docs say ON CONFLICT DO NOTHING skips the insert when the arbiter constraint is violated, and RETURNING returns only rows actually inserted or updated, so an empty result means the key exists. The 24 hours is only an example.

Queues and consumers: the idempotent producer is not enough

Kafka’s design docs describe the idempotent producer: the broker gives each producer an ID and uses per-message sequence numbers to drop duplicates from resends. With enable.idempotence=true, the config docs say exactly one copy of each message is written in the stream; it needs acks=all, retries above 0 and max.in.flight.requests.per.connection of 5 or less. Since Kafka 3.0 it defaults to true when no conflicting settings are present (KIP-679). The upgrade notes add two caveats: a bug kept the default from applying in 3.0.0 and 3.1.0 (fixed in 3.0.1, 3.1.1 and 3.2.0), and the 3.2.0 notes say Kafka Connect disables idempotence for its producers by default.

None of this protects your consumer. The design docs explain that a consumer which crashes after processing but before saving its position leads to reprocessing: at least once. Emails, payment calls and balance updates will repeat. Apply the same idea as in HTTP:

  • Give each event a business ID when the intent is created, so a re-publish carries the same ID.
  • In the consumer, insert that ID into a processed_events table with a primary key, in the same transaction as the side effect, with ON CONFLICT DO NOTHING. No row inserted means already done.
  • Pass the event ID downstream as the idempotency key for external calls.

Retries with backoff and jitter

Idempotency makes a retry safe; backoff and jitter make it gentle on a struggling server. Marc Brooker’s AWS Architecture Blog post “Exponential Backoff And Jitter” (4 March 2015, updated May 2023) simulates many clients updating the same row. Capped exponential backoff, sleep = min(cap, base * 2 ** attempt), still leaves clusters of calls; “Full Jitter”, sleep = random_between(0, min(cap, base * 2 ** attempt)), spreads them out. In that simulation, no jitter was the clear loser, “Equal Jitter” did worse than “Full Jitter”, and “Full” versus “Decorrelated” was less clear. The update says most AWS SDKs now support backoff with jitter in their standard and adaptive retry modes.

In practice: retry only what can succeed later (timeouts, connection errors, 5xx, 429, and 409 for an in-progress key); a 400 or 422 needs a changed request, which means a new key. Use full jitter with a cap, honour Retry-After and hints like Stripe’s Stripe-Should-Retry, and retry at one layer with few attempts.

Common mistakes

  • New key per retry. A key generated inside the retry loop makes every retry look new. Generate it once per intent and reuse it, as the AWS SDKs do with their tokens.
  • Caching every error forever. A stored 500 from before any work was done leaves the client stuck for the whole retention period.
  • Caching no errors at all. Dropping the key after a 500 that followed a real side effect lets the next retry charge again.
  • Keys not scoped to the caller. A global key table lets one customer collide with, or replay, another’s response.
  • TTL shorter than the retry window. Keys that expire in one hour do not help a retry that arrives after six.
  • Ignoring in-progress duplicates. A plain SELECT “does the key exist?” before inserting lets two requests both see nothing. Let the unique constraint decide.
  • Key outside the business transaction. Key in Redis and payment in PostgreSQL, with no plan for partial failure, brings back the problem the Builders’ Library warns about.

Checklist: eight questions before you ship a retryable POST

  1. Which of our POST and PATCH endpoints create money movement, orders, messages or other side effects that must not repeat?
  2. Is the key generated once per user intent, stored by the client, and reused across every retry layer?
  3. Is the key scoped to the authenticated account, and validated for format and length?
  4. Do we store a request fingerprint, and do we reject the same key with a different request?
  5. Are the key record and the business write committed in one transaction, or do we have an in_progress state with a lease and recovery for external calls?
  6. What exactly does a concurrent duplicate get: a wait and replay, or a 409 with Retry-After?
  7. Which results do we store and replay, which do we never store, and how do we handle an indeterminate 5xx?
  8. Is our published retention period longer than the longest retry window of every client and queue, and do we pass a key to every downstream API we call?

Drill: send the same POST three times at once

Run this against a sandbox copy of your service, never production or a real payment provider. It assumes a POST /payments endpoint backed by the tables above. In my own test of this sketch, a two-second sleep inside the transaction made the race easy to see.

Step 1. Fire three identical requests at the same moment with one key:

KEY=$(uuidgen)
for i in 1 2 3; do
  curl -s -i -X POST http://localhost:8080/payments \
    -H "Content-Type: application/json" \
    -H "Authorization: Bearer $SANDBOX_TOKEN" \
    -H "Idempotency-Key: $KEY" \
    -d '{"amount": 49900, "currency": "INR"}' > "out-$i.txt" &
done
wait
grep -h -E "^HTTP|payment_id" out-*.txt

Step 2. Check that only one payment exists for that key:

SELECT account_id, idem_key, count(*)
FROM payments
WHERE idem_key = '<the key>'
GROUP BY account_id, idem_key;

Step 3. Send the same key with a different amount. Expect 422 (or your documented mismatch error), and no new row.

Step 4. Send the same body with the same key from a second sandbox account. Expect a new, separate payment, because keys are scoped per account.

Step 5. Send the same request again after the retention period, or after manually expiring the row in the sandbox. Confirm the behaviour matches what your docs promise.

In my run on PostgreSQL 17, all three requests returned 201 with the same payment ID (two marked as replayed), the count was 1, the mismatched body got 422 and the second account got its own payment. If you see two rows, look first for a key check done with a SELECT outside the transaction, or a key generated inside the retry loop.

Sources

comments powered by Disqus

Releted Posts

MinIO and its alternatives: revisit your self-hosted S3 in 2026

For many teams, MinIO was the default answer to one simple need: “give us S3 on our own machines”. It sat behind CI pipelines, backup jobs, data lake experiments and almost every docker-compose file that needed a bucket.

Read more

Elasticsearch and OpenSearch: revisit the fork before you pick one

Many teams still talk about “our Elastic cluster” when the cluster is actually Amazon OpenSearch Service, or about “OpenSearch” when half the code paths still use an old Elasticsearch client.

Read more

Kafka and Redpanda: revisit the choice before you treat them as the same bus

Many teams say “we run Kafka” when they mean “our services talk to a Kafka-compatible broker”. That shortcut was useful when one open source broker dominated the protocol.

Read more