API Idempotency Keys
Stripe-style Idempotency-Key for safe retries — fingerprint, unique (account,key), in_progress/completed, response cache, ~24h TTL.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Why idempotency keys win for POST charges
Prefer
Client key + server claim + response cache
The client names one logical intent. The server executes it at most once in the TTL window and replays the stored HTTP outcome.
- Works for POST to a collection when the client does not know the resource id yet.
- Handles timeouts, 409 in-flight, and zombie workers with leases + fencing.
- Fingerprint mismatch (422) catches reused keys wired to a new payload.
Alternative
Just use PUT / rely on HTTP method semantics
PUT is idempotent for replace-by-URL. Creating a charge is not a replace.
- Clients rarely mint the charge id up front; POST to /v1/charges is the product API.
- Method idempotence does not stop concurrent in-flight duplicates.
- Without a key, a timeout retry is a second intent as far as the server can tell.
Happy path — one intent, many packets
Vertical flow for phones. Same story as the mermaid below, without a sideways SVG.
- 1
Mint the key once
UUIDv4 persisted with the outbound request. Never regenerate on retry. - 2
Atomic claim
INSERT (account_id, key) → in_progress with fingerprint + lease. Conflict → read state. - 3
Run the mutation once
Charge / write domain state. Prefer the same DB transaction or an outbox. - 4
Cache the HTTP outcome
Store status + body as completed or failed. Later retries replay it. - 5
If the worker dies
Expired lease + fencing token takeover. Never DELETE the row.
Overview
An operation is idempotent if applying it once or many times yields the same observable side effects as applying it once. For read-only GET / HEAD that is usually free. For payments, refunds, transfers, order placement, and other mutations, retries are inevitable: timeouts, load-balancer blips, flaky mobile networks, clients crashing mid-flight. Without a server-side contract, “retry the charge” becomes a double charge.
Idempotency keys (popularized by Stripe’s Idempotency-Key header) let a client label one logical intent. The server records that intent, executes at most once inside a TTL window, and replays the stored HTTP status plus body on later retries — even if the first attempt returned a 500.
Interviewers love this topic because it sits at the intersection of API design, distributed systems, and money correctness. It forces you to reason about races, crashes, storage, and client contracts — not just “add a UUID header.” Production systems must handle concurrent in-flight duplicates, fingerprint mismatches, and zombie / stuck-pending keys after process death.
Client vs server responsibilities
Client
- Generate a key once per logical operation — typically a UUIDv4 (or another high-entropy string, at most 255 characters). Do not regenerate on retry.
- Persist the key with the outbound request (DB row, outbox, local store) so retries reuse it.
- Send
Idempotency-Keyon mutating requests (POST, sometimesPUT/PATCH). - Keep request parameters stable across retries of the same key (same amount, currency, customer).
- On network uncertainty (timeout, connection reset): retry the same key with exponential backoff plus jitter.
- On 409 Conflict (in-flight): wait (
Retry-Afterif present) and retry the same key — do not invent a new key. - On 422 (fingerprint mismatch): treat as a client bug — do not reuse that key for a different payload.
- Never put PII (email, SSN, card PAN) in the key itself.
Server
- Scope keys per account / tenant (Stripe scopes per Stripe account). Unique constraint on
(account_id, key). - Compute a request fingerprint (hash of method + path + normalized body params that define the operation).
- Atomically claim the key into
in_progress(or reject a concurrent claim with 409). - Execute the business operation once under that claim.
- Persist status code + response body (and optionally headers) when the endpoint finishes — success or failure.
- Replay the cached response for later retries with a matching fingerprint.
- Enforce TTL (~24h for Stripe API v1; longer on some v2 surfaces) and reclaim / expire keys.
- Recover zombie pending keys (worker died after claim) via lease / TTL takeover with fencing.
Storage model
Logical schema
Treat this as the conceptual row you would put behind the unique key. Production migrations and types vary; the columns are the contract.
idempotency_records
account_id TEXT NOT NULL
key TEXT NOT NULL -- client Idempotency-Key
request_fingerprint TEXT NOT NULL -- hash of canonical request
state ENUM('in_progress','completed','failed')
response_status INT NULL
response_body BYTEA / JSONB NULL -- cached reply
response_headers JSONB NULL
recovery_point TEXT NULL -- for multi-step flows
lease_owner TEXT NULL -- worker id
lease_expires_at TIMESTAMPTZ NULL
fencing_token BIGINT NOT NULL DEFAULT 0
created_at TIMESTAMPTZ NOT NULL
updated_at TIMESTAMPTZ NOT NULL
expires_at TIMESTAMPTZ NOT NULL -- ~created_at + 24h
PRIMARY KEY / UNIQUE (account_id, key)
INDEX (expires_at) -- reaper
INDEX (state, lease_expires_at) -- zombie recoveryRequest fingerprint
Hash a canonical representation of the mutating inputs: sorted JSON keys, normalized amounts, exclude volatile headers. On retry:
| Same key + … | Server does |
|---|---|
| same fingerprint | Safe replay, or continue the in-flight claim |
| different fingerprint | 422 Unprocessable (misuse). Do not execute. |
Canonicalization is the sharp edge. JSON.stringify with an unsorted object, floating-point formatting, or extra headers in the hash will 422 a legitimate retry.
State machine
in_progress: claim acquired; work running (or zombie if the lease expired).completed: success response cached; always replay.failed: error response cached (including many 4xx / 5xx after the endpoint started); replay that outcome. Stripe caches failures too.
Validation failures before endpoint execution often do not create a durable idempotent result (Stripe behavior) — the client may fix the payload and retry.
States
- 1
Start → in progress
Start → in progress
atomic claim
- 2
in progress → completed
in progress → completed
success
- 3
in progress → failed
in progress → failed
error after start
- 4
in progress → takeover
in progress → takeover
lease expired
- 5
takeover → in progress
takeover → in progress
fencing token
- 6
completed → completed
completed → completed
replay cache
- 7
failed → failed
failed → failed
replay cache
Lesson map
API Idempotency Keys
Stripe-style Idempotency-Key for safe retries — fingerprint, unique (account,key), in_progress/completed, response cache, ~24h TTL.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB client["Client"] w1["Worker1"] w2["Worker2"] store["Store"] client -->|POST charge| w1 client -->|retry same key K| w2 w1 -->|INSERT ON| store store -->|claimed| w1 w2 -->|INSERT ON| store store -->|no row then| w2
Response cache and TTL
Store status + body for ~24 hours. After prune, the same key may start a new request. Design clients so keys are unique within the TTL window of an account. A TTL shorter than the client retry window is how you get a second charge.
Concurrent retries, 409, and zombie recovery
Concurrent in-flight
A unique constraint alone is not enough. Two requests can race: both miss, both insert. Prefer:
INSERT … ON CONFLICT DO NOTHING(or an equivalent atomic claim).- Winner proceeds; loser
SELECT … FOR UPDATE/ reads state. - If
in_progressand the lease is fresh → HTTP 409 Conflict +Retry-After. - If
completed/failed→ replay the cached response. - If the fingerprint differs → 422.
Do not block the second request until the first finishes. That ties up workers and connections. 409 plus client retry is the scalable contract.
Sequence
- 1
Client → Worker1
POST charge Idempotency-Key K
- 2
Client → Worker2
retry same key K
- 3
Worker1 → Store
INSERT ON CONFLICT DO NOTHING
- 4
Store → Worker1
claimed in_progress
- 5
Worker2 → Store
INSERT ON CONFLICT DO NOTHING
- 6
Store → Worker2
no row then SELECT FOR UPDATE
- 7
Worker2 → Client
409 Conflict Retry-After
- 8
Worker1 → Store
complete plus cached 200
- 9
Client → Worker1
retry same key K
- 10
Worker1 → Client
cached 200 Idempotent-Replayed true
Zombie / stuck pending
Crash after INSERT / in_progress but before completed leaves the key forever pending → perpetual 409s.
Recovery pattern:
- Attach a lease (
lease_owner,lease_expires_at) and optional heartbeat. - Before returning 409, if
nowis pastlease_expires_at, allow conditional takeover: bumpfencing_token, set a new owner / lease, re-enter execution only if the side effect is itself safe to retry (downstream idempotency, or reconcile an uncertain outcome). - Never blindly
DELETEthe row — that races with a slow-but-alive worker.
Fencing tokens ensure a stale worker cannot commit after takeover. The complete-write must check WHERE fencing_token = $mine (or compare-and-swap the token you received at claim time).
Deep dive · Why DELETE is the wrong zombie fix
A slow worker can still be executing after you decide it is dead. If you delete the row, a retry inserts a fresh in_progress claim and runs the charge again while the original worker also commits. Lease expiry plus an incremented fencing token lets a new worker continue, and the stale worker’s cache write no-ops.
System design / architecture
Components
- API gateway / edge — auth, rate limits, extract
Idempotency-Key, attachaccount_id. - Idempotency middleware — fingerprint, claim, lease, replay, 409 / 422 decisions.
- Idempotency store — Redis and/or Postgres (see tradeoffs).
- Domain service — charge / mutate business state; ideally participates in the same transaction or uses an outbox.
- Response interceptor — write
completed/failedplus cached body. - Reaper / sweeper — expire TTL rows; detect zombie leases for a takeover queue.
- Metrics and alerts — claim latency, 409 rate, zombie takeovers, fingerprint mismatches, store errors (fail closed).
Happy path
- Client sends
POST /v1/chargeswithIdempotency-Key: K. - Middleware hashes fingerprint
F; tries atomic insert of(account, K)intoin_progresswith a lease. - Domain runs payment; commits the business effect.
- Middleware stores status / body as
completed; returns the response. - Retry with the same
Kloads the row, checks fingerprint match, returns the cached body (Idempotent-Replayed: trueoptional).
Decisions
- 1
POST plus Idempotency-Key
- nextKnown account and key?
- ?
Known account and key?
- noAtomic claim plus lease
- yesRow status?
- 3
Atomic claim plus lease
- nextRun domain mutation
- 4
Run domain mutation
- nextCache completed response
- 5
Cache completed response
- nextReturn status and body
- 6
Return status and body
- ?
Row status?
- mismatch422 Unprocessable
- completed or failedReplay cache
- in_progressLease fresh?
- 8
422 Unprocessable
- 9
Replay cache
- ?
Lease fresh?
- yes409 Retry-After
- noTakeover plus fence
- 11
409 Retry-After
- 12
Takeover plus fence
- nextRun domain mutation
Failure modes
| Mode | What happened | Safe next step |
|---|---|---|
| Timeout before claim | No durable key | Retry; may create a new claim |
| Crash after claim, before side effect | Zombie pending | Lease expiry plus takeover |
| Crash after side effect, before cache write | Dangerous: charged but response lost | TX coupling, outbox, or reconcile by downstream id |
| Store unavailable | Uniqueness unknown | Fail closed (503). Do not silently allow duplicates |
| Concurrent duplicate | Two in-flight same key | 409 plus Retry-After |
| Payload changed under same key | Client reused key for a new intent | 422 |
The third row is the interview trap. If the charge committed in Postgres but the idempotency row is still in_progress, a takeover that “just runs the handler again” double-charges unless the domain write is itself idempotent or you look up the existing charge first.
Tradeoffs: DB vs Redis
| Store | Why you pick it | Cost |
|---|---|---|
| Postgres (or other durable DB) | Strong consistency with UNIQUE plus SELECT FOR UPDATE / advisory locks; can commit the idempotency row and the business write in one transaction | Higher latency |
| Redis | Sub-ms lookups, natural TTL | Risk of loss on failover unless AOF / cluster is carefully tuned; often paired with a durable fallback |
| Hybrid | Redis hot path; dual-write to Postgres; on Redis miss / failure, check Postgres | Complexity. Fail closed when neither store can confirm uniqueness |
Multi-step flows and recovery points
Long workflows (authorize, capture, ledger, notify) should not be one opaque black box under a single key without checkpoints.
Recovery points (inspired by Stripe phased execution):
- Persist
recovery_pointon the idempotency record after each durable phase, for exampleauth_created,funds_captured,ledger_posted,webhook_enqueued. - On retry / takeover, resume from the last committed recovery point instead of restarting from scratch.
- Each phase that calls an external system must either pass a derived downstream idempotency key, or query external state before acting.
- Emit an outbox event in the same DB transaction as the phase commit so notifications are at-least-once without losing the exactly-once side-effect story for money movement.
Teaching code: Postgres claim (Python)
Illustrative FastAPI-style sketch. Production needs a real pool, migrations, and JSON canonicalization. This fence is not runnable here — it talks to Postgres.
import hashlib, json, uuid
from datetime import datetime, timedelta, timezone
from enum import Enum
import asyncpg
from fastapi import FastAPI, Header, HTTPException, Request, Response
TTL = timedelta(hours=24)
LEASE = timedelta(seconds=30)
class State(str, Enum):
IN_PROGRESS = "in_progress"
COMPLETED = "completed"
FAILED = "failed"
def fingerprint(method: str, path: str, body: dict) -> str:
canonical = json.dumps(body, sort_keys=True, separators=(",", ":"))
raw = f"{method.upper()}|{path}|{canonical}".encode()
return hashlib.sha256(raw).hexdigest()
async def claim_or_replay(conn, account_id, key, fp, owner):
now = datetime.now(timezone.utc)
expires = now + TTL
lease_exp = now + LEASE
row = await conn.fetchrow(
"""
INSERT INTO idempotency_records
(account_id, key, request_fingerprint, state, lease_owner,
lease_expires_at, fencing_token, created_at, updated_at, expires_at)
VALUES ($1,$2,$3,'in_progress',$4,$5,1,$6,$6,$7)
ON CONFLICT (account_id, key) DO NOTHING
RETURNING *
""",
account_id, key, fp, owner, lease_exp, now, expires,
)
if row:
return None # we own the claim
existing = await conn.fetchrow(
"SELECT * FROM idempotency_records WHERE account_id=$1 AND key=$2 FOR UPDATE",
account_id, key,
)
if existing["request_fingerprint"] != fp:
raise HTTPException(422, "Idempotency key reused with different request")
if existing["state"] in (State.COMPLETED, State.FAILED):
return {
"status": existing["response_status"],
"body": existing["response_body"],
"replayed": True,
}
if existing["lease_expires_at"] and existing["lease_expires_at"] > now:
raise HTTPException(409, "Request in progress; retry later")
await conn.execute(
"""
UPDATE idempotency_records
SET lease_owner=$3, lease_expires_at=$4,
fencing_token = fencing_token + 1, updated_at=$5,
recovery_point = COALESCE(recovery_point, 'start')
WHERE account_id=$1 AND key=$2
""",
account_id, key, owner, lease_exp, now,
)
return None # takeover; re-enter execution
async def complete(conn, account_id, key, status, body, ok):
await conn.execute(
"""
UPDATE idempotency_records
SET state=$3, response_status=$4, response_body=$5::jsonb,
updated_at=NOW(), lease_expires_at=NULL
WHERE account_id=$1 AND key=$2
""",
account_id, key,
State.COMPLETED if ok else State.FAILED,
status, json.dumps(body),
)Wire the claim into the same transaction as the charge insert so a crash cannot leave “money moved, cache missing” — or persist a recovery_point and reconcile.
app = FastAPI()
@app.post("/v1/charges")
async def create_charge(
request: Request,
idempotency_key: str = Header(..., alias="Idempotency-Key"),
account_id: str = Header(..., alias="X-Account-Id"),
):
body = await request.json()
fp = fingerprint("POST", "/v1/charges", body)
owner = str(uuid.uuid4())
pool = request.app.state.pool
async with pool.acquire() as conn:
async with conn.transaction():
replay = await claim_or_replay(conn, account_id, idempotency_key, fp, owner)
if replay:
return Response(
content=json.dumps(replay["body"]),
status_code=replay["status"],
media_type="application/json",
headers={"Idempotent-Replayed": "true"},
)
charge_id = str(uuid.uuid4())
amount = body["amount"]
await conn.execute(
"INSERT INTO charges(id, account_id, amount) VALUES ($1,$2,$3)",
charge_id, account_id, amount,
)
await conn.execute(
"UPDATE idempotency_records SET recovery_point='charge_inserted' WHERE account_id=$1 AND key=$2",
account_id, idempotency_key,
)
result = {"id": charge_id, "amount": amount, "status": "succeeded"}
await complete(conn, account_id, idempotency_key, 200, result, True)
return resultTeaching code: Postgres claim (TypeScript)
Same semantics with Express + pg. Also a sketch — not runnable in the sandbox.
import crypto from "crypto";
import express from "express";
import { Pool } from "pg";
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
const TTL_MS = 24 * 60 * 60 * 1000;
const LEASE_MS = 30_000;
type Cached = { status: number; body: unknown; replayed: true };
function fingerprint(method: string, path: string, body: unknown): string {
const canonical = JSON.stringify(body, Object.keys(body as object).sort());
return crypto
.createHash("sha256")
.update(`${method.toUpperCase()}|${path}|${canonical}`)
.digest("hex");
}
async function claimOrReplay(
accountId: string,
key: string,
fp: string,
owner: string
): Promise<Cached | null> {
const client = await pool.connect();
try {
await client.query("BEGIN");
const now = new Date();
const expires = new Date(now.getTime() + TTL_MS);
const leaseExp = new Date(now.getTime() + LEASE_MS);
const inserted = await client.query(
`INSERT INTO idempotency_records
(account_id, key, request_fingerprint, state, lease_owner,
lease_expires_at, fencing_token, created_at, updated_at, expires_at)
VALUES ($1,$2,$3,'in_progress',$4,$5,1,$6,$6,$7)
ON CONFLICT (account_id, key) DO NOTHING
RETURNING id`,
[accountId, key, fp, owner, leaseExp, now, expires]
);
if (inserted.rowCount === 1) {
await client.query("COMMIT");
return null;
}
const { rows } = await client.query(
`SELECT * FROM idempotency_records
WHERE account_id=$1 AND key=$2 FOR UPDATE`,
[accountId, key]
);
const row = rows[0];
if (row.request_fingerprint !== fp) {
await client.query("ROLLBACK");
const err: { status?: number } & Error = new Error("fingerprint mismatch");
err.status = 422;
throw err;
}
if (row.state === "completed" || row.state === "failed") {
await client.query("COMMIT");
return { status: row.response_status, body: row.response_body, replayed: true };
}
if (row.lease_expires_at && new Date(row.lease_expires_at) > now) {
await client.query("ROLLBACK");
const err: { status?: number } & Error = new Error("in progress");
err.status = 409;
throw err;
}
await client.query(
`UPDATE idempotency_records
SET lease_owner=$3, lease_expires_at=$4,
fencing_token = fencing_token + 1, updated_at=$5
WHERE account_id=$1 AND key=$2`,
[accountId, key, owner, leaseExp, now]
);
await client.query("COMMIT");
return null;
} catch (e) {
await client.query("ROLLBACK");
throw e;
} finally {
client.release();
}
}
const app = express();
app.use(express.json());
app.post("/v1/charges", async (req, res) => {
const key = req.header("Idempotency-Key");
const accountId = req.header("X-Account-Id");
if (!key || !accountId) return res.status(400).json({ error: "missing headers" });
const fp = fingerprint("POST", "/v1/charges", req.body);
const owner = crypto.randomUUID();
try {
const replay = await claimOrReplay(accountId, key, fp, owner);
if (replay) {
res.setHeader("Idempotent-Replayed", "true");
return res.status(replay.status).json(replay.body);
}
const chargeId = crypto.randomUUID();
await pool.query(
`INSERT INTO charges(id, account_id, amount) VALUES ($1,$2,$3)`,
[chargeId, accountId, req.body.amount]
);
const body = { id: chargeId, amount: req.body.amount, status: "succeeded" };
await pool.query(
`UPDATE idempotency_records
SET state='completed', response_status=200,
response_body=$3::jsonb, lease_expires_at=NULL, updated_at=NOW(),
recovery_point='charge_inserted'
WHERE account_id=$1 AND key=$2`,
[accountId, key, JSON.stringify(body)]
);
return res.status(200).json(body);
} catch (e: any) {
const status = e.status ?? 500;
if (status === 409) res.setHeader("Retry-After", "1");
return res.status(status).json({ error: e.message });
}
});
app.listen(3000);In-memory store (run this)
Same claim / 409 / 422 / replay / zombie-lease / fencing-token rules, with a Map instead of Postgres. Advance now to expire a lease; a stale worker’s complete no-ops if the fencing token moved.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Why not just make the charge endpoint use HTTP PUT?
Answer
PUT idempotency is resource-replace semantics keyed by URL. Creating a new charge is naturally a POST to a collection; clients do not know the charge ID beforehand. Idempotency keys give POST the retry safety of PUT without forcing awkward client-generated resource IDs for every money movement.
Should you cache 500 responses under an idempotency key?
Answer
Stripe does: subsequent retries return the same result, including many failures after execution begins. That prevents “retry until success” from double-applying a partially ambiguous failure. Nuance: validation errors before execution often are not stored, so the client can fix and retry. Know the distinction in your design.
How do you handle two parallel requests with the same key?
Answer
Atomic claim into in_progress. The loser observes in-flight state and returns 409 Conflict (ideally with Retry-After), not a second execution and not a blocking wait that ties up workers. After completion, retries replay the cache.
What is a request fingerprint and why 422 on mismatch?
Answer
A hash of the canonical mutating parameters bound to the key. Reusing a key with a different amount or customer is almost always a client bug (or key-collision misuse). Rejecting with 422 protects against accidental cross-wiring of keys to payloads. Same key is the same intent — not a new charge.
How do you recover from a stuck in_progress (zombie) key?
Answer
Lease plus TTL. If the lease is expired, a new worker may take over with an incremented fencing token. The re-entered work must be safe: check the DB for an existing charge, pass downstream idempotency keys, or reconcile. Do not delete the row blindly.
Redis or Postgres for the idempotency store?
Answer
Postgres if you need to commit business writes and idempotency state atomically and prefer simpler consistency. Redis if you need very low latency and can accept dual-write / fallback complexity. Many high-scale systems use Redis hot path plus durable fallback and fail closed if uniqueness cannot be guaranteed.
How do multi-step payment flows use recovery points?
Answer
Persist named checkpoints (auth_ok, capture_ok, …) on the idempotency record. Retries resume at the last durable point. Each external call needs its own idempotency or read-before-write. Pair with a transactional outbox for async side effects.
What header signals a replay to the client?
Answer
Commonly Idempotent-Replayed: true (Stripe documents this in error / low-level guidance). Useful for metrics and client debugging; the body and status should still match the original outcome.
Who generates the idempotency key — client or server?
Answer
The client (or the first edge that owns the user intent) generates it and stores it with the intent so retries reuse the same key. A server-generated key on every request is useless — each retry is a new key. UUIDv4 is fine; derived keys from (user, cart, amount) can work if that tuple truly means one charge. Never use PII as the key.
How long should you keep an idempotency key?
Answer
At least as long as the client retry window plus downstream settlement lag — often 24h for payments, longer if webhooks can fire late. Too short: the key expires, a legitimate retry creates a second charge. Too long: storage cost and privacy. TTL the row, keep the business object forever.
What belongs in the request fingerprint?
Answer
Canonical mutating fields only: amount, currency, customer/destination, and any capture vs auth flag — not timestamps, Idempotency-Key itself, or volatile headers. Hash a stable JSON serialization. Mismatch → 422, not a new execution. That is how you stop “same key, different $500 vs $5” bugs. See request fingerprinting.
Is at-least-once delivery the same as an idempotent HTTP POST?
Answer
No. Broker delivery says how many times a message may arrive. HTTP idempotency says whether repeating a request changes observable state. You usually want both: at-least-once pipes plus an idempotent consumer (inbox / keys). See at-least-once vs exactly-once and the outbox/inbox.
Why not skip fingerprints if the client already sends a UUID key?
Answer
The UUID names the intent. The fingerprint binds that intent to a payload. Without it, a retried key with a different amount is a silent new charge.
How do retries interact with idempotency keys?
Answer
Backoff + jitter + Retry-After stop the storm. The key makes each retry safe. Neither is enough alone: keys without budgets still melt a down dependency; retries without keys double-charge. Retry storms.
Pitfalls
Draw two clients hitting POST /v1/charges with the same key while the first worker is still inside the processor. Label the atomic INSERT … ON CONFLICT, the 409 + Retry-After, and what happens if that worker dies before completed. Then add the lease expiry, fencing-token bump, and the check that prevents the dead worker from writing the cache.
Go Deeper
- Stripe Docs — Idempotent requests
- Stripe Engineering Blog — Designing robust and predictable APIs with idempotency
- Stripe Docs — Advanced error handling (retries and idempotency)
- YouTube — Designing Idempotent API Endpoints for Payments at Stripe
- FlowVerify — Idempotency keys: handling concurrent in-flight requests
- Sujeet Jaiswal — Stripe: Idempotency for Payment Reliability
- Stripe API v2 overview — Idempotency