Request Fingerprinting for Idempotent APIs
Canonical JSON (sorted keys, stable numbers) hashed with method+path; include/exclude map; Stripe-style 422 on mismatch; schema-evolution with null defaults.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Why canonical include-maps beat naive stringify
Prefer
Canonical JSON of an include-list, plus method and path
Hash only the fields that define the business intent. Key order, metadata, and timestamps must not change the fingerprint.
- sort_keys plus tight separators make two serializers agree.
- JSON pointers (amount, currency, customer.id) are the contract.
- Same key + same fingerprint replays. Same key + different fingerprint is 422.
- New optional fields default to null and stay out of the include map until you opt in.
Alternative
JSON.stringify the whole body as sent
Object key order, extra headers, client timestamps, and floating-point formatting all mutate the hash.
- A legitimate retry with keys in a different order 422s.
- request_id or timestamp in the body makes every retry a new fingerprint.
- POST /v1/payments and PUT /v1/payments/123 must not share a hash — method and path belong in the digest.
- No include map means every schema tweak is a production incident.
Happy path — fingerprint, then claim
Vertical cards for phones. Same story as the sequence diagram below.
- 1
Client sends one intent
POST /v1/payments with Idempotency-Key K and a JSON body. Key is minted once; body fields that matter stay stable across retries. - 2
Strip and sort
Drop ignored pointers (metadata, timestamp). Keep /amount, /currency, /customer/id. Sort object keys lexicographically. - 3
Hash METHOD|PATH|bytes
UTF-8 canonical JSON. SHA-256 (truncate only if you must). Store the hex next to the idempotency row. - 4
Lookup (account, key)
Unknown key: execute and persist entity + fingerprint. Known key + same fingerprint: replay the cached HTTP outcome. - 5
Mismatch
Known key + different fingerprint: 422 FingerprintMismatch with expected vs received. Do not execute.
Overview
An operation is idempotent only if repeating the same semantic intent does not produce a different outcome. The client names the intent with an Idempotency-Key. The server still has to answer: same key, but is this the same request?
Request fingerprinting is that answer. Compute a deterministic hash over a canonical representation of the request — HTTP method, resolved path, and the body fields that actually change money, inventory, or identity. Persist the fingerprint with the side effect. A later call with the same key:
| Same key + … | Server does |
|---|---|
| same fingerprint | Replay the cached status and body (or continue an in-flight claim) |
| different fingerprint | 422 Unprocessable — this key is already bound to another intent |
That 422 is a gift. A generic 500 hides a client bug (reused UUID wired to a new amount). A fingerprint-diff payload lets the caller fix the client, not “retry until it works.”
Why it matters
- Safe retries. Timeouts, load-balancer retries, and mobile radios will replay. The fingerprint decides “same charge” vs “new charge.”
- Observability. Log
request_fingerprint=<hash> outcome=hit|miss|mismatch. You can grep a double-submit across services. - UX. 422 with a diff is a product message (“this key was used for a different amount”), not a 500.
- Compliance. Audit wants exactly-once intent, not “we think the JSON was the same.”
Canonical JSON (RFC 8785)
Naive JSON.stringify(obj) is not a fingerprint. Key order is insertion order in modern engines, whitespace differs, 1.0 vs 1, Unicode normalization, and undefined dropping all change the bytes.
RFC 8785 (JCS) is the spec to name:
- Sort object keys lexicographically (UTF-16 code units in JCS; for interviews, “sorted Unicode keys” is enough).
- No insignificant whitespace — separators
","and":"with no spaces. - Numbers in shortest decimal form — no trailing
.0, no scientific notation surprises. Prefer integers in minor units (amount: 1200cents) so you never hash floats. - Unicode NFC if you accept free-text that can be composed two ways.
Production serializers: orjson.dumps(..., sort_keys=True) in Python; a key-sorting replacer in JS. Hash UTF-8 bytes, not a JS string you later re-encode.
Method + path
POST /v1/payments is not PUT /v1/payments/123. The verb and the resolved route (after path-parameter substitution) belong in the digest:
payload = METHOD.upper() + "|" + path + "|" + canonical_json_bytes
fingerprint = SHA-256(payload) // hex; truncate to 128 bits only if storage forces itHashing the body alone would let a GET and a POST collide, or a charge create and a charge capture.
Include / exclude map
A per-endpoint schema lists JSON pointers that define the operation. Everything else is stripped before canonicalization.
POST /v1/payments include:
/amount
/currency
/customer/id
ignored (must not affect the hash):
/metadata
/timestamp
/request_id
Idempotency-Key header itself
Authorization, User-Agent, X-Request-IdIf you hash timestamp, every retry is a mismatch. If you hash Idempotency-Key, you have circular nonsense. If you omit /amount, a client can reuse a key and change the charge.
Deep dive · Where to compute the hash
Do it at the service boundary, after auth, before business logic. A gateway that caches fingerprints will go stale when the include map evolves. Each service owns its schema. Keep payloads small (tens of KB); canonicalization is O(N) in bytes. Emit the hash on the tracing span.
Storage
Two boring options:
| Strategy | When |
|---|---|
Inline column on the primary row (payments.fingerprint) | One entity per intent; unique index is the collision detector |
Idempotency table (account_id, key) → fingerprint, response | Stripe-style; you already have this row from keys |
SHA-256 collisions are not a practical worry. If you somehow detect two different canonical payloads with the same hash, fail 409 and page yourself — do not execute.
Schema evolution
New optional fields must not 422 old clients.
Rules:
- The include map is versioned with the endpoint. Adding a pointer is a breaking change for old fingerprints unless you default the missing value to a stable sentinel (
null) and old clients omit it the same way. - Prefer: new field is not in the include map until you require it. Old and new clients keep hashing the same.
- When you must include it, default missing to
nullin the filtered object so{amount, currency}and{amount, currency, tax: null}match — only if you always insert that null in canonical form. Document it; this is easy to get backwards. - Never include fields whose values are generated server-side on first request (server timestamps, assigned ids).
422 mismatch payload
Stripe-style: do not run the handler. Return a body the client can show a developer:
{
"error": "FingerprintMismatch",
"expected": "a1b2…",
"received": "c3d4…"
}HTTP 422 Unprocessable Entity. This is a client bug (or a key collision). The fix is a new idempotency key for the new intent, not a retry of the old key.
Replay of a matching fingerprint should keep the original status — including a cached 500 if that is your policy — see the keys lesson.
Sequence
- 1
Client → Service
POST /v1/payments plus key K
- 2
Service
Strip includes, sort keys, SHA-256
- 3
Service → Service
Persist entity plus fingerprint
- 4
Service → Client
201 Created
- 5
Service → Client
200 replay cached body
- 6
Service → Client
422 FingerprintMismatch
Lesson map
Request Fingerprinting for Idempotent APIs
Canonical JSON (sorted keys, stable numbers) hashed with method+path; include/exclude map; Stripe-style 422 on mismatch; schema-evolution with null defaults.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB c["Client"] s["Service"] c -->|POST| s s -->|201 Created| c s -->|200 replay| c s -->|422| c
Architecture choices
- Stateless gateway, stateful service. Hash in the service so the include map ships with the code that knows the schema.
- Integer money. Fingerprint cents, not
19.99IEEE floats. - Observability.
outcome=hit|miss|mismatchplus the hash. Mismatch rate is a client-SDK bug dashboard. - Performance. Compiled JSON (
orjson, stable-stringify). Reject huge bodies before hashing.
In-memory fingerprint + 422 (run this)
No network. Canonical JSON with sorted keys. Include pointers only. Same payload, different key order → same hash. Extra ignored field → same hash. Amount change under the same key → 422.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Production uses SHA-256 of the UTF-8 bytes (djb2 here is a sandbox stand-in so we need no crypto). The equality behavior is what you are learning: order and ignored fields do not matter; amount does.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
How does request fingerprinting differ from an Idempotency-Key header?
Answer
The key is an opaque client token — the server stores key → result. The fingerprint is derived from the request and bound to that key so a retry cannot change amount or customer. Fingerprint-as-the-only-key is a different (usually worse) design: semantically identical POSTs from different intents collide.
What happens when you add a new optional field?
Answer
If it is not on the include map, old and new clients keep the same hash. If you add it to the map, missing values must canonicalize to a stable default (null) or you 422 every old SDK. Treat include-map changes as API breaks.
Why SHA-256 instead of MD5 or MurmurHash?
Answer
Financial APIs need collision resistance against a malicious client who wants two payloads to share a fingerprint (charge $1, replay as $10_000). MD5 is broken. Non-crypto hashes are fast and not a proof against crafted collisions. Truncate SHA-256 only if storage forces it, and still keep enough bits.
Why 422 and not 409 on fingerprint mismatch?
Answer
409 is the in-flight conflict (“this key is currently executing; wait”). 422 is “you reused a finished (or claimed) key with a different body.” Mixing them makes clients retry a bug. 409 + Retry-After; 422 + stop and mint a new key.
Should volatile headers be in the fingerprint?
Answer
No. Authorization rotates, User-Agent changes with app versions, X-Request-Id is unique per attempt by design. Hash method, path, and the include-list body. Auth still runs; it is just not part of intent identity.
Where do you store the fingerprint?
Answer
On the idempotency row next to account_id, key, state, and the cached response — unique (account_id, key), fingerprint as a column you compare. Optionally also on the domain row for audit.
What if two different canonical payloads hash the same?
Answer
Treat it as a 409 plus a pager, not as a replay. With SHA-256 this is not a capacity-planning scenario. Do not “just run the handler.”
Why include method and path?
Answer
Otherwise POST /v1/refunds and POST /v1/charges with a similar {amount, currency} body could share a digest if you ever mixed stores, and a capture vs a create on similar URLs would look like the same intent. The HTTP operation is part of the intent.
How do you fingerprint nested objects and arrays?
Answer
Pointers walk objects (/customer/id). Arrays are ordered — sorting array elements would change list semantics (two line items swapped is a different order). Sort keys, not arrays, unless the API declares a set. Canonicalize nested objects recursively.
Can the client send the fingerprint instead of the server computing it?
Answer
Do not trust it. The server computes. A client-supplied hash is a confused-deputy waiting to happen. The client supplies the key and a stable body; the server supplies the digest.
Pitfalls
Write two JSON bodies that a human calls “the same charge.” Put created_at on the retry. Hash the whole body — hashes differ. Then hash only /amount, /currency, /customer/id — hashes match. Add that diagram to your interview notes next to the 422 payload.