System design
Part 6 of 6 · Push NotificationsReliability, Dedup & Idempotency — Retries, Provider Receipts & Device State
Retries without idempotency create duplicate banners; provider receipts without token hygiene create silent loss. Persist send intent, keep notification_id stable, classify retryable vs terminal errors, and run a device state machine. Exactly-once bus theory lives in Kafka + idempotency clusters — cross-link, do not re-teach.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
How you retry a push
Prefer
Intent + idempotency key + terminal-error purge
Claim a worker lease so two pods do not double-send. Retry 5xx/timeout with the same nid. Purge Unregistered immediately. Collapse and a short client nid window catch anything that still slips through.
- Business key (order_id + template) or an explicit client key.
- Dead-letter poison payloads — ops inspect, not silent drop forever.
- Bind token to user_id; detect rare token reuse across users.
Alternative
Fire provider, hope, retry 4xx forever
No intent log means crash recovery double-sends. BadRequest retry storms. Dual writers without a lease duplicate critical alerts. Treating accept as durable user delivery lies to the funnel.
- 4xx BadRequest is a payload bug — fix it, do not backoff forever.
- Lease expiry without a claim protocol is a second-writer race.
- Doze delay is not a retryable provider error.
Reliability loop
Retry is a box, not a labeled self-loop on send — that collides edge labels on desktop.
- 1
Persist intent + idempotency key
At-least-once outbound. Ingress retries must not enqueue twice — see the idempotency cluster. - 2
Claim a worker lease
One in-flight send per nid. Crash recovery reclaims after expiry. - 3
Provider send → HTTP receipt
200 + id → accepted. 5xx/timeout → backoff. Unregistered or BadDeviceToken → invalidate that token. InvalidProviderToken or ExpiredProviderToken → fix credentials and keep device tokens. - 4
Resend same nid + key
New provider_message_id is fine. User-visible effect stays one banner via collapse. - 5
Dead-letter exhausted attempts
Poison payloads and ops, not an infinite 4xx loop.
Overview
Retries without idempotency create duplicate banners. Provider receipts without token hygiene create silent loss. This page ties retries, dedup keys, APNs/FCM error handling, and device state into one reliability story.
Deep exactly-once bus theory lives in Kafka delivery and idempotency keys — cross-link, do not re-teach.
You should be able to:
- Draw persist → claim → send → classify → retry box or purge.
- Name four dedup layers.
- Walk the device state machine without mass-deleting "suspect" tokens.
Dedup layers
| Layer | Job |
|---|---|
| Ingress idempotency | API retries do not enqueue twice — idempotency keys |
| Outbound claim | Worker lease so two pods do not double-send mid-flight |
| Provider collapse | Shade-level coalesce if duplicates slip through |
| Client display dedup | App ignores duplicate nid in a short window (last line of defense) |
At-least-once on the wire. Make the user-visible effect idempotent.
Device state
| State | Meaning | Action |
|---|---|---|
| active | Token fresh, last_seen recent | Send |
| suspect | Soft failures / no opens in cohort | Investigate; do not mass-delete |
| invalid | Device-token error (Unregistered, BadDeviceToken). Provider auth errors do not change device state | Stop sending that token |
| uninstalled / revoked | Permission off | Stop; optional in-app re-prompt |
Architecture (retry is a separate box)
Decisions
- 1
1 Persist intent + idem key
- next2 Provider send
- 2
2 Provider send
- next3 HTTP receipt
- 3
3 HTTP receipt
- next4 HTTP 200?
- ?
4 HTTP 200?
- yes5 Mark accepted
- no5 Terminal error?
- 5
5 Mark accepted
- ?
5 Terminal error?
- yes6 Invalidate token
- no6 Backoff retry same nid
- 7
6 Invalidate token
- next7 Purge registry
- 8
6 Backoff retry same nid
- 9
7 Purge registry
Doze and app standby delay delivery, not this classifier — Android Doze. If the user already got the in-app event, suppress the OS retry while the session is live — WebSockets & MQTT. InvalidProviderToken and ExpiredProviderToken are provider auth: fix credentials and keep device tokens. Only Unregistered and BadDeviceToken purge a token.
Diagrams - step by step
Three small diagrams for push reliability. Step numbers in the labels give the animation order. The lesson map under Diagram 1 plays those steps.
Diagram 1 - Happy path: persist, claim, send, classify
Decisions
- 1
Step 1 Persist intent with nid and idempotency key
- nextStep 2 Worker claims a lease on the nid
- 2
Step 2 Worker claims a lease on the nid
- nextStep 3 Send to the provider
- 3
Step 3 Send to the provider
- nextStep 4 Response class?
- ?
Step 4 Response class?
- 200 plus provider idStep 5a Mark accepted
- 5xx or timeoutStep 5b Backoff with jitter, retry the same nid
- Unregistered or BadDeviceTokenStep 5c Mark the token invalid and purge it
- 5
Step 5a Mark accepted
- 6
Step 5b Backoff with jitter, retry the same nid
- nextStep 3 Send to the provider
- retries exhaustedFailure path - dead-letter for ops review
- 7
Step 5c Mark the token invalid and purge it
- 8
Failure path - dead-letter for ops review
The nid stays the same across retries so dedup and analytics can join on it. Retryable errors back off; device-token errors are terminal for that token. A provider auth error such as InvalidProviderToken means fix your credentials, not purge device tokens.
Lesson map
Reliability, Dedup & Idempotency — Retries, Provider Receipts & Device State
Diagram 1 walks 7 steps from Step 1 Persist intent with nid and idempotency key through Step 5c Mark the token invalid and purge it.
Architecture. Step 1 Persist intent with nid and idempotency key Ready. Step 2 Worker claims a lease on the nid Ready. Step 3 Send to the provider Ready. Step 4 Response class? Ready. Step 5a Mark accepted Ready. Step 5b Backoff with jitter, retry the same nid Ready. Step 5c Mark the token invalid and purge it Ready. Failure path - dead-letter for ops review Ready
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB A["Step 1 Persist intent with nid and idempotency key Ready"] B["Step 2 Worker claims a lease on the nid Ready"] C["Step 3 Send to the provider Ready"] D["Step 4 Response class? Ready"] E["Step 5a Mark accepted Ready"] F["Step 5b Backoff with jitter, retry the same nid Ready"] G["Step 5c Mark the token invalid and purge it Ready"] X["Failure path - dead-letter for ops review Ready"] A -->|continues| B B -->|continues| C C -->|continues| D D -->|200 plus provider id| E D -->|5xx or timeout| F F -->|continues| C D -->|Unregistered or BadDeviceToken| G F -->|retries exhausted| X
Diagram 2 - Failure path: two workers send the same message
Sequence
- 1
Outbound queue → Worker 1
Step 1 deliver message nid 77
- 2
Worker 1 → Provider
Step 2 send is slow, ack not yet written
- 3
Outbound queue → Worker 2
Step 3 visibility timeout expires, nid 77 redelivered
- 4
Worker 2 → Provider
Step 4 second send for the same nid
- 5
Worker 1
Step 5 the user sees two banners
- 6
Worker 1
Fix - claim a lease on nid before sending, collapse id at the provider, client dedup on nid
At-least-once queues redeliver when a worker is slow. A lease on the nid stops the second worker; provider collapse and client dedup catch anything that still slips through.
Diagram 3 - Decision: which dedup layer stops which duplicate
Flow
- 1
1. Ingress idempotency key
- next2. Outbound lease on nid
- 2
2. Outbound lease on nid
- next3. Same nid plus collapse
- 3
3. Same nid plus collapse
- next4. Client nid window
- 4
4. Client nid window
- next5. Do not rely on the provider
- 5
5. Do not rely on the provider
Each layer covers a different source of duplicates. Exactly-once to the eyeball is not guaranteed, so the layers work together.
Sandbox: retry classifier (Python)
Terminal vs retryable vs dead-letter. No network.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same idea (TypeScript): idempotent enqueue
In-memory store + bus. Kafka publish semantics stay in the Kafka cluster. Replay returns the original mapping.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Failure modes
- Retry storms on 4xx BadRequest — fix the payload; do not infinite retry.
- Dual writers without a lease — duplicate critical alerts.
- Treating accept as durable user delivery — still need analytics honesty.
- Ignoring clock / lease expiry — second worker sends after crash recovery race.
- Token reuse across users (rare device bugs) — bind token to
user_id; detect hijack.
Retry rate and DLQ depth are SLIs — observability.
Pitfalls
Worker A claims nid-1, POSTs APNs, times out, dies. Worker B's lease wait ends. What do you send? What id stays stable? What happens if APNs actually accepted A's request? Draw collapse-id on the shade.
Interview Q&A
Is push at-least-once?
Answer
Yes on the wire. Make the user-visible effect idempotent via keys and collapse. Do not promise exactly-once to the eyeball.
Why keep nid stable across retries?
Answer
Analytics and client dedup join on one id. provider_message_id may change per attempt — store each receipt next to the same nid.
APNs 410 / Unregistered?
Answer
Purge the token immediately. Stop retrying that token. Fan-out to remaining active tokens for the user.
Exactly-once to the eyeball?
Answer
Not guaranteed. Design for at-most-one visible via collapse and idempotency. OS policy can still drop or delay.
What is the dead-letter for?
Answer
Poison payloads and exhausted retries. Ops inspect. Do not silent-drop forever, and do not infinite-retry 4xx BadRequest.
Where does Kafka fit?
Answer
Outbound topic acks and consumer retries. Semantics: Kafka delivery.
Where do idempotency keys fit?
Answer
Key design and store semantics: idempotency keys. This page applies them to push intent and worker claims.
In-app vs OS on retry?
Answer
If the user already got the WebSocket/MQTT event and the session is still live, suppress the OS retry. Depth: realtime hub.
Worker crash after a 200 you never saw?
Answer
Lease reclaim + same nid + collapse. You may send twice; the shade should still show one. Join both provider ids in analytics.
Suspect vs invalid?
Answer
Suspect is a cohort smell (no opens, soft failures) — investigate, do not mass-delete. Invalid is a device-token error (Unregistered, BadDeviceToken) — purge that token. Provider auth errors such as InvalidProviderToken and ExpiredProviderToken mean fix credentials and keep device tokens.