System design
Part 1 of 6 · Push NotificationsPush Notifications Architecture — APNs, FCM & Event-Driven Alerts
Mobile push is a distributed system: token registries, APNs/FCM provider APIs, priority and collapse, rich-content extensions, abuse-safe critical alerts, and a funnel from send to engagement. This hub maps the cluster. In-app realtime and durable buses are cousins — cross-link, do not re-teach.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
How you reach a phone that may be asleep
Prefer
OS push when backgrounded; in-app channel when live; durable bus for fan-out
APNs/FCM wake a killed app and own the shade. A connected session prefers WebSocket/MQTT. Kafka (or equivalent) holds the outbound intent so retries are durable — the provider message id is a receipt, not the product identity.
- Token registry + invalidation is the delivery control plane.
- Collapse keys coalesce stale state; notification_id joins analytics.
- No PII in payloads — opaque ids, fetch details after open.
Alternative
Always FCM, or a WebSocket that you pretend is push
Looks simple on a whiteboard. Killed apps never see the socket. Data-only payloads get deferred. Dual-send without collapse double-banners. Treating provider accept as 'the user saw it' lies to execs.
- Foreground-only sockets do not replace APNs/FCM.
- Everything 'high priority' burns OS and provider trust.
- Retries without an idempotency key spam the shade.
Failure path this cluster exists to name
Each hop is a later lesson. Interviews start at the silent drop and the duplicate banner, not at FCM trivia.
- 1
Enqueue a domain event
Persist notification_id before the provider call. Depth: reliability + Kafka delivery (acks only). - 2
Lookup tokens, check idempotency
Rotate, purge Unregistered. Same business key must not create a second banner. Depth: FCM tokens; idempotency keys. - 3
Send APNs or FCM
Priority, collapse, platform overrides. Depth: FCM fan-out; critical interruption levels. - 4
OS presents — or NSE times out
mutable-content is a best-effort enricher. Depth: NSE. - 5
Join the funnel honestly
Accept ≠ delivered ≠ opened. Depth: analytics + SLIs.
Overview
Interview prompt: design mobile push for a consumer app. Seniors are graded on tokens, provider semantics, and the funnel — not on pasting an FCM snippet.
OS push (APNs / FCM) wakes a backgrounded or killed app and owns banners, sounds, and channels. In-app realtime (WebSockets & MQTT) is live only while a session is open — lower latency for chat ticks, no OS banner unless you also local-notify. A durable outbound bus (Kafka delivery) is how you retry without losing the intent. Do not re-teach handshake frames, QoS, or ALO/EOS here.
This hub is the map. The five sibling pages are the whiteboard depth.
You should be able to:
- Draw compose → provider → OS → open, with token purge and a funnel dotted line.
- Say when you would refuse OS push (app already connected — prefer in-app).
- Contrast APNs direct vs FCM in one sentence — without folding NSE, critical entitlements, or exactly-once theory into this page.
OS push vs in-app realtime
| Path | Wins | Loses |
|---|---|---|
| OS push (APNs/FCM) | Wakes background/killed; OS-owned UX | Token churn; OS policy drops; not 100% telemetry |
| In-app WebSocket/MQTT | Live ticks while connected | No banner if the app is dead — realtime hub |
| Data-only payload | App-handled, quieter | May be deferred or killed (Doze / app death) |
| Alert / notification key | Requests user-visible presentation | Abuse trains users to disable everything |
| High vs normal priority | High wakes radios sooner | Abuse burns battery quotas and provider trust |
Rule of thumb: app foreground + connected → prefer in-app; background/killed → OS push; durable fan-out / audit → Kafka.
Architecture (two columns would shrink labels)
Split the path so the desktop SVG stays a single column. Compose is one chart; the device and funnel are another.
Outbound compose — identity is notification_id + idempotency key, not the provider message id:
Flow
- 1
1 Domain event / API
- next2 Outbound queue
- 2
2 Outbound queue
- next3 Dedup + personalize
- 3
3 Dedup + personalize
- next4 Token registry
- 4
4 Token registry
- next5 Idempotency store
- 5
5 Idempotency store
- next6 APNs HTTP/2 or FCM HTTP v1
- 6
6 APNs HTTP/2 or FCM HTTP v1
Device and funnel — NSE is best-effort; analytics joins on notification_id:
Flow
- 1
1 Provider accept receipt
- next2 OS shade / NSE
- next5 Invalid token purge
- 2
2 OS shade / NSE
- next3 User open / action
- 3
3 User open / action
- next4 Displayed opened engaged
- 4
4 Displayed opened engaged
- 5
5 Invalid token purge
Diagrams - step by step
Three small diagrams for the push notifications hub. Step numbers in the labels give the animation order. The lesson map under Diagram 1 plays those steps.
Diagram 1 - Happy path: event to banner
Decisions
- 1
Step 1 Domain event enqueued with a notification_id
- nextStep 2 Check the idempotency key
- 2
Step 2 Check the idempotency key
- nextStep 3 Look up device tokens and user prefs
- 3
Step 3 Look up device tokens and user prefs
- nextStep 4 Platform?
- ?
Step 4 Platform?
- iOSStep 5a Send to APNs over HTTP/2
- Android or webStep 5b Send to FCM HTTP v1
- 5
Step 5a Send to APNs over HTTP/2
- nextStep 6 Provider accepts - record accepted
- 410 UnregisteredFailure path - purge the token, stop retrying it
- 6
Step 5b Send to FCM HTTP v1
- nextStep 6 Provider accepts - record accepted
- UNREGISTEREDFailure path - purge the token, stop retrying it
- 7
Step 6 Provider accepts - record accepted
- nextStep 7 Device shows it and the app logs the open
- 8
Step 7 Device shows it and the app logs the open
- 9
Failure path - purge the token, stop retrying it
The notification service owns idempotency, token lookup and policy; APNs and FCM only deliver. Provider accept is recorded as accepted, not as seen. Terminal token errors come back in the send response and must purge the token.
Lesson map
Push Notifications Architecture — APNs, FCM & Event-Driven Alerts
Diagram 1 walks 8 steps from Step 1 Domain event enqueued with a notification_id through Step 7 Device shows it and the app logs the open.
Architecture. Step 1 Domain event enqueued with a notification_id Ready. Step 2 Check the idempotency key Ready. Step 3 Look up device tokens and user prefs Ready. Step 4 Platform? Ready. Step 5a Send to APNs over HTTP/2 Ready. Step 5b Send to FCM HTTP v1 Ready. Step 6 Provider accepts - record accepted Ready. Step 7 Device shows it and the app logs the open Ready. Failure path - purge the token, stop retrying it Ready
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB A["Step 1 Domain event enqueued with a notification_id Ready"] B["Step 2 Check the idempotency key Ready"] C["Step 3 Look up device tokens and user prefs Ready"] D["Step 4 Platform? Ready"] E["Step 5a Send to APNs over HTTP/2 Ready"] F["Step 5b Send to FCM HTTP v1 Ready"] G["Step 6 Provider accepts - record accepted Ready"] H["Step 7 Device shows it and the app logs the open Ready"] X["Failure path - purge the token, stop retrying it Ready"] A -->|continues| B B -->|continues| C C -->|continues| D D -->|iOS| E D -->|Android or web| F E -->|continues| G F -->|continues| G G -->|continues| H E -->|410 Unregistered| X F -->|UNREGISTERED| X
Diagram 2 - Failure path: retry without idempotency shows two banners
Sequence
- 1
Worker → APNs or FCM
Step 1 send the order shipped alert
- 2
APNs or FCM → Device
Step 2 provider delivers it
- 3
APNs or FCM → Worker
Step 3 the response times out on the way back
- 4
Worker → APNs or FCM
Step 4 retry with a new id and no collapse id
- 5
APNs or FCM → Device
Step 5 a second banner for the same event
- 6
Worker
Fix - stable notification_id plus idempotency key, and apns-collapse-id or FCM collapse_key
Push is at-least-once on the wire. A lost response looks like a failure, so the worker retries. Keeping one notification_id per business event and setting a collapse id makes the retry harmless.
Diagram 3 - Decision: OS push vs in-app channel, payload and priority
Decisions
- ?
Step 1 App in foreground and connected?
- yesIn-app WebSocket or MQTT
- no - background or killedStep 2 Must the user see it?
- 2
In-app WebSocket or MQTT
- ?
Step 2 Must the user see it?
- yesAlert payload via APNs or FCM
- no - sync hintData-only payload, may be deferred
- 4
Alert payload via APNs or FCM
- nextStep 3 Truly urgent?
- 5
Data-only payload, may be deferred
- ?
Step 3 Truly urgent?
- yesHigh priority, time-sensitive if allowed
- noNormal priority plus collapse key
- 7
High priority, time-sensitive if allowed
- Wrong pick for marketingBattery cost, throttling, users disable push
- 8
Normal priority plus collapse key
- 9
Battery cost, throttling, users disable push
This is the hub rule of thumb: in-app channel when the session is live, OS push when it is not. Alert payloads request display; data-only can be delayed by the OS. High priority is for real urgency only.
Depth: NSE, FCM fan-out, reliability, analytics. Ingress retries use idempotency keys — do not recap key stores here.
Sandbox: token registry (Python)
Educational in-memory registry. Upsert on launch / token refresh. Purge on provider Unregistered. The Doc sketch mangled __init__ and un-indented methods — this sandbox is a complete, runnable class.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same idea (TypeScript)
Enqueue is idempotent. Kafka/queue publish is an in-memory list — delivery semantics stay in the Kafka cluster. Record<string, string> replaces the Doc's mangled Record. String concat so MDX does not interpolate.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Privacy baseline
No email, phone, or message-body secrets in APNs aps.alert or FCM data. Opaque nid and type codes. Fetch details in-app after open with a real session. Hash tokens in logs (or last-4). Signed media URLs for NSE stay short-lived.
Cluster map
| Order | Page | You defend |
|---|---|---|
| 1 | This hub | OS push vs in-app; token registry; compose pipeline |
| 2 | NSE | mutable-content, time budget, fail-open |
| 3 | FCM fan-out | Tokens vs topics, collapse, priority |
| 4 | Critical alerts | Interruption levels, channels, abuse |
| 5 | Analytics | queued → accepted → delivered → opened |
| 6 | Reliability | Retries, receipts, device state |
Pitfalls
User places an order on the website. The iOS app is force-quit. Draw event → queue → token lookup → APNs → NSE → banner → open → funnel. Where does Unregistered land? Where does a retry with the same idempotency key land? If the app were foreground with a live socket, what do you suppress?
Interview Q&A
When OS push vs WebSocket?
Answer
Background or killed → OS push (APNs/FCM). Foreground live UI → WebSocket/MQTT. Often both, with a preference policy that suppresses the OS banner while the session is live. Depth: realtime protocols — do not recap ping/pong here.
Why does token churn break delivery?
Answer
Tokens rotate on reinstall and OS update. Stale tokens produce silent provider rejects until you purge Unregistered / NotRegistered / BadDeviceToken. Multi-device users need fan-out to every active token.
Data vs notification payload?
Answer
Notification keys request OS display. Data-only is app-handled and may be delayed or dropped when Doze or the process is dead. Quiet ticks can be data; user-visible alerts should request presentation.
What is a collapse key for?
Answer
Coalesce stale updates (latest score, latest order status) so a queued burst does not spam the shade. Collapse is not your analytics identity — that is notification_id.
Where does Kafka fit?
Answer
Durable outbound intent and analytics streams. Acks, consumer retries, and EOS live in Kafka delivery — cross-link only.
How do you make push idempotent?
Answer
Same business key (order + template, or an explicit client key) must not create a second user-visible alert after retries. Store the mapping to notification_id. Depth: idempotency keys and reliability.
Privacy baseline for payloads?
Answer
No PII. Opaque ids. Fetch details in-app after open. Hash tokens in logs. Short-lived signed URLs for rich media.
What SLIs would you put on this funnel?
APNs direct vs FCM for iOS?
Answer
FCM can pass through to APNs. Many stacks still speak APNs HTTP/2 direct for headers and error control. Depth: FCM fan-out.
What if NSE never runs?
Answer
Timeout or crash → OS presents the original payload. Design the alert to be useful without the image. Depth: NSE.