System design
Part 3 of 6 · Push NotificationsFCM & Multi-Platform Fan-out — Tokens, Topics, Collapse Keys & Priority
FCM is a common fan-out plane for Android (and often a pass-through toward APNs/web). Interviews test token lifecycle, topics vs device multicast, collapse keys, priority, and how your service abstracts APNs vs FCM behind one outbound pipeline.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
How you address devices
Prefer
Token path for account-critical; topics for coarse broadcast
Receipts, security, and OTP-adjacent UX go to the user's active tokens. Sports scores and ops incidents can ride topics. HTTP v1 carries android / apns / webpush overrides so one compose pipeline still speaks platform policy.
- Upsert tokens on launch and refresh; soft-delete Unregistered.
- Collapse per device so latest state wins in the shade.
- Never log raw tokens — hash or last-4.
Alternative
Topics for secrets, or everything high priority
Subscribe races leak the wrong cohort. High-everywhere trains Doze and users to disable the app. Dual-path APNs + FCM without collapse double-banners.
- Topics are not a per-user ACL.
- Normal priority will batch under Doze — design for delay.
- Partial multicast failure must retry only the failed tokens with the same nid.
Fan-out without a side-by-side platform row
Split token send and invalidation so desktop SVG labels stay 16px.
- 1
Compose once
notification_id, collapse key, priority, channel_id / interruption. Platform blocks in HTTP v1. - 2
OAuth to FCM HTTP v1
Service account. Legacy APIs are deprecated for a reason. - 3
Token or topic target
Transactional → tokens. Broadcast → topic/condition. Online session → in-app only. - 4
Partial failure
Multicast: retry only Unregistered-or-5xx tokens. Same nid. - 5
Purge terminal errors
NotRegistered / BadDeviceToken → invalidate. Depth: reliability.
Overview
FCM is a transport and policy adapter. Your product identity remains notification_id. Provider message ids are receipts, not substitutes.
Interviews test token lifecycle, topic messaging vs device multicast, collapse keys / event ids, and priority — plus how the service abstracts APNs vs FCM behind one outbound pipeline.
You should be able to:
- Draw compose → FCM HTTP v1 → Android / optional APNs / Web Push as separate single-column paths.
- Say when you would refuse a topic (per-user secrets).
- Contrast collapse key vs
nidin one sentence.
Fan-out model
| Mechanism | Wins | Loses |
|---|---|---|
| Device tokens | Precise; natural for transactional | Registry scale; rotate; invalidate |
| Topics | Cheap broadcast | Weak personalization; subscribe races |
| Conditions | Boolean topic combos for segments | Still not a secret channel |
| Multicast / batch | Fan-out to a token list | Partial failure semantics |
| Platform overrides | android / apns / webpush in HTTP v1 | Easy to dual-send if you also call APNs direct |
Architecture (split so platforms are not a row)
Token send — one compose, FCM HTTP v1, Android first:
Flow
- 1
1 Notification service
- next2 FCM HTTP v1 + OAuth
- 2
2 FCM HTTP v1 + OAuth
- next3 Android high/normal + collapse
- 3
3 Android high/normal + collapse
- next4 channel_id maps importance
- 4
4 channel_id maps importance
Optional APNs / Web and invalidation — do not draw three sibling sinks in one rank:
Flow
- 1
1 FCM project
- next2 Optional apns override
- 2
2 Optional apns override
- next3 Web Push / VAPID
- 3
3 Web Push / VAPID
- next4 Unregistered / BadDeviceToken
- 4
4 Unregistered / BadDeviceToken
- next5 Purge token registry
- 5
5 Purge token registry
Diagrams - step by step
Three small diagrams for FCM fan-out. Step numbers in the labels give the animation order. The lesson map under Diagram 1 plays those steps.
Diagram 1 - Happy path: token or topic send through FCM HTTP v1
Decisions
- 1
Step 1 Service gets an OAuth token for its service account
- nextStep 2 Audience?
- ?
Step 2 Audience?
- one userStep 3a Send to each device token
- broadcast segmentStep 3b Send to a topic or condition
- 3
Step 3a Send to each device token
- nextStep 4 Set priority, TTL, collapse key, platform overrides
- send returns UNREGISTEREDFailure path - purge the token from the registry
- 4
Step 3b Send to a topic or condition
- nextStep 4 Set priority, TTL, collapse key, platform overrides
- 5
Step 4 Set priority, TTL, collapse key, platform overrides
- nextStep 5 FCM fans out to Android, APNs and Web Push
- 6
Step 5 FCM fans out to Android, APNs and Web Push
- nextStep 6 Store the provider message id next to the nid
- 7
Step 6 Store the provider message id next to the nid
- 8
Failure path - purge the token from the registry
HTTP v1 uses OAuth and per-platform override blocks in one request. Tokens give per-user precision; topics give cheap broadcast. Invalid tokens are reported in the send response (UNREGISTERED), not by a later callback.
Lesson map
FCM & Multi-Platform Fan-out — Tokens, Topics, Collapse Keys & Priority
Diagram 1 walks 7 steps from Step 1 Service gets an OAuth token for its service account through Step 6 Store the provider message id next to the nid.
Architecture. Step 1 Service gets an OAuth token for its service account Ready. Step 2 Audience? Ready. Step 3a Send to each device token Ready. Step 3b Send to a topic or condition Ready. Step 4 Set priority, TTL, collapse key, platform overrides Ready. Step 5 FCM fans out to Android, APNs and Web Push Ready. Step 6 Store the provider message id next to the nid Ready. Failure path - purge the token from the registry Ready
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB A["Step 1 Service gets an OAuth token for its service account Ready"] B["Step 2 Audience? Ready"] C["Step 3a Send to each device token Ready"] D["Step 3b Send to a topic or condition Ready"] E["Step 4 Set priority, TTL, collapse key, platform overrides Ready"] F["Step 5 FCM fans out to Android, APNs and Web Push Ready"] G["Step 6 Store the provider message id next to the nid Ready"] X["Failure path - purge the token from the registry Ready"] A -->|continues| B B -->|one user| C B -->|broadcast segment| D C -->|continues| E D -->|continues| E E -->|continues| F F -->|continues| G C -->|send returns UNREGISTERED| X
Diagram 2 - Failure path: stale tokens never purged
Sequence
- 1
Notification service → FCM HTTP v1
Step 1 send to 1M tokens, 15 percent are stale
- 2
FCM HTTP v1 → Notification service
Step 2 404 UNREGISTERED for the stale tokens
- 3
Notification service → Token registry
Step 3 bug - error logged, token kept as active
- 4
Notification service → FCM HTTP v1
Step 4 next campaign resends to the same dead tokens
- 5
Notification service
Step 5 wasted quota, skewed accept rate, slower fan-out
- 6
Notification service
Fix - on UNREGISTERED delete the token from the registry right away
Tokens rotate on reinstall and OS updates. If terminal errors do not update the registry, every send pays for dead tokens and the accept-rate metric quietly drifts.
Diagram 3 - Decision: topics vs tokens, collapse vs distinct
Decisions
- ?
Step 1 Personalized or secret content?
- yesPer-user device tokens
- no - same for manyStep 2 Stable interest group?
- 2
Per-user device tokens
- nextStep 3 Only the latest state matters?
- ?
Step 2 Stable interest group?
- yesTopic or condition send
- noHybrid - topic for blasts, tokens for account alerts
- 4
Topic or condition send
- Wrong pick for per-user secretsEvery subscriber gets the same payload
- 5
Hybrid - topic for blasts, tokens for account alerts
- ?
Step 3 Only the latest state matters?
- yesSet a collapse key
- noSend distinct notifications
- 7
Set a collapse key
- 8
Send distinct notifications
- 9
Every subscriber gets the same payload
Topics are for coarse broadcasts, tokens for anything personal. A collapse key replaces pending messages with the newest one, which suits scores and status, not chat messages.
Queue + worker pool at scale. Durable outbound is Kafka delivery — acks only, no EOS recap.
Topics vs per-user send
- Topics — cheap broadcast (sports score, ops incident); hard to personalize; subscribe races.
- Per-user tokens — precise; need registry scale; transactional (receipt, OTP-adjacent UX — still rate-limit).
- Hybrid — topic for marketing blast; token path for account-critical.
Connected session → in_app_only. That preference lives next to WebSockets & MQTT.
Collapse keys and priority
- Collapse / event id — replace pending notifications of the same key (latest state wins).
- Android priority high vs normal — high may wake; normal can batch under Doze.
- APNs priority 10 vs 5 — similar wake tradeoff; misuse → throttling / user disable.
- TTL / time_to_live — expire stale before delivery (flash sale ended).
- Collapse ≠ nid — coalesces pending; nid is analytics/idempotency identity.
Token registry ops
- Upsert on every app launch / token refresh callback.
- Soft-delete on Unregistered; hard-delete after a retention window.
- Multi-device users: fan-out to all active tokens; collapse per device.
- Web tokens expire differently — refresh via service worker push subscription.
- Never log raw tokens in analytics — hash or last-4 for debug only.
Android channels (quick map)
Channel importance ≈ user-visible interruptiveness. Created in-app, not by FCM alone. Mis-mapped channel_id falls to default importance; users then disable broadly. Critical-style UX uses high-importance channels + full-screen intents carefully — policy on the critical alerts page.
Sandbox: FCM-shaped message (Python)
Opaque ids only. Platform overrides in one dict so you can talk through HTTP v1 without a network.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same idea (TypeScript): topic vs token router
If the user is in-app, skip OS push. Broadcast scores go to a topic. Everything else looks up a token at runtime (registry is the hub sandbox).
Press Run. Snippets must be self-contained — no network, files, or native modules.
Failure modes
- Token churn — reinstall creates a new token; old one 404s until purged.
- Partial multicast failure — some tokens succeed; retry only failures with the same nid.
- Topic fan-out storms — subscribe/unsubscribe lag → wrong cohort briefly.
- Priority inflation — everything "high" → OS + provider degrade trust.
- Dual-send APNs direct + FCM APNs — duplicate banners unless idempotent collapse.
Pitfalls
You send APNs HTTP/2 direct for iOS and FCM for Android. A new teammate also sets FCM apns overrides. What does the iPhone shade show? Name collapse-id, nid, and the idempotency key you would have used.
Interview Q&A
Why HTTP v1 over legacy?
Answer
OAuth service accounts, platform overrides, and a better error model. Legacy FCM APIs are deprecated.
Collapse key vs notification id?
Answer
Collapse coalesces pending shade rows so latest state wins. nid is your analytics and idempotency identity and stays stable across retries.
When topics?
Answer
Coarse broadcasts. Not per-user secrets, receipts, or security alerts.
FCM for iOS?
Answer
Optional. Many stacks still speak APNs direct for header and error control. If you use both, you will double-banner without collapse/idempotency.
How does Doze change the design?
Answer
Normal priority is deferred. High is still constrained. Design for delay: TTL, collapse, and in-app catch-up when the session opens.
How is Web Push different?
Answer
VAPID + service worker. Permission UX is distinct. Subscriptions expire on a different clock than APNs/FCM device tokens.
How do you scale fan-out?
Answer
Queue + worker pool. Durable outbound is Kafka delivery — cross-link only.
In-app vs FCM?
Answer
Connected session → WebSocket/MQTT. Else FCM/APNs. Depth: realtime hub.
What is channel_id?
Answer
Android channel created in-app. Importance is user-adjustable. Wrong id falls through to default and trains a global disable.
Partial multicast — retry strategy?
Answer
Retry only failed tokens, same nid, backoff. Terminal Unregistered → purge, do not retry that token. Depth: reliability.