EDA Failure Modes — Dual Writes, Schema Drift & Fan-out Blast Radius
Event-driven systems fail in known ways: dual-write loss or ghosts, schema changes that break a fleet, and one poison event amplified across every consumer. Containment is outbox, compatible rollout, and per-consumer isolation.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What are the top three event-driven failure modes?
Answer
Dual-write loss or ghosts, an incompatible schema rollout, and poison or fan-out amplification.
L2
How do you detect a lost event in production?
Answer
Outbox age and relay lag. Reconciliation compares the database, the log, and projections. CDC can audit the commit stream.
L3
Does a schema registry save you by itself?
Answer
No. You still need deploy order, tolerant readers, and a DLQ for payloads the registry did not imagine.
L4
What is a tolerant reader?
Answer
It accepts unknown fields, reads old and new names during a rename, and defaults a missing optional field.
L5
How do you keep one bad email template from looking like a platform outage?
Answer
Give email its own consumer and its own DLQ. Search and ledger keep consuming. Alert on that DLQ, then replay after the template fix.
L6
How do sagas fail when events fail?
Answer
A missing event stalls the state machine. A duplicate without idempotency double-compensates. The saga lesson owns the state machine. The outbox and inbox own the message.
L7
What is a useful blast-radius drill?
Answer
Inject a poison event in staging and show that one consumer dead-letters while the others stay inside their lag SLO.
Failure modes
Dual-write loss
The database has the row and no subscriber ever sees an event.
Ghost event
An event is published for a transaction that did not commit.
Schema break
A field change makes deserializers throw across every consumer of that stream.
Fan-out amplification
One bug is multiplied by every handler subscribed to the event.
Projection skew
Screens disagree because of lag or because older events applied out of version order.
Poison head-of-line block
Infinite retry pins lag to one partition.
Saga and event mismatch
Missing or duplicate events make compensations fire late or twice.
Invalidation storm
A burst of events deletes cache keys with no soft TTL and the origin falls over.
Misconceptions
The broker is the blast-radius boundary.
Consumer isolation, per-handler DLQs, and kill switches contain the damage. The topic only delivered the event.
A schema registry means rollouts cannot break readers.
The registry checks a compatibility mode. Deploy order, tolerant readers, and a DLQ still matter.
Interviewer traps
Listing Kafka ISR settings as the failure-mode answer.
Stay on dual-write, schema, and fan-out. Point at delivery semantics for poison and at the outbox lesson for the lost event.
Sending the interviewer to a saga failure page to avoid the event question.
Say what the missing event does to the saga, then point at the sagas hub for the state machine.
Design scenario
Same prompt for every reader.
Requirements
Search and ledger stay inside their lag SLO while email is fixed. The bad messages can be replayed after the template change. No dual-write on the producer.
Traffic / scale
About 5k payments per second at peak, fanned out to four consumers.
Latency
Ledger freshness stays under a few seconds. Email may pause. Checkout does not wait on email.
Consistency
PaymentCaptured is produced from an outbox row. Each consumer dedupes. Projections apply by version.
Availability
Email can fail closed into its own DLQ. The producer and the other consumers stay up.
Failure assumptions
- Two percent of email handlers throw.
- A schema change can ship to the producer before every consumer is tolerant.
- Cache invalidation subscribers can stampede origin if every event deletes keys at once.
Constraints
- Do not share one consumer thread across email and ledger.
- Do not drop the poison payload.
- Do not commit the payment and publish in two steps.
Prompt
Checkout emits PaymentCaptured. Search, email, ledger, and fraud subscribe. A bad email template throws on 2 percent of messages.
API
What does checkout return, and which consumer is allowed to be down?
Data
What does the DLQ record store so email can replay after the template fix?
Architecture
Where are the consumer boundaries, the kill switch, and the outbox?
One bad email template is not a down bus
Prefer
Per-consumer failure domains
Search, ledger, fraud, and email each have a handler, a lag SLO, and a DLQ. A template bug parks email. The others keep up. A kill switch stops the handler without stopping the producer.
- Checkout already returned from the write path.
- The poison payload is retained for replay.
- Deserializer errors page the schema owner, not every team.
Alternative
One consumer does every side effect
The throw in email stops the loop that also updates search and the ledger. The dashboard says events are down.
- Blast radius equals the fan-out you were proud of.
- Retries pin the shared partition.
- A rollback of the producer is the only kill switch you have.
Containment, not a new broker
Each failure already has a lesson. This page is the map of which one.
- 1
Make publish a consequence of the commit
Outbox or CDC. Dual-write loss and ghosts are the same missing boundary. - 2
Change schemas additively
Compatible checks in CI, a canary producer, tolerant readers, then the old field goes away. - 3
Split the fan-out
One group or subscription per handler. One DLQ per handler. A flag to disable a handler. - 4
Watch the disagreements
Outbox age, deserializer errors, per-consumer lag, projection freshness, DLQ depth. - 5
Drill the poison event
In staging, prove the other consumers stay inside their SLO.
Overview
Event-driven systems fail in predictable ways. Lost or ghost events come from dual-writes. Schema drift breaks a fleet of consumers. Fan-out multiplies one bad event across projections, email, and partners.
This page is the interview catalog. The fixes point backward:
- Dual-write: Transactional Outbox, Inbox & Consumer Idempotency and Transactional Outbox & Inbox Patterns.
- Commit stream as an audit: Change Data Capture.
- Poison and retry budgets: Ordering, Partitions, Poison Messages & Retry/DLQ Strategy and Kafka delivery semantics.
- The log and consumer groups: Apache Kafka — Topics, Partitions, Brokers & Consumer Groups.
- Multi-step undo: Sagas & Distributed Transactions.
- Cache storms: Cache Invalidation — TTL vs Event-Driven vs Versioned Keys.
Do not re-teach those pages. Name the symptom, the cause, and the lesson.
Failure catalog
| Failure | Symptom | Cause | Mitigation |
|---|---|---|---|
| Dual-write loss | The row exists and no event does | Commit, then produce, then crash | Outbox or CDC |
| Ghost event | An event with no durable state | Produce, then crash before commit | Outbox. Consumers tolerate a missing aggregate |
| Schema break | Deserializers throw | An incompatible field change | Compatibility checks, versioned contracts, DLQ |
| Fan-out amplification | One bug times N consumers | A shared poison event | Contract tests, staged rollout, a DLQ per consumer |
| Projection skew | Screens disagree | Lag, or an unordered apply | Version gates and a freshness SLO |
| Poison head-of-line | Lag climbs on one partition | Infinite retry | Bounded retries, DLQ, or a key park |
| Saga mismatch | Compensations fire oddly | A missing event or a duplicate | Outbox plus inbox. Undo stays on the saga hub |
| Invalidation storm | Origin meltdown after a burst of deletes | Event-driven invalidation with no soft TTL | The cache-invalidation lesson |
Dual-write, then the derived publish
Sequence
- 1
Service → DB
1. Commit with no outbox row
- 2
Service → Broker
2. Produce fails
- 3
DB
The row exists. Subscribers never hear it.
- 4
Service → DB
3. Commit the business row and the outbox row
- 5
Service → Broker
4. The relay publishes
- 6
Broker → Consumer
5. Deliver
- 7
Consumer
6. The inbox dedupes before the side effect.
Lesson map
Dual write, schema, and one throwing handler
The row and the outbox are committed. The broker has the fact. This consumer has an inbox key.
Architecture. Service Committed. Broker Has the fact. Consumer Deduped
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB service["Service Committed"] broker["Broker Has the fact"] consumer["Consumer Deduped"] service -->|Outbox relay| broker broker -->|Deliver| consumer consumer -->|This handler| broker service -->|Bad schema| consumer service -->|No outbox| broker
Interview line: publish is a derived effect of a committed outbox row, or of the commit log. It is not a second write in the request.
Change Data Capture is how you audit "the commit happened" when you do not trust the application to write an outbox row. It does not replace a domain event that is not the row image.
Schema drift
More than one consumer team means compatibility is not optional.
- A schema check in CI. The registry product is a Kafka concern. The public compatibility rules are in the Confluent Schema Registry docs linked below. This page does not re-teach wire formats. Kafka topics remains the log lesson, and delivery semantics remains the poison lesson when a deserializer throws.
- Additive changes first. Deprecate a field over a window.
- Tolerant readers ignore unknown fields.
- Expand, then contract for a rename: publish both names during the overlap, move readers, then drop the old name.
Decisions
- 1
1. Propose a schema change
- next2. Does CI accept the compatibility rule
- ?
2. Does CI accept the compatibility rule
- no3. Block the merge
- yes4. Canary the producer
- 3
3. Block the merge
- 4
4. Canary the producer
- next5. Watch the deserializer error rate
- 5
5. Watch the deserializer error rate
- next6. Do errors spike
- ?
6. Do errors spike
- yes7. Roll back and triage the DLQ
- no8. Finish the rollout
- 7
7. Roll back and triage the DLQ
- 8
8. Finish the rollout
A green registry check is not a finished rollout. Consumers that have not deployed the tolerant reader still throw. Those records belong on a DLQ with the payload intact, which is the previous lesson's poison path.
Fan-out blast radius
One OrderPlaced may wake search, email, fraud, analytics, partner webhooks, and a cache invalidator.
| Control | What it contains |
|---|---|
| A consumer boundary per handler | Lag and crashes stay local |
| A DLQ per handler | One bad handler does not freeze the others |
| A feature flag on the handler | A kill switch that does not stop the producer |
| Contract tests on golden events | Breaks show up before production |
| Quotas on webhooks | Partners do not take your whole retry budget |
The broker is not the blast-radius boundary. It delivered one fact to many groups. Isolation and kill switches contain the damage.
Cache subscribers are a special case of fan-out. A burst of invalidations without a soft TTL or a single-flight fill is an origin incident. Leave the patterns on Cache Invalidation — TTL vs Event-Driven vs Versioned Keys.
Tolerant reader (run this)
The v2 payload renamed the total field and kept a currency. The brittle reader assumes totalCents forever. The tolerant reader accepts either name and defaults currency.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Tolerant reads are not an excuse to rename a field every week. They are how the overlap window stays up.
One handler throws (run this)
Email throws. Search and analytics still return in the success list. The dead-letter record names the handler and the event id. A shared loop without this isolation would have stopped on email and never called analytics.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Production isolation is a separate consumer, not only a try/catch in one process. The sketch shows the policy: a failure is recorded for that handler, and the others proceed.
The email drill
Checkout emits PaymentCaptured. Search, email, ledger, and fraud subscribe. A bad template throws on 2 percent of messages.
Without a per-consumer DLQ, email lag climbs and the incident channel says events are down. With a separate consumer, a DLQ, and an alert, search and ledger stay inside their SLO. Email fixes the template and replays the parked ids. Checkout never waited on the template.
If those missing or duplicate events also drive a reserve-charge-ship saga, the state machine stalls or double-compensates. Say that, then hand the state machine to Sagas & Distributed Transactions. The message boundary stays the outbox and the inbox.
Interview Q&A
What are the top three event-driven failure modes?
Answer
Dual-write loss, an incompatible schema rollout, and poison or fan-out amplification.
How do you detect dual-write loss in production?
Answer
Outbox age and relay lag. Reconciliation compares database truth with the event log and the projections. CDC helps when you need the commit stream as an audit.
Can a schema registry alone save you?
Answer
No. You still need deploy order, tolerant readers, and a DLQ for unexpected payloads. A registry checks a rule. It does not restart old consumers.
How do sagas fail when events fail?
Answer
A missing event stalls the machine. A duplicate without idempotency can compensate twice. The message fix is the outbox and inbox. The state machine is the sagas hub.
What is a good blast-radius drill?
Answer
Inject a poison event in staging. Prove one consumer dead-letters and the others stay inside their lag SLO.
Why did the cache fall over after a healthy publish?
Answer
Every event deleted keys, and the origin refilled them at once. That is an invalidation storm. The controls live on cache invalidation.
What do you page on first?
Answer
Outbox age, deserializer error rate, DLQ depth per handler, and projection freshness. A single cluster-wide consumer lag hides which handler is actually stuck.
Ghost event or lost event?
Answer
Lost: the row exists, the event does not. Ghost: the event exists, the row does not. Both are a dual-write. The outbox removes the window. Consumers should still tolerate a missing aggregate, because a ghost can already be in the log from before the fix.
Pitfalls
Label columns lost event, bad schema, and poison email. Under each, write the user-visible symptom, the metric, and the sibling lesson you would open. If a column tempts you to redesign consumer groups, write the Kafka lesson title and stop.