Event-Driven Architecture — Sync vs Events, Patterns & Tradeoffs
Event-driven architecture publishes facts that already happened. This hub maps sync versus events, notification versus carried state versus sourcing, and where CQRS, the outbox, ordering, and failure modes sit.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
When does a sync checkout win over emitting OrderPlaced?
Answer
When the caller needs one authoritative answer now and the fan-out is tiny, such as an authz check, a price quote, or a hold confirmation.
L2
What is the payload tradeoff between a notification and event-carried state transfer?
Answer
A notification stays small and forces a fetch. A fat event lets the consumer act without a callback, and it duplicates state and schema.
L3
Why does a database commit followed by a Kafka produce lose events?
Answer
Those are two systems with no shared commit. A crash after the commit and before the produce drops the fact. That is a dual-write.
L4
How does the transactional outbox fix a dual-write?
Answer
The same local transaction writes the business row and an outbox row. A relay publishes the outbox row later. The event cannot vanish if the commit succeeded.
L5
What do you tell a product team about CQRS read-model lag?
Answer
The write path is the source of truth. The read path has a freshness SLO. Screens that need read-your-writes use the write model or wait for a version.
L6
How do you get per-order ordering without global order?
Answer
Partition or shard by orderId. That serializes one order and lets other orders run in parallel. Cross-partition views stay unordered.
L7
How do you contain fan-out when one bad schema ships?
Answer
Isolate consumers, give each a DLQ and a kill switch, and block incompatible changes in CI. One handler must not freeze the others.
Failure modes
Dual-write loss
The database commit succeeds and the later publish never happens, so subscribers never hear the fact.
Poison loop
A bad payload retries forever and stalls the partition for that key.
Schema drift
An incompatible field change makes a fleet of deserializers throw.
Unordered cross-partition view
Readers assume a global order the partition key never promised.
Projection lag mistaken for a bug
The command committed and the read model is simply behind.
Fan-out amplification
One bad event wakes search, email, fraud, and partners at once.
Misconceptions
Events give exactly-once for free.
Delivery is usually at-least-once. Idempotent consumers and an inbox make the effects safe.
Event-driven architecture replaces sagas.
Events carry the messages. Sagas still own multi-step business undo.
CQRS requires event sourcing.
CQRS splits commands from queries. Event sourcing is only one way to feed the read model.
Interviewer traps
Claiming exactly-once end to end.
Name broker limits, idempotent handlers, and deduped side effects. Do not stop at the producer.
Redrawing Kafka consumer groups when the question was the domain topology.
Point at the Kafka lessons for brokers and delivery. Stay on sync versus events, outbox, CQRS, and sagas.
Design scenario
Same prompt for every reader.
Requirements
Checkout confirms payment. Email and analytics may lag. No dual-write. Consumers are idempotent. Partition by orderId.
Traffic / scale
About 5k orders per second at peak.
Latency
The checkout UI needs a strong payment confirmation within 2 seconds. Email and analytics may lag 30 seconds.
Consistency
Payment confirmation is authoritative on the write path. Other order-status screens may be eventually consistent.
Availability
If payments are down for 90 seconds, accepted orders stay durable and consumers catch up from the outbox.
Failure assumptions
- The broker can be unavailable after the database commit.
- Consumers see at-least-once delivery, including duplicates.
- One bad event can fan out to many consumers.
Constraints
- No dual-write of a database commit and a later broker produce.
- Consumers are idempotent.
- Partition by orderId.
Prompt
Design the order lifecycle across inventory, payments, notifications, and analytics.
API
What does checkout return, and which confirmation stays on a synchronous path?
Data
Where do OrderPlaced and PaymentCaptured live, including the outbox row?
Architecture
Where do the CQRS order-status read model and the reserve, charge, ship saga sit?
Checkout needs a yes, and five other systems need the fact
Prefer
Sync confirmation, events for everything else
The user-critical answer stays on a request you control. Side effects publish one fact. Independent consumers project, notify, and analyze without holding the request open.
- Payment confirmation returns from the write path.
- OrderPlaced is durable in the same local transaction as the order.
- Email and analytics may lag. That lag is an SLO.
- A down consumer becomes backlog, not a failed checkout.
Alternative
One RPC chain for inventory, pay, email, and analytics
The caller waits on every downstream RTT. A 90-second payments outage, or a slow analytics hop, fails the user path.
- Latency is the sum of the chain.
- Temporal coupling means a deploy in one service stalls the others.
- Fan-out is N calls or a fragile in-request orchestrator.
- Partial failure has no durable fact to retry from.
Happy path — local commit, then reactions
Vertical cards for phones. The sequence diagram below is the same story.
- 1
Commit the order and the outbox row
One local transaction. The business row and the intent to publish succeed or fail together. - 2
Return accepted
The API does not wait for email, search, or analytics. - 3
Relay publishes OrderPlaced
A poller or a change stream reads the outbox. Broker downtime is lag, not a lost order. - 4
Consumers apply idempotently
An inbox keyed by eventId makes at-least-once delivery safe. Payments emits PaymentCaptured the same way. - 5
Project or dead-letter
Read models update by version. A poison payload stops after N retries and lands on a DLQ.
Overview
Event-driven architecture replaces request/response coupling with facts that happened. Services react, build read models, and coordinate workflows without holding each other's locks.
Interviews probe four things:
- Can you contrast sync RPC and async events with consistency and failure, not slogans?
- Do you know notification, event-carried state, and event sourcing?
- Can you place CQRS, the outbox, and ordering / DLQ on one board?
- Will you point at Kafka, sagas, CDC, and cache invalidation instead of re-teaching them?
Ask this out loud: if payments are down for 90 seconds, which user-visible paths still work, and which must wait?
Sync RPC versus events
| Dimension | Sync RPC | Event-driven |
|---|---|---|
| Coupling | Caller waits on the interface | Contract is the event. Time is decoupled |
| Latency | Sum of downstream RTTs | Producer returns after the local commit and the publish path it owns |
| Failure | Cascades unless you add timeouts and bulkheads | Consumers lag. The producer can stay up |
| Consistency | Often read-your-writes in one call | Eventual across services. Version the facts |
| Fan-out | N calls, or one fragile orchestrator | One publish, N independent consumers |
| Ops | Simpler per-request traces | Correlation ids, lag, DLQ depth |
| Prefer when | A strong immediate answer and a tiny fan-in | Many reactions, resilience, or an audit trail |
Rule of thumb: keep the user-critical confirmation on a sync path or a read model you control. Push side effects to events.
Within one service you still use a local transaction. Eventual consistency is the cross-service cost, not a reason to drop ACID on the aggregate that accepted the order.
Choreography is not a saga
Publishing events often implies choreography: peers react. A multi-step business outcome that must undo still needs a saga. Orchestration keeps a visible state machine. Choreography stays loosely coupled and can turn into an invisible event dance.
| Style | Strength | Weakness | Where it lives |
|---|---|---|---|
| Orchestrated saga | Observable timeouts and branches | Coordinator ownership | Sagas hub |
| Choreographed saga | Independent deploys | Event spaghetti | Same saga cluster |
| Pub/sub fan-out | Notifications and projections | No compensation policy | This cluster, plus the outbox |
Do not re-derive compensations here. Sagas & Distributed Transactions owns undo. This cluster owns the event topology and the reliability glue around it.
Decision path
Decisions
- 1
1. State change in service A
- next2. Caller needs an immediate answer
- ?
2. Caller needs an immediate answer
- yes3. Sync RPC or a local read
- no4. How much state is in the event
- 3
3. Sync RPC or a local read
- ?
4. How much state is in the event
- thin id5. Notification event
- enough fields6. Event-carried state transfer
- history is truth7. Event sourcing
- 5
5. Notification event
- next8. Outbox plus broker
- 6
6. Event-carried state transfer
- next8. Outbox plus broker
- 7
7. Event sourcing
- next8. Outbox plus broker
- 8
8. Outbox plus broker
- next9. Multi-step business undo
- ?
9. Multi-step business undo
- yes10. Saga on the sagas cluster
- no11. Idempotent consumers and projections
- 10
10. Saga on the sagas cluster
- 11
11. Idempotent consumers and projections
- next12. CQRS read models and cache invalidation
- 12
12. CQRS read models and cache invalidation
Lesson map
Commit, then the relay publishes
The order and the outbox row are committed. The broker has not published yet. The read model can still be old.
Architecture. Orders Committed. Broker Idle. Payments Idle. Read model Lagging
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB orders["Orders Committed"] broker["Broker Idle"] payments["Payments Idle"] read_model["Read model Lagging"] orders -->|OrderPlaced| broker broker -->|Deliver| payments broker -->|Project| read_model payments -->|PaymentCaptured| broker payments -->|DLQ| broker orders -->|Lost event| broker
Read it as a filter. If the caller needs the answer now, stop at sync. If later steps must undo, leave compensations on the saga lesson. Otherwise publish through the outbox and let idempotent consumers project.
What this cluster covers
- Event Types — Notification, Event-Carried State Transfer & Event Sourcing — payload richness and replay.
- CQRS — Commands, Queries, Projections & Consistency — write model versus query-shaped reads.
- Transactional Outbox, Inbox & Consumer Idempotency — how EDA uses the outbox. The relay encyclopedia stays on Transactional Outbox & Inbox Patterns.
- Ordering, Partitions, Poison Messages & Retry/DLQ Strategy — which order you actually need. Broker knobs stay on the Kafka lessons.
- EDA Failure Modes — Dual Writes, Schema Drift & Fan-out Blast Radius — the catalog, then back to this hub.
Adjacent lessons, by title only:
- Apache Kafka — Topics, Partitions, Brokers & Consumer Groups for the log, partitions, and consumer groups.
- Delivery Semantics — At-Least-Once, At-Most-Once & Exactly-Once for ordering, poison, and what the broker actually promises.
- Change Data Capture when the commit stream is the event source.
- Cache Invalidation — TTL vs Event-Driven vs Versioned Keys when an event should drop or version a cache key.
A delivery channel such as a WebSocket is not a domain event. Do not fold realtime protocols into this cluster.
Commit, publish, project
Sequence
- 1
Order API → Orders DB
1. Insert order and outbox row in one transaction
- 2
Orders DB → Order API
2. Commit OK and return accepted
- 3
Outbox relay → Orders DB
3. Poll the outbox
- 4
Outbox relay → Broker
4. Publish OrderPlaced
- 5
Broker → Payments
5. Deliver
- 6
Payments → Broker
6. Emit PaymentCaptured after an idempotent handle
- 7
Broker → Read model
7. Project order status
- 8
Payments
Poison payload exhausts retries
- 9
Payments → Broker
8. Send the original event to the DLQ
The client is done at step 2. Steps 3 through 7 are backlog. Step 8 keeps one bad payload from blocking the partition forever. Ordering and DLQ policy are the fifth lesson. Dual-write, schema, and fan-out are the sixth.
Dual-write versus outbox (run this)
No database and no broker. The first function writes the order and then publishes. The broker fails once, and the event is gone. The second function commits the order and the outbox row together. The relay retries. The fact survives.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Idempotent consumer (run this)
At-least-once delivery means the same eventId can arrive twice. Record it before the side effect. Production puts the inbox insert and the charge in one local transaction. This sketch keeps both in memory so you can see the duplicate skip.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The inbox pattern in an EDA flow is the fourth lesson. The existing Transactional Outbox & Inbox Patterns page owns relay variants. Do not rebuild that page here.
Interview Q&A
What is event-driven architecture in one sentence?
Answer
Systems publish immutable facts about state changes, and other components react asynchronously to those facts.
When is sync RPC still better?
Answer
When the client needs a single authoritative answer now, such as an authorization check, a price quote, or an inventory hold, and the fan-out is tiny.
Does event-driven architecture imply eventual consistency?
Answer
Across services, almost always. Inside one aggregate you still commit locally. The user-critical confirmation should read that local result, not a lagging projection.
Notification versus event-carried state versus event sourcing?
Answer
A notification is a thin signal and the consumer fetches. Event-carried state transfer puts the fields the consumer needs in the event. Event sourcing makes the log the system of record for the aggregate. The next lesson is the comparison.
How do you avoid a lost event after a database commit?
Answer
Write a transactional outbox row in the same transaction, or capture the commit with CDC. Never treat commit and produce as two non-atomic steps.
How do consumers survive at-least-once delivery?
Answer
An inbox or a natural business key, applied in the same local transaction as the side effect. Broker exactly-once does not cover email, charges, or shipments by itself. See delivery semantics.
Where do sagas fit?
Answer
Multi-step business processes that need compensations. Events are the messages. Sagas supply the undo policy.
Can you do CQRS without event sourcing?
Answer
Yes. Commands hit a write model that can still be an ordinary store. Queries hit a read model updated by events, projections, or CDC. CQRS is that split.
Which metrics matter?
Answer
Outbox age, consumer lag, DLQ depth, projection freshness, duplicate rate, and schema-compat failures. A freshness SLO is how you talk about lag with product.
What is the interview trap?
Answer
Saying exactly-once end to end without naming broker limits, idempotent handlers, and deduped side effects. The failure-modes lesson is the catalog of what that claim hides.
Pitfalls
Put checkout, the orders database, the outbox, payments, email, and the order-status read model on one column. Cross out payments for 90 seconds. Mark which boxes the user still gets a yes from, which boxes grow lag, and which lesson owns the undo if a later step must refund.
Go Deeper
- Martin Fowler — What do you mean by Event-Driven?
- Microsoft — Event-driven architecture style
- Uber Engineering — Reliable reprocessing
- microservices.io — Transactional outbox
- Apache Kafka documentation — delivery semantics
- Next: Event Types — Notification, Event-Carried State Transfer & Event Sourcing