Ordering, Partitions, Poison Messages & Retry/DLQ Strategy
Most event-driven flows need order per entity, not global order. Partitions buy parallelism. Bounded retries and a dead-letter path keep one poison message from stalling that entity.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Do orders need global ordering?
Answer
Usually no. Per orderId is enough for that aggregate and its projection.
L2
What does a partition key buy?
Answer
Events with the same key stay ordered and serialized. Different keys run in parallel. Consumers scale with the partition count, up to one active reader per partition in a group.
L3
What breaks if the key is too hot?
Answer
That key's partition becomes a queue of one. Split the key only when you can tolerate losing per-entity order, or fan out locally after a single ordered intake.
L4
Why do retries make an outage worse?
Answer
Synchronized retries add load. Poison retries never succeed and block the partition. Duplicates from retries still need idempotency.
L5
How is a dead-letter queue different from a retry topic?
Answer
A retry topic delays another attempt. A DLQ is where you stop and triage after the budget is spent.
L6
What must a dead-letter record carry?
Answer
The original payload plus correlation id, event id, error context, and headers so a human or a job can replay it on purpose.
L7
Where do Kafka exactly-once and consumer groups fit?
Answer
They are broker guarantees and assignment rules. This page still designs the key, the retry budget, and the poison path. Delivery semantics owns the broker chapter.
Failure modes
Head-of-line poison
The consumer never commits past a bad message, so every later event for that partition waits.
Retry storm
Transient errors retry in lockstep and knock over the dependency that was already failing.
Silent skip
A deserialization failure is ignored, so the business fact disappears with no DLQ record.
Misconceptions
We need total order for all orders.
The invariant is almost always per order, per account, or per aggregate. Global order is a single partition and a throughput ceiling.
A retry topic and a DLQ are the same queue.
Retry topics schedule another attempt. The DLQ means you gave up and someone must decide to fix, skip, or replay.
Interviewer traps
Answering with acks, ISR, and consumer-group protocol.
Name the order you need and the poison policy. Point at the Kafka lessons for broker knobs.
Infinite retry so we never lose data.
Infinite retry loses time. Bound it, dead-letter, and keep the original payload for a deliberate replay.
One bad barcode should not close the store
Prefer
Per-entity order, bounded retry, then a DLQ
The key is the order id. Transient errors back off with jitter. A poison payload moves aside with its event id, and the partition continues.
- Other orders stay parallel.
- The original payload is still available to replay.
- Duplicates from retries are an inbox problem, already designed.
Alternative
One global sequence and infinite retry
A single lane preserves a total order nobody asked for. The first poison message stops every later event.
- Throughput collapses to one partition.
- A validation error never succeeds without a code or data change.
- Synchronized retries amplify the outage.
Classify, then decide
Retry is a policy. It is not a while-loop around a handler.
- 1
Name the order you need
None, per entity, or global. Write the business rule that fails if two events for one order swap. - 2
Key the partition to that rule
orderId or accountId. A hot key is a capacity problem, not a reason to hash randomly and lose the order. - 3
Classify the error
Timeouts retry. Validation and bad payloads do not. A down dependency gets a budget, not an infinite loop. - 4
Dead-letter and move
Store the original event and the error. Commit past it so the key is not stuck. Alert on depth. - 5
Replay on purpose
Fix the code or the data first. Replay is a product decision, not an automatic second infinite loop.
Overview
Event-driven design cares about which order you need, how partitions buy parallelism, and how retries, poison messages, and a DLQ keep the system live.
Broker knobs stay next door. Apache Kafka — Topics, Partitions, Brokers & Consumer Groups owns topics, partitions, and groups. Delivery Semantics — At-Least-Once, At-Most-Once & Exactly-Once owns acks, duplicates, and exactly-once limits. This page stays on the architecture rule: align the key with the consistency boundary you project, and stop retrying poison.
Idempotency for the duplicate that a retry creates is the outbox and inbox lesson.
What you can promise
| Promise | How | Cost |
|---|---|---|
| No order | Any fan-out queue | Maximum parallelism. The reader must tolerate any interleaving |
| Per-entity order | Key equals the entity id | Hot keys serialize. That is often the correct trade |
| Global order | One partition, or a consensus log | A throughput ceiling |
| Causal order | Explicit happens-before | Extra metadata and a harder read model. Do not invent vector clocks unless the product needs them |
Interview line: "We need total order" is rarely true. Per orderId or per accountId is the usual invariant.
Cross-partition views are unordered. A read model that joins many orders must tolerate interleaving. That is normal. The CQRS lesson gates each entity by version so one order's own events still apply in order.
The key is an architecture lever
Decisions
- 1
1. Choose the partition key
- next2. Is that key hot
- ?
2. Is that key hot
- yes3. Split only if this entity can lose order
- no4. Key equals the entity id
- 3
3. Split only if this entity can lose order
- 4
4. Key equals the entity id
- next5. Parallelism follows the partition count
- 5
5. Parallelism follows the partition count
- next6. Cross-partition views stay unordered
- 6
6. Cross-partition views stay unordered
- next7. Read models tolerate interleaving across entities
- 7
7. Read models tolerate interleaving across entities
Lesson map
Park the poison, move the partition
The key is the entity id. The consumer is still inside the retry budget. The dead letter is empty.
Architecture. Partition In order. Consumer Retrying. Dead letter Empty
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB partition["Partition In order"] consumer["Consumer Retrying"] dlq["Dead letter Empty"] partition -->|Deliver| consumer consumer -->|Retry| partition consumer -->|Poison| dlq consumer -->|Commit offset| partition partition -->|Head blocked| consumer dlq -->|Alert| consumer
A hot key is a single busy entity, not a reason to drop the key. If one order must stay ordered, keep it on one partition and scale the others. Local fan-out after an ordered intake is a different design: you accepted order at the door, then parallelized work that is allowed to interleave.
Retry classes
| Class | When | Backoff | If you guess wrong |
|---|---|---|---|
| Transient | Network blip, brief 503 | Exponential backoff plus jitter | A thundering herd |
| Poison | Bad payload, deterministic crash | Do not retry forever | The partition stalls |
| Downstream outage | Dependency down | A budget, plus a circuit if you have one | Lag cascades |
| Schema incompatible | Deserializer rejects the record | DLQ and an alert | A silent drop if you swallow the error |
Jitter matters because many consumers otherwise retry on the same clock. The classification matters more: a validation error is not a timeout.
Poison path
Sequence
- 1
Partition → Consumer
1. Deliver the message
- 2
Consumer → Consumer
2. The handler throws
- 3
Consumer → Partition
3. Retry with backoff up to a budget
- 4
Consumer
4. Still failing. Classify it as poison.
- 5
Consumer → Dead letter
5. Publish the original payload and the error context
- 6
Consumer → Partition
6. Commit the offset so the partition moves
- 7
Dead letter → Oncall
7. Alert on depth
- 8
Partition
Never committing leaves the entity blocked at the head.
Architecture rules:
- Bound retries. Classify poison versus transient before you spend the budget.
- Prefer per-entity blocking over freezing every consumer. A key-level park lets other keys continue when the platform supports it. A partition-level DLQ is the simpler default.
- The DLQ carries correlation id, event id, and error context. Without those, replay is a guess.
- Replay is a product decision: fix the data, skip, or ship a code fix first.
Block, skip, or park
| Policy | Strength | Cost |
|---|---|---|
| Never commit | No silent loss | Head-of-line blocking for that partition |
| Skip to a DLQ after N tries | The pipeline stays healthy | The DLQ is now real work. Ignoring it loses the fact |
| Park one key | Other keys continue | You built a key-level pause, which is more machinery |
There is no policy that is both "never look at failures" and "never lose work."
Classify an error (run this)
Validation never retries. A timeout retries until the budget, then dead-letters. Backoff grows and adds jitter so clones do not wake together.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Per-key order in one worker (run this)
Messages arrived interleaved. Sorting inside each orderId restores the sequence that key promised. It does not create an order between a and b. That is the point.
Press Run. Snippets must be self-contained — no network, files, or native modules.
A partition key is this map, enforced by the log instead of by an in-process sort. If two events for order a land on different partitions, the sort never sees them together.
Interview Q&A
Do we need global ordering for orders?
Answer
Usually no. Per orderId is enough for that aggregate's projection and its business rules.
Why can retries make things worse?
Answer
Synchronized retries amplify load. Poison retries block a partition. Every retry can duplicate a side effect, so the consumer still needs the inbox from the previous lesson.
How is a DLQ different from a retry topic?
Answer
A retry topic is delayed staging for another attempt. A DLQ is triage after you give up. Automatic replay from the DLQ without a fix recreates the poison loop.
How does this link to Kafka exactly-once?
Answer
Broker exactly-once reduces duplicate produces inside a transaction scope. You still design consumer idempotency and a poison path. Delivery semantics is that chapter.
What is a hot key in this design?
Answer
One entity id that takes a disproportionate share of events, so its partition becomes the bottleneck. Splitting it trades away the per-entity order you chose the key for.
What do you alert on?
Answer
DLQ depth, retry rate by error class, and lag on the partitions that are not parked. Age of the oldest unprocessed event for a key beats a raw message count.
Where do schema failures go?
Answer
They are poison, not timeouts. Dead-letter them and alert. Compatibility and rollout are the failure-modes lesson. Deserializer internals stay with the Kafka topics lesson.
Can two consumers in one group order a single partition together?
Answer
No. One partition is read by at most one member of a group. That is why extra consumers sit idle past the partition count. The assignment rules live on Kafka topics and consumer groups.
Pitfalls
For payments, email, and a per-order status projection, write the key you would use and the event pair that must not swap. Then name one error you retry and one you dead-letter on the first classification.