At-Least-Once vs Exactly-Once Delivery
At-most-once loses; at-least-once duplicates; exactly-once needs transactional producer+consumer+sink; effectively-once is idempotent consumers + dedup.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Why effectively-once is the default winner
Prefer
At-least-once + idempotent consumer + dedup
The broker may deliver twice. The handler makes the second delivery a no-op. Business state looks like one process.
- Works with Kafka, RabbitMQ confirms, SQS, HTTP retries — any at-least-once channel.
- Dedup key is a natural id (order_id, msg_id) or a payload hash in a durable store.
- Pairs with an outbox so the producer side does not dual-write.
- Throughput stays high; you pay storage and a unique constraint, not a distributed transaction.
Alternative
True exactly-once across producer, broker, consumer, sink
Kafka EOS can atomic-commit produce + offset. Logical exactly-once still needs a transactional sink.
- Requires idempotent producer, unique transactional.id, isolation read_committed, and a sink in the same transaction (or an idempotent write).
- Cuts throughput, adds coordinator load, and is broker-specific.
- A Postgres INSERT outside the Kafka transaction still duplicates on retry.
- Reserve for ledgers and money movement — not analytics or cache invalidation.
Happy path — one record, one process (effectively-once)
Vertical cards for phones. The mermaid sequence below is the same story.
- 1
Producer assigns identity
Serialize the event. Stamp a natural key (order_id) and, on Kafka, a producer id plus sequence number. - 2
Broker appends and replicates
acks=all waits for the in-sync replica set. The producer retries until it sees that ack. - 3
Consumer processes idempotently
Look up the key in a durable dedup store. New key: apply the side effect. Duplicate: skip. - 4
Commit offset after the sink
Manual or transactional commit only after the domain write succeeds. Auto-commit before process is at-most-once loss. - 5
Crash before commit
The broker redelivers. Dedup makes the replay a no-op. This is at-least-once on the wire, effectively-once in the business.
Overview
Delivery semantics are the contract between a producer, a broker (or any middleware), and a consumer. They answer one interview question: if a process dies, a packet is lost, or a leader fails over, does the business see zero, one, or many applications of the same logical event?
Four names get used as if they were interchangeable. They are not.
| Contract | Promise | What you actually get |
|---|---|---|
| At-most-once | Zero or one delivery | No duplicates. Loss is allowed. Fire-and-forget (acks=0), or commit the offset before processing. |
| At-least-once | One or more deliveries | No silent loss if retries and durability are configured. Duplicates are allowed. The default of Kafka with acks=all plus manual commit-after-process, RabbitMQ with publisher confirms plus consumer acks, SQS, HTTP retries. |
| Exactly-once | One process of each logical message | Requires a transactional boundary across producer, broker, consumer, and sink. Kafka EOS is the famous instance. Rare, expensive, and still only as strong as the sink. |
| Effectively-once | Looks like once from the business | At-least-once on the wire + idempotent handler and/or a dedup store. This is what most production systems actually ship. |
Interviewers love this because it sits under payments, analytics, cache invalidation, and IoT — and because “we use Kafka exactly-once” is often a lie about the sink.
Why the wrong guarantee hurts
| Scenario | Wrong contract | Business impact |
|---|---|---|
| Financial debit / credit | Duplicate debit, or a lost debit | Double-spend or missing revenue; audit and trust |
| User-activity analytics | Duplicate events | Inflated funnels, bad product bets |
| Cache invalidation | Missed invalidation | Stale reads; “it works on my box” incidents |
| IoT / sensor pipelines | Dropped readings | Blind spots; in safety domains, hazard |
A senior engineer names the semantic, prices the cost (latency, throughput, storage), and designs the surrounding pieces — idempotent handlers, dedup tables, transactional boundaries — to the SLA.
The producer–broker–consumer triangle
Every delivery debate is three axes, not one broker checkbox.
| Axis | Producer | Broker | Consumer |
|---|---|---|---|
| Ack / durability | acks=0 / 1 / all (Kafka). When does the send count as success? | Replication factor + ISR (in-sync replicas). Did the write survive a leader death? | Auto-commit vs manual vs transactional commit. When is the message “processed”? |
| Retry | Exponential backoff, max attempts. Retry after timeout is how duplicates are born. | Idempotent writes (Kafka 0.11+: producer id + sequence). | Idempotent processing or transactional consume. |
| Idempotence | Monotonic sequence per partition. | Broker drops duplicates with producerId + seq <= last. | Store a processed-message key; skip repeats. |
Decisions
- 1
Producer
- send plus seqBroker log
- 2
Broker log
- replicateISR
- ack if acks=allProducer
- poll from offsetConsumer
- 3
ISR
- nextBroker log
- 4
Consumer
- check dedup storeSeen key?
- ?
Seen key?
- noWrite sink
- yesSkip and still commit
- 6
Write sink
- then commit offsetBroker log
- 7
Skip and still commit
- nextBroker log
Lesson map
At-Least-Once vs Exactly-Once Delivery
At-most-once loses; at-least-once duplicates; exactly-once needs transactional producer+consumer+sink; effectively-once is idempotent consumers + dedup.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB prod["Producer"] brk["Broker log"] con["Consumer"] sink["Sink"] prod -->|Send record| brk brk -->|Ack acks=all| prod con -->|Poll from| brk con -->|Apply side| sink con -->|Commit offset| brk con -->|Commit offset| brk
Producer acks
acks=0— fire and forget. Producer never waits. Fast, at-most-once. A timeout looks like success and the broker may never have the bytes.acks=1— leader wrote to its local log. Faster thanall. If the leader dies before followers catch up, the record can vanish.acks=all(or-1) — every in-sync replica has the record. This is the durability you want for at-least-once. Combined with retries, a producer crash after the broker persisted but before the ack arrived will resend → duplicate unless the producer is idempotent.
ISR (in-sync replicas)
The ISR is the set of followers that have caught up within replica.lag.time.max.ms. acks=all only waits for the current ISR, not for replication.factor if some replicas are out of sync. A leader election with ISR smaller than you thought is how “we set acks=all” still loses data. min.insync.replicas is the floor: refuse the produce rather than ack with a tiny ISR.
Idempotent producer (not the same as EOS)
Kafka enable.idempotence=true attaches a producer id and a monotonically increasing sequence per partition. The broker stores the highest sequence it has accepted. A retry with the same pid+seq is dropped. That kills producer-retry duplicates. It does not stop consumer redelivery, and it does not make your Postgres insert exactly-once.
EOS includes the idempotent producer, then adds transactions.
Offset commit: auto vs manual vs transactional
This is the consumer half of the triangle. Get it wrong and you silently change the semantic.
| Mode | When the offset moves | Semantic if the process dies |
|---|---|---|
| Auto-commit | On a timer, often before your handler returns | Crash after commit, before process → loss (at-most-once). Crash after process, before commit → duplicate. You get a mix. |
| Manual commit after sink | You call commit only after the domain write succeeds | Crash before commit → redelivery (at-least-once). Correct default. |
| Transactional commit | Offset + output records (or offset + sink) in one Kafka transaction | Read-process-write atomicity inside Kafka. Still needs the sink to participate or be idempotent. |
Exactly-once semantics (EOS)
Kafka’s version of EOS is a transaction that spans:
- Transactional producer — a batch of records is wrapped in a
transactional.id. The broker makes them visible only aftercommitTransaction. Aborted transactions never becomeread_committeddata. - Transactional consumer — reads a transactional topic, writes outputs, and commits offsets in the same transaction. Read-process-write is atomic from Kafka’s point of view.
- Requirements
enable.idempotence=true- Unique
transactional.idper producer instance (fencing: a new instance with the same id fences the old one) - Consumers set
isolation.level=read_committedso they skip open / aborted tx - An exactly-once sink — a DB that can take an idempotent write, or a write path in the same transaction
Even with all of that, logical exactly-once holds only if downstream sinks are transactional with the broker (or themselves idempotent). A INSERT INTO charges over a separate Postgres session after a Kafka consume is still effectively-once at best: crash after insert, before offset commit, and you will see the message again.
Sequence
- 1
Producer
1. Serialize, pid plus seq
- 2
Producer → Broker
Send record
- 3
Broker
2. Append, replicate to ISR
- 4
Broker → Producer
Ack acks=all
- 5
Consumer → Broker
Poll from committed offset
- 6
Consumer
3. Dedup check
- 7
Consumer → Sink
Apply side effect
- 8
Consumer → Broker
Commit offset after sink
- 9
Consumer → Consumer
Skip
- 10
Consumer → Broker
Commit offset
read_committed is the consumer half of EOS: without it, a consumer can see records from an open transaction that later aborts. That looks like a phantom event.
Effectively-once via idempotent consumers
This is the pattern you should be able to whiteboard in five minutes.
- Deterministic keys — prefer a natural key (
order_id,sensor_id + timestamp, Stripeevent.id) over a random UUID minted at consume time. A UUID at consume time creates duplicates instead of detecting them. - Idempotent operations
INSERT … ON CONFLICT DO NOTHINGUPDATE … SET … WHERE version = expected- A dedup table keyed by message id (or payload hash) with a TTL
- Stateless vs stateful — a pure function of the event can rely on upstream dedup. A stateful service keeps a small durable set of recent keys (DB table, Redis with HA, compacted topic, embedded LSM). In-memory LRU dies on restart and you re-process.
The producer still uses acks=all and retries. The consumer still commits after the write. Duplicates are expected; they are safe.
Deep dive · Compacted topics are not a dedup store
Log compaction keeps the latest value per key. A consumer that is mid-rebalance, or that reads from the earliest offset, still sees the same key more than once. Compaction is a storage optimization, not an exactly-once consumer. If you need “process this key once,” you still need a processed-keys table or an idempotent sink.
Failure modes
Walk this table in an interview. Each row is a different semantic unless you designed for it.
| Failure point | What goes wrong | Resulting semantic |
|---|---|---|
| Producer crash after send, before ack | Broker may already have the record. Producer retries. | Duplicate (at-least-once) unless idempotent producer. |
| Broker leader election | In-flight writes vanish if ISR is weaker than you thought. | Loss (at-most-once) or later duplicate on retry, depending on acks. |
| Consumer crash before offset commit | Group rebalance; the same offset is delivered again. | Duplicate unless the handler / sink is idempotent. |
| Network partition | Producer saw an ack the new majority never replicated — or never saw an ack for a write that landed. | Loss or duplicate. This is why min.insync.replicas and idempotent producers exist. |
| Sink failure (DB down mid-write) | Offset not committed (good) or committed (loss). Partial writes. | Breaks EOS unless the sink is in the transaction or the write is idempotent. |
Mitigations you should name: retry with idempotence, write-ahead / outbox, fencing tokens on transactional ids, DLQ for poison pills.
Architecture choices
| Decision | Options | Trade-off |
|---|---|---|
| Broker | Kafka (log, EOS), Pulsar (segments, tiered storage), RabbitMQ (ack/NACK, confirms) | Kafka: strongest EOS, more ops. RabbitMQ: simpler, no native EOS — you build effectively-once. |
| Producer idempotence | Broker-native vs hand-rolled sequence numbers | Native is less code; needs a new enough Kafka. |
| Offset management | Auto / manual / transactional | Auto is a footgun. Manual is at-least-once. Transactional is EOS-shaped and slower. |
| Dedup store | Process memory, Redis SET + TTL, DB unique table, compacted topic | Memory is volatile. Redis is fast, extra HA. DB is durable, extra latency. Compacted topic is cheap and eventual. |
| EOS vs effectively-once | Full Kafka transactions + transactional sink vs idempotent consumer + outbox | EOS for money. Effectively-once for almost everything else. |
| Errors | Retry, circuit breaker, DLQ | Retries create duplicates. DLQ isolates poison pills; unmonitored DLQs are silent back-pressure. |
A senior design mixes them: at-least-once everywhere, idempotent downstream writes, outbox on the produce side, and EOS only on the financial spine.
In-memory broker (run this)
No Kafka, no Postgres. A list is the log. A Set is the dedup table. Advance the story: producer retry without idempotence duplicates; consumer crash before commit redelivers; dedup makes the second process a skip; commit-before-process loses the message.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Read the logs in order:
- A timeout retry without producer idempotence appends twice.
- A crash after process, before offset commit, redelivers.
- The dedup set makes the second delivery a skip — effectively-once.
- The same pid+seq on an idempotent producer is dropped at the broker.
- Commit-before-process loses the message — accidental at-most-once.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Explain the difference between at-least-once and effectively-once.
Answer
At-least-once is a delivery contract: the message will show up one or more times; loss is not allowed if durability and retries are set. Effectively-once is a business contract built on top: the consumer is idempotent (or a dedup store drops repeats), so the observable side effect happens once. The wire still duplicates.
Why is true exactly-once rare in practice?
Answer
It needs a transactional boundary that spans producer, broker, consumer, and sink — two-phase commit or Kafka transactions plus a sink that participates. That costs latency, throughput, and ops. Most databases and caches are not in the broker’s transaction, so you fall back to effectively-once.
How does Kafka’s idempotent producer actually dedupe?
Answer
Each record carries a producer id and a monotonic sequence per partition. The broker remembers the highest sequence accepted for that pid+partition. A retry with seq less than or equal to that high-water mark is ignored. This stops producer-timeout duplicates. It does not stop consumer redelivery.
What does isolation.level=read_committed do?
Answer
The consumer skips records that belong to open or aborted transactions. Only committed transactional data is visible. Without it, EOS consumers can observe records that later disappear — which is not exactly-once.
How does the outbox pattern give you effectively-once on the produce side?
Answer
The service writes the domain row and an outbox row in one DB transaction. A poller or CDC publishes to the broker. The event cannot exist without the business write, and cannot be lost if the process dies after commit. Duplicates on publish are handled by a unique outbox key plus consumer idempotence. See transactional outbox.
When would you deliberately choose at-most-once?
Answer
When a duplicate is worse than a gap. Real-time telemetry where a missed sample is fine, but a duplicated page-alert pages the on-call twice. Cache busts sometimes prefer miss-and-refresh over double-invalidation storms. Write it down as a product decision, not an accident of acks=0.
Redis SET vs a compacted Kafka topic for dedup?
Answer
Redis: sub-millisecond SISMEMBER / SET + TTL, extra HA and failover story. Compacted topic: durable with the broker, scales with partitions, but you consume to rebuild state and still see a key more than once during rebalance. A DB unique table is the boring, correct default for money.
How do you test that a consumer is actually idempotent?
Answer
Publish one message. Crash (or skip commit) after the side effect. Restart. Assert the domain row exists once and the dedup store has one row for that msg_id. Then publish the same msg_id again from a second producer and assert still one side effect.
acks=all plus retries — is that exactly-once?
Answer
No. That is durable at-least-once. Without an idempotent producer you get broker-side duplicates on retry. Without an idempotent consumer you get handler-side duplicates on redelivery. Exactly-once still needs transactions plus a transactional (or idempotent) sink.
What happens if the sink is not in the Kafka transaction?
Answer
EOS at the broker does not extend to that sink. Crash after INSERT and before offset commit → redelivery → second insert, unless the insert is ON CONFLICT DO NOTHING or keyed by message id. You have effectively-once only if you designed the sink that way.
Pitfalls
Draw producer, broker, consumer, sink. Label acks=all, ISR, and “commit offset after sink.” Then X out the producer after send-before-ack, and X out the consumer after process-before-commit. For each X, write loss, duplicate, or skip under three designs: auto-commit, manual + dedup, Kafka EOS with a non-transactional INSERT. The third column is the trap.