Messaging
Part 1 of 5 · Kafka messagingApache Kafka — Topics, Partitions, Brokers & Consumer Groups
Append-only partitioned logs; consumer group assigns each partition to at most one member; durability=ISR; scale=partitions; order=within partition.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Why Kafka is a commit log you can fan out
Prefer
Partitioned log + independent consumer groups
Producers append. Replication is ISR. Each group assigns a partition to at most one member. Other groups keep their own offsets and replay the same topic.
- Ordering is total only inside a partition (sticky key).
- Retention is time/size — replay is first-class.
- Scale consumers in one group up to the partition count.
- Durability knobs are acks, min.insync.replicas, unclean.leader.election.
Alternative
Classic broker queue (delete-on-ack)
RabbitMQ/SQS shine at routing and simple jobs. They are the wrong backbone for multi-TB replayable CDC or many independent readers of the same stream.
- SQS FIFO caps throughput; no true log replay.
- Rabbit fan-out needs extra exchanges/bindings; retention is TTL/plugins.
- Kafka for tiny RPC work queues is overkill ops — use AMQP.
- Pulsar wins when you need multi-tenancy plus tiered storage out of the box.
Happy path — produce, replicate, consume, fan out
Vertical cards for phones. Same story as the mermaid below.
- 1
Produce with a key
hash(key) pins the record to one topic partition. Null keys sticky/round-robin for throughput. - 2
Leader appends, ISR replicates
acks=all waits for in-sync followers. min.insync.replicas refuses a too-small ISR instead of silently acking. - 3
One group splits partitions
Each partition is consumed by at most one member of that group. Extra members sit idle. - 4
Another group is independent
A second group.id has its own offsets. Same topic, full fan-out — billing and search do not steal messages. - 5
Process, then commit offset
Commit-before-process is at-most-once loss. Commit-after is at-least-once. See delivery semantics.
Overview
Kafka stores append-only logs partitioned across brokers. Producers write records to a topic partition. A consumer group divides those partitions so each partition is consumed by at most one member of that group. Durability comes from replication (ISR + min.insync.replicas). Scale comes from more partitions. Ordering is guaranteed only within a single partition.
Think of Kafka as a distributed commit log you can fan out to many independent consumer groups — not a classic message broker that deletes messages on ack.
You should be able to:
- Draw producer → leader → ISR followers → two consumer groups with separate offsets.
- Say why order is per-partition and why
group.idis the fan-out knob. - Name what
acks=allactually waits for (current ISR, not every replica forever). - Choose Kafka vs a queue without cargo-culting “we use Kafka.”
Architecture
1 Produce
- 1
Producer
- key hashTopic partition
- 2
Topic partition
- nextLeader broker
2 Durability
- 3
Leader broker
- replicateFollower
- replicateFollower
- nextISR ack if acks=all
- nextConsumer group
- nextOther group
- 4
Follower
- 5
Follower
- 6
ISR ack if acks=all
3 Consume
- 7
Consumer group
- assignConsumer A: p0, p2
- assignConsumer B: p1
- 8
Consumer A: p0, p2
- 9
Consumer B: p1
4 Independent fan-out
- 10
Other group
- nextOwn offsets
- 11
Own offsets
Lesson map
Apache Kafka — Topics, Partitions, Brokers & Consumer Groups
Append-only partitioned logs; consumer group assigns each partition to at most one member; durability=ISR; scale=partitions; order=within partition.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB prod["Producer"] lead["Partition leader"] fol["Followers ISR"] ca["Consumer A: p0, p2"] prod -->|produce(key,| lead lead -->|replicate| fol fol -->|ack| lead lead -->|offset N| prod ca -->|fetch p0 from| lead ca -->|commit offset| lead
Kafka vs RabbitMQ vs SQS vs Pulsar
| Dimension | Kafka | RabbitMQ | AWS SQS | Pulsar |
|---|---|---|---|---|
| Core model | Append-only partitioned log | Smart broker + queues/exchanges | Managed queue | Segmented log + brokers + BookKeeper |
| Ordering | Per partition (key sticky) | Per queue (weaker at scale) | FIFO queues only (limited TPS) | Per partition / key |
| Replay / retention | First-class (time/size) | Limited (TTL / plugins) | Short retention; no true replay | First-class (tiers) |
| Fan-out | Many consumer groups on same topic | Fan-out exchanges / bind | SNS + SQS pattern | Multi-subscription |
| Ops burden | High (KRaft/ZK, disks, ISR) | Medium | Near-zero managed | Medium–high |
| Best fit | Event streaming, CDC, high throughput | Complex routing, low-latency RPC-ish | Simple async jobs | Multi-tenant streaming |
What fails if you choose wrong
- Pick RabbitMQ for multi-TB replayable CDC → you fight retention and lose cheap fan-out.
- Pick SQS for strict key ordering at millions of messages/s → FIFO caps throughput; no log replay.
- Pick Kafka for tiny RPC work queues with complex routing → overkill ops; better AMQP.
- Ignore Pulsar when you need true multi-tenancy + tiered storage out of the box.
Topics, partitions, brokers
Topics and partitions
- A topic is a named stream. A partition is an ordered, immutable sequence of records with contiguous offsets.
- Partition count ≈ max parallel consumers in one group. Too few → lag under load. Too many → metadata/ISR churn, small files, long rebalances.
- Replication factor (RF) copies each partition to RF brokers. Writes go to the leader; followers in the ISR stay caught up.
Partition keys, hot shards, and what happens when you change N are the next lesson: ordering vs throughput.
Brokers and durability knobs
| Knob | Safer default | What breaks if weaker |
|---|---|---|
acks | all | Leader ack only → data loss on crash |
min.insync.replicas | ≥ 2 (with RF ≥ 3) | Under-replicated writes accepted |
unclean.leader.election | false | Stale follower can become leader → truncate committed data |
| Disk / log retention | time + size policy | Blind forever growth or surprise truncation |
acks=all waits for the current ISR, not for every replica that is supposed to exist. If ISR shrinks below min.insync.replicas, producers get errors instead of a silent ack. That is the interview sentence.
Consumer groups
group.idis the unit of competing consumers. Same group → partitions split. Different groups → each gets the full topic (fan-out).- Offset commits mark progress. Commit before process → at-most-once. After → at-least-once. The full contract, Kafka EOS, and why HTTP/Stripe still need keys: delivery semantics, plus at-least-once vs exactly-once, outbox/inbox, and API idempotency keys.
- Membership changes trigger a rebalance. Eager vs cooperative sticky, lag math, and
pause: rebalancing, lag, backpressure.
Sequence
- 1
Producer
1 Produce with acks=all
- 2
Producer → Partition leader
produce(key, value)
- 3
Partition leader → Followers ISR
replicate
- 4
Followers ISR → Partition leader
ack
- 5
Partition leader → Producer
offset N
- 6
Consumer A
2 Group consumes partitions
- 7
Consumer A → Partition leader
fetch p0 from offset
- 8
Consumer B → Partition leader
fetch p1 from offset
- 9
Consumer A
3 Process then commit
- 10
Consumer A → Partition leader
commit offset N+1
Partition assignment model (run this)
No broker. Same key sticks to one partition. Null keys round-robin. A two-member group splits three partitions; a third member in that group would idle.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pros / cons
| Choice | Pros | Cons | Prefer when |
|---|---|---|---|
| More partitions | Parallelism, throughput | Rebalance cost, metadata | Sustained lag with saturated consumers |
| Higher RF | Durability | Disk + network | Money movement / audit logs |
| Many consumer groups | Independent fan-out | Broker read amplification | Analytics + billing + search on same events |
| Kafka vs SQS | Replay, ordering at scale | Ops | Platform event backbone |
Interview Q&A
Where does Kafka guarantee order?
Answer
Only within a partition. Cross-partition order is undefined. Same key → same partition → total order for that key, until you change partition count.
How many consumers in a group can read one partition?
Answer
At most one active member per partition per group. Extra members sit idle. Scale past partition count does nothing until you add partitions (and pay affinity costs).
Why not 10,000 partitions just in case?
Answer
Controller/KRaft metadata pressure, longer rebalances, more open files, weaker batching. Partition count is a capacity plan, not a lucky charm.
What is ISR?
Answer
In-Sync Replicas: followers that have caught up within replica.lag.time.max.ms. acks=all waits for the current ISR. Needed for durable produces.
Kafka vs RabbitMQ in one line?
Answer
Kafka = durable replayable log + consumer-group fan-out. Rabbit = smart routing broker for work queues.
How do two teams both process the same topic without stealing messages?
Answer
Different group.id values — each maintains its own offsets. That is fan-out, not competing consumers.
What happens if the leader dies with acks=all and min.insync.replicas=2?
Answer
A remaining ISR follower becomes leader. If ISR shrinks below min, producers get errors instead of silently losing data (with unclean election off).
Pitfalls
Draw one topic with three partitions, RF=3, two groups (billing and search). Assign two consumers in billing. Then add a third billing consumer. Write which members are idle. Then kill the leader with acks=all and min.insync.replicas=2 vs acks=1 + unclean election. Label error, loss, or continue.