Distributed systems
Part 3 of 5 · Raft consensusLeader Election Deep Dive — Timeouts, Randomized Election, Split Votes
Heartbeat silence → new-term election via RequestVote; exclusive votes + up-to-date log; randomized timeouts; pre-vote reduces disruption.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
How you pick a leader
Prefer
Raft randomized timeouts + RequestVote
First timeout in the randomized window campaigns. Exclusive votes + majority + up-to-date log. Simple enough to operate in etcd.
- Breaks symmetry after split votes.
- Does not depend on synced wall clocks.
- Pre-vote (etcd) stops partitioned nodes from disrupting a healthy leader.
- Timeout must be several heartbeats, not one missed packet.
Alternative
Bully, Zab ranking, or an external lease
Deterministic ID or zxid ranking, or a cloud lease. Different failure modes.
- Bully: flaky high-ID node wins forever.
- Zab: epoch + zxid — right inside ZooKeeper, not a Raft substitute.
- k8s lease: familiar, not full log consensus.
- Identical timeouts on all nodes: synchronized split votes.
Overview
Raft elections convert heartbeat silence into a new term's leader with high probability and without split-brain. The levers are election timeout, randomized backoff, RequestVote with the log up-to-date rule, and majority.
This lesson explains why fixed identical timeouts cause perpetual split votes, how etcd-style timeout ranges sit relative to heartbeat RTT, and what happens under asymmetric partitions.
Election strategies
| Strategy | How it picks a leader | Pros | Cons / failure mode |
|---|---|---|---|
| Raft randomized timeouts | First timeout wins the term | Simple; breaks symmetry | Needs timeout >> heartbeat RTT |
| Bully algorithm | Highest ID wins | Deterministic | Bad under a flaky high-ID node |
| ZooKeeper Zab election | Epoch + zxid ranking | Strong ordering tie-break | ZK-specific ops model |
| External registrar (k8s lease) | apiserver lease holder | Familiar in cloud | Not full log consensus |
What fails if you choose wrong
- Set election timeout ≤ heartbeat interval → constant spurious elections.
- Use identical timeouts on all nodes → synchronized elections → split votes forever under load.
- Ignore log up-to-date and elect the empty restarting node → truncate committed history.
Step-by-step election
Sequence
- 1
Follower A
Step 1 heartbeats stop
- 2
Follower A → Follower A
Step 2 election timeout random
- 3
Follower A → Follower A
Step 3 candidate term plus 1
- 4
Follower A → Follower B
RequestVote
- 5
Follower A → Follower C
RequestVote
- 6
Follower B → Follower A
Step 4 grant
- 7
Follower C → Follower A
Step 4 grant
- 8
Follower A
Step 5 majority becomes leader
- 9
Follower A → Follower B
Step 6 AppendEntries heartbeat
- 10
Follower A → Follower C
AppendEntries heartbeat
Lesson map
Leader Election Deep Dive — Timeouts, Randomized Election, Split Votes
Heartbeat silence → new-term election via RequestVote; exclusive votes + up-to-date log; randomized timeouts; pre-vote reduces disruption.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB f1["Follower A"] f2["Follower B"] f3["Follower C"] f1 -->|RequestVote| f2 f1 -->|RequestVote| f3 f2 -->|Step 4 grant| f1 f3 -->|Step 4 grant| f1 f1 -->|Step 6| f2 f1 -->|AppendEntries| f3
Timeouts & randomization
Typical guidance (order-of-magnitude, LAN etcd):
- Heartbeat interval: small multiple of network RTT (e.g. 50–100ms).
- Election timeout: several heartbeats (e.g. 150–300ms randomized range) so a single missed heartbeat does not flip leadership, but dead leaders are noticed quickly.
- Randomization window must be wide enough that candidates rarely timeout together.
Raft uses local timers, not synced absolute time. Clock skew does not replace randomization.
Architecture
Step 1 — Stable
- 1
Leader heartbeats
- nextFollowers reset timers
- 2
Followers reset timers
Step 2 — Silence
- 3
No heartbeat
- nextCandidates fire at random
- 4
Candidates fire at random
- nextMajority?
Step 3 — Resolve
- 5
Majority?
- yesNew leader
- split voteNew term plus random
- 6
New leader
- 7
New term plus random
Split votes
With three candidates starting in the same millisecond, each may get one vote (self) and none reaches majority. Raft's answer: increase term again after another randomized timeout. Probability of repeated ties drops quickly if randomness is good.
Vote granting rules (checklist)
- Candidate term ≥ voter
currentTerm(else reject). - Voter has not voted for a different candidate in that term.
- Candidate log is at least as up-to-date: compare last entry term, then index.
- On higher term, voter steps down and clears
votedFor.
A node with an empty log wins only if others also lack more up-to-date logs. Otherwise the vote rule blocks it — that protects committed entries.
Partitions & disruption
- Minority partition with old leader: old leader cannot commit; majority elects a new leader; old leader demotes on higher term.
- Candidate storm after a network blip: mitigated by pre-vote (etcd) — check electability before incrementing term.
- Disruption: a lagging server starting elections can annoy a healthy leader; pre-vote / check quorum before disrupting.
Split-brain? Two leaders in different terms can briefly exist; only the higher term is accepted. Same-term dual leaders are prevented by majority exclusive votes.
Sandbox: randomized vs identical timeouts (Python)
Toy Monte Carlo: unique minimum timeout vs a three-way tie. Identical timeouts always "split" in this model.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Vote grant helper (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Why randomize election timeouts?
Answer
To desynchronize candidates so one usually obtains a majority before others time out. Identical timeouts correlate elections.
Relationship between heartbeat and election timeout?
Answer
Election timeout should be meaningfully larger than heartbeat period and message RTT variance. Too small → spurious elections. Too large → slow failover.
What is pre-vote?
Answer
An optimization: ask if you could win before bumping currentTerm, reducing disruption from partitioned or lagging nodes. etcd uses it; the Raft paper's core safety does not require it.
Can a node with an empty log win?
Answer
Only if others also lack more up-to-date logs; otherwise the vote rule blocks it — protecting committed entries. This is the "restarted empty disk" interview trap.
Is split-brain possible?
Answer
Two leaders in different terms can briefly exist; the lower term steps down. Same-term dual leaders are prevented by exclusive majority votes.
What resets a follower's election timer?
Answer
Valid AppendEntries / RequestVote from the current or newer term (heartbeats included). Garbage or stale-term RPCs should not reset you forever.
Why vote only once per term?
Answer
Election safety — exclusive votes make dual majorities in one term impossible, given 2f+1 intersection.
How does etcd avoid election storms?
Answer
Tuned timings + pre-vote + stable leadership under load. Operators still must set heartbeat/election suitably for RTT (LAN defaults on WAN = pain).
Pitfalls
Three logs: A [1,1,1], B [1,1,2], C [1,1]. Term 4 election. Who can gather a majority, and why does C lose both votes? Then shrink B's last index and show a split.