Concurrency
Part 3 of 6 · Distributed locksZooKeeper / etcd Lock Recipes — Ephemeral Nodes, Sessions & Watches
Consensus locks use ephemeral/lease-bound keys, sequential recipes, and watches; session timeout is the fence boundary. Prefer Curator / etcd concurrency APIs over hand-rolled nodes.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Coordinator for a correctness-critical lock
Prefer
ZooKeeper or etcd with a battle-tested recipe
Ephemeral/lease-bound keys disappear when the session dies. Sequential children plus predecessor watches avoid a thundering herd. Session timeout is the failure detector and fence boundary.
- Quorum log (Zab / Raft) survives coordinator crash from the log.
- Curator InterProcessMutex or etcd Mutex.Lock — not copy-paste from a blog.
- Txn compare-revision is still the better tool for short single-key updates — see CAS.
Alternative
Hand-rolled ZK nodes, or Redis multi-primary
Wrong watch target wakes everyone. Session loss mid-critical-section is easy to miss. Inventing multi-primary Redis is how you get split-brain.
- Redis SET NX is faster and simpler — and not a quorum lock.
- Unlock can fail while the session still lives; always defer unlock and keep a TTL lease.
- Cross-datacenter etcd is possible; RTT and partitions usually argue for regional owners.
ZooKeeper exclusive lock recipe
Lowest sequence holds. One watch wakes one waiter.
- 1
Create ephemeral sequential child
Under /locks/resource/. The node dies with the session — no polite unlock required. - 2
List children; lowest sequence wins
Total order of waiters. Fair-ish FIFO under the recipe. - 3
Else watch the predecessor
Largest sequence lower than yours — not the parent. Parent watches are O(n) herd. - 4
On watch fire, re-check
NodeDeleted → getChildren again. You hold the lock only if you are now lowest. - 5
Session death releases
Ephemeral node gone → next waiter wakes. If pause > session timeout, you lost the lock — stop writing.
Overview
ZooKeeper and etcd give locks backed by consensus: ephemeral/lease-bound keys disappear when the session dies, sequential recipes avoid herd effects, and watches notify the next waiter. Session timeout is your failure detector and fence boundary — stronger than Redis TTL guesses, at higher ops cost.
Use Curator InterProcessMutex or etcd's concurrency lock API instead of hand-rolling node gymnastics.
You should be able to:
- Fill the Redis vs ZK/etcd comparison table from memory.
- Draw predecessor watches, not a parent watch.
- Point at Raft for quorum internals and stay on recipes here.
Why consensus locks beat Redis for critical paths
| Axis | Redis SET NX | ZK / etcd |
|---|---|---|
| Agreement | Single primary | Quorum log (Zab / Raft) |
| Crash of coordinator | Lock state may vanish or resurrect oddly | New leader continues from the log |
| Client death | TTL only | Session expiry deletes ephemeral / lease keys |
| Split-brain risk | High if you invent multi-primary | Designed to avoid dual leaders |
| Latency | Sub-ms local | ms–tens of ms (quorum) |
| Ops | Simple | Ensembles, snapshots, watches load |
Interview one-liner: Redis = fast best-effort; ZK/etcd = consensus locks for correctness-critical exclusion. Depth on Redis: SET NX EX vs Redlock.
ZooKeeper recipe (classic)
- Create ephemeral sequential child under
/locks/resource/. - List children; if your sequence is lowest → you hold the lock.
- Else watch the largest sequence lower than yours (not the parent — avoids thundering herd).
- On watch fire, re-check.
- Session death → ephemeral node gone → next waiter wakes.
Sequence
- 1
1 Client A → 2 ZooKeeper
create /locks/r/lock-0000000001 EPHEMERAL_SEQUENTIAL
- 2
3 Client B → 2 ZooKeeper
create /locks/r/lock-0000000002 EPHEMERAL_SEQUENTIAL
- 3
1 Client A → 2 ZooKeeper
getChildren — A is lowest
- 4
2 ZooKeeper → 1 Client A
acquired
- 5
3 Client B → 2 ZooKeeper
watch lock-0000000001
- 6
1 Client A
session expires / close
- 7
2 ZooKeeper → 3 Client B
NodeDeleted watch
- 8
3 Client B → 2 ZooKeeper
getChildren — B lowest
- 9
2 ZooKeeper → 3 Client B
acquired
Lesson map
ZooKeeper / etcd Lock Recipes — Ephemeral Nodes, Sessions & Watches
Consensus locks use ephemeral/lease-bound keys, sequential recipes, and watches; session timeout is the fence boundary. Prefer Curator / etcd concurrency APIs over hand-rolled nodes.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Client A"] zk["2 ZooKeeper"] b["3 Client B"] a -->|create| zk b -->|create| zk a -->|getChildren - A| zk zk -->|acquired| a b -->|watch| zk zk -->|NodeDeleted| b
Ephemeral nodes vanish when the session ends, releasing the lock without relying on polite unlock. Sequential children give a total order of waiters. Watch the predecessor so one deletion wakes one waiter, not O(n).
Apache Curator InterProcessMutex
Hand-rolled ZK locks miss edge cases (session loss mid-critical-section, wrong watch targets). Curator recipes:
- InterProcessMutex — reentrant exclusive lock
- InterProcessSemaphoreV2 — leases/count
- Handles cleanup, retries, and recipe details
Prefer Curator (or battle-tested clients) over copy-paste from a blog.
etcd lock API
etcd v3 concurrency package:
- Create a lease with TTL; keep-alive in background.
concurrency.NewMutex(session, "/locks/resource").Lock(ctx)- Lock key is bound to the lease; lease expiry → key gone → lock released.
- Transactions (
Txn) can compare revisions for CAS-style critical sections.
Flow
- 1
1 Session + lease keep-alive
- next2 Mutex.Lock
- keep-alive fail5 Lease expires / keys deleted
- 2
2 Mutex.Lock
- next3 Critical section
- 3
3 Critical section
- next4 Mutex.Unlock
- 4
4 Mutex.Unlock
- 5
5 Lease expires / keys deleted
- next6 Process must stop writing
- 6
6 Process must stop writing
If pause > session timeout, you lose the lock — treat it like fencing: stop writing. Keep-alive is the heartbeat.
If unlock fails but the session lives, you still hold until unlock or session end — always use finally / defer unlock; defend with a TTL lease anyway.
etcd (Raft, gRPC, Kubernetes ecosystem) is the usual greenfield choice; keep ZK when the estate already depends on it. Locking across datacenters is possible but RTT and partition behavior hurt; often prefer regional owners.
Acquires go through the leader log (linearizable enough to say in an interview); clients must still obey session loss. Observe hold duration, wait queue depth (children count), session expire rate, keep-alive failures.
Sandbox: sequential lock, predecessor wait (Python)
In-memory children list. Lowest sequence holds. A waiter watches only the predecessor.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same idea (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
Ten waiters under /locks/r/. You watch the parent. The holder deletes. How many clients wake? Now each waiter watches only its predecessor. How many wake? If the holder GC-pauses past session timeout, what must the old holder do before any write?
Interview Q&A
Why ephemeral nodes?
Answer
They vanish when the session ends, releasing the lock without relying on polite unlock.
Why sequential?
Answer
Total order of waiters; lowest wins; fair-ish FIFO under the recipe.
Why watch the predecessor, not the parent?
Answer
Parent watches wake everyone → O(n) herd; predecessor wakes one.
What is Curator InterProcessMutex?
Answer
Production Java recipe implementing the ZK lock protocol with edge cases handled.
How does etcd lock relate to leases?
Answer
Mutex key is attached to a lease; keep-alive renews; expiry deletes key.
Session timeout vs GC pause?
Answer
If pause > session timeout, you lose the lock — treat like fencing: stop writing.
ZK vs etcd for new systems?
Answer
etcd (Raft, gRPC, Kubernetes ecosystem) is the usual greenfield choice; keep ZK when the estate already depends on it.
Can you lock across datacenters with etcd?
Answer
Possible but RTT and partition behavior hurt; often prefer regional owners.
What happens if unlock fails but session lives?
Answer
You still hold until unlock or session end — always use finally / defer unlock; defend with TTL lease anyway.
How do you observe lock health?
Answer
Hold duration, wait queue depth (children count), session expire rate, keep-alive failures.
Is a ZK lock linearizable?
Answer
Acquires go through the leader log; clients must still obey session loss.
Redis vs ZK one-liner for interviews?
Answer
Redis = fast best-effort; ZK/etcd = consensus locks for correctness-critical exclusion.