Concurrency
Part 2 of 6 · Distributed locksRedis SET NX EX vs Redlock — Single-Instance Safety & the Kleppmann Debate
Atomic SET NX EX is best-effort on one Redis; Redlock’s majority story is contested (Kleppmann). Use single Redis + fencing for soft work; etcd/ZK/DB CAS for correctness-critical paths.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Where the lock lives
Prefer
Single Redis SET NX EX + fencing for soft work; etcd/ZK/CAS when safety matters
Atomic acquire-if-absent with a TTL. Unlock compares the random token. Critical sections still fence at storage. Correctness-critical exclusion does not ride independent Redis clocks.
- NX and EX together avoid the set-then-expire crash window.
- Lua GET+DEL is atomic so you cannot delete the next owner's lock.
- Redlock is the wrong interview answer for ledgers, inventory, or unique constraints.
Alternative
Redlock as consensus, or DEL without a token
Majority of independent Redis masters looks like a quorum. Kleppmann: pauses and clock assumptions break safety. Bare DEL after expiry steals the next lock.
- Independent Redis nodes are not a replicated log.
- Async replica promotion can resurrect a lock you thought was gone.
- Validity math from client wall clocks is the Redlock critique in one line.
Single-instance acquire and release
One Redis. Token identity is the whole unlock protocol.
- 1
SET key token NX PX 10000
One round-trip. OK means you hold the lease for about 10s. nil means someone else holds it. - 2
Do the work under the lease
Finish before TTL, or renew with compare-and-expire. GC can still exceed TTL — fence at storage. - 3
Unlock with Lua if GET == token
Compare-and-delete. Bare DEL after expiry deletes the next holder's lock. - 4
Next client SET NX PX succeeds
Deadlock freedom: crashes expire. Mutual exclusion was only as strong as this one Redis.
Overview
SET key token NX EX ttl is an atomic best-effort lock on a single Redis: acquire if absent, unlock only with a matching token via Lua. Redlock spreads that idea across N independent Redis nodes for majority quorum — but clocks, GC pauses, and correlated failures make it a poor choice for correctness-critical mutual exclusion.
Prefer single Redis + fencing for soft coordination. Prefer etcd / ZooKeeper / DB CAS when safety matters.
You should be able to:
- Explain why NX and EX must be one command.
- Contrast single-instance Redis with a consensus lock without re-teaching Raft.
- Refuse Redlock for money, inventory, or unique constraints.
Atomic SET NX EX
Redis SET with NX (only if Not eXists) and EX/PX (TTL) is one round-trip and atomic on one server:
SET lock:ledger <random-token> NX PX 10000- Success → you hold the lease for ~10s.
- Unlock must be compare-and-delete: only delete if
value == your token(Lua), or you delete someone else's lock after expiry/reacquire.
Sequence
- 1
1 Client A → 2 Redis
SET lock:x tokenA NX PX 10000
- 2
2 Redis → 1 Client A
OK
- 3
3 Client B → 2 Redis
SET lock:x tokenB NX PX 10000
- 4
2 Redis → 3 Client B
nil not acquired
- 5
1 Client A → 2 Redis
EVAL unlock if value==tokenA
- 6
2 Redis → 1 Client A
1
- 7
3 Client B → 2 Redis
SET lock:x tokenB NX PX 10000
- 8
2 Redis → 3 Client B
OK
Lesson map
Redis SET NX EX vs Redlock — Single-Instance Safety & the Kleppmann Debate
Atomic SET NX EX is best-effort on one Redis; Redlock’s majority story is contested (Kleppmann). Use single Redis + fencing for soft work; etcd/ZK/DB CAS for correctness-critical paths.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Client A"] r["2 Redis"] b["3 Client B"] a -->|SET lock:x| r r -->|OK| a b -->|SET lock:x| r r -->|nil not acquired| b a -->|EVAL unlock if| r r -->|1| a
Unlock with token + Lua
Never DEL blindly:
-- KEYS[1]=lock key, ARGV[1]=token
if redis.call("GET", KEYS[1]) == ARGV[1] then
return redis.call("DEL", KEYS[1])
else
return 0
endA random token exists so unlock only clears your lock if the key was reacquired after expiry by someone else. GET+DEL must be atomic; otherwise two clients can interleave and delete the wrong lock.
Clock and process assumptions (single instance)
| Assumption | Reality |
|---|---|
| Client finishes work before TTL | GC / STW / SIGSTOP can exceed TTL |
| Redis clock for TTL is fine | Redis uses an approximate LRU clock; TTL expiry is server-side OK |
| Network is timely | Delayed unlock after reacquire races without a token check |
| One Redis is enough | Failover / async replica promotion can resurrect locks |
Single-instance Redis lock is not consensus. It is useful for cache stampede coalescing and best-effort jobs if the shared resource is fenced or the work is idempotent. Stampede waiter behavior is single-flight.
What is Redlock?
Antirez's Redlock algorithm:
- Get time t1.
- Sequentially
SET NX PXon N independent Redis masters (often 5). - Count successes; need majority (
N/2+1). - Validity time ≈ TTL − (t2−t1) − clock drift.
- Unlock: release on all nodes (best-effort).
Goal: survive minority Redis failures without a single coordinator.
Decisions
- 1
1 Read time t1
- next2 SET NX PX on N masters
- 2
2 SET NX PX on N masters
- next3 Majority N/2+1?
- ?
3 Majority N/2+1?
- no4 Fail — unlock all
- yes5 Validity = TTL minus elapsed minus drift
- 4
4 Fail — unlock all
- 5
5 Validity = TTL minus elapsed minus drift
Kleppmann: pauses and clock assumptions break safety; a majority of independent Redis is not the same as consensus; use fencing or a real consensus service for critical sections.
Defense: improves availability vs one node for soft locks; many workloads are best-effort; fencing can be added.
When is Redlock the wrong answer? Any time losing mutual exclusion corrupts money, inventory, or unique constraints — interviewers expect etcd/ZK/DB.
Do not "fix" multi-node Redis by inventing another Redlock. Put the lock in etcd/ZK or use DynamoDB conditional writes / Postgres row versions. Depth: ZK/etcd, CAS.
Leader election for soft leadership (cache warmer) maybe; for metadata ownership use Raft-based systems — see Raft hub, not this page.
Redis Cluster
Hash-tag the key to one slot. During resharding/failover expect weaker guarantees — another reason fencing matters.
Renew
PEXPIRE only if GET still equals your token (Lua). On renew failure: stop writing immediately. Depth: lease renewal.
Export acquire latency, acquire failures, hold duration, unlock mismatches, renew failures, fence rejects downstream.
Sandbox: token unlock vs bare DEL (Python)
In-memory Redis. After A's TTL, B acquires. A's late DEL would steal B's lock; token compare-and-delete would not.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same idea (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
Five independent Redis masters. You get OK from three, then a 15s GC pause on a 10s TTL. Who holds the lock on each node? What does storage do if you never sent a fencing token? Now repeat with one Redis plus a monotonic fence at the DB.
Interview Q&A
Why NX and EX together?
Answer
NX gives acquire-if-free; EX bounds the lease so crashes do not deadlock forever. Atomic together avoids "set then expire" races.
Why a random token?
Answer
So unlock only clears your lock if the key was reacquired after expiry by someone else.
What does Lua buy you on unlock?
Answer
GET+DEL must be atomic; otherwise two clients can interleave and delete the wrong lock.
Is single Redis locking CP?
Answer
No. One node; async replicas can lose locks on failover; not a consensus quorum.
Summarize Kleppmann's Redlock critique.
Answer
Pauses and clock assumptions break safety; majority of independent Redis is not the same as consensus; use fencing or a real consensus service for critical sections.
Summarize the defense of Redlock.
Answer
Improves availability vs one node for soft locks; many workloads are best-effort; fencing can be added.
When is Redlock the wrong answer?
Answer
Any time losing mutual exclusion corrupts money, inventory, or unique constraints — interviewers expect etcd/ZK/DB.
How do you renew a Redis lock?
Answer
PEXPIRE only if GET still equals your token (Lua). On renew failure: stop writing immediately.
Redis Cluster and locks?
Answer
Hash-tag the key to one slot; during resharding/failover expect weaker guarantees — another reason fencing matters.
Alternative to Redlock for multi-node Redis?
Answer
Don't. Put the lock in etcd/ZK or use DynamoDB conditional writes / Postgres row versions.
Can you use Redlock just for leader election?
Answer
Leader election for soft leadership (cache warmer) maybe; for metadata ownership use Raft-based systems.
What metrics do you export?
Answer
Acquire latency, acquire failures, hold duration, unlock mismatches, renew failures, fence rejects downstream.