Concurrency
Part 5 of 6 · Distributed locksLease Renewal, Clock Skew & Heartbeat Failure Modes
Renew on monotonic timers; on renew failure stop writing immediately. Clock skew and GC pauses force lease-length tradeoffs—pair renewals with fencing.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
How you know you still hold the lease
Prefer
Monotonic renew + fail-closed + fencing
Heartbeat on time.monotonic / performance.now. Token must still match. Unknown renew → stop writes. Storage still rejects a stale fence after a pause the heartbeat missed.
- T/3 gives three chances under moderate jitter; adjust with data.
- Hard expiry in the lock service is authoritative; soft TTL is a UX hint.
- Single failed renew is not loss if TTL remains — retry inside the remaining budget.
Alternative
Wall-clock sleep, or write until you feel the TTL
NTP steps, client clocks ahead of the service, and GC pauses past T all produce zombies or false loss. Redlock validity math from client clocks is the same bug.
- Do not invent a soft TTL that lets you write after the service granted the lock elsewhere.
- Do not block stopping writes on unlock success.
- If GC regularly exceeds T, shrink the critical section or fix the runtime.
Heartbeat renewal
Acquire, renew, or lose. Losing means stop — then fence.
- 1
Acquire with TTL T plus a fence
The grant is time-bounded. Depth of the token: fencing lesson. - 2
Renew every T/3 if token matches
Compare-and-expire. Redis Lua PEXPIRE; etcd keep-alive. - 3
Retry a blip inside remaining budget
Cap retries so you cannot renew after another client acquired (token check). Distinguish timeout vs NOT_OWNER. - 4
Renew fail or TTL expiry → lost
Set a volatile lost flag or cancel a context. Stop writing immediately. Best-effort unlock; do not wait on it.
Overview
Distributed locks in production are leases: you must heartbeat/renew before TTL, and on any renew failure you must stop writing immediately. Clock skew, GC pauses, and network blips create false expiries or zombie holders; soft vs hard TTL and lease-length vs pause-budget are the tuning knobs.
Pair renew discipline with fencing tokens — renewal alone does not make TTL locks safe.
You should be able to:
- Draw acquiring → holding (renew OK) → lost (stop writes).
- Recite the TTL inequality from p99 work, RTT, and GC budget.
- Refuse client wall clocks as a safety mechanism (the Redlock critique).
Lease lifecycle
Holding is not "renew forever." A successful heartbeat returns you to Holding; failure is Lost; polite unlock is Releasing.
Decisions
- 1
1 Acquiring
- acquire OK + fence2 Holding
- 2
2 Holding
- next3 Renew result
- unlock5 Releasing
- ?
3 Renew result
- OK2 Holding
- fail / TTL expiry4 Lost — stop writes
- 4
4 Lost — stop writes
- 5
5 Releasing
Lesson map
Lease Renewal, Clock Skew & Heartbeat Failure Modes
Renew on monotonic timers; on renew failure stop writing immediately. Clock skew and GC pauses force lease-length tradeoffs—pair renewals with fencing.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB c["1 Client"] l["2 Lock service"] c -->|acquire TTL=30s| l c -->|renew if token| l l -->|OK| c l -->|FAIL| c
Renew OK returns to Holding. Fail or TTL expiry is Lost — stop writes. Unlock is a separate exit.
Heartbeat renewal
Pattern:
- Acquire with TTL T.
- Renew every T/3 (rule of thumb) via compare-and-expire (token must still match).
- If renew fails or times out beyond budget → treat lock as lost.
- Critical section code checks a volatile lost flag or cancels a context.
Sequence
- 1
1 Client → 2 Lock service
acquire TTL=30s
- 2
1 Client → 2 Lock service
renew if token matches
- 3
2 Lock service → 1 Client
OK
- 4
1 Client → 2 Lock service
renew if token matches
- 5
2 Lock service → 1 Client
FAIL
- 6
1 Client
cancel context — stop writes
Soft TTL vs hard TTL
| Axis | Soft TTL | Hard TTL |
|---|---|---|
| Meaning | Hint to renew / refresh | Absolute expiry in the lock service |
| Client | May continue until hard expiry | Must stop at loss detection |
| Storage | Should still fence | Must fence |
| Example | Cache soft-refresh | etcd lease, Redis PX |
Never invent a soft TTL that lets you write after the service already granted the lock elsewhere.
Clock skew failure modes
| Skew | Symptom | Mitigation |
|---|---|---|
| Client clock ahead | Thinks lease still valid; service expired | Use service TTL only; monotonic renew deadlines via time.monotonic / performance.now |
| Client clock behind | Renews too early (OK) or mis-schedules | Monotonic timers for scheduling |
| NTP step backwards | Sleep schedules wake late | Monotonic clocks; avoid wall clock for intervals |
| Redis vs client | Rare for server-side TTL | Do not compute safety from client wall clock (Redlock critique) |
Rule: schedule heartbeats with a monotonic clock; let the lock service own expiry.
Network blips
- Single failed renew ≠ lost lock if TTL still valid — retry within remaining budget.
- Cap retries so you cannot renew after another client acquired (token check).
- Distinguish transport timeout vs definitive
NOT_OWNER. - If unknown, prefer fail-closed for critical writes.
Multi-region renew latency: RTT eats the budget — prefer regional locks.
Lease length vs GC pause
| Longer lease | Shorter lease |
|---|---|
| Survives longer STW / blips | Faster failover after crash |
| Slower recovery when holder dies | More renew traffic |
| Larger zombie window without fencing | More false loss under load |
Size T from:
T > work_slice_p99 + renew_rtt_p99 + gc_budget + safety_marginIf GC regularly exceeds T, fix the runtime or the design (smaller critical sections). A lease much longer than work is OK for low-churn leaders; still renew and fence; document failover SLO.
Should you unlock on renew failure? Best-effort; do not block stopping writes on unlock success.
Sandbox: lost flag stops writes (Python)
Monotonic ticks, not wall clock. A missed renew past budget sets lost and further writes are skipped.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same idea (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
Work p99 is 200ms, renew RTT p99 is 20ms, GC budget 150ms. Pick T and a T/3 heartbeat. Now inject a 400ms STW. Did you lose the lock? What does the lost flag do, and what must storage do with the old fence?
Interview Q&A
Why renew at T/3?
Answer
Three chances to renew before expiry under moderate jitter; adjust with data.
Wall clock or monotonic for renew timers?
Answer
Monotonic for intervals; never NTP-affected wall clock.
Soft vs hard TTL?
Answer
Hard expiry is authoritative in the service; soft is a client UX hint only.
Renew timeout but lock still valid?
Answer
Retry until safety deadline; if unknown, prefer fail-closed for critical writes.
GC pause longer than TTL?
Answer
You lost the lock; fencing must save you; also reduce pause or critical section size.
Leader election heartbeats vs locks?
Answer
Same shape: lease + renew + step down on failure (Raft leadership is the gold-standard story — see the Raft hub).
Multi-region renew latency?
Answer
RTT eats the budget — prefer regional locks.
Should unlock on renew failure?
Answer
Best-effort; do not block stopping writes on unlock success.
Clock skew between Redis nodes (Redlock)?
Answer
Core critique — avoid depending on synchronized client validity math.
How to test renew paths?
Answer
Fault-inject renew errors; assert no writes after loss; assert fence rejects late writes.
Lease much longer than work?
Answer
OK for low churn leaders; still renew and fence; document failover SLO.
One-liner?
Answer
Renew on monotonic timers; on failure cancel the critical section; fence at storage.
Go Deeper
- Kleppmann — distributed locking
- etcd leases and keep-alive
- Redis SET / PEXPIRE
- Related heartbeat shape: Raft leader election
- Next: CAS and optimistic coordination