Distributed systems
Part 2 of 6 · Time, Clocks & OrderingPhysical Clocks - NTP/PTP, Drift & Skew, Wall vs Monotonic Time & Leap Seconds
How clocks are kept in sync: oscillator drift (ppm), NTP four-timestamp offset/delay math and the delay/2 error bound (runnable), slew vs step, NTP vs PTP vs cloud time (ClockBound); wall vs monotonic APIs per language; runnable lease bug under an NTP step; leap seconds step vs smear; failure catalog (LWW, leases, TTL/JWT, negative durations).
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is skew, and what is drift?
Answer
Skew is the offset at one instant. Drift is the rate the clock gains or loses time, in ppm.
L2
Why is NTP's error bound delay/2?
Answer
The offset formula assumes equal path times. Asymmetry can be as large as half the round trip.
L3
When should NTP step instead of slew?
Answer
Step large offsets, often at boot, when slewing would take too long. Slew small offsets so time stays smooth.
L4
Which APIs are monotonic?
Answer
CLOCK_MONOTONIC, time.monotonic, performance.now, System.nanoTime, and Go's monotonic reading inside time.Time since Go 1.9.
L5
What does a backwards NTP step do to a wall-clock lease?
Answer
Elapsed wall time shrinks, so the holder can believe the lease is still valid after it expired.
L6
What is leap smearing?
Answer
Spread the extra second over a long window by running clocks slightly slow, so no second repeats. Google's published smear is 24 hours, noon to noon UTC.
L7
What still saves you across machines?
Answer
A safety margin for drift plus a fencing token, because a pause can outlast the lease.
Failure modes
Backwards NTP step extends a wall-clock lease
In the example, 12 seconds truly elapse, the wall clock steps back 3 seconds, and wall elapsed reads 9000 ms against a 10 second lease.
Negative duration around a repeated second
A wall clock that repeats a second yields a negative subtraction. That is how the 2017 Cloudflare DNS panic started.
Last-writer-wins on a slow node
A node with a slow clock produces writes that always lose when conflicts are resolved by timestamp.
Misconceptions
Synced to 1 ms means correctness.
That is a typical case, not a bound. Asymmetry, a bad source, or a VM pause can blow it.
Disabling steps makes time safe.
A badly wrong clock then takes hours to slew back, and every stamp written meanwhile is wrong.
A monotonic clock can be stored as a timestamp.
The value is meaningless after reboot and across machines.
Interviewer traps
Quoting a vendor's typical NTP accuracy as a correctness bound.
State the delay/2 bound and say the path can be asymmetric.
Fixing the local lease measurement and forgetting the other machine.
Monotonic time is local. The grantor still needs a margin and a fencing token.
Design scenario
Same prompt for every reader.
Requirements
A holder must not keep acting after the grantor has expired the lease. Durations on one process must ignore wall-clock steps.
Traffic / scale
Example lease: 10 seconds. Example step: 3 seconds backwards in the middle, with 12 seconds of true elapsed time.
Latency
Slewing is capped, so a large offset is stepped instead of waiting hours.
Consistency
The protected resource must reject a stale holder even if that holder still believes the lease is valid.
Availability
Refusing steps entirely leaves a wrong clock wrong until slew catches up.
Failure assumptions
- NTP can step the wall clock backwards.
- A process can pause longer than the lease.
- The holder and grantor drift a few milliseconds between syncs.
Constraints
- Do not compute the lease from Date.now or time.time.
- Do not rely on a 1 ms sync claim as a bound.
Prompt
A lock service grants 10 second leases. NTP sometimes steps a holder backwards by a few seconds.
API
What does the holder renew, and which clock does it read?
Data
What fencing token does the resource store and compare?
Architecture
Where is the monotonic measurement, and where is the cross-machine check?
Which clock measures the lease
Prefer
Monotonic elapsed time, plus a fence
On one machine, CLOCK_MONOTONIC, time.monotonic, or performance.now ignores an NTP step. Across machines, add a margin and a fencing token.
- A backwards step no longer extends the lease.
- A forward step no longer expires it early on that host.
- A pause longer than the lease still needs a fence.
Alternative
Subtract two wall-clock reads
Date.now, time.time, and CLOCK_REALTIME jump when NTP steps or a leap second repeats.
- The example 3 second step backwards leaves a 12 second lease looking unexpired.
- A repeated second produces a negative duration.
- Two machines still disagree after the local measurement is fixed.
From oscillator to a number you can trust
Physical time only. Logical clocks start on the next page.
- 1
Estimate offset
offset = ((t2 - t1) + (t3 - t4)) / 2 and delay = (t4 - t1) - (t3 - t2). The error is up to delay/2. - 2
Prefer the lowest-delay sample
In the example, the congested sample estimates 2.5 ms instead of about 24.5 ms. The minimum-delay sample stays near the true +25 ms. - 3
Slew small errors, step large ones
ntpd steps above 128 ms by default and limits slewing to 500 ppm. A step can move time backwards. - 4
Measure leases monotonically
The example backwards step of 3 seconds leaves wall elapsed at 9000 ms and monotonic elapsed at 12000 ms.
Overview
A computer's wall clock is a quartz oscillator plus a correction service. The oscillator drifts (tens of ppm is normal, more with temperature), NTP or PTP nudges it back, and sometimes the correction steps the clock, including backwards. On top of that, UTC occasionally inserts a leap second. The practical rules are simple: measure durations with a monotonic clock, treat cross-machine wall-clock comparisons as approximate with a known error bound, never let correctness depend on two machines agreeing on the time, and know how your platform handles leap seconds.
How a clock is kept in sync
- Oscillator: a crystal ticks at a nominal frequency, but real frequency varies with manufacturing and temperature. A 50 ppm error is about 4.3 seconds per day.
- Time source: GPS receivers or atomic references feed stratum-1 servers; clients sync from those (directly or via more strata).
- Measurement: the client exchanges timestamped packets and estimates offset and round-trip delay.
- Correction: small offsets are slewed (run the clock slightly fast or slow until it catches up); large ones are stepped (jump). The reference ntpd steps when the offset exceeds 128 ms by default and limits slewing to 500 ppm; chrony's
makestepsetting controls when it is allowed to step. - Discipline: the daemon also learns the oscillator's frequency error so the clock drifts less between polls.
Sequence
- 1
"Client"
"Step 1: record t1 and send request"
- 2
"Client" → "NTP server"
"request (t1)"
- 3
"NTP server"
"Step 2: record t2 on arrival"
- 4
"NTP server"
"Step 3: record t3 on reply"
- 5
"NTP server" → "Client"
"reply (t1, t2, t3)"
- 6
"Client"
"Step 4: record t4 on arrival"
- 7
"Client"
"Step 5: offset = ((t2 - t1) + (t3 - t4)) / 2, delay = (t4 - t1) - (t3 - t2)"
- 8
"Client"
"Step 6: prefer low-delay samples, then slew (small error) or step (large error)"
Lesson map
Physical Clocks - NTP/PTP, Drift & Skew, Wall vs Monotonic Time & Leap Seconds
How clocks are kept in sync: oscillator drift (ppm), NTP four-timestamp offset/delay math and the delay/2 error bound (runnable), slew vs step, NTP vs PTP vs cloud time (ClockBound); wall vs monotonic APIs per language; runnable lease bug under an NTP step; leap seconds step vs smear; failure catalog (LWW, leases, TTL/JWT, negative durations).
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB c["Client"] s["NTP server"] c -->|request (t1)| s s -->|reply (t1, t2, t3)| c
The offset math, drift and smearing (runnable)
The offset formula assumes the network path is symmetric. If the return path is slower, the estimate is wrong by up to half the round-trip delay, which is why NTP prefers the lowest-delay samples.
# How NTP estimates clock offset from four timestamps, why the error bound is delay/2,
# how much an unsynced crystal drifts, and what a 24h linear leap smear looks like.
# Timestamps are EXAMPLE values in milliseconds.
def offset_delay(t1, t2, t3, t4):
"""t1 client send, t2 server receive, t3 server send, t4 client receive."""
offset = ((t2 - t1) + (t3 - t4)) / 2 # how far the server is ahead of the client
delay = (t4 - t1) - (t3 - t2) # round trip minus server processing time
return offset, delay
# Several samples against the same server; the true offset is +25 ms.
# Sample 3 hit a congested return path (asymmetric delay), which skews its estimate.
samples = [
(0, 35, 36, 22),
(1000, 1037, 1038, 1026),
(2000, 2036, 2037, 2068), # slow way back
(3000, 3034, 3035, 3020),
]
print("sample offset_ms delay_ms error_bound(+/-delay/2)")
results = []
for i, s in enumerate(samples, 1):
off, d = offset_delay(*s)
results.append((d, off))
print(f"{i:>6} {off:>9.1f} {d:>8.1f} {d/2:>8.1f}")
# NTP's clock filter prefers the lowest-delay sample: least room for asymmetry error.
best_delay, best_off = min(results)
print(f"chosen (min delay) offset = {best_off:+.1f} ms (true +25.0)\n")
# Drift: a quartz oscillator off by N parts per million gains/loses N microseconds per second.
for ppm in (10, 50, 200):
per_day_ms = ppm * 86_400 / 1_000
print(f"drift {ppm:>3} ppm -> {per_day_ms:>7.1f} ms/day unsynced, {ppm*60/1000:>5.2f} ms after a 60 s sync gap")
# Leap smear: instead of inserting 23:59:60, spread the extra second over 24 hours
# (noon to noon UTC, the published Google approach). Clocks run ~11.6 ppm slow meanwhile.
print("\nhours into 24h smear -> smeared clock lags UTC-without-leap by (ms)")
for h in (0, 6, 12, 18, 24):
print(f" {h:>2}h -> {1000 * h / 24:6.1f}")
print("smear rate =", round(1e6 / 86_400, 2), "ppm")Output:
sample offset_ms delay_ms error_bound(+/-delay/2)
1 24.5 21.0 10.5
2 24.5 25.0 12.5
3 2.5 67.0 33.5
4 24.5 19.0 9.5
chosen (min delay) offset = +24.5 ms (true +25.0)
drift 10 ppm -> 864.0 ms/day unsynced, 0.60 ms after a 60 s sync gap
drift 50 ppm -> 4320.0 ms/day unsynced, 3.00 ms after a 60 s sync gap
drift 200 ppm -> 17280.0 ms/day unsynced, 12.00 ms after a 60 s sync gap
hours into 24h smear -> smeared clock lags UTC-without-leap by (ms)
0h -> 0.0
6h -> 250.0
12h -> 500.0
18h -> 750.0
24h -> 1000.0
smear rate = 11.57 ppmNTP vs PTP vs cloud time services
| Option | Typical accuracy (example ranges) | How | Good for | Watch out for |
|---|---|---|---|---|
| Public NTP over the internet | Milliseconds to tens of ms | UDP packets, software timestamps | General servers, laptops | Asymmetric paths, untrusted sources |
| Datacenter / cloud NTP (for example Amazon Time Sync, Google internal NTP) | Sub-millisecond to low ms | Local stratum-1 sources, short paths | Most production fleets | Know whether it smears leap seconds |
| PTP (IEEE 1588) with hardware timestamping | Sub-microsecond in good networks | NIC and switch timestamps, boundary/transparent clocks | Trading, telecom, databases that exploit tight bounds | Needs supporting NICs and switches |
| GPS / atomic references per datacenter | Tightest bounds | Local reference clocks | Spanner-style TrueTime | Cost, operational expertise |
Meta has written publicly about moving parts of its fleet from NTP to PTP, and AWS publishes the open-source ClockBound daemon, which exposes an error bound (earliest/latest) instead of a single "now". That interval idea is exactly what TrueTime uses later in this cluster.
Wall clock vs monotonic clock
| Wall clock | Monotonic clock | |
|---|---|---|
| Meaning | Time of day (Unix epoch) | Time since an arbitrary point (often boot) |
| Can go backwards? | Yes (NTP step, admin change, leap second handling) | No |
| Comparable across machines? | Approximately | No |
| Use for | Timestamps humans read, calendar logic, expiry dates stored in data | Timeouts, latency, rate limits, lease countdowns, retry backoff |
| Linux | CLOCK_REALTIME | CLOCK_MONOTONIC (CLOCK_BOOTTIME also counts suspend) |
| Python | time.time() | time.monotonic(), time.perf_counter() |
| JavaScript | Date.now() | performance.now() |
| Java | System.currentTimeMillis() | System.nanoTime() |
| Go | time.Now() wall part | Monotonic reading inside time.Time (since Go 1.9) |
Why measuring a lease with the wall clock is a bug (runnable)
// Wall clock vs monotonic clock for measuring elapsed time, under an NTP step.
// We simulate both clocks so the run is deterministic; real code would use
// Date.now() (wall) vs performance.now() (monotonic).
class SimClocks {
private trueMs = 0; // real elapsed time since start
private wallOffset = 0; // adjustments applied by NTP/admins to the wall clock only
advance(ms: number): void { this.trueMs += ms; }
stepWall(deltaMs: number): void { this.wallOffset += deltaMs; } // NTP step / manual change
wall(): number { return 1_700_000_000_000 + this.trueMs + this.wallOffset; }
monotonic(): number { return this.trueMs; } // only ever moves forward at a steady rate
}
const LEASE_MS = 10_000; // example lease: 10 s
function scenario(label: string, stepMs: number): void {
const c = new SimClocks();
const wallStart = c.wall();
const monoStart = c.monotonic();
c.advance(6_000); // 6 s pass
c.stepWall(stepMs); // NTP steps the wall clock
c.advance(6_000); // 6 more seconds pass: 12 s truly elapsed, lease SHOULD be expired
const wallElapsed = c.wall() - wallStart;
const monoElapsed = c.monotonic() - monoStart;
console.log(label);
console.log(` true elapsed 12000 ms`);
console.log(` wall elapsed ${String(wallElapsed).padStart(9)} ms -> lease ${wallElapsed < LEASE_MS ? "STILL HELD (bug)" : "expired"}`);
console.log(` mono elapsed ${String(monoElapsed).padStart(9)} ms -> lease ${monoElapsed < LEASE_MS ? "still held" : "expired (correct)"}`);
}
scenario("A) wall clock stepped BACK 3 s mid-lease:", -3_000);
scenario("B) wall clock stepped FORWARD 5 s mid-lease:", 5_000);
// A negative duration is how Cloudflare's 2017 leap-second bug started: code computed
// a duration from wall-clock reads that went backwards, then passed it to a function
// that requires a positive value (Go's rand.Int63n), which panicked.
const before = 1_483_228_800_000; // 2017-01-01T00:00:00Z
const after = before - 1_000; // wall clock repeats a second around the leap
console.log(`\nwall-clock 'duration' across a repeated second: ${after - before} ms (negative!)`);Output:
A) wall clock stepped BACK 3 s mid-lease:
true elapsed 12000 ms
wall elapsed 9000 ms -> lease STILL HELD (bug)
mono elapsed 12000 ms -> lease expired (correct)
B) wall clock stepped FORWARD 5 s mid-lease:
true elapsed 12000 ms
wall elapsed 17000 ms -> lease expired
mono elapsed 12000 ms -> lease expired (correct)
wall-clock 'duration' across a repeated second: -1000 ms (negative!)ExpectedA) wall clock stepped BACK 3 s mid-lease: true elapsed 12000 ms wall elapsed 9000 ms -> lease STILL HELD (bug) mono elapsed 12000 ms -> lease expired (correct) B) wall clock stepped FORWARD 5 s mid-lease: true elapsed 12000 ms wall elapsed 17000 ms -> lease expired mono elapsed 12000 ms -> lease expired (correct) wall-clock 'duration' across a repeated second: -1000 ms (negative!)
Press Run. Snippets must be self-contained — no network, files, or native modules.
A backwards step makes a lease holder think it still owns the lease after it has expired everywhere else; a forward step makes it give up early. The monotonic clock gets both right on one machine. Across machines you still need a safety margin (the grantor and holder count time at slightly different rates) and fencing tokens, because a process can be paused (GC, VM migration) for longer than the lease.
Leap seconds: step, repeat, or smear
Since 1972, UTC has inserted 27 leap seconds to stay aligned with Earth's rotation. There are three ways a system can experience one:
- Step / repeat: the kernel shows 23:59:59 twice (or a 23:59:60). Code that assumes time never repeats sees zero or negative durations.
- Smear: the time service runs clocks slightly slow over a window (Google smears linearly over 24 hours, noon to noon UTC) so no second is ever repeated. Clocks are up to about half a second off "true" UTC during the window.
- Ignore and resync: the clock is simply wrong by one second until NTP steps it.
Mixing smeared and non-smeared sources in one fleet is the worst option: machines disagree by up to a second for a day. Pick one policy for the whole fleet. The CGPM voted in 2022 to stop adding leap seconds by 2035, but the code paths still have to handle existing data and the remaining years.
Why timestamps lie: the failure catalog
- Last-writer-wins data loss: a node with a slow clock produces writes that always "lose". Cassandra resolves cell conflicts by timestamp, so client or coordinator clock skew directly decides which write survives.
- Lease and lock expiry bugs: a process pauses or its clock jumps, it keeps acting as leader after the lease expired; without fencing, two leaders write.
- TTL and cache expiry: a server whose clock is a minute fast expires sessions or tokens early; one a minute slow accepts expired JWTs (allow a small, bounded leeway for
expandnbf). - Negative durations: metrics, rate limiters and backoff code that subtract two wall-clock reads.
- Out-of-order logs: merged logs from many machines sorted by timestamp show effects before causes; use trace IDs and span parent links for causality.
What happens if you choose otherwise
- Disable NTP stepping entirely to avoid backwards jumps: a badly wrong clock then takes hours to slew back, during which every timestamp it writes is wrong.
- Rely on "our clocks are synced to 1 ms" for correctness: that is a typical case, not a bound. Network asymmetry, a bad time source, or a VM pause can blow it, and you will not see it until data is lost.
- Use a monotonic clock for persisted timestamps: values are meaningless after reboot and across machines.
Interview Q&A
What is the difference between clock skew and clock drift?
Answer
Skew (offset) is how far apart two clocks are at one instant; drift is the rate at which a clock gains or loses time. Drift causes skew to grow between synchronizations.
Why does NTP's error bound depend on round-trip delay?
Answer
NTP assumes the outbound and return paths take equal time. If they don't, the offset estimate is off by up to half the round trip, so a 20 ms round trip means up to 10 ms of uncertainty even with perfect timestamps.
When should NTP slew instead of step?
Answer
Slew small offsets so time stays monotonic and smooth; step only large offsets (typically at boot or after a long outage) when slewing would take too long. Many teams allow steps only at startup.
How would you make a distributed lease safe despite clock problems?
Answer
Measure the lease locally with a monotonic clock, have the holder give it up before the grantor's expiry with a margin for drift, and attach a monotonically increasing fencing token that the protected resource checks, so a stale holder's writes are rejected.
What is leap smearing and what are its trade-offs?
Answer
Spreading the extra second over a long window by slowing clocks slightly, so time never repeats. The trade-off is that smeared clocks differ from strict UTC by up to about half a second during the window, so every machine and dependency should agree on the same smear.
How large is a 50 ppm drift?
Answer
A 50 ppm oscillator gains or loses 50 microseconds per second, which is 4320 ms per day unsynced, and 3.00 ms after a 60 second sync gap in the example.
Why does mixing smeared and non-smeared time sources hurt?
Answer
Machines disagree by up to about a second for a day. Pick one leap-second policy for the whole fleet and its dependencies.
What does the lowest-delay NTP sample buy you?
Answer
Less room for path asymmetry. In the example the minimum-delay sample has a 19 ms delay and a 9.5 ms error bound, and its offset is +24.5 ms against a true +25 ms.
Check yourself
Find one timeout or lease in a service you know. Say whether it subtracts wall-clock reads. If the clock stepped back by the example 3 seconds, would the holder keep acting, and what fence would reject it?
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Lease Renewal, Clock Skew & Heartbeat Failure Modes, Distributed Locks — Correctness, Leases & Fencing Tokens, Stream Processing — Event Time, Windows, State & Exactly-Once, Replication Lag & Session Guarantees — Read-Your-Writes, Monotonic Reads & Consistent Prefix.
Go Deeper
- RFC 5905: Network Time Protocol Version 4
- Cloudflare: How and why the leap second affected Cloudflare DNS
- Google Public NTP: Leap Smear
- Meta Engineering: How Precision Time Protocol is being deployed at Meta
- AWS ClockBound (error-bounded clock daemon)
- Go time package: Monotonic Clocks
- Martin Kleppmann, How to do distributed locking