TopicsDistributed systems
Distributed systems
Raft consensus, replication, consistent hashing, saga-style distributed transactions, two-phase commit, and conflict-free replicated data types.
Common tags: raft, consensus, replication, sagas, two-phase-commit, crdt
- Distributed systems
Vector Clocks vs Version Vectors - Detecting Concurrent Writes, Siblings & Dotted Version Vectors
Cluster · Time, Clocks & Ordering
Vector clock rules and four-way compare; runnable Dynamo-style sibling store with context and merge; vector clocks vs version vectors vs client-id vclocks vs dotted version vectors; runnable LWW vs per-server VV (lost write) vs DVV (siblings); size growth and pruning; LWW vs siblings vs CRDTs vs consensus; Dynamo, Riak 2.0, Cassandra.
Open study →- distributed-systems
- clocks
- time
- ntp
- ptp
- monotonic-clock
- leap-seconds
- lamport-clocks
- happens-before
- vector-clocks
- version-vectors
- hlc
- truetime
- spanner
- cockroachdb
- snowflake
- uuidv7
- ulid
- interview
- Distributed systems
Time, Clocks & Ordering in Distributed Systems - Physical Clocks, Lamport, Vector Clocks, HLC & TrueTime
Cluster · Time, Clocks & Ordering
Interview hub: why no machine knows the real time; wall vs monotonic; the ladder from NTP wall clocks to Lamport, vector clocks, HLC and TrueTime with a decision flow; runnable LWW-on-skewed-clocks data loss vs Lamport vs vector clocks; comparison table incl. timestamp oracles; Cloudflare 2017 leap second, Spanner, CockroachDB, Dynamo, Snowflake/UUIDv7 anchors.
Open study →- distributed-systems
- clocks
- time
- ntp
- ptp
- monotonic-clock
- leap-seconds
- lamport-clocks
- happens-before
- vector-clocks
- version-vectors
- hlc
- truetime
- spanner
- cockroachdb
- snowflake
- uuidv7
- ulid
- interview
- Distributed systems
Physical Clocks - NTP/PTP, Drift & Skew, Wall vs Monotonic Time & Leap Seconds
Cluster · Time, Clocks & Ordering
How clocks are kept in sync: oscillator drift (ppm), NTP four-timestamp offset/delay math and the delay/2 error bound (runnable), slew vs step, NTP vs PTP vs cloud time (ClockBound); wall vs monotonic APIs per language; runnable lease bug under an NTP step; leap seconds step vs smear; failure catalog (LWW, leases, TTL/JWT, negative durations).
Open study →- distributed-systems
- clocks
- time
- ntp
- ptp
- monotonic-clock
- leap-seconds
- lamport-clocks
- happens-before
- vector-clocks
- version-vectors
- hlc
- truetime
- spanner
- cockroachdb
- snowflake
- uuidv7
- ulid
- interview
- Distributed systems
Lamport Clocks - Happens-Before, Logical Timestamps & Total Order
Cluster · Time, Clocks & Ordering
Happens-before precisely; Lamport clock rules and total order with (ts, pid) tie-break (runnable); runnable proof that the clock condition holds but L(a)<L(b) does not imply causality; Lamport vs wall vs vector vs HLC vs consensus log index; Raft terms and fencing tokens as logical clocks; pitfalls (tie-break, persistence).
Open study →- distributed-systems
- clocks
- time
- ntp
- ptp
- monotonic-clock
- leap-seconds
- lamport-clocks
- happens-before
- vector-clocks
- version-vectors
- hlc
- truetime
- spanner
- cockroachdb
- snowflake
- uuidv7
- ulid
- interview
- Distributed systems
Hybrid Logical Clocks & TrueTime - Commit Wait, Uncertainty Intervals & External Consistency
Cluster · Time, Clocks & Ordering
HLC algorithm (l, c) with skewed nodes and a max-offset guard (runnable); CockroachDB max-offset self-termination, MongoDB cluster time, YugabyteDB; TrueTime intervals and commit wait (runnable) plus CockroachDB-style uncertainty restarts; external consistency vs serializability vs SI; TrueTime vs HLC vs timestamp oracle vs single leader.
Open study →- distributed-systems
- clocks
- time
- ntp
- ptp
- monotonic-clock
- leap-seconds
- lamport-clocks
- happens-before
- vector-clocks
- version-vectors
- hlc
- truetime
- spanner
- cockroachdb
- snowflake
- uuidv7
- ulid
- interview
- Distributed systems
Durable Objects Real-Time - WebSocket Hibernation, Alarms, Chat, Presence & Collaborative Editing
Cluster · Durable Objects
Real-time: WebSocket Hibernation API (acceptWebSocket, tags, attachments, auto-response), real chat room with presence and history, alarms, collaborative editing via CRDT or server ordering, fan-out batching and backpressure, hibernation cost math ($20.65 vs $420.65 docs example), transport decision chart.
Open study →- distributed-systems
- cloudflare
- durable-objects
- actor-model
- workers
- edge
- sqlite
- websockets
- hibernation
- alarms
- sharding
- rate-limiter
- leases
- fencing-tokens
- concurrency
- testing
- interview
- Distributed systems
Durable Objects Storage - SQLite vs Legacy KV, Transactions, Write Coalescing & Point-in-Time Recovery
Cluster · Durable Objects
Storage: SQLite backend vs legacy KV, SQL API, sync KV API, transactionSync, write coalescing, in-memory cache, allowUnconfirmed/noCache, PITR bookmarks with ctx.abort, limits quoted from docs; output-gate and PITR simulations; storage decision chart.
Open study →- distributed-systems
- cloudflare
- durable-objects
- actor-model
- workers
- edge
- sqlite
- websockets
- hibernation
- alarms
- sharding
- rate-limiter
- leases
- fencing-tokens
- concurrency
- testing
- interview
- Distributed systems
Durable Objects Scaling - Sharding by ID, Hot Objects, Location Hints, Limits & Cost
Cluster · Durable Objects
Scaling: per-object throughput soft limits, sharding by ID, hot objects and sharded counters, location hints and jurisdictions, cold starts and eviction, pricing; real rate limiter, lease with fencing token and counter run under wrangler dev; shard sizing and global-vs-per-key simulations.
Open study →- distributed-systems
- cloudflare
- durable-objects
- actor-model
- workers
- edge
- sqlite
- websockets
- hibernation
- alarms
- sharding
- rate-limiter
- leases
- fencing-tokens
- concurrency
- testing
- interview
- Distributed systems
Durable Objects in Production - Resets, Deploys & Class Migrations, Observability, Testing & Interview Q&A
Cluster · Durable Objects
Production: failure catalog (deploys, eviction, write failure resets, blockConcurrencyWhile throws, overloaded vs retryable errors), version skew, declarative exports class lifecycle (create, rename three-deploy alias, transfer, delete), at-least-once alarms, observability, testing with @cloudflare/vitest-plugin (4 real passing tests), interview Q&A.
Open study →- distributed-systems
- cloudflare
- durable-objects
- actor-model
- workers
- edge
- sqlite
- websockets
- hibernation
- alarms
- sharding
- rate-limiter
- leases
- fencing-tokens
- concurrency
- testing
- interview
- Distributed systems
Durable Objects Concurrency - Single Thread, Input & Output Gates, blockConcurrencyWhile & the Races That Remain
Cluster · Durable Objects
Concurrency: single thread plus input and output gates, where interleaving still happens (fetch, timers, other objects), blockConcurrencyWhile cost, the external-call oversell race reproduced under wrangler dev (reserved 6 of 3) and fixed with claim-first and version check.
Open study →- distributed-systems
- cloudflare
- durable-objects
- actor-model
- workers
- edge
- sqlite
- websockets
- hibernation
- alarms
- sharding
- rate-limiter
- leases
- fencing-tokens
- concurrency
- testing
- interview
- Distributed systems
Distributed ID Generation - Snowflake, UUIDv4 vs UUIDv7, ULID, KSUID & Sequences
Cluster · Time, Clocks & Ordering
ID schemes compared (sequences, hi-lo, Flickr ticket servers, Snowflake, Instagram, UUIDv4, UUIDv7/RFC 9562, ULID, KSUID); runnable Snowflake with sequence overflow and clock-rollback handling; runnable UUIDv7 with monotonic counter plus B-tree right-edge locality vs UUIDv4; ordering vs locality vs coordination; hot-partition twist; JS 2^53 pitfall.
Open study →- distributed-systems
- clocks
- time
- ntp
- ptp
- monotonic-clock
- leap-seconds
- lamport-clocks
- happens-before
- vector-clocks
- version-vectors
- hlc
- truetime
- spanner
- cockroachdb
- snowflake
- uuidv7
- ulid
- interview
- Distributed systems
Cloudflare Durable Objects - Single-Instance Actors, Edge State & When to Use Them
Cluster · Durable Objects
Hub: a Durable Object is a single-threaded actor with its own SQLite, one live instance per ID worldwide; routing with getByName/idFromName/newUniqueId, stubs and RPC; DO vs KV, D1, R2, Queues, Redis, Postgres row locks, Orleans and Akka with a decision chart; runnable routing simulation; what happens if you pick the alternative.
Open study →- distributed-systems
- cloudflare
- durable-objects
- actor-model
- workers
- edge
- sqlite
- websockets
- hibernation
- alarms
- sharding
- rate-limiter
- leases
- fencing-tokens
- concurrency
- testing
- interview
- Distributed systems
CRDTs — Conflict-Free Types, Convergence & When Consensus Wins
Cluster · CRDTs
Interview map for conflict-free replicated data types: a join that converges, the state versus op versus delta split, and the invariants that still belong on Raft or one ledger.
Open study →- distributed-systems
- crdt
- eventual-consistency
- convergence
- collaborative-editing
- interview
- Distributed systems
State-based vs Op-based vs Delta-CRDTs & Compaction
Cluster · CRDTs
The same abstract type can ship as full state, as operations, or as deltas. The channel, the fresh replica, and the tombstone pile follow from that choice.
Open study →- distributed-systems
- crdt
- eventual-consistency
- convergence
- collaborative-editing
- interview
- Distributed systems
Sets & Maps — G-Set, 2P-Set, OR-Set, OR-Map
Cluster · CRDTs
G-Set only adds. 2P-Set removes forever. An observed-remove set tags each add with a dot so a later add can win. An OR-Map nests a CRDT under each key.
Open study →- distributed-systems
- crdt
- eventual-consistency
- convergence
- collaborative-editing
- interview
- Distributed systems
Sequences & Collaborative Text — RGA, LSEQ, Yjs/Automerge
Cluster · CRDTs
Sequence CRDTs give each insert a stable identity so two people typing at the same place converge. RGA, LSEQ, Yjs, and Automerge are that idea with different identifiers. A last-writer-wins string is not.
Open study →- distributed-systems
- crdt
- eventual-consistency
- convergence
- collaborative-editing
- interview
- Distributed systems
Production CRDTs & When NOT to Use Them (Riak, Redis CRDT, collab apps, vs Raft/linearizability)
Cluster · CRDTs
Riak data types, Redis Active-Active, and Yjs or Automerge are production CRDTs. Unique names, non-negative money, and exactly-once effects still belong on a linearizable store.
Open study →- distributed-systems
- crdt
- eventual-consistency
- convergence
- collaborative-editing
- interview
- Distributed systems
Counters & Registers — G-Counter, PN-Counter, LWW & MV-Register
Cluster · CRDTs
G-Counter and PN-Counter merge per-replica counts with a component-wise max. Last-writer-wins drops a concurrent value. The multi-value register keeps it.
Open study →- distributed-systems
- crdt
- eventual-consistency
- convergence
- collaborative-editing
- interview
- Distributed systems
XA & Database 2PC - Postgres PREPARE, MySQL XA & Heuristic Decisions
Cluster · Two-Phase Commit — Protocol, Coordinator & Participants
X/Open XA, Postgres PREPARE TRANSACTION, MySQL XA, and heuristic commit or rollback.
Open study →- distributed-systems
- two-phase-commit
- 2pc
- xa
- atomic-commit
- interview
- Distributed systems
2PC vs 3PC vs Consensus Commit vs Sagas
Cluster · Two-Phase Commit — Protocol, Coordinator & Participants
Why 3PC splits under partition, Paxos Commit as F>0 2PC, and when to use the existing saga series.
Open study →- distributed-systems
- two-phase-commit
- 2pc
- xa
- atomic-commit
- interview
- Distributed systems
Two-Phase Commit — Protocol, Coordinator & Participants
Cluster · Two-Phase Commit — Protocol, Coordinator & Participants
Interview hub: 2PC coordinator, votes, forced logs, blocking vs 3PC, Paxos Commit, and sagas.
Open study →- distributed-systems
- two-phase-commit
- 2pc
- xa
- atomic-commit
- interview
- Distributed systems
Optimizations - Presumed Abort, Presumed Commit & Read-Only Votes
Cluster · Two-Phase Commit — Protocol, Coordinator & Participants
Presumed abort, presumed commit, and read-only votes without shrinking the uncertainty window.
Open study →- distributed-systems
- two-phase-commit
- 2pc
- xa
- atomic-commit
- interview
- Distributed systems
Prepare & Commit - Votes, Logging & the Decision
Cluster · Two-Phase Commit — Protocol, Coordinator & Participants
Prepare and commit votes, force-before-send logging, and the decision record.
Open study →- distributed-systems
- two-phase-commit
- 2pc
- xa
- atomic-commit
- interview
- Distributed systems
Failures & Recovery - Blocking, Coordinator Crash & Participant Crash
Cluster · Two-Phase Commit — Protocol, Coordinator & Participants
Coordinator and participant crashes, blocking, termination rules, and heuristic decisions.
Open study →- distributed-systems
- two-phase-commit
- 2pc
- xa
- atomic-commit
- interview
- Distributed systems
Two-Phase Commit vs Sagas — Why 2PC Breaks at Scale
Cluster · Sagas & Distributed Transactions
Why 2PC blocks at scale, how that compares with sagas, and when a shared-database ACID transaction still wins.
Open study →- distributed-systems
- sagas
- transactions
- orchestration
- choreography
- compensation
- consistency
- interview
- Distributed systems
Sagas & Distributed Transactions — Orchestration, Choreography & Compensations
Cluster · Sagas & Distributed Transactions
Interview hub on sagas versus 2PC, orchestration versus choreography, compensations, deadlines, and reconciliation for multi-service business transactions.
Open study →- distributed-systems
- sagas
- transactions
- orchestration
- choreography
- compensation
- consistency
- interview
- Distributed systems
Saga State Machines — Timeouts, Retries & Deadlines
Cluster · Sagas & Distributed Transactions
Saga state machines with per-step timers, whole-saga deadlines, bounded retries, and durable state.
Open study →- distributed-systems
- sagas
- transactions
- orchestration
- choreography
- compensation
- consistency
- interview
- Distributed systems
Saga Failure Modes — Poison Steps, Partial Failure & Reconciliation
Cluster · Sagas & Distributed Transactions
Poison steps, partial failure, dead-letter quarantine, and idempotent reconciliation for sagas that miss their SLA.
Open study →- distributed-systems
- sagas
- transactions
- orchestration
- choreography
- compensation
- consistency
- interview
- Distributed systems
Orchestration vs Choreography — Central Coordinator vs Event Dance
Cluster · Sagas & Distributed Transactions
Central coordinator versus an event dance for sagas, with decision guidance and pointers to outbox, inbox, and deadlines.
Open study →- distributed-systems
- sagas
- transactions
- orchestration
- choreography
- compensation
- consistency
- interview
- Distributed systems
Compensating Transactions — Idempotent Undo & Semantic Rollback
Cluster · Sagas & Distributed Transactions
Idempotent semantic undo for sagas: reverse-order compensations, void versus refund, and reversing ledger entries.
Open study →- distributed-systems
- sagas
- transactions
- orchestration
- choreography
- compensation
- consistency
- interview
- Distributed systems
Raft vs Multi-Paxos vs Zab — When to Choose What
Cluster · Raft consensus
Same CFT RSM goal, different models/ops/ecosystems; prefer Raft/etcd/Consul greenfield; keep ZK when watches own estate; custom Multi-Paxos only with deep expertise.
Open study →- raft
- paxos
- zab
- zookeeper
- etcd
- consul
- kraft
- Distributed systems
Quorums & Majority — Why 2f+1, Read Quorums & Stale Reads
Cluster · Raft consensus
N=2f+1 tolerates f crashes; majority intersection prevents conflicting commits; linearizable reads need ReadIndex/lease not blind follower reads.
Open study →- raft
- quorums
- 2f+1
- read-quorum
- stale-reads
- Distributed systems
Log Replication & Commit Index — Matching, Conflict Resolution & Safety
Cluster · Raft consensus
AppendEntries + prevLog match; nextIndex backoff; truncate divergent suffixes; commitIndex on majority current-term matchIndex (Figure 8); safety sketch.
Open study →- raft
- log-replication
- commit-index
- appendentries
- figure-8
- safety
- Distributed systems
Leader Election Deep Dive — Timeouts, Randomized Election, Split Votes
Cluster · Raft consensus
Heartbeat silence → new-term election via RequestVote; exclusive votes + up-to-date log; randomized timeouts; pre-vote reduces disruption.
Open study →- raft
- leader-election
- timeouts
- split-votes
- pre-vote
- Distributed systems
Raft Consensus — Leader Election, Log Replication & Safety
Cluster · Raft consensus
Single leader, append-only log, commit after majority; terms/roles/heartbeats/log matching vs Multi-Paxos; used in etcd/Consul/TiKV/K8s metadata.
Open study →- raft
- consensus
- distributed-systems
- leader-election
- log-replication
- etcd
- quorums
- Distributed systems
Vnode Rebalancing & Membership: Stream ~K/N Without Split-Brain
Cluster · Consistent hashing
Adding a node only steals ~1/N of keys, but streaming those keys still needs throttling, versioned membership, and hinted handoff so clients and replicas do not split-brain.
Open study →- distributed-systems
- sharding
- Distributed systems
Topology-Aware Replica Placement: Racks, AZs, and Honest Quorums
Cluster · Consistent hashing
RF walks must skip the same host and prefer different racks/AZs. Quorum R+W>RF is not enough if all copies share a failure domain.
Open study →- distributed-systems
- sharding
- Distributed systems
Rendezvous Hashing (HRW): Highest Random Weight
Cluster · Consistent hashing
HRW scores hash(key, node) and picks the max. No ring to maintain; membership change remaps about 1/N; lookup is O(N) unless approximated. Weights fold into the score.
Open study →- distributed-systems
- sharding
- Distributed systems
Jump Consistent Hash: Dense Buckets, Almost No Memory
Cluster · Consistent hashing
Jump hash maps a key onto 0..N-1 with almost no memory and ~K/N movement. Buckets must be a dense integer range — no arbitrary node ids, weights, or AZ walks.
Open study →- distributed-systems
- sharding
- Distributed systems
Hot Keys & Bounded Loads: When Consistent Hashing Is Not Enough
Cluster · Consistent hashing
Consistent hashing balances key cardinality, not QPS. Salt hot partitions, coalesce, or use bounded-load / power-of-two-choices so one viral key does not melt a shard.
Open study →- distributed-systems
- sharding
- Distributed systems
Consistent Hashing: Rings, Virtual Nodes & Replica Placement
Cluster · Consistent hashing
Modulo remaps ~all keys on membership change; consistent hashing remaps ~K/N via a clockwise hash ring. Vnodes fix skew and fan out failures; RF walks collect distinct physical nodes (topology-aware).
Open study →- distributed-systems
- consistent-hashing
- virtual-nodes
- replication
- system-design
- interview
- sharding