High-level design
Part 1 of 3 · SlotWiseSlotWise Reservation Service Incident - HLD Debug & Fix Path
A reservation service starts double-booking slots under load. Reproduce, trace isolation and locking choices, fix with the right transaction boundaries, and ship a regression test.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is the symptom?
Answer
Two 201 responses for one slot_id, and two active rows.
L2
What is the suspect code?
Answer
Autocommit sessions and if not taken: insert.
L3
Which isolation still loses?
Answer
Read Committed does not make that check-then-act safe.
L4
What is the default fix?
Answer
A partial unique index on slot_id where status is active, and IntegrityError mapped to 409.
L5
When do you add FOR UPDATE?
Answer
When the slot row is the thing you want to serialize. It is optional next to the index.
L6
Why not UNIQUE(slot_id, user_id)?
Answer
Two different users both succeed. The invariant is one active reservation per slot.
L7
What does the regression test assert?
Answer
Under concurrent clients, one 201 and the rest 409.
Failure modes
Double 201
Both transactions saw no active row and both inserted.
Wrong unique key
A constraint on slot and user still allows two users.
Blocking migration
A unique index on a large table without a concurrent build locks writers for too long.
Misconceptions
Read Committed hides the race.
Both transactions can see zero active rows and both can insert.
Serializable replaces the constraint.
It can work with retries. The unique index documents the invariant and fails as 409.
This is an engineering-practices page because it is an incident.
The topic is the reservation design and the isolation choice. The debug method is a linked study.
Interviewer traps
Test with one thread.
The bug needs overlapping check-then-act.
Hold the HTTP call open with no timeout.
A lock you forget to release becomes a stuck slot.
Design scenario
Same prompt for every reader.
Requirements
Reproduce on Postgres. Name the isolation level. Ship one fix and a regression test. Keep cancelled reservations possible.
Traffic / scale
One popular slot, many clients, low create rate everywhere else.
Latency
The conflict path should fail fast with 409, not wait on a retry storm.
Consistency
At most one active reservation per slot.
Availability
A taken slot is 409. A missing slot is not a 500.
Failure assumptions
- Two transactions read taken=false before either inserts.
- A user cancels and another user books the same slot later.
Constraints
- Do not use a unique key that includes user_id.
- Do not put the page topic on engineering-practices.
Prompt
Stop SlotWise from giving one slot to two users under 50 concurrent clients.
API
Which status codes distinguish created, taken, and missing?
Data
What does the partial unique index include, and what does it exclude?
Architecture
Where does the transaction start relative to the insert?
Where the invariant lives
Prefer
Partial unique index on the active slot
The database rejects the second insert even if the application forgets a lock.
- Cancelled rows stay allowed.
- IntegrityError becomes 409.
- A load test is part of the fix.
Alternative
Raise isolation and keep the Python check
Serializable can help and still needs retries. The check-then-act remains a footgun.
- Read Committed definitely loses.
- A missed retry becomes a 500.
- The business rule is clearer as a constraint.
Overview
A reservation service starts double-booking slots under load. You must reproduce, trace isolation and locking choices, fix with the right transaction boundaries, and ship a regression test. This is an HLD/incident narrative that bridges to LLD debugging and database isolation studies.
Incident brief
SlotWise lets users POST /reservations {slot_id, user_id}. Under 50 concurrent clients for one popular slot, two 201 responses appear for the same slot_id. Dashboard shows reservations rows with the same slot and different users. Autocommit sessions and a Python if not taken: insert path are suspects.
Step-by-step response
- Freeze deploys; capture DB isolation (
SHOW transaction_isolation), SQLAlchemy session scope, and the reserve endpoint code. - Reproduce with a small Python thread/async flood against a local Postgres.
- Hypothesize: read-then-write race under Read Committed.
- Prove with logging of both transactions seeing
taken=false. - Fix options (pick and defend):
- Unique constraint on
slot_idwhere active + catch IntegrityError SELECT ... FOR UPDATEon slot row inside a transaction- Serializable + retry on serialization failure
- Unique constraint on
- Add migration + regression test that fails on the old code.
- Roll forward; watch conflict metrics.
From double 201 to one winner
Diagram 1. Read Committed lets both inserts land. A constraint or a row lock leaves one winner.
- 1
Symptom: two 201s
One popular slot, two active rows, two users. - 2
Reproduce
Flood a local Postgres with concurrent clients. One thread will not show it. - 3
Trace check then insert
Both transactions read no active row. - 4
Read Committed does not save you
The phantom insert is allowed. Both commits succeed. - 5
Fix
Partial unique index or SELECT FOR UPDATE inside one transaction. - 6
Regression
The load test expects one 201 and conflicts for the rest.
Decisions
- 1
1. Symptom: double 201
- next2. Reproduce with N workers
- 2
2. Reproduce with N workers
- next3. Trace: check then insert
- 3
3. Trace: check then insert
- next4. Isolation prevents phantom?
- ?
4. Isolation prevents phantom?
- No under RC5. Failure path: both insert
- Yes with lock or constraint6. One winner one conflict
- 5
5. Failure path: both insert
- next7. Fix: uniqueness or FOR UPDATE
- 6
6. One winner one conflict
- 7
7. Fix: uniqueness or FOR UPDATE
- next8. Regression test under load
- 8
8. Regression test under load
Lesson map
SlotWise Reservation Service Incident - HLD Debug & Fix Path
A reservation service starts double-booking slots under load. Reproduce, trace isolation and locking choices, fix with the right transaction boundaries, and ship a regression test.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Symptom: double 201"] b["2. Reproduce with N workers"] c["3. Trace: check then insert"] d["4. Isolation prevents phantom?"] e["5. Failure path: both insert"] f["6. One winner one conflict"] g["7. Fix: uniqueness or FOR UPDATE"] h["8. Regression test under load"] a -->|continues| b b -->|continues| c c -->|continues| d d -->|No under RC| e d -->|Yes with lock or constraint| f e -->|continues| g g -->|continues| h
Concepts used, learn more
Read the underlying idea on its own study page. This lesson applies it. It does not replace those pages.
- Interview Debugging & Implementation — Reproduce, Trace, Fix, Ship
- MVCC, Snapshot Isolation & Write Skew
- Isolation Levels Deep Dive
- SSI vs Snapshot Isolation
- Distributed Locks — Correctness, Leases & Fencing Tokens
- Mutexes, Condition Variables, Deadlocks & Happens-Before
- When Locks Win — Contention, Fairness & Hybrid Designs
- Low-Level Design Under Time — Interfaces, State & Tradeoffs
- API Design — Naming, Paths, Routing & Contracts
| Concept | Role | Study page |
|---|---|---|
| Interview debug method | Reproduce, trace, fix, ship | Interview Debugging & Implementation — Reproduce, Trace, Fix, Ship |
| MVCC / snapshot isolation | Why RC allows the race | MVCC, Snapshot Isolation & Write Skew |
| Isolation levels | RC vs RR vs Serializable | Isolation Levels Deep Dive |
| Distributed locks / fencing | When app-level locks appear | Distributed Locks — Correctness, Leases & Fencing Tokens |
| Mutex intuition | Local analog of critical section | Mutexes, Condition Variables, Deadlocks & Happens-Before |
| LLD interfaces | Service invariants for reserve | Low-Level Design Under Time — Interfaces, State & Tradeoffs |
Comparative: fix strategies
| Fix | Pros | Cons | Verdict |
|---|---|---|---|
| UNIQUE(slot_id) active | Simple, DB enforced | Need partial unique index for cancels | Strong default |
| SELECT FOR UPDATE | Explicit critical section | Must lock the right row; easy to miss | Good when slot row exists |
| Serializable + retry | General | Latency under contention; app retry logic | Use when multi-row invariants |
| Advisory lock | Flexible | Easy to forget; not row-backed | Situational |
Pitfalls
- Testing only with one thread.
- Unique constraint on
(slot_id, user_id)that still allows two users. - Holding HTTP connections open across slow commits without timeouts.
- Migrating a unique index on a huge table without
CREATE INDEX CONCURRENTLYdiscipline.
Interview Q&A
Why do two requests both get 201?
Answer
Each transaction saw no active row under Read Committed and each inserted.
Why not UNIQUE(slot_id, user_id)?
Answer
Two different users both succeed. The rule is one active reservation for the slot.
Why a partial index?
Answer
A cancelled row must not block the next active reservation for that slot.
When is FOR UPDATE enough on its own?
Answer
When every reserve path locks the same slot row inside the transaction. A forgotten path still needs the unique index.
What does the regression test assert?
Answer
Concurrent clients produce one 201 and 409 for the rest.
Why is the topic hld?
Answer
The lesson is the reservation invariant and the isolation choice. Interview debugging is a linked study, not the page topic.
Related
The series pager also walks these pages.