High-level design
Part 3 of 3 · SlotWiseSlotWise - Load Repro & Isolation Decision Chart
How to reproduce the double-book under load, which isolation level still loses without a constraint, and a decision chart for constraint versus row lock versus serializable retries.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is the repro setup?
Answer
Docker Postgres, default Read Committed, one seeded slot, the old SELECT then INSERT.
L2
How many clients?
Answer
Fifty threads. One thread will not double-book.
L3
What do you observe before the fix?
Answer
Two active reservations for that slot.
L4
What do you observe after the index?
Answer
One 201 and IntegrityError-shaped 409s.
L5
When is the partial unique index the whole fix?
Answer
The invariant is one active row per slot, a single-row rule.
L6
When do you add serializable?
Answer
The rule spans several rows and a constraint cannot say it. Budget the retries.
L7
Why not only raise isolation?
Answer
The constraint documents the rule and fails fast. Serializable without it still needs a correct retry.
Failure modes
Repro with one client
The race never opens. You ship the old code.
Serializable without retries
A serialization failure becomes a 500 instead of a 409.
Isolation change without a metric
You cannot tell conflicts from errors. Count reservation_conflicts_total.
Misconceptions
Repeatable Read makes check-then-act safe here.
Two inserts of a new active row can still both commit if nothing they read conflicts. The unique index is the conflict.
SSI replaces a primary business rule.
SSI versus snapshot isolation explains write skew. One active reservation is still clearest as a constraint.
An advisory lock is the same as the index.
It is easy to forget on one code path and it is not row-backed.
Interviewer traps
Tune isolation in production before the local repro.
Fifty threads on a laptop already show two active rows.
Log only the 201s.
The conflict counter is the page you operate.
Design scenario
Same prompt for every reader.
Requirements
A load repro. A decision between partial unique, row lock, and serializable. A metric for conflicts.
Traffic / scale
Fifty concurrent clients on one slot. The rest of the slots are quiet.
Latency
Conflicts fail fast. Serializable retries need a budget so they do not stampede.
Consistency
One active reservation per slot is true after every commit.
Availability
A conflict is a 409 the client can retry against a different slot or a later cancel.
Failure assumptions
- Postgres is Read Committed.
- The application counts active rows and then inserts.
Constraints
- Do not call the bug fixed because the isolation level string changed.
- Do not skip the load test in CI.
Prompt
Choose the isolation and locking story for SlotWise after you have reproduced the double book.
API
What does the client see on the losing request?
Data
Which rows does the partial index cover after the rerun?
Architecture
When is serializable the branch you take instead of the unique index?
What you change first
Prefer
Partial unique index, then rerun the load
The second run is the proof. One 201, and conflicts you can count.
- Read Committed stays the default.
- The invariant is in the schema.
- FOR UPDATE is an optional extra on the slot row.
Alternative
Serializable alone
Possible with retries, and a weak place to hide a check-then-act.
- Serialization failures need a budget.
- The business rule is less obvious.
- A quiet load test can miss it.
Overview
How to reproduce the double-book under load, which isolation level still loses without a constraint, and a decision chart for constraint vs row lock vs serializable retries. Ties the incident back to Study pages on isolation and debug method.
Minimal repro outline
- Docker Postgres with default Read Committed.
- Seed one slot.
- Old code path:
SELECTcount active; if zeroINSERT. - 50 threads; observe two actives.
- Apply unique index; rerun; observe one 201 and IntegrityError-driven 409s.
Decision chart
Constraint, lock, or serializable
Diagram 1. A single-slot rule takes the unique index. A multi-row rule takes serializable or a careful lock, then a retry budget.
- 1
One slot or many rows?
A single active reservation is not a multi-row invariant. - 2
Partial unique index
Index slot_id where status is active. - 3
Map the conflict
IntegrityError becomes 409. - 4
Multi-row branch
Serializable or explicit locking when one index cannot say the rule. - 5
Retry budget
Count conflicts. Do not loop forever. - 6
Ship the load test
CI runs the flood. One winner is the assertion.
Flow
- 1
1. Multi-row invariant?
- No single slot2. Partial UNIQUE on slot
- Yes complex3. Serializable or careful locking
- 2
2. Partial UNIQUE on slot
- next4. Map IntegrityError to 409
- 3
3. Serializable or careful locking
- next5. Retry budget and metrics
- 4
4. Map IntegrityError to 409
- next6. Optional FOR UPDATE on slot row
- 5
5. Retry budget and metrics
- next7. Ship load test in CI
- 6
6. Optional FOR UPDATE on slot row
- next7. Ship load test in CI
- 7
7. Ship load test in CI
Lesson map
SlotWise - Load Repro & Isolation Decision Chart
How to reproduce the double-book under load, which isolation level still loses without a constraint, and a decision chart for constraint versus row lock versus serializable retries.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Multi-row invariant?"] b["2. Partial UNIQUE on slot"] c["3. Serializable or careful locking"] d["4. Map IntegrityError to 409"] e["5. Retry budget and metrics"] f["6. Optional FOR UPDATE on slot row"] g["7. Ship load test in CI"] a -->|No single slot| b a -->|Yes complex| c b -->|continues| d c -->|continues| e d -->|continues| f f -->|continues| g e -->|continues| g
When isolation alone is not enough
Read Committed does not make check-then-act safe. Repeatable Read / SSI help some write-skew shapes but a uniqueness constraint remains the clearest invariant for "one active reservation per slot."
Concepts used, learn more
Read the underlying idea on its own study page. This lesson applies it. It does not replace those pages.
- Interview Debugging & Implementation — Reproduce, Trace, Fix, Ship
- MVCC, Snapshot Isolation & Write Skew
- Isolation Levels Deep Dive
- SSI vs Snapshot Isolation
- Distributed Locks — Correctness, Leases & Fencing Tokens
- Mutexes, Condition Variables, Deadlocks & Happens-Before
- When Locks Win — Contention, Fairness & Hybrid Designs
- Low-Level Design Under Time — Interfaces, State & Tradeoffs
- API Design — Naming, Paths, Routing & Contracts
Interview Q&A
Why not only raise isolation to Serializable?
Answer
It can work with retries, but the unique constraint documents the invariant and fails fast with a clear 409. Use both if the business rule is sacred.
Where do you log?
Answer
Correlation id from the gateway, transaction ids if available, and a reservation_conflicts_total counter.
Why does Repeatable Read still lose without a constraint?
Answer
Two transactions can each see zero active rows and each insert a new one. Nothing they updated conflicts until the unique index makes the second insert fail.
What do you seed?
Answer
One slot, default Read Committed, and the old count-then-insert path.
What do you log?
Answer
A correlation id and reservation_conflicts_total. Transaction ids if you have them.
Why mention CREATE INDEX CONCURRENTLY?
Answer
On a large reservations table the unique index should not take an exclusive lock for the whole build.
Related
The series pager also walks these pages.