Low-level design
Part 3 of 3 · Feature Flag ServiceFeature Flag Service - Evaluation Order & Concurrency
Why the kill switch short-circuits before the percentage, why get copies the flag, and how this round connects to progressive delivery.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is the precedence?
Answer
Missing flag, kill switch, disabled, segment miss, percentage hit or miss.
L2
Why must the kill switch be first among stored flags?
Answer
A dying feature has to go dark for everyone, including users under the percentage.
L3
Does evaluate hash a killed flag?
Answer
No. It returns KILL_SWITCH before sticky_bucket.
L4
What does get copy?
Answer
segments and variants. The caller receives a new Flag.
L5
Is the sticky bucket lock-free?
Answer
The function is pure and takes no lock. evaluate still calls it while holding the RLock.
L6
Why is one RLock enough?
Answer
upsert and evaluate are the writers and the reader. The critical section is the map.
L7
What stays on other pages?
Answer
Progressive delivery, stale-flag debt, and the segment rule model.
Failure modes
Percentage before kill switch
Buckets under the percentage stay on during an incident.
Live set returned
The caller adds a segment and the next evaluate sees it without an upsert.
Cached variant
A process that keeps the last Evaluation misses the kill switch until the cache is dropped.
Misconceptions
The hash has to run under the lock because it is shared state.
The hash reads two strings. The lock exists so the Flag fields do not tear.
An RWLock is required because evaluate is a read.
The critical section is a dictionary lookup and a few field reads. A short exclusive lock is the coding-round choice.
Most-specific match should beat this order.
This store has one order. A segment does not outrank a kill switch.
Interviewer traps
Move the hash above the kill switch to save a branch.
The branch is the product. The hash is the cheap part.
Treat the segment gate as an authorization check.
It selects a behavior. Permission is a different study.
Design scenario
Same prompt for every reader.
Requirements
Kill switch before percentage. Defensive get. One RLock. A pure sticky bucket.
Traffic / scale
Readers in evaluate, a writer in upsert.
Latency
The hash can run after the early returns. It does not need its own lock.
Consistency
A kill switch hides the percentage for every user.
Availability
Missing, disabled, and segment miss are Evaluations.
Failure assumptions
- kill_switch is true and percentage is 100.
- A caller mutates the Flag returned by get.
Constraints
- Do not reorder the kill switch after the percentage.
- Do not return the live segments set.
Prompt
Show the evaluation order and the copy that keeps callers out of the store.
API
Which reasons return before sticky_bucket is called?
Data
Which fields does get copy?
Architecture
What does the RLock cover, and what is pure?
What has to win
Prefer
Kill switch, then the rest
An incident turns the flag off for every bucket.
- DISABLED is next.
- A segment miss is next.
- The hash is last.
Alternative
Percentage first
Users under the rollout stay on while the kill switch is set.
- The reason code lies.
- The incident is not instant.
- The progressive-delivery page has nothing to call.
Evaluation precedence
- Missing flag returns FLAG_NOT_FOUND.
- Kill switch returns KILL_SWITCH.
- A disabled flag returns DISABLED.
- A segment miss returns SEGMENT_MISS.
- Percentage hit or miss.
Never reorder the kill switch after the percentage. A dying feature must go dark for everyone immediately.
def evaluate(self, key: str, user_id: str, segment: Optional[str] = None) -> Evaluation:
with self._lock:
f = self._flags.get(key)
if f is None:
return Evaluation(key, False, "FLAG_NOT_FOUND", False, None)
if f.kill_switch:
return Evaluation(key, False, "KILL_SWITCH", f.variants.get("off", False), None)
if not f.enabled:
return Evaluation(key, False, "DISABLED", f.variants.get("off", False), None)
if f.segments:
if not segment or segment not in f.segments:
return Evaluation(key, False, "SEGMENT_MISS", f.variants.get("off", False), None)
bucket = self.sticky_bucket(key, user_id)
if bucket < f.percentage:
return Evaluation(key, True, "PERCENTAGE_HIT", f.variants.get("on", True), bucket)
return Evaluation(key, False, "PERCENTAGE_MISS", f.variants.get("off", False), bucket)Concurrency
- One RLock for upsert and evaluate is enough for the coding round.
- get returns a new Flag, with a new segment set and a new variant dict.
- sticky_bucket is pure. It does not take a lock. evaluate in this module calls it before the lock is released, after the earlier returns.
Early returns, then the hash
Diagram 1. KILL_SWITCH never computes a bucket.
- 1
Lock and load
evaluate reads one flag under the RLock. - 2
Missing
FLAG_NOT_FOUND does not hash. - 3
Kill switch
KILL_SWITCH returns the off variant. - 4
Segment
A non-empty set rejects a missing or unknown segment. - 5
Bucket
Only a flag that passed the earlier checks is hashed.
Decisions
- 1
1. Lookup flag
- next2. Exists?
- ?
2. Exists?
- No3. Failure: FLAG_NOT_FOUND
- Yes4. Kill switch?
- 3
3. Failure: FLAG_NOT_FOUND
- ?
4. Kill switch?
- Yes5. Failure: KILL_SWITCH deny
- No6. Enabled?
- 5
5. Failure: KILL_SWITCH deny
- ?
6. Enabled?
- No7. DISABLED
- Yes8. Segment OK?
- 7
7. DISABLED
- ?
8. Segment OK?
- No9. SEGMENT_MISS
- Yes10. Bucket under percentage?
- 9
9. SEGMENT_MISS
- ?
10. Bucket under percentage?
- Yes11. PERCENTAGE_HIT
- No12. PERCENTAGE_MISS
- 11
11. PERCENTAGE_HIT
- 12
12. PERCENTAGE_MISS
Lesson map
Feature Flag Service - Evaluation Order & Concurrency
Why the kill switch short-circuits before the percentage, why get copies the flag, and how this round connects to progressive delivery.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Lookup flag"] b["2. Exists?"] c["3. Failure: FLAG_NOT_FOUND"] d["4. Kill switch?"] e["5. Failure: KILL_SWITCH deny"] f["6. Enabled?"] g["7. DISABLED"] h["8. Segment OK?"] i["9. SEGMENT_MISS"] j["10. Bucket under percentage?"] k["11. PERCENTAGE_HIT"] l["12. PERCENTAGE_MISS"] a -->|continues| b b -->|No| c b -->|Yes| d d -->|Yes| e d -->|No| f f -->|No| g f -->|Yes| h h -->|No| i h -->|Yes| j j -->|Yes| k j -->|No| l
Flow
- 1
Caller holds a Flag?
- getCopy segments and variants
- 2
Copy segments and variants
- 3
Hash inputs?
- flag key and userPure sticky bucket
- 4
Pure sticky bucket
Concepts used, learn more
Read the underlying idea on its own study page. This lesson applies it. It does not replace those pages.
- Feature Flags — Targeting, Experimentation & Kill Switches
- Flag Types & Evaluation — Boolean, Multivariate, Percentage & Sticky Buckets
- Targeting & Context — Attributes, Segments, Rules & Precedence
- Experimentation & A/B — Exposure, Metrics, Guardrails & Peeking
- Progressive Delivery & Kill Switches — Canary, Ramp, Instant Rollback
- Flag Architecture & Hygiene — SDK Placement, Consistency, Stale Flags & Debt
- Mutexes, Condition Variables, Deadlocks & Happens-Before
- Mutex vs RWLock
- Low-Level Design Under Time — Interfaces, State & Tradeoffs
- Hashing, Frequency Maps & Counting — Two Sum Family & Anagrams
Interview Q&A
Why is first match this fixed list?
Answer
The store has one precedence. A segment and a percentage can both look true. The kill switch still wins, and the segment wins over the percentage, because that is the order of the returns.
Why is the hash described as lock-free if evaluate holds the lock?
Answer
The function does not touch the map and does not take a lock of its own. The call site in evaluate is still inside the RLock, after the early returns. Moving the call outside the lock is safe for the hash. It is not required for the test.
What is the torn read get prevents?
Answer
A caller who removes a segment from the returned set and changes a later evaluate. The copy keeps that edit local.
Why not an RWLock by default?
Answer
The critical section is a dictionary lookup and a handful of field reads. A short exclusive lock is the usual coding-round choice until a profile says the read is long.
Does a cached Evaluation see the kill switch?
Answer
Not until the cache is dropped. That staleness is the architecture study, not a reason to hash before the kill switch.
What does the concurrent test on the solution page prove?
Answer
Workers that upsert and evaluate the same 100 percent flag do not raise. It does not pin one interleaving.
Related
The series pager also walks these pages.