Low-level design
Part 2 of 3 · Feature Flag ServiceFeature Flag Service - Sticky Bucket + Kill Switch Solution & Tests
Runnable FlagStore with upsert, sticky SHA-256 bucketing, segment gates, kill-switch precedence, and a concurrent upsert/evaluate test.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
How do you run the tests?
Answer
From the feature_flags directory, python3 test_flags.py -v. Six tests pass.
L2
What does test_kill_switch_wins assert?
Answer
A flag with percentage 100 and kill_switch set evaluates to enabled false and reason KILL_SWITCH.
L3
What does test_sticky_bucket_stable assert?
Answer
Two calls with the same flag key and user id match, and the bucket is in 0 to 99.
L4
What does test_percentage_rollout assert?
Answer
Percentage 0 is disabled for the user. Percentage 100 is enabled.
L5
What does test_segment_gate assert?
Answer
segment external is SEGMENT_MISS. segment internal is enabled.
L6
What does test_missing_flag assert?
Answer
Reason FLAG_NOT_FOUND.
L7
What does the concurrent test assert?
Answer
Forty threads upsert or evaluate. The error list stays empty.
Failure modes
Kill switch checked too late
test_kill_switch_wins would see the on variant. The reason would not be KILL_SWITCH.
Unstable bucket
Two calls for user-42 on dark-mode would differ. The test expects equality.
Shared segments object
upsert copies the set. A later edit of the caller's set must not change the store.
Misconceptions
The test should assert bucket 17.
It asserts equality and the range. The digest is not a fixture number.
Percentage 100 ignores the kill switch.
The kill switch returns before the comparison.
Hash the user id alone.
sticky_bucket hashes the flag key and the user id. flags.py is the module the tests import.
Interviewer traps
Paste a shorter stub than the tests import.
The module and the unittest file on this page are the full source.
Assert a particular thread order.
The concurrent test asserts that no worker raised.
Design scenario
Same prompt for every reader.
Requirements
Stdlib only. Kill switch over 100 percent. Sticky equality. Segment gate. Missing flag. Concurrent upsert and evaluate.
Traffic / scale
Forty threads on one flag whose percentage is 100.
Latency
The hash is SHA-256 of two strings.
Consistency
One flag key and one user id keep one bucket.
Availability
A missing flag is an Evaluation, not an exception. A blank key on upsert raises INVALID_KEY.
Failure assumptions
- The kill switch and the percentage are both set.
- The caller passes a segment that is not in the set.
Constraints
- Do not use the language hash of the user id.
- Do not skip the kill switch when the rollout is complete.
Prompt
Implement FlagStore and the six tests.
API
Which reason does a killed full rollout return?
Data
What does upsert copy into the map?
Architecture
Which test would fail if evaluate hashed before the kill switch?
What the six tests pin
Prefer
Reasons and equality
The kill switch, the segment, and the sticky pair are the product.
- Percentage 0 and 100 are the easy bounds.
- The bucket number is not a fixture.
- Workers must not raise.
Alternative
Skip the kill switch when the rollout is complete
An incident cannot force the off variant without editing the percentage.
- The on variant still returns.
- The reason lies.
- test_kill_switch_wins fails.
Put flags.py and test_flags.py on the path and run python3 test_flags.py -v. You want a kill switch over a full rollout, a stable bucket, a segment miss, and a concurrent run that records no errors.
Overview
FlagStore is the evaluator. sticky_bucket is a pure SHA-256. evaluate reads one flag under the lock and walks the order. The tests do not pin a magic bucket number.
Run
python3 test_flags.py -vSolution
"""Feature Flag Service - targeting, sticky bucketing, kill switch.
Sandbox: python3 -c "from flags import FlagStore; ..."
Tests: python3 test_flags.py
stdlib only. Provisional teaching implementation.
"""
from __future__ import annotations
import hashlib
import threading
from dataclasses import dataclass, field
from typing import Any, Optional
@dataclass
class Flag:
key: str
enabled: bool = False
kill_switch: bool = False
percentage: int = 0 # 0..100 sticky rollout
segments: set[str] = field(default_factory=set)
variants: dict[str, Any] = field(default_factory=lambda: {"on": True, "off": False})
@dataclass(frozen=True)
class Evaluation:
key: str
enabled: bool
reason: str
variant: Any = False
bucket: Optional[int] = None
class FlagStore:
"""In-memory flag store with sticky percentage bucketing."""
def __init__(self) -> None:
self._lock = threading.RLock()
self._flags: dict[str, Flag] = {}
def upsert(self, flag: Flag) -> None:
if not flag.key or not flag.key.strip():
raise ValueError("INVALID_KEY")
if not 0 <= flag.percentage <= 100:
raise ValueError("INVALID_PERCENTAGE")
with self._lock:
self._flags[flag.key] = Flag(
key=flag.key,
enabled=flag.enabled,
kill_switch=flag.kill_switch,
percentage=flag.percentage,
segments=set(flag.segments),
variants=dict(flag.variants),
)
def get(self, key: str) -> Optional[Flag]:
with self._lock:
f = self._flags.get(key)
return None if f is None else Flag(**{**f.__dict__, "segments": set(f.segments), "variants": dict(f.variants)})
@staticmethod
def sticky_bucket(flag_key: str, user_id: str) -> int:
h = hashlib.sha256(f"{flag_key}:{user_id}".encode()).hexdigest()
return int(h[:8], 16) % 100
def evaluate(self, key: str, user_id: str, segment: Optional[str] = None) -> Evaluation:
with self._lock:
f = self._flags.get(key)
if f is None:
return Evaluation(key, False, "FLAG_NOT_FOUND", False, None)
if f.kill_switch:
return Evaluation(key, False, "KILL_SWITCH", f.variants.get("off", False), None)
if not f.enabled:
return Evaluation(key, False, "DISABLED", f.variants.get("off", False), None)
if f.segments:
if not segment or segment not in f.segments:
return Evaluation(key, False, "SEGMENT_MISS", f.variants.get("off", False), None)
bucket = self.sticky_bucket(key, user_id)
if bucket < f.percentage:
return Evaluation(key, True, "PERCENTAGE_HIT", f.variants.get("on", True), bucket)
return Evaluation(key, False, "PERCENTAGE_MISS", f.variants.get("off", False), bucket)Tests
"""Tests for feature flag service."""
from __future__ import annotations
import threading
import unittest
from flags import Evaluation, Flag, FlagStore
class FlagStoreTests(unittest.TestCase):
def setUp(self) -> None:
self.s = FlagStore()
def test_kill_switch_wins(self):
self.s.upsert(Flag("x", enabled=True, kill_switch=True, percentage=100))
e = self.s.evaluate("x", "u1")
self.assertFalse(e.enabled)
self.assertEqual(e.reason, "KILL_SWITCH")
def test_sticky_bucket_stable(self):
a = FlagStore.sticky_bucket("dark-mode", "user-42")
b = FlagStore.sticky_bucket("dark-mode", "user-42")
self.assertEqual(a, b)
self.assertTrue(0 <= a < 100)
def test_percentage_rollout(self):
self.s.upsert(Flag("roll", enabled=True, percentage=0))
self.assertFalse(self.s.evaluate("roll", "u").enabled)
self.s.upsert(Flag("roll", enabled=True, percentage=100))
self.assertTrue(self.s.evaluate("roll", "u").enabled)
def test_segment_gate(self):
self.s.upsert(Flag("beta", enabled=True, percentage=100, segments={"internal"}))
self.assertEqual(self.s.evaluate("beta", "u", segment="external").reason, "SEGMENT_MISS")
self.assertTrue(self.s.evaluate("beta", "u", segment="internal").enabled)
def test_missing_flag(self):
e = self.s.evaluate("nope", "u")
self.assertEqual(e.reason, "FLAG_NOT_FOUND")
def test_concurrent_upsert_evaluate(self):
self.s.upsert(Flag("c", enabled=True, percentage=100))
errors = []
def worker(i: int):
try:
if i % 2 == 0:
self.s.upsert(Flag("c", enabled=True, percentage=100))
else:
assert self.s.evaluate("c", f"u{i}").enabled
except Exception as e:
errors.append(e)
threads = [threading.Thread(target=worker, args=(i,)) for i in range(40)]
for t in threads:
t.start()
for t in threads:
t.join()
self.assertEqual(errors, [])
if __name__ == "__main__":
unittest.main()The order the tests lock in
Diagram 1. test_kill_switch_wins never reaches the percentage compare.
- 1
Kill switch
percentage 100 and kill_switch still return KILL_SWITCH. - 2
Sticky
dark-mode and user-42 hash to the same bucket twice. - 3
Bounds
Percentage 0 is off. Percentage 100 is on. - 4
Segment
external misses. internal is enabled. - 5
Concurrency
Forty threads leave the error list empty.
Decisions
- 1
1. Lookup flag
- next2. Exists?
- ?
2. Exists?
- No3. Failure: FLAG_NOT_FOUND
- Yes4. Kill switch?
- 3
3. Failure: FLAG_NOT_FOUND
- ?
4. Kill switch?
- Yes5. Failure: KILL_SWITCH deny
- No6. Enabled?
- 5
5. Failure: KILL_SWITCH deny
- ?
6. Enabled?
- No7. DISABLED
- Yes8. Segment OK?
- 7
7. DISABLED
- ?
8. Segment OK?
- No9. SEGMENT_MISS
- Yes10. Bucket under percentage?
- 9
9. SEGMENT_MISS
- ?
10. Bucket under percentage?
- Yes11. PERCENTAGE_HIT
- No12. PERCENTAGE_MISS
- 11
11. PERCENTAGE_HIT
- 12
12. PERCENTAGE_MISS
Lesson map
Feature Flag Service - Sticky Bucket + Kill Switch Solution & Tests
Runnable FlagStore with upsert, sticky SHA-256 bucketing, segment gates, kill-switch precedence, and a concurrent upsert/evaluate test.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Lookup flag"] b["2. Exists?"] c["3. Failure: FLAG_NOT_FOUND"] d["4. Kill switch?"] e["5. Failure: KILL_SWITCH deny"] f["6. Enabled?"] g["7. DISABLED"] h["8. Segment OK?"] i["9. SEGMENT_MISS"] j["10. Bucket under percentage?"] k["11. PERCENTAGE_HIT"] l["12. PERCENTAGE_MISS"] a -->|continues| b b -->|No| c b -->|Yes| d d -->|Yes| e d -->|No| f f -->|No| g f -->|Yes| h h -->|No| i h -->|Yes| j j -->|Yes| k j -->|No| l
Concepts used, learn more
Read the underlying idea on its own study page. This lesson applies it. It does not replace those pages.
- Feature Flags — Targeting, Experimentation & Kill Switches
- Flag Types & Evaluation — Boolean, Multivariate, Percentage & Sticky Buckets
- Targeting & Context — Attributes, Segments, Rules & Precedence
- Experimentation & A/B — Exposure, Metrics, Guardrails & Peeking
- Progressive Delivery & Kill Switches — Canary, Ramp, Instant Rollback
- Flag Architecture & Hygiene — SDK Placement, Consistency, Stale Flags & Debt
- Mutexes, Condition Variables, Deadlocks & Happens-Before
- Mutex vs RWLock
- Low-Level Design Under Time — Interfaces, State & Tradeoffs
- Hashing, Frequency Maps & Counting — Two Sum Family & Anagrams
Interview Q&A
Why copy on upsert and on get?
Answer
The caller's set and dict must not be the objects in the map. upsert stores new containers. get returns new containers.
What does percentage 0 return?
Answer
enabled false and PERCENTAGE_MISS, after the flag has passed kill switch, enabled, and segment. The bucket is still computed.
What does a missing segment argument do when segments is non-empty?
Answer
SEGMENT_MISS. The caller did not name a member of the set.
What does the concurrent test refuse to prove?
Answer
A particular interleaving. It proves the workers did not raise while upsert and evaluate shared the flag.
Which import does the test need?
Answer
from flags import Evaluation, Flag, FlagStore.
Does this page replace the feature-flag series?
Answer
No. Those pages own targeting, experiments, ramps, and hygiene. This page is the coding-round module.
Related
The series pager also walks these pages.