Low-level design
Part 1 of 3 · Feature Flag ServiceFeature Flag Service LLD - Targeting, Sticky Buckets & Kill Switches
Provisional machine-coding LLD for an in-memory flag service: kill switch, disabled, segment gate, then a sticky SHA-256 percentage.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does evaluate return?
Answer
key, enabled, reason, variant, and an optional bucket.
L2
What is the order?
Answer
Missing, kill switch, disabled, segment, then percentage.
L3
What is the bucket input?
Answer
flag key and user id, hashed with SHA-256 and reduced modulo 100.
L4
When does a kill switch win?
Answer
Whenever kill_switch is set, including when the percentage is 100. The hash is not consulted.
L5
When does a segment miss?
Answer
When the flag lists segments and the caller segment is missing or not in that set.
L6
What does get return?
Answer
A new Flag whose segments and variants are copies, or None.
L7
What do you not rebuild here?
Answer
Remote config, exposure analytics, and the stale-flag lifecycle. Those studies already exist.
Failure modes
Random each request
The percentage can look right overall and flip for one user on refresh.
Percentage before kill switch
A killed flag still returns the on variant for buckets under the percentage.
Hash of the user alone
Unrelated flags assign the same users together.
Misconceptions
This cluster replaces the feature-flag series.
It links into that series. Targeting, experiments, ramps, and hygiene stay on those pages.
Any hash is fine.
Both services must use the same function. A language hash of the object is not stable.
The live set can be returned.
A caller who mutates it edits the store. get copies the set and the variant dict.
Interviewer traps
Start with a vendor SDK tour.
Write upsert, evaluate, and the reason codes.
Treat a segment match as authorization.
The flag exposes a behavior. Permission is a different study.
Design scenario
Same prompt for every reader.
Requirements
RLock, copied flag, SHA-256 bucket from flag key plus user id, reason code.
Traffic / scale
Many readers, rare upserts.
Latency
The lock covers the map. The hash is pure.
Consistency
The same flag key and user id always land in the same bucket.
Availability
A missing flag returns FLAG_NOT_FOUND. It does not raise.
Failure assumptions
- A writer replaces the flag while readers evaluate.
- The kill switch is set on a flag whose percentage is 100.
Constraints
- Do not draw a fresh random number per request.
- Do not run the percentage ahead of the kill switch.
Prompt
Evaluate a flag so a kill switch wins, a segment can exclude a user, and everyone else sticks to a bucket.
API
Which reason codes can evaluate return?
Data
Which two strings go into the bucket hash?
Architecture
What does get copy, and why?
What this coding round ships
Prefer
In-memory FlagStore
A kill switch, a segment gate, and a sticky percentage. Reasons name which rule won.
- Percentage is 0 to 100.
- The hash includes the flag key.
- get does not return the live set.
Alternative
The feature-flag product series
Targeting, experiments, ramps, and hygiene. Already published.
- This cluster does not replace those pages.
- Exposure math stays there.
- Stale-flag debt stays there.
Overview
This cluster is the evaluator you implement in one process. upsert stores a flag. evaluate walks a fixed order. sticky_bucket is SHA-256 of the flag key and the user id, modulo 100. Remote config, analytics, and the stale-flag lifecycle are already studies. Link them. Do not rewrite them.
Spec summary
| Operation | Rules |
|---|---|
| upsert | Reject a blank key and a percentage outside 0 to 100 |
| evaluate | Return enabled, reason, variant, and bucket |
| kill switch | Always deny with KILL_SWITCH |
| enabled | A false flag returns DISABLED |
| segments | If the set is non-empty, require membership |
| percentage | Sticky bucket under the percentage is PERCENTAGE_HIT |
Step-by-step design
- Keep boolean flags and a simple on/off variant map. Multivariate experiments are out of this round.
- Store flags under an RLock. Copy on write and on get.
- Draw the evaluation order before writing the hash.
- Make the bucket a pure function of the flag key and the user id.
- Test the kill switch, a segment miss, sticky equality, and a concurrent upsert with evaluate.
Kill switch before the hash
Diagram 1. The percentage is the last check.
- 1
Lookup
A missing key returns FLAG_NOT_FOUND. - 2
Kill switch
A set kill switch returns KILL_SWITCH and the off variant. - 3
Enabled
A disabled flag returns DISABLED without hashing. - 4
Segment
A non-empty segment set requires the caller segment. - 5
Percentage
The sticky bucket is compared with the percentage.
Decisions
- 1
1. Lookup flag
- next2. Exists?
- ?
2. Exists?
- No3. Failure: FLAG_NOT_FOUND
- Yes4. Kill switch?
- 3
3. Failure: FLAG_NOT_FOUND
- ?
4. Kill switch?
- Yes5. Failure: KILL_SWITCH deny
- No6. Enabled?
- 5
5. Failure: KILL_SWITCH deny
- ?
6. Enabled?
- No7. DISABLED
- Yes8. Segment OK?
- 7
7. DISABLED
- ?
8. Segment OK?
- No9. SEGMENT_MISS
- Yes10. Bucket under percentage?
- 9
9. SEGMENT_MISS
- ?
10. Bucket under percentage?
- Yes11. PERCENTAGE_HIT
- No12. PERCENTAGE_MISS
- 11
11. PERCENTAGE_HIT
- 12
12. PERCENTAGE_MISS
Lesson map
Feature Flag Service LLD - Targeting, Sticky Buckets & Kill Switches
Provisional machine-coding LLD for an in-memory flag service: kill switch, disabled, segment gate, then a sticky SHA-256 percentage.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Lookup flag"] b["2. Exists?"] c["3. Failure: FLAG_NOT_FOUND"] d["4. Kill switch?"] e["5. Failure: KILL_SWITCH deny"] f["6. Enabled?"] g["7. DISABLED"] h["8. Segment OK?"] i["9. SEGMENT_MISS"] j["10. Bucket under percentage?"] k["11. PERCENTAGE_HIT"] l["12. PERCENTAGE_MISS"] a -->|continues| b b -->|No| c b -->|Yes| d d -->|Yes| e d -->|No| f f -->|No| g f -->|Yes| h h -->|No| i h -->|Yes| j j -->|Yes| k j -->|No| l
Flow
- 1
Need a stable percent?
- YesHash flag key and user id
- 2
Hash flag key and user id
- 3
Need an instant off?
- YesKill switch before the hash
- 4
Kill switch before the hash
Pitfalls
- Hashing the user id alone, so unrelated flags move together.
- Checking the percentage before the kill switch.
- Returning the live segments set from get.
Concepts used, learn more
Read the underlying idea on its own study page. This lesson applies it. It does not replace those pages.
- Feature Flags — Targeting, Experimentation & Kill Switches
- Flag Types & Evaluation — Boolean, Multivariate, Percentage & Sticky Buckets
- Targeting & Context — Attributes, Segments, Rules & Precedence
- Experimentation & A/B — Exposure, Metrics, Guardrails & Peeking
- Progressive Delivery & Kill Switches — Canary, Ramp, Instant Rollback
- Flag Architecture & Hygiene — SDK Placement, Consistency, Stale Flags & Debt
- Mutexes, Condition Variables, Deadlocks & Happens-Before
- Mutex vs RWLock
- Low-Level Design Under Time — Interfaces, State & Tradeoffs
- Hashing, Frequency Maps & Counting — Two Sum Family & Anagrams
Interview Q&A
Why is this not the feature-flag series?
Answer
Those pages are the product: targeting, experiments, ramps, and hygiene. This cluster is the in-memory evaluator those ideas assume. It is provisional and it does not replace them.
Why include the flag key in the hash?
Answer
So two flags do not assign the same users together. The salt is the flag key in this round.
Why does a 100 percent flag still go dark?
Answer
Because the kill switch is checked first. The percentage is never read.
What is a sticky result?
Answer
The same flag key and user id always produce the same bucket. Changing either string moves the bucket.
Why copy the segment set?
Answer
The caller can add or remove a segment on the object they received. A copy keeps that edit off the store.
What does this round omit?
Answer
Remote config, exposure analytics, and the stale-flag cleanup. Those are the linked studies.
Related
The series pager also walks these pages.