Flag Types & Evaluation — Boolean, Multivariate, Percentage & Sticky Buckets
Flag types and evaluation mechanics decide whether your rollout is trustworthy. Booleans gate on or off. Multivariate flags return named variants. Percentage rollouts need a sticky hash so the same subject stays in the same bucket. This lesson implements those sketches, contrasts hash choices, and shows why a fresh random draw breaks experiments and UX.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does a boolean flag return?
Answer
True or false, after rules, with a coded default when the SDK has no rules or the flag is archived.
L2
What does a multivariate flag return?
Answer
One variant key, such as control, treatment_a, or treatment_b, and sometimes a payload attached to that key.
L3
How do you assign 50, 25, and 25 without reshuffling?
Answer
Hash the unit into a bucket from 0 to 99. Walk cumulative percent ranges. The same hash hits the same range until the salt or the ranges change.
L4
Why is the flag key inside the hash?
Answer
So assignment to flag A is independent of flag B. Omit it and the same users bunch up across experiments.
L5
What breaks if you rotate the salt during the experiment?
Answer
Users move between variants. Metrics mix two experiences, and the UI flips under people who were already treated.
L6
MurmurHash3 or SHA-256?
Answer
MurmurHash3 is what flag SDKs usually ship: fast and well distributed, not a cryptographic hash. SHA-256 is the portable demo. A language hashCode is not stable across services.
L7
When does evaluation stop early?
Answer
A kill or force variant wins before segments and percentages. An archived flag returns the default and does not pretend the old rules still exist.
Failure modes
Fresh random each request
The percentage is right in aggregate and wrong for every user who refreshes.
Salt rotated mid-experiment
The population reshuffles. Lift and the UX both become uninterpretable.
Nested booleans
a and b and not c explode the matrix. A named variant or a layered flag is the smaller model.
Two services, two seeds
The UI shows the treatment and the API still runs the control.
Misconceptions
Any hash is fine if the modulo looks uniform.
The hash has to be the same function in every service. A runtime hashCode is not.
random() is acceptable for a user-facing canary.
It is a load-test coin flip at best. User-facing features use a sticky unit.
A JSON payload flag is just config, so it needs no default.
Invalid payloads fail to the coded default. The schema is part of the flag.
Interviewer traps
Implement MurmurHash3 on the whiteboard.
State the contract: hash(seed, flag, unit) modulo 100, sticky, flag key included. Name MurmurHash3 as the production choice and stop.
Explain every vendor's rule engine.
Give the short-circuit order and send segments, reason codes, and PII to the targeting page.
Design scenario
Same prompt for every reader.
Requirements
One bucketing unit. One seed shared by both services. Ranges that sum to 100. A default when the flag is missing. No reshuffle unless you version the experiment.
Traffic / scale
Five thousand accounts in the demo, and the real population is every logged-in account plus a stable anonymous id.
Latency
Evaluation is local over a snapshot. The hash is cheap next to a remote rules fetch.
Consistency
Web and API share flag key, seed, and unit. Anonymous users keep a stable device id until login, and you document the identity merge.
Availability
Offline evaluation uses the last snapshot or the coded default. The risk class of the flag picks fail-open or fail-closed.
Failure assumptions
- Someone will change a range without versioning the salt.
- Anonymous and logged-in ids can disagree if you do not stitch them.
- A payload can arrive malformed.
Constraints
- Do not reteach segment precedence. That is the next page.
- Do not turn the hash into a cryptography lecture.
Prompt
Checkout needs control, express, and one_page at 50, 25, and 25. The same account must stay on one variant across the web app and the API.
API
What does assign return for one unit, and what is the default?
Data
Which three strings go into the hash?
Architecture
Which process is allowed to use a different seed, and what do you do at login?
You need three checkout treatments
Prefer
Multivariate flag
One flag returns control, express, or one_page. Ranges sit on a sticky bucket so a user does not hop.
- The split is visible in one place.
- Changing ranges without a new salt moves people. Version the experiment if you must.
- The binary still contains every branch.
Alternative
A pile of booleans
new_nav and express and not one_page. The combinations are the product of the flags.
- Each flag looks simple.
- Precedence is implicit and easy to get wrong.
- Experiments cannot name a treatment.
What the evaluator does
Vendors differ in the details. The order is the idea worth defending.
- 1
Missing or archived
Return the coded default. Do not invent a variant the SDK cannot see. - 2
Kill or force
An ops override wins before any percentage. This is the kill switch's slot. - 3
Individual target
An allowlist or denylist for a user key. Fine for dogfood, not for a million rows. - 4
Segment, then percentage
Reusable audiences, then a sticky split on whoever is left. Segments are the next page. - 5
Fall through
The default off or on. Prerequisites that require another flag to be on couple the two lifecycles. Use them rarely.
Overview
Types decide whether the rollout is something you can trust. A boolean gates on or off. A multivariate flag returns a name. A percentage flag needs a sticky hash so the same subject stays in the same bucket. A payload carries copy or a threshold, and it still needs a schema.
new_nav as a boolean is not checkout_flow with the variants control, express, and one_page. Experiments need variants. Kill switches need a force-off. Progressive delivery needs a percentage that can ramp without reshuffling people who were already treated.
A percentage flag is not a new container image. Pod replacement is Rolling, Blue-Green & Canary. The digest you shipped is CI/CD Pipelines.
Decisions
- 1
1. Evaluation context
- next2. Evaluator
- 2
2. Evaluator
- next3. Which flag type
- ?
3. Which flag type
- on or variant4. Branch in code
- percent or JSON5. Bucket or payload
- 4
4. Branch in code
- next6. Return the decision
- 5
5. Bucket or payload
- next6. Return the decision
- 6
6. Return the decision
Lesson map
Flag Types & Evaluation — Boolean, Multivariate, Percentage & Sticky Buckets
Flag types and evaluation mechanics decide whether your rollout is trustworthy. Booleans gate on or off. Multivariate flags return named variants. Percentage rollouts need a sticky hash so the same subject stays in the same bucket. This lesson implements those sketches, contrasts hash choices, and shows why a fresh random draw breaks experiments and UX.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB ctx["1. Evaluation context"] eval["2. Evaluator"] kind["3. Which flag type"] app["4. Branch in code"] ctx -->|1. Evaluation context| eval eval -->|2. Evaluator to 3. Which flag type| kind kind -->|on or variant| app
Boolean flags
The simplest release or ops toggle. Evaluation returns true or false after the rules run. Always define the default used when the SDK has no rules or the flag is archived.
The pitfall is nesting. Many booleans combined with and, or, and not create a combinatorial explosion. Prefer one multivariate flag, or layered flags with an explicit precedence. Precedence is the targeting page.
Multivariate flags
Return a variant key and an optional payload. Example names: control, treatment_a, treatment_b. Percentage splits allocate traffic by mapping a sticky bucket onto ranges.
| Bucket range | Variant |
|---|---|
| 0-49 | control |
| 50-74 | treatment_a |
| 75-99 | treatment_b |
Changing those ranges mid-flight moves users unless you version the experiment key or the salt. The sandbox below uses control at 50, express at 25, and one_page at 25. Same idea.
Percentage and the sticky unit
A fresh random draw each request puts the same user in and out of the treatment. The UX flickers. The statistics are not a comparison of two stable groups.
Sticky means: hash the seed, the flag, and the unit, take modulo 100, and compare with the percent. The unit is usually a user id or an account id. Anonymous users get a stable cookie or device id. When they log in, document the identity merge. Holdout groups and identity stitching are the advanced version of that sentence, not a second system you invent in the interview.
Hash function notes
- MurmurHash3 (32 or 128 bit) is the common SDK choice. Fast, good distribution, not cryptographic. Bucketing does not need a cryptographic hash.
- SHA-256 is the portable demo and the test oracle. Slower. Fine at evaluation volume if you are not hashing on the hottest line without a cache, and these sandboxes are not that line.
- Never use a language runtime
hashCodeacross services. It is not stable across platforms, and two processes will disagree. - Include the flag key and the seed so flag A and flag B do not correlate. Independence is the reason, not tradition.
Sandbox: multivariate sticky split
The Python snippet uses hashlib. The TypeScript snippet inlines the same SHA-256 because the browser cannot import Node's crypto module. Seed exp1 is part of the material. Two calls for user-99 return the same name. Five thousand users land near 50, 25, and 25.
ProblemVariants control, express, and one_page at 50, 25, and 25. One user is stable. A population of 5,000 lands near those weights.
Expectedassign is idempotent for user-99. The counts stay within five points of 50, 25, and 25 percent.
Edge cases
- A split that does not sum to 100 is rejected.
- The same unit and seed never move.
- A new seed is a different experiment.
- Test: same user is stable
assign('checkout', u, split) == assign('checkout', u, split) - Test: control is near half
abs(c['control'] / 5000 - 0.5) < 0.05 - Test: express is near a quarter
abs(c['express'] / 5000 - 0.25) < 0.05 - Test: one_page is near a quarter
abs(c['one_page'] / 5000 - 0.25) < 0.05
Press Run. Snippets must be self-contained — no network, files, or native modules.
ProblemMatch the Python ranges: control 50, express 25, one_page 25, sticky for one unit.
Expecteduser-99 is stable. Counts stay within five points of the target mix.
Edge cases
- The last variant catches a bucket that never matched.
- A second flag key is an independent assignment.
- Test: same user is stable
assign('checkout', u, split) === assign('checkout', u, split) - Test: control is near half
Math.abs(counts.control / 5000 - 0.5) < 0.05 - Test: express is near a quarter
Math.abs(counts.express / 5000 - 0.25) < 0.05 - Test: one_page is near a quarter
Math.abs(counts.one_page / 5000 - 0.25) < 0.05
Press Run. Snippets must be self-contained — no network, files, or native modules.
Defaults, prerequisites, and payloads
| Approach | Strength | Cost |
|---|---|---|
| Boolean plus a code branch | Simple | A new behavior still needs a deploy |
| Multivariate plus code | Treatments have names | Every path still ships |
| JSON payload | Tune copy and limits remotely | Easy to invent a half-language. Validate the schema and fail to the default |
Payload flags are powerful for copy and thresholds. They are a poor place to hide business logic you cannot test. Reject an invalid payload. Do not half-apply it.
Prerequisites mean flag B evaluates only when flag A is on. That couples their delete dates. Prefer a single multivariate flag when the treatments are really one decision.
Two services disagree when they use different seeds, a buggy context, or stale snapshots. Standardize the SDK config and the bucketing key. The architecture page is where split-brain gets its own lesson.
Interview Q&A
Boolean versus multivariate?
Answer
A boolean is on or off. A multivariate flag returns one of several named treatments. A/B/n needs the second one. A kill switch can be a boolean force-off, or a force onto the safe variant name.
Why include the flag key in the hash?
Answer
So assignment to flag A is independent of flag B. If you hash only the user, the same people land in the treatment of every experiment, and the experiments are no longer separate.
What breaks if you change the salt mid-experiment?
Answer
Users reshuffle across variants. Metrics mix experiences. People see the UI flip. Version the experiment key when you truly want a new population.
Is random() acceptable for a canary?
Answer
For an ephemeral load test, maybe. For a user-facing feature, no. Use a sticky assignment on a stable unit.
What is the default when the SDK is offline?
Answer
The last-known snapshot, or the coded default. Pick fail-open or fail-closed from the flag's risk, and write it down. The hub states that choice. This page assumes the default exists.
Can two services disagree on a flag?
Answer
Yes, if the seeds differ, the context is built differently, or one snapshot is stale. Share the SDK configuration and the bucketing key. A UI that says on and an API that says off is a support ticket with a cause.
When is a payload flag the wrong tool?
Answer
When the payload is really a program: branching rules, unbounded maps, or a value you cannot validate. Keep a schema. Fall back to the default on reject. Permanent pricing tables belong in a settings store.
What is a prerequisite flag?
Answer
Flag B runs only if flag A is on. It looks like reuse. It ties two lifecycles together, so deleting A silently changes B. Use it sparingly.
Pitfalls
- Drawing a new random number per request and calling the chart "10 percent."
- Leaving
hashCodein one language and SHA-256 in another. - Editing bucket ranges in place during a live experiment.
- Modeling three treatments as three booleans.
- Accepting a payload that failed schema validation because the UI would otherwise be blank.
- Forgetting the anonymous-to-login identity story and then debugging "the variant changed when I signed in."
Pick user-99 and the 50/25/25 split. Say which strings enter the hash, what a second call must return, and what you would version if product asked to move express from 25 to 40 tomorrow.
Go deeper
- Martin Fowler on feature toggles separates the types this page implements.
- LaunchDarkly variations and Unleash activation strategies are the product versions of boolean, multivariate, and percentage.
- The OpenFeature specification is the evaluation API those call sites should sit behind.
Next: Targeting & Context.