Feature Flags — Targeting, Experimentation & Kill Switches
Feature flags (feature toggles) decouple deploy from release: you ship dark code safely, then turn behavior on for segments, percentages, or experiments, and turn it off instantly when metrics or incidents demand it. This hub teaches release, experiment, ops, and permission toggles, sticky bucketing, targeting, A/B guardrails, progressive delivery, kill switches, and hygiene so flags do not become permanent debt.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What problem does a feature flag solve that a deploy does not?
Answer
It changes who experiences a behavior without shipping a new binary. Rollback of behavior is a rule change, not a pipeline.
L2
When is static config the right tool?
Answer
Non-behavioral knobs such as timeouts, pool sizes, and URLs. Permanent product settings belong in config or a settings store, not an eternal flag.
L3
What are the four toggle types?
Answer
Release, experiment, ops, and permission. Release is short-lived. Experiment is sticky and measured. Ops is a kill switch. Permission overlaps authorization and still needs a server-side policy.
L4
Why must percentage assignment be sticky?
Answer
A fresh random draw each request flips the same user across variants. That breaks the UX and invalidates the experiment.
L5
Fail-open or fail-closed when the flag service is down?
Answer
Security and entitlement flags fail closed. Availability-critical paths may serve the last-known snapshot. Document the choice per flag.
L6
What is sample ratio mismatch?
Answer
Observed assignment percentages diverge from the intended split. Treat it as a stop-ship for the experiment until bucketing or exposure is fixed.
L7
What has to be true before you call the flag done?
Answer
The winning behavior is the code default, dead branches are gone, the flag is archived, and an owner would have noticed if the cleanup SLA slipped.
Failure modes
Big-bang release
Every deploy is a launch. The only rollback is another deploy.
Config sprawl
Environment variables act as flags with no targeting, audit, or sticky bucket.
Experiment theater
An A/B has no exposure event, no guardrail, and a user who changes variants between requests.
Permanent release toggle
The flag stays after general availability and the dead branch never leaves the binary.
Misconceptions
A client-side flag can enforce security.
Client flags are UX. Entitlements and dangerous operations are server-side, and authorization policy is a different system.
A Kubernetes canary and a flag canary are the same control.
A workload canary shifts traffic across versions. A flag shifts behavior inside a version.
Random assignment is fine if the percentage is right on average.
The average can be right while one user flickers. Experiments and user-facing rollouts need a sticky key.
Interviewer traps
Walk the entire Kubernetes rollout, probe, and PDB model.
Say the flag gates behavior inside the pods. Point at the Kubernetes workloads hub and the canary page for the binary.
Use the flag as the only check on an admin API.
Flags manage exposure of a capability the policy already allows. Authorization stays server-side.
Design scenario
Same prompt for every reader.
Requirements
Ship dark. Ramp with a stable user key. Log exposure if the split is an experiment. Name the kill switch, who can flip it, and the safe variant. Say what you will not use a flag for.
Traffic / scale
One checkout path. Internal staff first, then a sticky percentage, not a fresh coin flip per request.
Latency
Evaluation uses a local snapshot. A remote call on every request is the wrong default.
Consistency
The same flag key, bucketing key, and salt return the same variant across services until the rules change.
Availability
If the flag service is down, the last snapshot still evaluates. Entitlement flags fail closed.
Failure assumptions
- The new flow can be wrong without crashing the process.
- A guardrail can breach after the binary is healthy.
- Someone will be tempted to leave the flag in place after launch.
Constraints
- Do not reteach Deployment rollouts, probe math, or error-budget burn formulas.
- Do not treat a schema cutover flag as a member of this cluster.
Prompt
Checkout is about to ship a new flow. You can deploy the binary, ramp a sticky percentage, or run a measured experiment. An incident must be able to force the old flow in seconds.
API
What does the caller pass as context, and what variant comes back?
Data
Which key is hashed, and what is logged the first time the surface renders?
Architecture
Where does evaluation run, and which sibling page owns targeting, stats, the ramp, or cleanup?
You need to change who sees the new checkout
Prefer
Feature flag
The binary is already in production. Rules pick a variant in seconds, with an audit trail and a sticky key.
- Target a segment, a percentage, or an experiment.
- A kill switch forces the safe variant without a new build.
- Right for progressive release, ops toggles, and short-lived experiments.
Alternative
Redeploy or restart on config
Environment variables and a previous artifact change everyone on that version, on pipeline time.
- Right for timeouts, URLs, and a correctness bug the flag cannot hide.
- Awkward as a per-user kill switch.
- Wrong when you need instant off without a rollout.
Evaluate, assign, observe, kill
Same flag key, bucketing key, and salt must return the same variant until you change the rules on purpose.
- 1
Build context
User, tenant, device, country, build. The bucketing key is a stable user or account id, not an IP address. - 2
Evaluate the snapshot
The SDK applies rules locally and refreshes in the background. A remote call on every request is the slow path. - 3
Assign a sticky variant
Hash the seed, the flag key, and the unit. The bucket decides the percentage or the named treatment. - 4
Log exposure once
Experiment metrics need the people who actually hit the surface. Assignment alone dilutes the result. - 5
Kill or graduate
A guardrail breach or an incident forces the safe variant. A finished release flag becomes the default and is deleted.
Overview
Code is already in production. Behavior is gated by rules evaluated against context. That split is the whole point of a flag. Deploy and release stop being the same event.
Without that control plane you get five familiar failures:
- Big-bang releases. Every deploy is a user-visible launch. Rollback means another deploy.
- Long-lived branches. Feature branches rot while they wait for the big merge.
- Config sprawl. YAML and environment variables pretend to be flags, with no targeting, no audit, and no sticky bucket.
- Experiment theater. A random A/B has no exposure event, no guardrail, and no consistent assignment.
- Incident paralysis. You cannot disable a bad path without a hotfix pipeline.
Flags fix the control plane. They do not fix a segfault, a bad schema migration, or a missing authorization check.
Flags, config, deploys, and experiment platforms
| Dimension | Feature flag service | Static config or env | Code deploy or rollback | Dedicated experiment platform |
|---|---|---|---|---|
| Change latency | Seconds, remote | Redeploy or restart | Minutes to hours | Seconds to minutes |
| Targeting | Rules, segments, percentage, sticky | Usually global or coarse | All-or-nothing per version | Strong stats and metrics |
| Kill switch | First-class | Awkward | Redeploy the previous artifact | Possible, not the ops primary |
| Who flipped it | Usually audited | Git blame or a ticket | Deploy history | Experiment ACL |
| Best fit | Progressive release, ops toggles, permissions | Timeouts, URLs, non-behavioral knobs | Binary correctness fixes | Long-running A/B with guardrails |
| Avoid when | Permanent product settings | Per-user personalization | You need instant off without a redeploy | A simple boolean release (overkill) |
Rule of thumb: use flags to control who sees what behavior when. Use config for non-behavioral knobs (pool size, timeout). Use deploys to ship code and to fix correctness bugs a flag cannot paper over. Use an experiment platform, or a flag service plus a stats stack, when you need exposure, metrics, and peeking discipline. That stack is the experimentation page.
Progressive pod rollouts still live in Kubernetes. Flags gate product behavior inside those pods. The workload map is Kubernetes Workloads. How a canary shifts traffic across versions is Rolling, Blue-Green & Canary.
Toggle taxonomy
Pete Hodgson's categories, via Martin Fowler, are the interview vocabulary:
- Release toggles hide unfinished work. They are short-lived. Remove them after general availability.
- Experiment toggles return A/B/n variants. They are sticky, and they are tied to metrics and exposure.
- Ops toggles are kill switches, degrade modes, and circuit-adjacent behavior. They may live longer, and they need a named owner.
- Permission toggles encode entitlement, plan, or beta access. They overlap authorization. The policy engine is Authorization. Do not rebuild RBAC on this page.
Client, server, and the snapshot
| Mode | Strength | Cost | Typical use |
|---|---|---|---|
| Server-side | Secrets stay on the server. APIs stay consistent. | Latency if you call out on every request. | Pricing, entitlements, dangerous paths |
| Client-side | Instant UI. Offline-friendly once bootstrapped. | Rules leak. A user can tamper with the UX. | Copy and layout experiments |
| Edge or BFF | Low latency with a place to keep control. | An extra hop you have to design. | CDN personalization, server-rendered UI |
SDKs usually download a rules snapshot, evaluate locally, and refresh in the background. That is stale-while-revalidate. When the flag service is down:
- Fail closed for security-sensitive flags. Default is deny the new capability.
- Fail open carefully for availability. Serve the last-known-good snapshot, and write that choice down per flag.
Timeouts and a breaker around the flag service itself belong with Resilience Patterns. Do not invent a second circuit-breaker lecture here.
Consistency contract
For a given flag key, bucketing key, and salt, evaluation returns the same variant across devices and services for the life of the experiment, unless the rules change on purpose. Sticky bucketing is not optional if you want to trust an A/B.
Decisions
- 1
1. Context arrives
- next2. SDK evaluates
- 2
2. SDK evaluates
- next3. Rules and sticky hash
- ?
3. Rules and sticky hash
- variant4. Serve the behavior
- 4
4. Serve the behavior
- next5. Log one exposure
- 5
5. Log one exposure
- next6. Read guardrails
- 6
6. Read guardrails
- breach7. Kill or ramp down
- healthy8. Graduate the flag
- 7
7. Kill or ramp down
- 8
8. Graduate the flag
Lesson map
Feature Flags — Targeting, Experimentation & Kill Switches
Feature flags (feature toggles) decouple deploy from release: you ship dark code safely, then turn behavior on for segments, percentages, or experiments, and turn it off instantly when metrics or incidents demand it. This hub teaches release, experiment, ops, and permission toggles, sticky bucketing, targeting, A/B guardrails, progressive delivery, kill switches, and hygiene so flags do not become permanent debt.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Context arrives"] b["2. SDK evaluates"] c["3. Rules and sticky hash"] d["4. Serve the behavior"] a -->|1. Context arrives| b b -->|2. SDK evaluates to 3. Rules and sticky hash| c c -->|variant| d
Change the seed only when you intentionally reshuffle the population. A casual salt change breaks experiment continuity and flips users who were already treated.
The next page implements boolean, multivariate, and percentage assignment. The hash below is the contract those types share.
Sticky percentage, in the sandbox
Production SDKs often use MurmurHash3. It is fast, it distributes well, and it is not a cryptographic hash, which is fine for bucketing. These demos use SHA-256 so Python and the browser agree. Take the first eight hex characters as an integer, then modulo 100. The bucket is in 0–99. The user is in the percentage when the bucket is below the percent.
Include the flag key and the seed in the material so two flags do not share an assignment. Never use a language runtime's hashCode across services. It is not stable across platforms.
The browser sandbox has no Node crypto module. The TypeScript block inlines SHA-256. The material string is the same one Python hashes with hashlib.
ProblemSame user and flag always land in the same bucket. About 10 percent of 10,000 users are enabled at a 10 percent rollout.
Expecteduser-42 is stable. The empirical rate sits between 0.08 and 0.12.
Edge cases
- Percent 0 enables nobody.
- Percent 100 enables every bucket.
- A different flag key is an independent draw.
- Test: same user is stable
enabled('new_checkout', 'user-42', 25) == enabled('new_checkout', 'user-42', 25) - Test: about 10 percent
0.08 < rate < 0.12 - Test: zero percent is off
enabled('new_checkout', 'user-42', 0) == False - Test: full percent is on
enabled('new_checkout', 'user-42', 100) == True
Press Run. Snippets must be self-contained — no network, files, or native modules.
ProblemMatch the Python contract: sticky bucket 0-99, enabled when the bucket is below the percent.
Expecteduser-42 is stable. The empirical rate sits between 0.08 and 0.12.
Edge cases
- A changed seed is a reshuffle.
- Percent outside 0-100 throws.
- Two flag keys do not share buckets.
- Test: same user is stable
enabled('new_checkout', 'user-42', 25) === enabled('new_checkout', 'user-42', 25) - Test: about 10 percent
rate > 0.08 && rate < 0.12 - Test: zero percent is off
enabled('new_checkout', 'user-42', 0) === false - Test: full percent is on
enabled('new_checkout', 'user-42', 100) === true
Press Run. Snippets must be self-contained — no network, files, or native modules.
Progressive delivery and kill switches, in one pass
Ramp 1 percent, then 5, 25, 50, and 100, with a gate on error rate, latency, and a business guardrail such as checkout conversion. A kill switch is an ops toggle that forces the safe variant for the blast radius you named. It is faster than rolling back a Deployment when the bug is behavioral.
Coordinate the gate with SLOs, Error Budgets & Distributed Tracing. The binary you ramped still has an artifact identity on CI/CD Pipelines. Schedules, abort criteria, and dual-control live on Progressive Delivery & Kill Switches.
Architecture you should not negotiate away
- One bucketing key, usually a stable user or account id. Never an IP alone for an A/B.
- An OpenFeature-style abstraction so call sites do not hard-bind to one vendor SDK.
- Exposure logging on the first evaluation that reaches the experiment surface. Without it, metrics lie.
- Change control. Name who can flip a production kill switch. Dual-control for a high blast radius.
- A cleanup SLA. A stale flag after general availability is debt. Track age and owners.
- Do not put secrets in a client bundle. Do not use a flag as the sole authorization layer.
SDK placement, split-brain evaluation, and the stale-flag lifecycle are Flag Architecture & Hygiene.
What this cluster covers
- Flag Types & Evaluation — boolean, multivariate, percentage, payload, and the sticky hash.
- Targeting & Context — attributes, segments, rule order, and reason codes.
- Experimentation & A/B — exposure, guardrails, peeking, and sample ratio mismatch.
- Progressive Delivery & Kill Switches — ramps, abort gates, and instant rollback of behavior.
- Flag Architecture & Hygiene — where the SDK sits, consistency, and deleting the flag.
What this cluster leaves alone
- Pod rollouts. Surge, probes, and a traffic shift between versions are Kubernetes Workloads and Rolling, Blue-Green & Canary. A flag cannot repair a crashing binary.
- Artifact identity. What you deployed, and how you know the digest, is CI/CD Pipelines.
- Overload tools. Breakers, bulkheads, and shedding are Resilience Patterns. A kill switch is a cousin of those tools, not a replacement for them.
- Error budgets. Ramp gates should watch burn. The math is SLOs, Error Budgets & Distributed Tracing.
- Permission. Beta access still needs a server-side decision. That system is Authorization.
- Schema cutover. A read flag that flips a migration path is Rollback, Feature Flags & Cutover. It stays in the zero-downtime database migrations series. It is not a member of this cluster.
Interview Q&A
Feature flag versus configuration?
Answer
Config sets non-behavioral parameters: timeouts, endpoints, pool sizes. Flags gate behavior variants with targeting, audit, and often a sticky assignment. Permanent product settings belong in settings or config, not in an eternal flag.
Why sticky bucketing?
Answer
Without it, one user can flip variants across requests. That poisons the UX, invalidates A/B statistics, and creates support tickets that say the UI changed under them.
Fail-open or fail-closed when the flag service is down?
Answer
Security and entitlement flags fail closed: deny the new capability. Availability-critical paths often serve the last-known-good snapshot. Document the choice on the flag. Pair the client with a timeout and a breaker.
Can a client-side flag enforce security?
Answer
No. Client flags are UX. A user can change them. Enforce entitlements and dangerous operations on the server, in authorization policy.
Flag canary versus Kubernetes canary?
Answer
A Kubernetes canary shifts traffic across versions. A flag shifts behavior inside a version. Use both: canary the binary, flag the feature. Probes, surge, and Pod disruption budgets stay on the Kubernetes pages.
What is sample ratio mismatch?
Answer
Observed assignment percentages diverge from the intended split, for example 48/52 when you asked for 50/50, at a large N. Usual causes are a bucketing bug, a filter that drops one variant, or exposure that is not comparable. Treat it as a stop-ship. Do not interpret lift.
When should you not use a flag?
Answer
Permanent settings, a one-off data migration that belongs in an offline job, a security boundary that needs a policy engine, or a team that will never clean the flag up. A schema cutover is the migrations lesson, not a percentage of rows.
Why care about OpenFeature?
Answer
It is a vendor-neutral evaluation API. You can swap LaunchDarkly, Unleash, Flagsmith, Split, or a homegrown provider without rewriting every call site. Tests can use an in-memory provider.
Pitfalls
- Treating an environment variable as a flag and then wondering why there is no audit or sticky bucket.
- Hashing an IP address for an experiment and calling the split random-but-fine.
- Rotating the salt mid-experiment because a dashboard looked noisy.
- Shipping a client flag as the only gate on a dangerous operation.
- Leaving a release toggle wrapped around dead code for two quarters.
- Analyzing an experiment on assignment when only some users reached the UI.
Out loud: what a flag decides, what config is for, why the hash includes the flag key, and which page you open for targeting, for peeking, for the kill switch, and for deleting the flag. Then name the Kubernetes page you will not reteach.
Go deeper
- Martin Fowler on feature toggles is the taxonomy and the debt warning. Pete Hodgson's categories live in that article.
- The OpenFeature specification is the vendor-neutral evaluation API.
- LaunchDarkly targeting and Unleash feature flags show rules, percentage rollouts, and ops toggles in real products.
- Overlapping Experiment Infrastructure is the paper for running many experiments without letting them collide. The experimentation page is the interview version.
Next: Flag Types & Evaluation.