Flag Architecture & Hygiene — SDK Placement, Consistency, Stale Flags & Debt
Flags fail quietly when the architecture is wrong, and loudly when hygiene is ignored. This lesson places SDKs so services do not split-brain a decision, adopts an OpenFeature-style seam, and sets a lifecycle from create to ramp to delete so stale flags do not become a permanent settings database.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does application code call?
Answer
A small API such as getBooleanValue(flag, default, context). The provider adapter talks to the vendor.
L2
Why put OpenFeature, or something like it, in the middle?
Answer
You can swap providers, log exposures in one hook, and unit-test with an in-memory provider.
L3
What causes split-brain flags?
Answer
Different seeds, different rules versions, or a BFF snapshot that is not the API's snapshot.
L4
Where should a sensitive decision evaluate?
Answer
On the server. Treat client rules as public. Never ship a secret in the bundle.
L5
How do you retire a flag?
Answer
Default the winning variant in code, delete the dead branch, archive the flag, and ship those together. Then watch the metric that proved the winner.
L6
What is flag debt?
Answer
Toggles that outlived their purpose and still complicate code, tests, and reasoning. Age, a missing owner, and an always-on wrapper are the symptoms.
L7
What must not be a flag?
Answer
Permanent product configuration, secrets, the only authorization check, and a one-time backfill. Those have stores, policy engines, and jobs.
Failure modes
UI on, API off
Two tiers hashed different seeds or read different snapshots. The user sees a button the API rejects.
Always-on for six months
The release flag still wraps both branches. Deleting it feels riskier than leaving it.
Two sources of truth
Git and the console both claim to be canonical. An incident flip and a later deploy disagree.
Exposure logged in three places
The metric double-counts. The wrapper hook was the one place that should have emitted.
Misconceptions
An in-memory test provider is a toy, so production must call the vendor SDK directly.
The call site stays on the abstraction. The provider changes. Tests and production use the same method names.
A flag left at 100 percent is free.
It still forks the test matrix and hides the dead branch. Graduate it.
Client evaluation is fine for entitlements if the UI hides the button.
Hiding a button is UX. The API still authorizes. Client rules are public.
Interviewer traps
Pick a vendor and design their console.
Design the seam: one API, one provider, one seed, one exposure hook. The vendor is replaceable.
Store the price list in a flag forever.
That is a settings service. Say so, and point permanent config out of the flag system.
Design scenario
Same prompt for every reader.
Requirements
One evaluation API. Shared flag key, seed, unit, and rules version. An in-memory provider for tests. An owner and an expiry on every new flag. A delete plan for flags past general availability.
Traffic / scale
Evaluation is local on a snapshot. The consistency check is a checksum of the seed, not a remote call per request.
Latency
A disagreement between BFF and API is a correctness bug, not a latency optimization.
Consistency
Same inputs, same variant. A seed mismatch fails a contract test before it fails a user.
Availability
Code defaults still evaluate when the snapshot is missing. The emergency overlay is audited and is not a second permanent store.
Failure assumptions
- Someone will copy a vendor SDK call into a new service.
- A flag at 100 percent will look done and stay in the code.
- Git and the console will both be edited during an incident.
Constraints
- Do not rebuild the ramp gates or the experiment statistics.
- Do not move the migration cutover page into this series.
Prompt
Web, a BFF, and a checkout API all decide new_checkout. You are about to add a second vendor, and the flag list is full of year-old release toggles.
API
What does getBoolean return for a missing flag?
Data
Which four values must match across BFF and API?
Architecture
Where is the one exposure hook, and what is the source of truth after the incident overlay expires?
The UI and the API must agree on new_checkout
Prefer
One API, one seed, one snapshot generation
Call sites share an evaluation interface. BFF and services use the same flag key, seed, unit, and rules version.
- A contract test fails when the seeds diverge.
- The vendor sits behind a provider.
- Exposure is logged in one hook.
Alternative
Each tier imports a vendor SDK
Web, BFF, and API each evaluate with whatever config that deploy happened to ship.
- Fast to start.
- UI on and API off becomes a support ticket.
- Swapping vendors means editing every call site.
Lifecycle of one flag
A flag without an owner and an expiry is debt on the day you create it.
- 1
Create
Name the owner, the expiry, and the default. Release flags get a short SLA. Ops flags get a runbook. - 2
Ramp
The previous page owns the percentage and the kill switch. This page owns that both tiers see the same decision. - 3
Graduate
The winning variant becomes the code default. Dead branches go away in the same change. - 4
Archive and delete
Remove the vendor flag and the code references together. A dashboard of age and last flip shows what you missed.
Overview
Flags fail quietly when two services disagree, and they fail loudly when the codebase is a museum of toggles. This page is placement, the seam in front of the vendor, and the delete step. The hub is the map back to types, targeting, experiments, and ramps.
Flow
- 1
1. Web and mobile SDKs
- next2. Rules snapshot
- 2
2. Rules snapshot
- 3
3. BFF and services
- next4. OpenFeature API
- 4
4. OpenFeature API
- next5. One provider
- 5
5. One provider
- next2. Rules snapshot
Lesson map
Flag Architecture & Hygiene — SDK Placement, Consistency, Stale Flags & Debt
Flags fail quietly when the architecture is wrong, and loudly when hygiene is ignored. This lesson places SDKs so services do not split-brain a decision, adopts an OpenFeature-style seam, and sets a lifecycle from create to ramp to delete so stale flags do not become a permanent settings database.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB web["1. Web and mobile SDKs"] rc["2. Rules snapshot"] bff["3. BFF and services"] of["4. OpenFeature API"] web -->|1. Web and mobile SDKs| rc bff -->|3. BFF and services| of
Web and mobile read a rules snapshot. The BFF and the services call one abstraction. That abstraction talks to LaunchDarkly, Unleash, Flagsmith, Split, or an in-house provider. The provider is fed by the same remote config the clients snapshot.
The consistency rule: every path that influences the same user-visible decision shares the flag key, the seed, the bucketing unit, and the rules version. A BFF that evaluates differently from the API creates the ticket "the button was there and the API said no."
The seam
Application code asks for a boolean with a default and a context. It does not import the vendor. Provider adapters talk to the service. You get three things: a vendor you can swap, one place to log exposures, and an in-memory provider for tests.
Exposure belongs in that wrapper, once. Logging it again in the handler double-counts. The event shape is the experimentation page.
ProblemgetBoolean returns the stored value, and a missing flag returns the default the caller passed.
Expectednew_checkout is true. A missing key is the default false.
Edge cases
- A stored false must not fall through to the default.
- The targeting key is accepted and unused by this in-memory demo.
- Test: stored flag wins
client.getBoolean('new_checkout', false, { targetingKey: 'u1' }) === true - Test: missing flag uses the default
client.getBoolean('missing', false, { targetingKey: 'u1' }) === false - Test: stored false stays false
off.getBoolean('new_checkout', true, { targetingKey: 'u1' }) === false
Press Run. Snippets must be self-contained — no network, files, or native modules.
Seeds have to match
A checksum across services is enough to catch the boring bug: one deploy still on seed v1, another on v2. Same seed, same variant. Different seeds can disagree, but a 50 percent gate can also land both buckets on the same side by chance. The demo uses the unit user, where those two seeds fall on opposite sides of 50. A contract test has to pick a case that can diverge. This is not the full sticky hash lesson. That is the types page. This is the check that both binaries use that hash.
ProblemTwo callers with seed v1 agree. v1 and v2 do not.
Expectedagree is true for a repeated seed and false for a mismatched pair.
Edge cases
- Percent 0 agrees on off for every seed.
- Percent 100 agrees on on.
- A different flag key is a different decision even with one seed.
- Test: matching seeds agree
agree(['v1', 'v1'], 'f', 'user', 50) is True - Test: mismatched seeds disagree
agree(['v1', 'v2'], 'f', 'user', 50) is False - Test: zero percent agrees
agree(['v1', 'v2'], 'f', 'user', 0) is True
Press Run. Snippets must be self-contained — no network, files, or native modules.
The zero-percent case agrees because both seeds are off. That is a reminder to test a percent that can actually diverge, not only the trivial gates.
One source of truth
| Pattern | Strength | Cost |
|---|---|---|
| Remote console only | Fast | Drift from the repo, weaker review |
| Definitions in Git | Review and env promotion | Slow kills unless there is a break-glass path |
| Dual-write, Git and remote, both canonical | Looks redundant | Ambiguous source of truth. Avoid it unless sync is automatic and one side is derived |
| Code defaults plus a remote overlay | Safe when the network is gone | Write down which one wins, and expire the overlay |
Prefer one source of truth and a promotion pipeline. The pipeline is CI/CD Pipelines. The emergency overlay is allowed, audited, and temporary. Who may use it is the kill switch page.
Stale flags
Symptoms:
- The flag has been always-on for six months and still wraps a dead branch.
- Nobody knows the owner, so deletion feels unsafe.
- Evaluation cost and the cognitive load both rise.
- A forgotten permission toggle is still in the client.
Hygiene that keeps the list short:
- Owner and expiry at creation, on the ticket, not in someone's memory.
- A cleanup SLA. A reasonable release-flag rule is to remove it within two sprints of general availability.
- A dashboard of age, last flip, and evaluation volume.
- Graduate: set the winning default in code, delete the flag, delete the dead branch.
- Archive in the vendor and delete the code references in the same series of changes.
The interview cost model is qualitative and enough: each stale flag adds a test matrix and a question in review. Hundreds of them slow delivery. You do not need a dollar figure to refuse the next permanent toggle.
What is not a flag
- Permanent product configuration, theme defaults, and a pricing table that is not an experiment. That is a settings service or a database.
- Secrets and credentials. They do not belong in a rules snapshot.
- The only authorization check. Policy is Authorization.
- A one-time backfill. Use a job and the expand/contract plan. The cutover flag for a migration is Rollback, Feature Flags & Cutover, still in the databases series.
- A feature whose branch you will never remove. That is conditional debt with a console.
When the flag service is down, the resilience choices from the hub still apply: last snapshot, or fail closed for a sensitive flag. The breaker around that dependency is Resilience Patterns. Guardrail numbers that tell you a ramp should have stopped are SLOs, Error Budgets & Distributed Tracing.
How you test it
- Unit-test both variants with the in-memory provider. Do not click the console as the only test.
- Contract-test that the BFF and the API share seed and rules version. The Python sketch is the shape of that test.
- In staging, shadow-evaluate with production-like segments before the ramp.
- Assert the exposure hook fires once. A second log line in the handler is a bug.
Client SDKs stay on a hostile-client footing. Rules are public. Sensitive decisions stay on the server even if the button is hidden.
Interview Q&A
Why OpenFeature?
Answer
It keeps the evaluation API stable while the provider changes. Tests use an in-memory provider. A migration off one vendor does not rewrite every if-statement. The same hook can log exposure once.
What causes split-brain flags?
Answer
Different seeds, stale snapshots, or two tiers applying different rules to the same user. The visible bug is UI on and API off. The fix is one flag key, one seed, one unit, and one rules version, checked in a contract test.
How do you retire a flag?
Answer
Make the winning variant the default in code, remove the other branch, archive or delete the flag, and confirm the metric that justified the winner. Do the code delete and the vendor delete in the same series of changes so you do not leave a dangling key.
What is flag debt?
Answer
Toggles that have outlived their purpose and still complicate the code, the tests, and the review. An always-on wrapper with no owner is the usual specimen. The cost is cognitive load and a larger test matrix, and it compounds.
What is the client SDK security posture?
Answer
Treat the downloaded rules as public. Do not ship secrets. Do not put a raw allowlist of emails in the payload. Enforce sensitive decisions on the server.
Where do exposures get logged?
Answer
In one wrapper or hook, the first time the surface runs for that unit. A second log in the client and a third in the API double-count the experiment. The experimentation page owns the event. This page owns putting it in one place.
GitOps versus the console for a kill?
Answer
Git for reviewed defaults and promotion. A console or API for the incident, with IAM and an audit trail. After the incident, the overlay expires and Git is the source of truth again. Two lasting sources of truth will diverge.
Which decisions should leave the flag system?
Answer
Permanent settings, secrets, the sole authorization boundary, and one-time data backfills. If you will never delete the branch, it is not a rollout. It is a conditional you forgot to finish.
Pitfalls
- Importing the vendor SDK in a new service because the wrapper was "just for the web app."
- Shipping a client bundle that contains a server-only rule set.
- Calling a flag done because it sits at 100 percent.
- Editing Git and the console during one incident and keeping both edits.
- Logging exposure in the provider and again in the handler.
- Using a flag as the backfill switch for a schema change.
Pick an always-on release flag from six months ago. Say the default you will hard-code, the branch you will delete, the contract test that checks the seed, and the dashboard row that should disappear. Then name the one emergency overlay you will not leave behind.
Go deeper
- The OpenFeature specification is the evaluation API and the provider model this page sketches.
- Martin Fowler on feature toggles is the debt and lifecycle warning.
- LaunchDarkly code references are one way to find flags that no longer appear in code.
- Unleash and Flagsmith document lifecycle and architecture from the open-source side.
Back to the hub.