Experimentation & A/B — Exposure, Metrics, Guardrails & Peeking
An experiment is a sticky multivariate flag plus exposure logging, a success metric, and guardrail metrics, run with statistical discipline. This lesson covers assignment versus exposure, sample ratio mismatch, peeking, a one-sentence view of CUPED, and when a dedicated experiment platform earns its place next to the flag.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is assignment?
Answer
The variant the evaluator returned for this unit. It is not proof the user saw the change.
L2
What is exposure?
Answer
An analytics event that the user hit the experiment surface. Log it once per experiment and unit, when the surface renders or the code path runs.
L3
Why analyze exposures instead of every assignment?
Answer
People who never reached the UI dilute the effect. If one variant is harder to reach, analyzing assignments also biases it.
L4
What is sample ratio mismatch?
Answer
The observed split is not the intended split at a large N. Pause, fix bucketing or instrumentation, and do not interpret lift.
L5
What is peeking?
Answer
Checking significance on a schedule you did not precommit and stopping when the chart looks good. That inflates false positives unless the method allows continuous looks.
L6
What can veto a winning primary metric?
Answer
A guardrail: error rate, latency, refunds, crashes. A win that burns the error budget is a loss.
L7
What is CUPED in one sentence?
Answer
Variance reduction that uses pre-experiment behavior as a covariate so you detect an effect sooner. It is not a substitute for a correct exposure.
Failure modes
Analysis on assignment
Most assigned users never saw the surface. The effect shrinks, or it biases when reach differs by variant.
Daily peek
Someone stops the test the first afternoon p looks small. The false-positive rate is no longer the one you think you set.
Ignored SRM
The split is 46/54 and the team still reads lift. The bug is in the pipe, not in the product.
Double exposure
The client and the API both emit the event. The unit is counted twice and the metric moves for the wrong reason.
Misconceptions
A sticky flag is an experiment.
A sticky flag is the assignment. The experiment also needs exposure, a primary metric, guardrails, and a stopping rule.
CUPED fixes a biased sample.
CUPED reduces variance. It does not repair a missing exposure or a sample ratio mismatch.
Bayesian dashboards mean you can stop whenever you like.
You still need a decision rule you wrote down, and guardrails that can veto.
Interviewer traps
Derive the CUPED estimator.
One sentence: pre-experiment covariates reduce variance. Offer the book or the platform, and come back to exposure.
Ship because the primary metric is up and latency is worse.
Guardrails veto. Point at the error budget instead of redefining success.
Design scenario
Same prompt for every reader.
Requirements
Log exposure once when checkout renders. Precommit the primary metric and the horizon or a sequential method. Name the guardrails that pause the ramp. Pause if the sample ratio is wrong.
Traffic / scale
A 50/50 assignment on a stable account id. Analysis uses exposures, not every assignment.
Latency
p99 on the checkout path is a guardrail, not the success metric, unless you precommitted it as primary.
Consistency
One salt, one exposure stream. Web and API must not emit two exposures for the same unit.
Availability
A guardrail breach pauses or kills the treatment. The kill switch itself is the progressive-delivery page.
Failure assumptions
- The first look at the dashboard will be early.
- Bots, caches, or a redirect can drop one variant.
- A second experiment may be running on the same surface.
Constraints
- Do not teach a full sequential-testing proof.
- Do not recompute error-budget burn. Name the SLO page.
Prompt
Checkout v2 is a 50/50 sticky split. You can see conversion, p99, and the 5xx rate. Someone wants to stop on day two because the chart looks good.
API
When do you emit exposure, and what do you refuse to emit twice?
Data
Which metric decides, and which metrics can veto?
Architecture
What do you pause first when the split is 46/54?
The checkout bet needs a number, not a vibe
Prefer
Exposure, a primary metric, guardrails
Sticky assignment plus one exposure per unit, a metric you precommitted, and a veto if latency or errors move.
- Analysis matches the people who saw the surface.
- Stopping follows a horizon or a sequential method.
- A kill switch still exists if the treatment is harmful.
Alternative
Refresh the dashboard and ship
The split is a flag. The decision is whichever day the chart first looks good.
- Fast to narrate.
- False positives pile up. That is peeking.
- A 46/54 bug looks like a product win.
One request through an experiment
Assignment can happen on the server. Exposure happens when the surface actually runs.
- 1
Evaluate
The app asks the SDK for experiment_x. The sticky variant comes back. That is assignment. - 2
Expose once
Emit exposure for this experiment and this user the first time checkout renders or the code path runs. - 3
Render
The user sees that treatment, not a second draw. - 4
Record metrics
Conversion, revenue, and the guardrails. Tie the guardrails to the error budget. - 5
Decide on the rule
Fixed horizon or a sequential method. An unplanned peek is not a rule. SRM pauses the whole thing.
Overview
An experiment is a sticky multivariate flag plus measurement. The types page makes the assignment stable. The targeting page decides the audience. This page is why the metric is allowed to change your mind.
You can run an experiment in a dedicated tool and still need the same three properties: consistent assignment, an exposure event, and a way to turn a bad treatment off. The off switch is the next page.
Flow
- 1
1. User request
- next2. Evaluate the flag
- 2
2. Evaluate the flag
- next3. Treatment comes back
- 3
3. Treatment comes back
- next4. Emit exposure once
- 4
4. Emit exposure once
- next5. Render that treatment
- 5
5. Render that treatment
- next6. Record metric events
- 6
6. Record metric events
Lesson map
Experimentation & A/B — Exposure, Metrics, Guardrails & Peeking
An experiment is a sticky multivariate flag plus exposure logging, a success metric, and guardrail metrics, run with statistical discipline. This lesson covers assignment versus exposure, sample ratio mismatch, peeking, a one-sentence view of CUPED, and when a dedicated experiment platform earns its place next to the flag.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB u["1. User request"] app["2. Evaluate the flag"] flag["3. Treatment comes back"] exp["4. Emit exposure once"] u -->|1. User request to 2. Evaluate the flag| app app -->|2. Evaluate the flag| flag flag -->|3. Treatment comes back| exp
Assignment versus exposure
Assignment is the variant the evaluator returned.
Exposure is the analytics event that the user actually experienced the surface: they saw the checkout step, or they opened the paywall. Log it once per experiment and unit, when that surface renders or that code path runs.
If you analyze everyone who was assigned, and only some of them reached the UI, you dilute the effect. If one variant is harder to reach, you also bias it. The analysis key is the exposure, not the assignment list.
The artifact that contained the code is still a CI concern. Identity of the build is CI/CD Pipelines. The experiment does not replace that.
Metrics
| Class | Examples | Role |
|---|---|---|
| Primary | Checkout conversion, activation rate | The decision |
| Secondary | Revenue per user, time to complete | Interpretation |
| Guardrails | p99 latency, 5xx rate, refund rate, crash-free sessions | Veto or auto-stop |
| Debugging | Click positions, funnel step counts | Not a ship decision by themselves |
A "winning" experiment that burns the error budget loses. Burn rate and the budget live on SLOs, Error Budgets & Distributed Tracing. Name them. Do not re-derive the multi-window formula here.
Peeking
The naive workflow is to refresh the dashboard every day and stop when the p-value first looks small. Repeated looks inflate false positives. That is peeking.
Mitigations, at interview depth:
- Fixed horizon. Precommit the sample size. Do not stop early because the chart is flattering.
- Sequential tests or always-valid inference. Methods that allow continuous monitoring with a corrected error rate.
- Bayesian platforms. Still need a decision rule you wrote down, and guardrails.
The sentence to say: we precommit the metrics, the horizon or the sequential method, and the guardrails. We do not ship on a single unplanned peek.
Sample ratio mismatch
You asked for 50/50. At a large N you observe something like 46/54. A rough chi-square check for one degree of freedom treats a statistic above about 3.84 as the "this is not 50/50" side of a 0.05 line. The sandbox is that sketch. It is not a license to skip a real stats review.
Causes that show up in production:
- A bucketing bug, or a different salt in one service.
- Filters that drop exposures differently per variant.
- Bots, or a cache that keeps serving one variant.
- A client redirect that loses a variant.
Action: pause the experiment, fix the instrumentation, and do not interpret lift. The 46/54 result is not a product insight.
ProblemA near split stays under 3.84. A 4000/6000 split does not.
Expected5000 and 5050 pass the rough check. 4000 and 6000 fail it.
Edge cases
- Equal counts score 0.
- The check assumes you intended 50/50.
- A small N can look wild without being the same bug.
- Test: healthy split passes
srm_chi2(5000, 5050) < 3.84 - Test: broken split fails
srm_chi2(4000, 6000) > 3.84 - Test: equal counts are quiet
srm_chi2(1000, 1000) == 0
Press Run. Snippets must be self-contained — no network, files, or native modules.
ProblemThe second call for the same experiment and user must not emit again.
ExpectedOne event is stored. A different user still emits.
Edge cases
- A second experiment for the same user still emits.
- A different variant string does not create a second event once the key is seen.
- Test: one exposure for one user
events.length === 1 - Test: a second user emits
eventsAfterSecond.length === 2 - Test: the first event is exposure
events[0].type === 'exposure'
Press Run. Snippets must be self-contained — no network, files, or native modules.
CUPED, in one sentence
CUPED (Controlled-experiment Using Pre-Experiment Data) adjusts the metric with a covariate from before the experiment, which cuts variance. You get a tighter interval or you notice an effect with less traffic. The interview sentence is: we use pre-experiment behavior as a covariate so detection is faster and cleaner. It is not a substitute for a correct exposure. Leave the formula to a stats review.
Overlapping experiments
Several experiments at once can interact. Google's overlapping experiment infrastructure uses layers and domains so a user is not in conflicting treatments of the same surface.
Practical version:
- Put UI chrome, pricing, and ranking in different layers when they would fight over one screen.
- Use the platform's mutual-exclusion groups.
- Watch shared guardrails. Two "harmless" experiments can burn one error budget together.
Flag-native A/B versus a dedicated platform
| Need | Flag A/B | Dedicated stack |
|---|---|---|
| Sticky split and exposure | Yes | Yes |
| CUPED, sequential tests | Basic to medium | Stronger |
| Metric warehouse | Often do-it-yourself | Often built in |
| Ops kill switch | Strong | Medium. Pair it with flags |
A common split of labor: flags for delivery and the kill switch, an experiment platform (or the flag vendor's experiment product) for the stats-heavy bet. Split, Optimizely, an internal stack, and the vendor's own experiment product are the usual names. The choice is about the statistics and the warehouse, not about whether assignment should be sticky. Sticky is not optional on either side.
Interview Q&A
Assignment versus exposure?
Answer
Assignment is the variant decision. Exposure is the event that the user hit the experiment surface. Analysis should key off exposures. Assignment-only analysis dilutes the effect and biases it when the variants are not equally reachable.
What is peeking?
Answer
Repeatedly checking significance and stopping when you get lucky. That inflates the false-positive rate unless you are using a sequential method that priced in the looks. A fixed horizon means you wait for the sample you precommitted.
What is sample ratio mismatch?
Answer
The observed traffic split is not the split you configured. At a large N, 46/54 against a 50/50 plan is a bug in bucketing, filtering, caching, or redirects. Pause and fix it. Do not read lift off a broken sample.
Guardrail versus primary metric?
Answer
The primary metric is the decision. Guardrails can veto a ship even when the primary metric wins. Latency, error rate, refunds, and crashes are the usual vetoes. A win that spends the error budget is not a win.
Why does an experiment need sticky assignment?
Answer
A user who flickers between variants does not have one experience. The behavior is biased and the causal read is gone. Sticky is the types-page contract, applied here so the metric means something.
Can you experiment without a flag system?
Answer
Yes. Dedicated tools assign and measure on their own. You still need consistent assignment, an exposure event, and an ops off-switch for a bad treatment. The off-switch is often still a flag.
What is CUPED in one sentence?
Answer
Variance reduction that uses pre-experiment covariates so you detect effects with less noise. It does not repair a wrong exposure or a sample ratio mismatch.
How do you run two experiments on one screen?
Answer
Put them in layers or a mutual-exclusion group so one user is not in two treatments of the same surface. Watch the shared guardrail. Interaction is a property of the surface, not a footnote.
Pitfalls
- Counting every assignment as if it were a view of the new checkout.
- Emitting exposure from both the client and the API for the same unit.
- Stopping on day two because the p-value dipped.
- Explaining a 46/54 split as "users preferred the treatment."
- Treating CUPED as permission to skip instrumentation.
- Letting two experiments share a button and a guardrail with no layer between them.
In three sentences: what you precommitted, why today's chart is a peek, and which guardrail would pause the ramp even if conversion is up. Then say where the error budget is documented, without re-deriving it.
Go deeper
- Overlapping Experiment Infrastructure is the Google paper on layers and domains.
- Trustworthy Online Controlled Experiments is Kohavi, Tang, and Xu on peeking, sample ratio mismatch, and metrics.
- LaunchDarkly experimentation is one vendor path from a flag to a measured split.
- The Microsoft Experimentation Platform group is the public face of a dedicated stack.