Engineering practices
Part 3 of 6 · Interview Debugging & Implementation — Reproduce, Trace, Fix, ShipReproduce & Hypothesize — Narrow, Bisect & Local Repro
Tighten the window and the tenant, change one variable, and keep a hypothesis log. Git bisect is named at a high level and taught in the Git series.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What must you know before a patch?
Answer
Since when, who, what exactly fails, and what still works. Plus one failing request id if you can get it.
L2
What is in a hypothesis row?
Answer
The prediction, the evidence that would confirm it, the evidence that would kill it, and the result once you look.
L3
Why one variable?
Answer
If the flag, the restart, and the deploy change together, you cannot say which one moved the symptom.
L4
Local, staging, or logs?
Answer
Local when the inputs are known and the bug is logic. Staging when flags and wiring matter. Logs when the sandbox has no shell.
L5
When do you bisect instead of reading the latest pull request?
Answer
Read the latest change when the deploy correlation is strong and the diff is small. Bisect when the window is wide or many commits landed.
L6
You cannot reproduce locally. Then what?
Answer
Capture one payload, flag, build, and region from the logs. Say the gap. Continue on staging or on the log trail.
L7
How do you show you are not guessing?
Answer
Speak the prediction before the next query. Update the result column. Narrate a kill without defending the dead guess.
Failure modes
Three knobs at once
The flag flips, the process restarts, and a new build rolls. The symptom vanishes and nobody knows why.
Blank result column
Predictions accumulate and nothing is marked confirm or kill. The next edit is still a guess.
Local green, prod config ignored
The unit test uses the happy payload. Production sends a tenant the flag only enables there.
Bisect with a flaky check
The script sometimes returns 500 for unrelated reasons. The commit it names is noise.
Misconceptions
A hypothesis is a hunch you keep private.
If you did not say what would kill it, you do not have a hypothesis yet.
Bisect replaces reading the deploy.
A small diff at the correlated deploy is the faster read. Bisect is for a wide window.
No local repro means stop.
Log-only and staging are reproductions. Name which one you have.
Interviewer traps
Teach rebase versus merge while you are still narrowing the tenant.
Say you would bisect between the last green deploy and HEAD, then point at the Git series.
Explain error-budget burn math because the symptom mentions latency.
Use the SLO page for the budget. Here, since-when and who still come first.
Design scenario
Same prompt for every reader.
Requirements
Fill one hypothesis row before the next query. Change one variable. Say whether the repro is local, staging, or log-only.
Traffic / scale
One cohort is enough. Do not boil the ocean to 'all users'.
Latency
If the failure is slowness rather than 500, still narrow the window before any profiler.
Consistency
The failing request id, the build id, and the flag value must be from the same event.
Availability
A shared staging env can be flaky. Say so if you use it, and do not treat a single green retry as proof.
Failure assumptions
- The first query returns many error types.
- Local may miss IAM or a flag default.
- The regression window may contain more than one commit.
Constraints
- Do not reteach git object model or interactive rebase.
- Do not patch until since-when, who, and what-fails have answers.
Prompt
Checkout 500s started sometime after a morning deploy. Several tenants are mentioned. You may use logs, a staging flag, or a local test once you have a payload.
API
What is the one-request script you would bisect with?
Data
Which fields go in the hypothesis row for H1?
Architecture
Which single variable do you flip first, and which two do you leave alone?
How a row gets a result
Prefer
Prediction, then one knob
H1 says PaymentTimeout dominates after 14:12. You filter that window only. Other tenants and the flag stay put until this row is confirm or kill.
- The kill criterion is spoken before the query.
- A dead row is written down and left dead.
- The next edit has one cause attached.
Alternative
Change everything, then look
Flag off, restart, and redeploy, then decide the bug is gone.
- You cannot tell a restart from a fix.
- The table never gains a result.
- The next regression has no smaller search space.
Narrow, then mark the row
If you cannot answer since-when, who, and what fails, you are not ready to patch.
- 1
Since when, who, what
Deploy or config time, tenant or region or canary, the endpoint and error code, and what still works. - 2
Write H1
Prediction, evidence that confirms, evidence that kills. Leave result blank until you look. - 3
One variable
Time range, or tenant, or flag, or build id. Not two of those. - 4
Mark confirm or kill
Update the row out loud. A kill is progress. - 5
Only then the fix
The next page owns the smallest diff. This page stops at a confirmed row.
Narrowing, out loud
- Since when? A deploy, a config change, or a traffic spike.
- Who? Tenant, region, plan, canary versus everyone.
- What exactly fails? Endpoint, error code, or a latency burn. The budget math is the SLO lesson, not this one.
- What still works? Reads versus writes, or a neighboring service.
- Can you hold one failing request id?
If (1) through (3) are blank, do not patch.
Hypothesis log
| # | Prediction | Evidence that confirms | Evidence that kills | Result |
|---|---|---|---|---|
| H1 | Payment client timeout after the 14:12 deploy | PaymentTimeout dominates, and the trace leaf is payments | No timeouts, and database errors dominate | |
| H2 | Pool exhaustion | Pool metrics saturated, acquire latency high | Pool idle, and the payments span fails first | |
| H3 | Bad flag for tenant T | Only T fails, and the flag is on for T | Other tenants fail too |
Interviewers like seeing the result column move. A killed H1 is a good minute.
Decisions
- 1
1 Symptom is clarified
- next2 Write H1 prediction
- 2
2 Write H1 prediction
- next3 Narrow time and tenant
- 3
3 Narrow time and tenant
- next4 Gather the evidence
- 4
4 Gather the evidence
- next5 Confirm or kill?
- ?
5 Confirm or kill?
- kill6 Next hypothesis
- confirm7 Proceed to the fix
- 6
6 Next hypothesis
- next4 Gather the evidence
- 7
7 Proceed to the fix
Lesson map
Reproduce & Hypothesize — Narrow, Bisect & Local Repro
Tighten the window and the tenant, change one variable, and keep a hypothesis log. Git bisect is named at a high level and taught in the Git series.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s["1 Symptom is clarified"] h["2 Write H1 prediction"] n["3 Narrow time and tenant"] e["4 Gather the evidence"] s -->|1 Symptom is clarified| h h -->|2 Write H1 prediction| n n -->|3 Narrow time and tenant| e
One variable
Flip time range, or tenant, or flag, or build id. Pros: you can name the cause. Cons: it feels slow in a 45-minute room. Skipping it buys false confidence and a "fix" that was a restart.
Where the repro lives
| Approach | When it wins | What it misses | If you pick the other |
|---|---|---|---|
| Local repro | The inputs are known and the bug is logic | Prod config, data shape, IAM | You step a case the sandbox will never send |
| Staging | Flags and integration | Shared flaky state, slower loops | You wait on an env while logs already have the id |
| Log or trace only | AWS-style rooms, no shell, prod-only | You cannot step | You misread a line and never notice |
Say which one you are in. "I cannot reproduce locally" is fine if the next sentence is the payload, the flag, and the build you captured, or an explicit log-only continuation.
Bisect, high level
If the failure is a regression and you know a good revision, binary-search the commits with a minimal script. In the room: "I would bisect between the last green deploy and HEAD with a one-request script that asserts the status is not 500." The commands, the worktree so you do not stash over the bisect, and rebase versus merge live on Git and on branching, bisect, and worktrees. Read the latest pull request first when the deploy correlation is strong and the diff is small. Bisect when many commits landed or the window is wide. A flaky script names the wrong commit.
Flags
If a flag gates the new path, force it off for one request or one tenant. If the symptom disappears, that path is implicated. Pros: the blast radius shrinks immediately. Cons: flag debt and bad defaults. The rollback story for that flag is the fix lesson.
Sandbox
Encode prediction, evidence, and kill. The sample logs confirm H1.
ProblemH1 predicts at least two PaymentTimeout events. A second prediction, that every error is a database pool error, should die.
Expectedh1 is confirm. h2 is kill.
Edge cases
- An empty log kills a dominates prediction.
- Confirm is not a license to skip the result column.
- Test: timeouts confirm H1
h1['result'] == 'confirm' - Test: pool errors kill H2
h2['result'] == 'kill' - Test: both rows keep their text
h1['hypothesis'].startswith('PaymentTimeout')
Press Run. Snippets must be self-contained — no network, files, or native modules.
ProblemTwo predictions against the same three lines. Only the timeout prediction should confirm.
Expectedh1.result is confirm and h2.result is kill.
Edge cases
- A predicate that throws is a broken row, not a kill.
- Order of confirms does not matter.
- Test: H1 confirms
h1.result === 'confirm' - Test: H2 is killed
h2.result === 'kill'
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
The first query shows many error types. What do you do?
Answer
Stats count by error, pick the mode, and follow one id for that mode. Do not read the window chronologically until a pattern appears by luck. That query shape is the signal lesson.
You cannot reproduce locally.
Answer
Capture one failing payload and the env around it: flag, build, region. Write a characterization test that locks today's wrong behavior, or keep going on staging or logs and say which gap remains. The characterization itself is the fix lesson.
When do you bisect versus read the latest pull request?
Answer
Latest pull request when the deploy lines up and the diff is small. Bisect when the regression window is wide. The command sequence is branching, bisect, and worktrees, under the Git series.
How do you show you are not guessing?
Answer
Say the prediction before the next query. Update the table. A kill is a sentence you can say without embarrassment.
What does one-variable discipline cost?
Answer
Minutes. What it buys is a cause. Mixing a flag flip with a restart makes a green result unusable.
The symptom is latency, not a 500. Do you profile yet?
Answer
Not until since-when and who are known. If you need a budget or a burn rate, that is SLOs, error budgets, and tracing. A flame graph is still a different page.
A feature flag makes the bug vanish for one tenant. What is confirmed?
Answer
The flagged path is implicated for that tenant. It is not yet a line-level cause. The next query is inside that path, still one variable.
What do you hand the next slice?
Answer
A confirmed row: the prediction, the evidence, and the request id or build you will use as the proof after the diff.
Pitfalls
- Patching before since-when, who, and what-fails.
- A hypothesis with no kill criterion.
- Flag plus restart plus deploy in one step.
- Bisecting a flaky script.
- Reteaching rebase while the tenant is still unknown.
Using the table above, say the query that would kill pool exhaustion. Then say the single variable you would flip for H3, and the two you would not touch in the same minute.
Go deeper
- git-bisect for the command. The workflow around it stays on the Git series.
- OpenFeature for the idea of a flag evaluation, not a vendor tour.
Next: Fix, verify, and roll back.