Engineering practices
Part 1 of 6 · Interview Debugging & Implementation — Reproduce, Trace, Fix, ShipInterview Debugging & Implementation — Reproduce, Trace, Fix, Ship
Timed interview loop: clarify, hypothesize, evidence, fix, verify, then a feature or one-component LLD. Logs versus a debugger versus reading code cold.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is the interviewer scoring?
Answer
A loop you can narrate: symptom, prediction, evidence, one change, proof, then a contract. Vocabulary without that trail does not pass.
L2
What belongs in the first five minutes?
Answer
Restate the symptom, who is hit, since when, and what fixed means. Do not open a log group yet.
L3
What makes a hypothesis useful?
Answer
It predicts a specific signal. If that signal is absent, the hypothesis is dead and you say so.
L4
Logs, a debugger, or reading the code?
Answer
Logs and traces when the prompt is a distributed sandbox. A debugger for a local single-process crash. Cold reading when a small diff or a stack is already in hand.
L5
How do you avoid drowning in logs?
Answer
Filter time, service, severity, and one correlation id. Aggregate to find the mode, then follow that one id.
L6
The interviewer pivots to a feature before the fix is done. What do you say?
Answer
Give a thirty-second status and ask which slice they want scored. Prefer closing the fix when fixed was the success criterion.
L7
How is the low-level design different from a system sketch?
Answer
Low-level design is one component's types, invariants, state, and failure behavior. Boxes for the whole company are a different interview.
Failure modes
Fix before evidence
Three files change before any prediction. The room cannot tell a repair from a restart.
Alarm treated as the cause
The page says checkout 5xx rose at 14:12. Scaling the service hides a bad deploy or a dependency timeout.
Unfiltered day of logs
A query with no time, service, or id returns a wall of text. The mode never gets named.
Feature starts as a framework
A new retry library begins before the contract says one retry, an idempotency key, and a 503 when attempts are exhausted.
Misconceptions
The person who pastes the most log lines is the most thorough.
Thorough means a prediction, a filter, and an update when the prediction dies.
Logs Insights syntax is the skill being scored.
The syntax is a tool. The skill is narrowing to one request and one code path.
Low-level design means drawing every service.
This loop wants interfaces and state for one component after the bug is under control.
Interviewer traps
Explain metrics versus logs versus traces from first principles.
Pick the cheapest signal that can kill the current hypothesis, then point at the observability triad.
Design a multi-region architecture while the 500 is still open.
Finish the smallest fix or explicitly park it. Keep the design inside one component.
Design scenario
Same prompt for every reader.
Requirements
Name the symptom and the success criterion. Confirm one hypothesis with a filtered query and one trace. Ship the smallest diff and rerun that query. State the retry contract, including what you will not build.
Traffic / scale
A 45 to 60 minute room. One failing cohort, not the whole product.
Latency
The interesting latency is the payment leaf, if the trace says so. A flame graph is a different page.
Consistency
A retried charge must not double-bill. Say the idempotency key and stop before designing a ledger.
Availability
Rollback is a flag, a revert, or the previous artifact. Say who is still exposed.
Failure assumptions
- The first hypothesis can be wrong.
- The console may be CloudWatch, Cloud Logging, or kubectl.
- Time will run out if the fix becomes a rewrite.
Constraints
- Do not reteach the observability triad, SLO burn math, or git rebase.
- Do not open a separate high-level design series.
Prompt
Checkout returns 500s for some users after 14:12. You have a log dump or a staging console. Find the bug, fix it, then add a retry with idempotency on the payment call.
API
What does fixed mean, and what is the retry contract if time remains?
Data
Which id do you follow from the log line to the span?
Architecture
Which single code path do you open, and what stays out of scope?
What the room should hear
Prefer
One prediction, then one signal
Hypothesis: payment client timeout after the 14:12 deploy. If true, checkout logs show PaymentTimeout after that time and the trace leaf is payments. If false, the next guess is pool exhaustion.
- The symptom, the cohort, and the meaning of fixed are restated first.
- One query, one id, one code path.
- The same query is the proof, and the feature starts as a contract.
Alternative
Five tabs and a maybe
Unfiltered logs, three files edited before any prediction, and a Redis guess with nothing that would kill it.
- The first fifteen minutes land in unrelated modules.
- A restart and a code change get mixed, so the result has no cause.
- The feature never starts, or it starts as a new framework.
The 45-minute script
Say the slice out loud. The sibling pages hold the query, the repro, the diff, the contract, and the types.
- 1
Clarify
Restate the symptom, the blast radius, since when, who is affected, and what fixed means. - 2
Hypothesize
Rank two or three causes. Pick the top one. Say the evidence that would kill it. - 3
Find one signal
Filter logs or traces. Follow one correlation id. Open one code path. - 4
Smallest fix
Change one thing. A characterization test if there is time. No drive-by rewrite. - 5
Verify
Rerun the same query or the same request. Done means that signal is quiet for the repro. - 6
Feature or low-level design
Contract, edges out of scope, then interfaces or a state machine, with a happy path and one failure path.
What good sounds like
You are in a 45 to 60 minute loop. The prompt looks like checkout returning 500s, with a log dump or a staging URL, and then a request to add a retry with idempotency on the payment call. Or it looks like designing the types for a request-scoped log correlator and implementing the missing piece. The score is the trail, not a syntax recital.
| Slice | Minutes in a 45-minute room | What you say |
|---|---|---|
| Clarify | 0-5 | Symptom, blast radius, since when, who, and the meaning of fixed |
| Hypothesize | 5-10 | Two or three causes, the top one, and the evidence that would kill it |
| Evidence | 10-22 | Filtered logs or traces, one correlation id, one code path |
| Fix | 22-32 | Smallest correct diff, a characterization test if time allows, same query again |
| Feature or LLD | 32-45 | Contract, non-goals, interfaces or state, tests for the happy path and one failure |
A useful sentence sounds like this. Hypothesis: payment client timeout after the deploy at 14:12. If true, the checkout log group shows elevated PaymentTimeout for that service after that time, and the trace shows the payment span as the leaf error. If false, the next hypothesis is database pool exhaustion.
Decisions
- 1
1 Clarify the symptom
- next2 One hypothesis
- 2
2 One hypothesis
- next3 Find one signal
- 3
3 Find one signal
- next4 Evidence matches?
- ?
4 Evidence matches?
- no2 One hypothesis
- yes5 Smallest fix
- 5
5 Smallest fix
- next6 Same query is quiet
- 6
6 Same query is quiet
- next7 Feature contract
- 7
7 Feature contract
- next8 LLD interfaces
- 8
8 LLD interfaces
- next9 Ship or stop
- 9
9 Ship or stop
Lesson map
Interview Debugging & Implementation — Reproduce, Trace, Fix, Ship
Timed interview loop: clarify, hypothesize, evidence, fix, verify, then a feature or one-component LLD. Logs versus a debugger versus reading code cold.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Clarify the symptom"] b["2 One hypothesis"] c["3 Find one signal"] d["4 Evidence matches?"] a -->|1 Clarify the symptom| b b -->|2 One hypothesis to 3 Find one signal| c c -->|3 Find one signal| d d -->|no| b
The no branch is an update, not a failure. Write the kill in the hypothesis table and move. Interviewers score that update.
How you hunt
| Mode | When it wins | What you give up | If you pick it anyway |
|---|---|---|---|
| Logs and traces | Distributed sandboxes, a CloudWatch or kubectl prompt, on-call shaped bugs | You cannot see locals. Bad ids leave you blind | You match the setup they handed you |
| Attach a debugger | Local unit bug, single-process crash, a clear stack | Rarely available against prod, and it can hide systemic thinking | You look like you ignored a log prompt |
| Read the code cold | A small recent diff, a typo, a config flag, a stack already in hand | Env-only failures stay invisible | You spend the first fifteen minutes in the wrong module |
Starting with a debugger when the prompt is "use CloudWatch" signals you skipped the setup. Reading cold without a hypothesis does the same. Logs without ever opening code leave you unable to name a diff.
What this cluster does not reteach
Metrics tell you something is wrong and since when. Logs give the message and the fields. Traces connect services for one request. Start with the cheapest signal that can kill the hypothesis. The pillar lesson is the observability triad. Urgency against a budget is SLOs, error budgets, and tracing. When the bug is latency rather than a hard error, the flame graph lives on performance engineering. git bisect for "when did this start" lives on Git. Safe retries on writes live on API idempotency keys. This series does not rewrite those pages and does not open a separate high-level design or pattern catalog. Low-level design in the last lesson is the in-interview component, not a second series.
Series map
- This hub, the loop and the rubric.
- Finding the signal in CloudWatch, traces, and correlation ids, with the GCP and kubectl twins.
- Reproduce and hypothesize: narrow, bisect at a high level, local repro.
- Fix, verify, and roll back: the smallest change and the proof.
- Implement the feature in the same session: a contract under a clock.
- Low-level design under time: interfaces, state, and tradeoffs for one component.
Sandbox
The helper is a rubric, not production code. A fix recorded before evidence is a thrash.
ProblemMark which interview phases you actually did. A fix with no evidence is a thrash even if other boxes are ticked.
ExpectedThe partial run is a thrash and not a pass. The full run passes with no missing phase.
Edge cases
- Verify without a fix is still a missing phase.
- A full board is a pass only when thrash is false.
- Test: partial run missed evidence
'Evidence' in partial['missing'] - Test: fix without evidence is a thrash
partial['thrash_risk'] is True - Test: partial run does not pass
partial['pass'] is False - Test: full board passes
full['pass'] is True - Test: full board is not a thrash
full['thrash_risk'] is False
Press Run. Snippets must be self-contained — no network, files, or native modules.
ProblemThe checklist is identical. A fix without evidence fails the room even when clarify and hypothesis are done.
Expectedpartial.thrash_risk is true and partial.pass is false. full.pass is true.
Edge cases
- An empty board is all missing and not a thrash.
- Evidence without Fix is incomplete, not a thrash.
- Test: partial run missed evidence
partial.missing.includes('Evidence') - Test: fix without evidence is a thrash
partial.thrash_risk === true - Test: partial run does not pass
partial.pass === false - Test: full board passes
full.pass === true
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Walk me through a production bug you debugged.
Answer
Symptom, time window, who was hit, the first hypothesis, the exact query or trace that confirmed it, the small fix, how you proved the signal went quiet, and the follow-up (an alert, a test, or a runbook). Name CloudWatch Logs Insights, X-Ray, or kubectl without turning the answer into a product tour.
How do you avoid drowning in CloudWatch?
Answer
Filter by time, service, severity, and a correlation id first. Use stats and sort to find the mode, then drill into one request. A bare message dump over a full day is how the room stalls.
Logs, traces, or metrics for this interview?
Answer
Metrics say something is wrong and since when. Logs give the message and the fields. Traces connect services for one request. Start with the cheapest signal that can kill the current hypothesis. The pillar split itself is the observability triad.
What if the first hypothesis is wrong?
Answer
Say so. Write prediction, evidence, and result. Move to the next cause. The update is the score, not a lucky first guess.
They ask for a feature before the fix is finished. What do you do?
Answer
Offer a thirty-second status. You have a likely cause and a one-line fix. You can land the fix and then design, or park the fix and design now. Ask which one they want scored. Prefer closing the fix when fixed was the success criterion.
Design the classes for a request-scoped log correlator.
Answer
A short version here, and the full walk on the low-level design page. A correlation context holds a request id and maybe a tenant id, stored in request scope. Middleware creates or accepts it. A logger adapter injects the fields. Outbound calls copy the bound id into headers. Never mint a second id mid-request, and never put personal data in the id.
How is low-level design different from a system sketch here?
Answer
A system sketch is services, stores, and request flow across the company. Low-level design is one component's types, state, and failure behavior. This loop usually wants the second one after a bug fix.
When does performance profiling enter this loop?
Answer
When the symptom is latency rather than a hard error, and a filtered trace already names the slow span. The capture and the flame graph live on performance engineering. Do not start there while 500s are still unexplained.
Pitfalls
- Opening the log group before saying what would kill the hypothesis.
- Changing a flag, restarting, and redeploying in one step.
- Treating the alarm threshold as the root cause.
- Starting a retry framework before the one-line fix is proven.
- Drawing every service when the prompt asked for one component.
Out loud, in under thirty seconds: the symptom, the deploy time, the log field you expect, the span you expect to be the leaf, and the next hypothesis if that field is absent. Then name which sibling page you would open for the query, the repro, the diff, the retry contract, or the types.
Go deeper
- CloudWatch Logs Insights query syntax for the filter, stats, and sort shape.
- AWS X-Ray and OpenTelemetry traces for one request across services.
- The SRE workbook incident-response chapter for the structure of a timeline, not a tool.
Next: Finding the signal.