Engineering practices
Part 5 of 6 · Interview Debugging & Implementation — Reproduce, Trace, Fix, ShipImplement the Feature in the Same Session
Turn a vague also-add-X into a contract: inputs, errors, idempotency, and the edges you will not build in the remaining time.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What do you do with 'also add X'?
Answer
Turn it into inputs, outputs, errors, and non-goals before typing.
L2
What is the default retry slice?
Answer
One extra attempt on 408 or 504, with an idempotency key, then 503 if it still fails.
L3
How do you prevent a double charge?
Answer
The same idempotency key reaches payments. You do not design a new ledger in this session.
L4
Sync retry or a queue?
Answer
Sync once, unless the work is multi-minute or the load story is the thing they want scored.
L5
Which tests do you write?
Answer
Success after one failure, a non-retryable 409, a missing key, and exhausted retries becoming 503.
L6
They expand scope mid-way. Then what?
Answer
Re-open the non-goals. Ask what to swap out. Keep one shippable slice.
L7
Where does authz go?
Answer
One sentence: same caller checks as checkout. Do not redesign OAuth.
Failure modes
Retry without a key
The second attempt charges again. The demo looks green and the ledger does not.
Retrying a 409
A conflict is treated like a timeout. The client loops on a request that will never succeed unchanged.
Hidden non-goals
Circuit breakers, hedging, and a new metrics stack appear as comments and then as half-written types.
Feature reopens the bug
The new retry path maps every failure back to a generic 500 and undoes the status fix.
Misconceptions
More retries is more reliable.
Unbounded retries amplify a down dependency. One attempt, then a 503, is the interview default.
Idempotency means the handler is pure.
It means the same key returns the same outcome and does not repeat the side effect. The store that guarantees that is a different page.
A queue is always more senior.
A queue is more surface. Senior is noticing you cannot finish that surface in the time left.
Interviewer traps
Design the full idempotency store, TTL, and fingerprint mismatch.
Require the key, name the 409 on a body mismatch if you must, and point at API idempotency keys for the store.
Add a new authorization model for the retry.
Same caller as checkout. Stop.
Design scenario
Same prompt for every reader.
Requirements
Say trigger, max attempts, backoff, the idempotency key, success, exhausted failure, and non-goals. Implement happy path and one failure. Keep 409 non-retryable.
Traffic / scale
One checkout request, one extra attempt. Not a fleet-wide hedge.
Latency
The retry waits at most a short backoff you state, or it is immediate. Do not invent a multi-minute delay inside the request.
Consistency
The same key must not create two charges. A different body with the same key is a conflict, not a second success.
Availability
When attempts are exhausted, checkout returns 503 and logs the request id. It does not hang.
Failure assumptions
- The first attempt returns 504 and the second returns 200.
- A 409 can arrive on the first attempt.
- The caller can forget the key.
Constraints
- No new HTTP client.
- No multi-region hedge and no breaker dashboard.
Prompt
After the status fix, add retries when payments times out. You have the rest of the session.
API
What are the status codes for success, conflict, missing key, and exhausted timeout?
Data
Which field is the idempotency key, and who generates it?
Architecture
Why is this attempt inline instead of on a queue?
The slice you can finish
Prefer
One retry, one key, one exhausted status
408 and 504 retry once. 409 does not. A missing key is 400. Two failures become 503 PaymentTimeout. The key is required on the charge.
- The table exists before the loop.
- Non-goals are spoken: no breaker UI, no hedge, no new auth.
- Tests cover the four rows you named.
Alternative
A platform, then maybe the retry
A queue, a worker, a new client, and a metrics pipeline, with the timeout behavior still implicit.
- The clock dies in scaffolding.
- Double-charge is untested.
- The status fix from the previous slice gets lost.
Contract, then a thin implementation
If they widen scope, edit the non-goal list before you edit the code.
- 1
Write the table
Trigger, attempts, backoff, key, success, exhausted failure. - 2
Name non-goals
Breaker dashboard, multi-region hedge, a new permission model, a new client library. - 3
Happy path
One failure, then 200 and a payment id for the same key. - 4
One failure path
Missing key, or a 409, or two timeouts becoming 503. Say which one you coded. - 5
Stop
Ask what to deepen. Do not silently start the queue.
From "add retries" to a contract
Example prompt: add retries when payments times out.
| Field | Decision |
|---|---|
| Trigger | Upstream 504 or 408, or the client timeout |
| Max attempts | 2 total, meaning one retry |
| Backoff | Immediate, or 50ms. Say which |
| Idempotency | Require an idempotency key on the charge |
| Success | 200 with a payment id |
| Exhausted | 503 and Retry-After, logged with the request id |
| Non-goals | Breaker UI, multi-region hedging, a new auth model |
Authorization is one sentence: the same caller checks as checkout. You are not changing the permission model.
Flow
- 1
1 Vague also-add-X
- next2 Write the contract
- 2
2 Write the contract
- next3 Name the non-goals
- 3
3 Name the non-goals
- next4 Code the happy path
- 4
4 Code the happy path
- next5 One failure path
- 5
5 One failure path
- next6 Tests for both
- 6
6 Tests for both
- next7 Stop or deepen
- 7
7 Stop or deepen
Lesson map
Implement the Feature in the Same Session
Turn a vague also-add-X into a contract: inputs, errors, idempotency, and the edges you will not build in the remaining time.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Vague also-add-X"] b["2 Write the contract"] c["3 Name the non-goals"] d["4 Code the happy path"] a -->|1 Vague also-add-X| b b -->|2 Write the contract| c c -->|3 Name the non-goals| d
Pick one tradeoff and say why
| Choice | Pros | Cons | If you choose the other |
|---|---|---|---|
| Sync retry in the request | Simple, and the user path stays in this process | Holds a worker and can amplify load | A queue needs more surface than this room |
| Fail-fast, client retries | Protects payments | The client must be correct, and duplicates return without a key | Server retry hides that complexity and spends your attempts |
| Queue and worker | Room to back off, and the request stays short | New failure modes and a larger design | Overkill for one retry in forty minutes |
The default in this session is the sync retry. Mention the queue if the work is long or they ask about load.
Edges to name, even if you skip the code
- The same key with a different body is a 409, or first-write-wins. Say which. The store that enforces it is API idempotency keys.
- A retry after a success, because the response was late, must not charge twice.
- Charged, then checkout crashed, is a reconciliation case. Do not invent a second commit protocol in this sitting.
- A deadline that expires mid-retry means stop and return. Do not start the extra attempt.
What you will not build
Say it. Not a new metrics platform, not a new HTTP client, not multi-tenant rate limits. The fix page already separated timeout from conflict. This feature must not collapse them again.
Sandbox
One retry on timeout, a required key, and no retry on 409.
ProblemThe first 504 for key k1 is retried into a 200. An empty key is 400. A 409 is returned as-is. Two 504s become 503.
Expectedok_result.payment_id is pay_k1 after 2 calls. missing.status is 400. conflict.status is 409. exhausted.status is 503.
Edge cases
- A 200 on the first attempt never increments past one call.
- 409 does not become 503.
- Test: retry then pay
ok_result.payment_id == 'pay_k1' - Test: two attempts were used
ok_calls == 2 - Test: missing key is 400
missing.status == 400 - Test: conflict is not retried
conflict.status == 409 and conflict_calls == 1 - Test: two timeouts become 503
exhausted.status == 503 and exhausted.error == 'PaymentTimeout'
Press Run. Snippets must be self-contained — no network, files, or native modules.
Problemk1 fails once and then succeeds. An empty key is 400. A 409 returns immediately. Two 504s become 503.
ExpectedokResult.payment_id is pay_k1, missing.status is 400, conflict.status is 409, exhausted.status is 503.
Edge cases
- 408 is retryable the same way 504 is.
- A 200 does not turn into 503.
- Test: retry then pay
okResult.payment_id === 'pay_k1' - Test: two attempts
okCalls === 2 - Test: missing key
missing.status === 400 - Test: conflict stops
conflict.status === 409 && conflictCalls === 1 - Test: exhausted timeout
exhausted.status === 503
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
How do you prevent a double charge on retry?
Answer
Pass an idempotency key through to payments. The same key returns the original outcome. Fingerprints, TTLs, and the in-flight row are API idempotency keys. Do not invent a ledger unless they ask.
Sync retry or a queue?
Answer
Sync once for this scope. Mention a queue if the work is long or they want the load story. Do not start the worker in the same breath as the contract.
What tests do you write?
Answer
Success after one failure. A non-retryable 409. A missing key. Exhausted retries become 503. Those four match the table.
They expand scope mid-way.
Answer
Re-open the non-goals. Ask which item to swap out. Keep one slice you can ship. The hub script is the same status check you used when the fix was still open.
Where does authorization go?
Answer
The retry uses the caller checkout already authenticated and authorized. One sentence. A new permission model is a non-goal.
What if the body changes under the same key?
Answer
Say 409, or say the first body wins. Pick one. The page on idempotency keys is where the fingerprint check lives.
The payment succeeded and the response was lost. Then what?
Answer
The retry must hit the same key and return the stored success. If you did not build the store, say that is the dependency and stop. Do not add a second charge "just in case."
How does this stay compatible with the status fix?
Answer
504 and 408 are the only retryable upstream codes. 409 stays 409. A generic 500 must not return. That mapping was the previous slice.
Pitfalls
- Retrying every status, including 409.
- A key that is optional in the demo and required in the story.
- Starting a queue because it sounds more complete.
- Dropping the request id on the exhausted log line.
- Redesigning login as part of "also add retries."
Trigger, attempts, backoff, key, success, exhausted, and three non-goals. Then say the four tests. If you cannot name the non-goals, you are not ready to type.
Go deeper
- Stripe idempotent requests for the key header. The Study page owns the store.
- AWS retry behavior for backoff guidance you can cite without implementing a full client.
- Deadline propagation is the "stop when the budget is gone" sentence. A client timeout that starts a second attempt after the caller has already given up is not a retry.
Next: Low-level design under time.