Engineering practices
Part 2 of 6 · Interview Debugging & Implementation — Reproduce, Trace, Fix, ShipFinding the Signal — CloudWatch, Traces & Correlation IDs
CloudWatch Logs Insights query craft, alarms as pointers, and X-Ray or OpenTelemetry correlation. The same loop on Cloud Logging and kubectl.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What do you constrain before you read a message?
Answer
Time, service, and an identity such as status or request id.
L2
What is a Logs Insights query made of?
Answer
fields, filter, optional stats, sort, and limit. Filter before you pull the raw message.
L3
What is an alarm for?
Answer
It says look here, starting then. A metric filter that counts a named error is a better pointer than a naked 5xx rate.
L4
How do you follow one user?
Answer
Take one request id from the dominant error. Find that id in every service log. Open the trace with the same id.
L5
Which span do you trust?
Answer
The leaf error span. The root 500 is often just the propagation of that leaf.
L6
There is no request id in the logs. Then what?
Answer
Stats by error, status, or path. Use a trace id if the tracer minted one. Propose a structured id as follow-up, not as a reason to stop.
L7
How does kubectl change the loop?
Answer
It does not. --since, a label, and a request-id filter replace Insights. Multi-pod still needs an aggregator.
Failure modes
Day-long scan
fields @message with no time or service filter. The UI becomes the investigation.
Alarm as root cause
The 5xx alarm fired, so the diff scales the deployment. The leaf span was a payment timeout after a bad build.
Two ids for one request
The log request id and the trace id diverge mid-hop. You chase two stories and call them two bugs.
Span with no parent
A broken propagator looks like an unrelated service. It is one broken header.
Misconceptions
Knowing every CloudWatch feature is the interview.
The interview is the narrow-then-follow loop. The feature list is optional color.
A metric filter explains why.
It counts a class of lines. Why is the code path behind one of those lines.
GCP or Kubernetes is a different method.
The store changes. Time, identity, one id, then the code path do not.
Interviewer traps
Derive a multi-window burn alert before you have a request id.
Use the alarm as a start time. Burn math stays on the SLO page.
Refuse the AWS console because you prefer another vendor.
Use the tool they handed you, and say the loop is the same on Cloud Logging or kubectl.
Design scenario
Same prompt for every reader.
Requirements
Show a filter for 5xx, a stats-by-error query, one request id, and the leaf span you would open.
Traffic / scale
Thousands of lines in thirty minutes. You read dozens, not thousands.
Latency
Only if the leaf span's duration is the story. Otherwise status and error code come first.
Consistency
One id crosses logs and traces. A second minted id mid-request is a bug in propagation, not a second incident.
Availability
If Insights is slow, the same filter on a shorter window still counts as progress.
Failure assumptions
- Some lines lack a request id.
- The alarm has no error class on it.
- Payments and checkout log to different groups.
Constraints
- Do not design a new telemetry pipeline.
- Do not re-teach sampling policy.
Prompt
Checkout 5xx rose at 14:12. You have Logs Insights on the checkout and payments log groups, plus a trace backend.
API
Which query finds the mode, and which query follows one id?
Data
Which field joins the log line to the trace?
Architecture
Which service owns the leaf span, and which file do you open?
Where the first query should land
Prefer
Mode, then one id
Stats by error code after 14:12, then a single request id, then the trace for that id.
- Time and service are in the filter before the message body.
- The dominant error names the hypothesis.
- The leaf span picks the service whose code you open.
Alternative
Read the stream
Sort the last day of checkout logs and scroll until a 500 appears.
- Volume hides the mode.
- You never learn whether PaymentTimeout or pool errors dominate.
- The trace, if you open one, is a random success.
From alarm to one code path
Each step throws information away on purpose.
- 1
Read the pointer
The alarm gives a start time and a service. It does not give a cause. - 2
Find the mode
Stats by error code, status, or path inside that window. - 3
Take one id
Filter to the dominant error and copy one request id. - 4
Open the trace
The leaf error span names the service and the exception. A missing parent is a propagation bug. - 5
Open one path
That service's handler for that error. Stop there until the hypothesis is confirmed or killed.
Query shape
A log group is a stream of events. Insights queries @timestamp, @message, and fields you parsed or that the app emitted as JSON.
Typical interview pattern for recent checkout failures:
fields @timestamp, @message, request_id, status, latency_ms, service
| filter service = "checkout" and status >= 500
| sort @timestamp desc
| limit 50Put the time window in the console picker or add a filter on @timestamp. When you do not have an id yet, aggregate:
fields @timestamp, error_code
| filter service = "checkout"
| stats count(*) as n by error_code
| sort n desc
| limit 20Then filter to the dominant error_code, take one request_id, and follow only that id. The payments group should show the same id if propagation worked.
| Construct | Use | Antipattern |
|---|---|---|
| filter early | Cut volume before you read messages | A full-text scan of a day |
| stats by | Find the mode | Ten thousand raw lines and a hope |
| sort and limit | Newest, or top N | An unlimited result set |
| exact request id | One journey | Mixing unrelated ids in one story |
Alarms are pointers
An alarm on 5xx above a threshold says look here, starting then. It does not say why. A metric filter that counts PaymentTimeout in the checkout group is a sharper pointer because it names the class. In the room: "The alarm says checkout errors rose at 14:12. Next I stats by error type after that timestamp."
If you treat the alarm as the cause, you scale the pods while the real change is a bad deploy or a dependency timeout. How that 5xx rate spends an error budget is SLOs, error budgets, and tracing. Which of metrics, logs, or traces you open first is the observability triad.
One request across services
Prefer one id that propagates. The log field is request_id. The span carries trace_id. Outbound headers should carry the same value when you control them. A span without a parent is a broken propagation story. The leaf error span is usually closer to the fault than the root 500.
Open the trace for that id, note the failing subsegment's service, duration, and exception, then open that service's code path. Sampling policy and W3C header layout stay on the SLO and triad pages. This page only needs the join.
Decisions
- 1
1 Alarm says when
- next2 Stats by error code
- 2
2 Stats by error code
- next3 Pick one request id
- 3
3 Pick one request id
- next4 Follow that id in logs
- 4
4 Follow that id in logs
- next5 Trace shows a leaf?
- ?
5 Trace shows a leaf?
- yes6 Open that code path
- no7 Next hypothesis
- 6
6 Open that code path
- 7
7 Next hypothesis
Lesson map
Finding the Signal — CloudWatch, Traces & Correlation IDs
CloudWatch Logs Insights query craft, alarms as pointers, and X-Ray or OpenTelemetry correlation. The same loop on Cloud Logging and kubectl.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Alarm says when"] b["2 Stats by error code"] c["3 Pick one request id"] d["4 Follow that id in logs"] a -->|1 Alarm says when| b b -->|2 Stats by error code| c c -->|3 Pick one request id| d
The diagram is a column on purpose. A five-actor sequence overflows a narrow card. The story is still checkout calling payments, both writing the same request id, and the trace backend holding the leaf.
Same loop, other stores
| Platform | Entry | Narrow | Follow |
|---|---|---|---|
| AWS | Log group and Logs Insights | filter on time, service, status | request id, then the X-Ray or OTel trace |
| GCP | Cloud Logging | resource labels, jsonPayload, time range | the trace field, then Cloud Trace |
| Kubernetes | kubectl logs with --since | a label, then a request-id filter | stern or the mesh tracer if many pods |
Staying vendor-agnostic shows the skill transfers. If they gave you an AWS console, use it. Refusing the provided tool looks stubborn.
Sandbox
Filter, count, then one id. Do not print the successful requests.
ProblemKeep checkout events at status 500 or above, count by error, and follow the first request id of the dominant error across every service.
ExpectedPaymentTimeout has count 2. The followed id is r2. The trail has the checkout line and the payments line.
Edge cases
- A service with no errors returns an empty mode.
- The trail includes payments even though the filter service was checkout.
- Test: dominant error is PaymentTimeout
result['stats']['PaymentTimeout'] == 2 - Test: follow the first id of that error
result['follow_request_id'] == 'r2' - Test: trail crosses services
len(result['trail']) == 2 - Test: payments line is on the trail
any(event['service'] == 'payments' for event in result['trail'])
Press Run. Snippets must be self-contained — no network, files, or native modules.
ProblemCheckout 5xx lines only. The dominant error's first request id should pull the payments line too.
Expectedstats.PaymentTimeout is 2, follow_request_id is r2, and the trail length is 2.
Edge cases
- Missing error becomes the key unknown.
- A 200 never enters the hit list.
- Test: two PaymentTimeout lines
result.stats.PaymentTimeout === 2 - Test: sample id is r2
result.follow_request_id === 'r2' - Test: trail has both hops
result.trail.length === 2 - Test: payments is on the trail
result.trail.some((event) => event.service === 'payments')
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
Write a Logs Insights query for checkout 5xx in the last 30 minutes.
Answer
fields @timestamp, request_id, status, @message, then filter service = "checkout" and status >= 500, then sort @timestamp desc, then limit 50. Put the 30 minutes in the time picker or in a timestamp filter. Do not select the message until the filter is on.
You have no request id in the logs. What now?
Answer
Stats by error, status, or path to find the mode. Check whether traces still carry a generated id. Adding a structured request id is the follow-up, not a reason to stop the loop.
Alarm, log, or trace: which first?
Answer
The alarm for when and how big. Logs for what. The trace for where across services. Start with whichever one can kill the current hypothesis fastest. The triad is the longer version of that sentence.
How is kubectl different from CloudWatch here?
Answer
Same loop, different store. --since, a label selector, and a request-id filter replace Insights. Many pods need an aggregator. You still end on one id and one code path.
Why is the leaf span more useful than the root 500?
Answer
The root status is often the mapping of a downstream failure. The leaf names the service, the duration, and the exception you can search for in that service's code.
What is a metric filter allowed to conclude?
Answer
That a named string became more frequent after a time. It cannot conclude the deploy, the pool, or the client. Those are hypotheses the filter helps you order.
Two ids appear on one request. What do you call that?
Answer
A propagation bug until proven otherwise. Do not file two incidents because checkout logged one id and payments logged another.
When do you leave this page?
Answer
When you have a confirmed or killed hypothesis and a code path. Reproduction and the one-variable rule are the next lesson. The diff and the proof are the fix lesson.
Pitfalls
fields @messageacross a day with no filter.- Scaling because the alarm fired.
- Following a new id on every hop.
- Opening five services because five log groups exist.
- Treating a missing parent span as a second root cause.
Write the stats query for checkout after 14:12. Circle the top error code. Write the second query that keeps only that code and prints twenty request ids. Pick the first id and say which span you open. Throw the other nineteen away until that one story is confirmed or killed.
Go deeper
- Logs Insights query syntax and metric filters.
- W3C Trace Context for the header that should match the log id.
- Cloud Logging query language and kubectl logs as the same loop.
Next: Reproduce and hypothesize.