Probes & Pod Lifecycle — Liveness, Readiness, Startup & PreStop
Startup probes buy slow boots. Readiness gates Service traffic. Liveness restarts stuck processes — never point it at a dependent database or you restart-storm yourself. PreStop plus terminationGracePeriodSeconds is how you drain cleanly.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
What a failed check is allowed to do
Prefer
Local liveness, dependency-aware readiness
Liveness answers only: is this process wedged? Readiness answers: should this Pod take traffic right now?
- A slow database sheds load without killing every Pod.
- Startup covers a long boot so liveness does not fire early.
- The process stays up while it is merely unready.
Alternative
One heavy health URL for every probe
The same handler checks the database, the cache, and disk. Liveness and readiness fail together.
- Dependency slowness restarts the whole fleet.
- Restarts add load to the thing that is already slow.
- You lose in-flight work that a NotReady Pod could have drained.
Shutdown order
The grace period includes the hook and the process exit. Endpoint removal is not instant, which is why a short preStop exists.
- 1
Terminating
The Pod is marked deleted. EndpointSlice removal starts. - 2
preStop
Optional hook. A brief sleep can cover a slow endpoint update. It is not a substitute for app drain. - 3
SIGTERM
Sent after preStop finishes. The app stops accepting work and finishes in-flight calls. - 4
SIGKILL
If the container is still running when the grace period ends, the kubelet kills it.
Overview
Wrong probes are a top cause of restart storms and of "half the Pods are Ready and still broken." Pasting /healthz into every field does not design a failure domain.
This page is kubelet probes and Pod shutdown. Load-balancer health checks, slow start, and connection draining at a proxy are Health Checks, Slow Start & Connection Draining. North-south routing is Private Networking. Do not merge those into a probe answer.
Three probes, three jobs
startupProbe. Slow boots (JVM, migrations, huge classpaths) must not be killed by liveness. While the startup probe is failing, liveness and readiness are not run. After it succeeds, the other probes take over. Size it so failureThreshold * periodSeconds covers the worst cold start, plus margin. If it never succeeds and the failure threshold is hit, the kubelet restarts the container. "Still starting" is not infinite.
livenessProbe. Detect a deadlock or a wedged process. Failure past the threshold restarts the container, subject to the Pod restart policy. The check must be local. Do not call a database or a cache. If that dependency is slow, every Pod restarts, and the dependency gets worse.
readinessProbe. This is Service membership. Not Ready means the Pod leaves the Endpoints or EndpointSlice. The process keeps running. Use it for warm-up and for dependencies where the right move is to shed traffic, not to die. A readiness failure does not restart the container. If every Pod fails readiness, the Service has zero backends. That is an outage you chose.
Mechanisms and fields
httpGeton a path and port. Prefer a cheap in-process handler.tcpSocketonly proves the port is open. That is weak semantic health.execruns a command in the container. It forks. At large replica counts that cost shows up. It can also contend on a lock or a GIL if the command is heavy.grpcuses the gRPC health protocol on modern kubelets. Prefer it over an exec wrapper aroundgrpc_health_probewhen the kubelet supports it.
Shared fields: initialDelaySeconds, periodSeconds, timeoutSeconds, successThreshold, failureThreshold. Readiness is allowed a successThreshold above 1 so a flapping Pod must pass several times before it returns to the Service. Liveness successThreshold must be 1.
Lifecycle when the Pod is deleted
- The Pod is Terminating and endpoint removal begins. It is not guaranteed to finish before the process is signaled.
preStopruns if you defined it. A short sleep is a common bridge so kube-proxy or a mesh finishes removing the endpoint.- The kubelet sends SIGTERM after preStop completes. Both the hook and the process share
terminationGracePeriodSeconds. - The app must stop accepting work, finish in-flight requests, and exit 0.
- If it is still alive when the grace period ends, SIGKILL.
A long preStop sleep with no application drain is an anti-pattern. The sleep only buys time. The process still has to handle SIGTERM. If preStop consumes the whole grace period, the process gets little time before SIGKILL.
CPU throttling can make a probe miss its timeout even when the process is logically fine. That interaction is Requests, Limits & QoS.
Decisions
- 1
1. Containers start
- next2. Startup probe
- ?
2. Startup probe
- failing3. Hold other probes
- ok5. Readiness and liveness
- 3
3. Hold other probes
- next4. Retry until success
- 4
4. Retry until success
- next5. Readiness and liveness
- 5
5. Readiness and liveness
- next6. Ready for traffic
- ?
6. Ready for traffic
- yes7. Join EndpointSlice
- no8. Stay out of Service
- 7
7. Join EndpointSlice
- next9. Process still healthy
- 8
8. Stay out of Service
- next9. Process still healthy
- ?
9. Process still healthy
- no10. Restart container
- yes11. Serve until delete
- 10
10. Restart container
- 11
11. Serve until delete
- next12. preStop then SIGTERM
- 12
12. preStop then SIGTERM
- next13. Exit or SIGKILL
- 13
13. Exit or SIGKILL
Lesson map
Probes & Pod Lifecycle — Liveness, Readiness, Startup & PreStop
Startup probes buy slow boots. Readiness gates Service traffic. Liveness restarts stuck processes — never point it at a dependent database or you restart-storm yourself. PreStop plus terminationGracePeriodSeconds is how you drain cleanly.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Containers start"] b["2. Startup probe"] c["3. Hold other probes"] d["4. Retry until success"] a -->|1. Containers start| b b -->|failing| c c -->|3. Hold other probes| d
The diagram is linear on purpose. In the cluster, a failed startup check loops until success or the failure threshold restarts the container. A failed liveness check restarts and the cycle begins again.
Restart-storm anti-patterns
- Liveness hits a shared database. The database slows down. Every Pod fails liveness. Mass restart makes the database slower.
- Liveness and readiness share one heavy endpoint.
failureThresholdis tiny and a GC pause exceeds it.- No startup probe on a slow boot, so liveness kills the process during init.
- The process ignores SIGTERM, so every deploy ends in SIGKILL mid-request.
- Readiness checks a dependency with no cache and no timeout, so one blip empties the Service.
Probe decisions, in memory
Liveness ignores dependency_ok. Readiness includes it. Startup only cares that the process has finished booting. Crossing the failure threshold restarts on startup and liveness, and only removes traffic on readiness.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
While the app is still booting, startup fails under its threshold and the other probes are not the ones you would trust yet. After boot, readiness stays failed until the dependency is ok. When the event loop wedges, liveness reaches RESTART_CONTAINER. Readiness reaches REMOVE_FROM_SERVICE. Those are different outcomes.
Interview Q&A
What is the difference between readiness and liveness?
Answer
Readiness turns traffic on or off by Service membership. Liveness restarts the container. Readiness failure leaves the process running. Liveness failure does not.
When do you need a startup probe?
Answer
When a correct boot is slower than the liveness deadline. Cold JVM, migrations, or a large classpath. Without it, liveness kills the container during init, and the restart loop never finishes booting.
Why is liveness-on-database dangerous?
Answer
The dependency is shared. When it slows, every Pod fails the same check, every Pod restarts, and the database takes a thundering herd. Shed that failure with readiness, or with a local "am I wedged" check that does not call the database.
What happens on SIGTERM?
Answer
The process should stop accepting new work and drain in-flight calls, then exit. If it is still running after terminationGracePeriodSeconds, the kubelet sends SIGKILL. Fail readiness as soon as you begin shutdown so new requests stop. A preStop hook can buy a moment for endpoint removal.
Can a readiness failure restart a container?
Answer
No. It removes the Pod from the Service endpoints. Only startup and liveness, past their failure thresholds, ask the kubelet to restart the container.
How do you size startup failureThreshold?
Answer
failureThreshold * periodSeconds must be at least the p99 cold start, plus margin. If the product is shorter than boot, you restart a healthy-but-slow process.
exec probe or httpGet?
Answer
httpGet hits an in-process handler and is usually the right default. exec forks a process on every period, on every Pod. That is expensive, and a slow exec can time out for reasons that are not "the app is dead."
What should a good health endpoint do?
Answer
Stay cheap and local. No authentication redirect loop. Liveness does not fan out to dependencies. Readiness may check a critical dependency if that check is fast, cached, and an intentional load-shed. It must not be a full synthetic transaction on every second.
Why a preStop sleep at all?
Answer
Endpoint removal races the SIGTERM. A short preStop gives kube-proxy time to stop sending new connections. The sleep is not the drain. The application still has to finish requests after SIGTERM. A 60-second sleep that eats the grace period just moves the SIGKILL earlier in the app's shutdown.
How is this different from a load balancer health check?
Answer
The kubelet probe decides Pod restart and EndpointSlice membership inside the cluster. A load balancer health check is outside that loop. Slow start and proxy draining are Health Checks, Slow Start & Connection Draining. Publishing the Service past the node is Ingress Controllers & North-South.
Pitfalls
- Copying one probe block into startup, liveness, and readiness.
timeoutSecondsshorter than a normal GC pause, withfailureThresholdof 1.- Readiness that depends on a Pod that depends on you. You empty both Services.
- Ignoring SIGTERM and discovering it only on the next deploy, as client errors.
- Using a mesh retry policy to hide a probe that restarts healthy Pods. Fix the probe.
The handler pings Postgres and Redis and is wired to all three probes with failureThreshold 1 and periodSeconds 1. Say what you change for each probe, and what you expect during a 2-second database stall.