Kubernetes Control Plane - API Server Request Path, etcd, Watches, Informers, Optimistic Concurrency & CRDs
Control plane: API server request pipeline (authn, API Priority and Fairness, RBAC, mutating/validating admission, schema validation), etcd limits and compaction, resourceVersion, watches, 410 Gone and relist, 409 optimistic concurrency, shared informers and workqueues, CRDs and operators vs aggregated APIs.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What order does the API server process a write in?
Answer
Authentication, flow control, authorization, mutating admission, schema validation, validating admission, then etcd.
L2
What is a resourceVersion?
Answer
An opaque version, in practice the etcd revision of the object's last write, used for watches, read consistency and optimistic concurrency.
L3
What does 409 Conflict mean?
Answer
Your write was based on a stale resourceVersion. Re-read, recompute and retry.
L4
What does 410 Gone on a watch mean?
Answer
The resourceVersion you resumed from was compacted. Relist and start a new watch.
L5
Why do controllers use informers?
Answer
One LIST plus one WATCH feeds a local cache, so reads cost the API server almost nothing and work is de-duplicated per key.
L6
Webhook or ValidatingAdmissionPolicy?
Answer
CEL policies run in-process with no network hop. Use webhooks only when you need external data or complex logic, with tight timeouts and selectors.
L7
CRD or aggregated API server?
Answer
A CRD for almost all declarative config stored in etcd. An aggregated API for non-declarative APIs, custom storage or very high churn.
Failure modes
Webhook outage blocks writes
A failurePolicy Fail webhook whose backend is down rejects or times out every matching request.
etcd quota exceeded
Large objects or high churn fill the backend and writes stop with database space exceeded.
Lost updates from blind writes
Patches without resourceVersion preconditions silently overwrite another automation's change.
Misconceptions
Admission can hide data from readers.
Admission only sees writes. Reads are controlled by RBAC.
resourceVersions can be compared across resource types.
They are only orderable within one resource type. Treat them as opaque.
Informer caches are always current.
They lag slightly, so writes must carry a resourceVersion and conflicts must requeue.
Interviewer traps
Saying controllers poll the API server.
They list once and watch, through shared informers and a work queue of keys.
Recommending failurePolicy Ignore as the fix for webhook outages.
Ignore keeps writes flowing but silently skips policy. Prefer in-process CEL policies and scoped webhooks.
Design scenario
Same prompt for every reader.
Requirements
Declarative spec per database, status conditions, safe upgrades, and no lost updates when the HPA, CI and the operator touch related objects.
Failure assumptions
- The operator restarts mid-reconcile.
- Its watch falls behind the compaction window.
- Two automations update the same object.
Constraints
- etcd default quota of 2 GiB.
- API Priority and Fairness is on with default levels.
Prompt
Design an operator that manages 2,000 Postgres clusters as a custom resource on one Kubernetes cluster without overloading the API server.
API
What does the CRD schema, status subresource and versioning look like?
Data
What is stored in etcd vs elsewhere, and how big is each object?
Architecture
How do informers, the work queue, finalizers and leader election fit together?
Overview
The control plane is a REST API in front of a replicated key-value store, with a change feed. The API server is the only component that talks to etcd. It authenticates and authorizes each request, runs admission plugins that can rewrite or reject it, validates it against a schema, and stores it with a monotonically increasing resourceVersion. Every other component, including the scheduler and all controllers, is just a client that lists objects once and then watches for changes, keeping a local cache. Concurrency control is optimistic: a write that was computed from a stale read fails with 409 Conflict instead of silently overwriting someone else's change.
Get this layer right and the rest of Kubernetes stops being magic. Get it wrong and you see the classic incidents: a webhook outage that blocks every deploy, an etcd database that hits its quota, controllers that hammer the API server with full LISTs, or two automations that keep undoing each other.
The API server request path
Flow
- 1
Step 1: HTTPS request arrives with a client cert, bearer token or OIDC token
- nextStep 2: authentication maps the request to a user and groups
- 2
Step 2: authentication maps the request to a user and groups
- nextStep 3: API Priority and Fairness picks a flow and priority level from that identity
- nextFailure path: 401 Unauthorized
- 3
Step 3: API Priority and Fairness picks a flow and priority level from that identity
- nextStep 4: authorization (Node, RBAC, webhook) checks verb, resource, namespace
- 4
Step 4: authorization (Node, RBAC, webhook) checks verb, resource, namespace
- nextStep 5: mutating admission: built-in plugins, MutatingAdmissionPolicy, mutating webhooks in series
- nextFailure path: 403 Forbidden
- 5
Step 5: mutating admission: built-in plugins, MutatingAdmissionPolicy, mutating webhooks in series
- nextStep 6: schema validation against the OpenAPI schema or CRD structural schema
- nextFailure path: every matching create or update fails cluster-wide
- 6
Step 6: schema validation against the OpenAPI schema or CRD structural schema
- nextStep 7: validating admission: built-in plugins, ValidatingAdmissionPolicy, validating webhooks in parallel
- 7
Step 7: validating admission: built-in plugins, ValidatingAdmissionPolicy, validating webhooks in parallel
- nextStep 8: etcd transaction compares the expected resourceVersion and writes
- 8
Step 8: etcd transaction compares the expected resourceVersion and writes
- nextStep 9: watch cache receives the event and fans it out to watchers
- nextFailure path: 409 Conflict, client must re-read and retry
- 9
Step 9: watch cache receives the event and fans it out to watchers
- 10
Failure path: 401 Unauthorized
- 11
Failure path: 403 Forbidden
- 12
Failure path: every matching create or update fails cluster-wide
- 13
Failure path: 409 Conflict, client must re-read and retry
Lesson map
Kubernetes Control Plane - API Server Request Path, etcd, Watches, Informers, Optimistic Concurrency & CRDs
Control plane: API server request pipeline (authn, API Priority and Fairness, RBAC, mutating/validating admission, schema validation), etcd limits and compaction, resourceVersion, watches, 410 Gone and relist, 409 optimistic concurrency, shared informers and workqueues, CRDs and operators vs aggregated APIs.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB r["Step 1: HTTPS request arrives with a client cert, bearer token or OIDC token"] n["Step 2: authentication maps the request to a user and groups"] f["Step 3: API Priority and Fairness picks a flow and priority level from that identity"] z["Step 4: authorization (Node, RBAC, webhook) checks verb, resource, namespace"] ma["Step 5: mutating admission: built-in plugins, MutatingAdmissionPolicy, mutating webhooks in series"] sv["Step 6: schema validation against the OpenAPI schema or CRD structural schema"] va["Step 7: validating admission: built-in plugins, ValidatingAdmissionPolicy, validating webhooks in parallel"] et["Step 8: etcd transaction compares the expected resourceVersion and writes"] wc["Step 9: watch cache receives the event and fans it out to watchers"] e1["Failure path: 401 Unauthorized"] e2["Failure path: 403 Forbidden"] e3["Failure path: every matching create or update fails cluster-wide"] e4["Failure path: 409 Conflict, client must re-read and retry"] r -->|continues| n n -->|continues| f f -->|continues| z z -->|continues| ma ma -->|continues| sv sv -->|continues| va va -->|continues| et et -->|continues| wc n -->|continues| e1 z -->|continues| e2 ma -->|continues| e3 et -->|continues| e4
| Stage | Built-in examples | What you can plug in | Typical failure |
|---|---|---|---|
| Authentication | X.509 client certs, service account tokens, OIDC | Webhook token authenticators | 401 from expired tokens or wrong audience |
| Flow control | API Priority and Fairness (APF), keyed on the authenticated user and request | FlowSchema and PriorityLevelConfiguration objects | A noisy client gets 429s instead of starving everyone |
| Authorization | Node authorizer, RBAC | Authorization webhooks | 403: missing Role or RoleBinding |
| Mutating admission | DefaultStorageClass, ServiceAccount, LimitRanger, Priority | MutatingAdmissionPolicy (CEL), mutating webhooks | Objects that differ from what you sent; webhook timeouts |
| Schema validation | OpenAPI for built-ins, structural schemas for CRDs | CRD schema, CEL validation rules on CRDs | 400 Bad Request; unknown fields warned or rejected |
| Validating admission | ResourceQuota, PodSecurity, NamespaceLifecycle | ValidatingAdmissionPolicy (CEL), validating webhooks, policy engines | Denials, or a down webhook blocking writes |
Facts worth remembering, all from the Kubernetes docs:
- Admission only sees writes. Reads (get, list, watch) bypass admission, so you cannot use admission to hide data; that is RBAC's job.
- Mutating webhooks run in series; validating webhooks run in parallel. Mutations happen before schema validation, so a mutation cannot produce an invalid object that slips through.
- Side effects need reconciliation. An admission plugin cannot know whether a later plugin will reject the request, so anything it changes elsewhere (quota counters, external records) needs a cleanup loop.
- The defaults matter. In Kubernetes 1.37 the default admission plugins include
MutatingAdmissionPolicy,ValidatingAdmissionPolicy, both webhook plugins,PodSecurity,ResourceQuota,DefaultStorageClassandStorageObjectInUseProtection.
In-process CEL policies vs webhooks. A ValidatingAdmissionPolicy runs inside the API server, so it adds no network hop and cannot be "down". A webhook can call external systems and run arbitrary code, but it is a remote dependency on your critical write path. Prefer CEL policies for simple rules, and give webhooks tight timeouts, narrow rules and namespaceSelectors, and a deliberate failurePolicy.
Flow control. APF divides the API server's concurrency (by default --max-requests-inflight 400 plus --max-mutating-requests-inflight 200 are summed into the total when APF is on) into priority levels, and queues requests fairly per flow inside each level. The point is isolation: one controller stuck in a LIST loop gets throttled with 429s while node heartbeats and leader election keep flowing. Clients must honor 429 and Retry-After.
Polling the API or using informers?
Prefer
Shared informers and a work queue
List once, watch, cache locally, and reconcile each key once.
- 1 LIST and 1 open WATCH instead of 200,000 reads.
- 1,000 events collapsed to 20 reconciles.
- Stale cache reads are safe because writes carry a resourceVersion.
Alternative
Polling with LIST
Ask the API server for everything on every loop.
- API load grows with cluster size and loop rate.
- Clients hit 429s from API Priority and Fairness.
- etcd and the watch cache take the pressure.
One write through the control plane
Diagram 1 condensed: the request path and its failure exits.
- 1
Authenticate and classify
Identity first, then API Priority and Fairness picks a priority level. - 2
Authorize and admit
RBAC, mutating admission, schema validation, validating admission. - 3
Compare and store
etcd writes the object if the resourceVersion still matches. - 4
Fan out to watchers
The watch cache streams the change to informers. - 5
Conflict or compaction
A stale write gets 409; a watcher behind compaction gets 410 and relists.
etcd: the store behind it
etcd is a Raft-replicated key-value store with MVCC: every write creates a new revision of the whole keyspace, and old revisions are kept until compacted. The API server maps each object to a key and uses the etcd revision of the last write as the object's resourceVersion. If you know Raft from the Raft series, the consequences follow directly.
| Property | What it means for Kubernetes | Number to know (etcd and Kubernetes docs) |
|---|---|---|
| Raft quorum | Writes need a majority; run an odd number of members (3 or 5) | 3 members tolerate 1 failure, 5 tolerate 2 |
| Every write goes through the leader and is fsynced | Disk latency is API latency | Use fast SSDs; etcd is sensitive to disk and network I/O |
| Request size limit | Large objects fail outright | Default max request 1.5 MiB |
| Storage quota | Writes stop when the backend is full (alarm: database space exceeded) | Default 2 GiB, 8 GiB suggested maximum |
| MVCC history | Watches can resume from an old revision | Kubernetes keeps about 5 minutes of history by default |
| Compaction | Old revisions dropped; resuming a watch from before it fails | --etcd-compaction-interval default 5m in kube-apiserver |
| Defragmentation | Compaction frees space logically; defrag returns it to the filesystem | Run defrag per member, never all at once |
Why etcd is the bottleneck. Every component's writes funnel through one Raft leader, every object lives in one keyspace, and every watcher depends on its revision stream. Kubernetes' published scale envelope (no more than 5,000 nodes, 150,000 pods, 300,000 containers, 110 pods per node) is largely a statement about what the API server and etcd can keep up with. The usual ways teams hurt it: storing megabytes in ConfigMaps or Secrets, creating thousands of short-lived custom resources or Events, and controllers that LIST the world on every loop.
Watches, resourceVersion and optimistic concurrency
The API server keeps an in-memory watch cache per resource type. Watches and most list requests are served from it rather than etcd. In Kubernetes 1.37, even "most recent" consistent reads can be served from the watch cache (with etcd 3.4.31+ or 3.5.13+), using etcd progress notifications to prove the cache is fresh.
| Read | resourceVersion parameter | Consistency | Cost |
|---|---|---|---|
| get or list | unset | Most recent (consistent) | Cache with freshness check, or quorum read |
| get or list | "0" | Any: may be stale | Cheapest, from cache |
| list | N with resourceVersionMatch=NotOlderThan | At least as new as N | From cache |
| list | N with resourceVersionMatch=Exact | Exactly N, or 410 Gone if compacted | May fall back to etcd |
| watch | N | Events strictly after N | 410 Gone if N is older than retained history |
Writes use the same number for optimistic concurrency. A PUT that carries a resourceVersion succeeds only if the object has not changed since; otherwise 409 Conflict. This is compare-and-swap, the same idea as in CAS and optimistic coordination, implemented with an etcd transaction.
"""Simulation: etcd-style revisions, optimistic concurrency and watch history.
Models three documented behaviours:
* every write bumps one global revision; an object's resourceVersion is the
revision of its last write
* an update carrying a stale resourceVersion gets 409 Conflict
* watches can resume only inside the retained history window; after
compaction an old resourceVersion gets 410 Gone and the client must relist
"""
class Conflict(Exception): pass
class Gone(Exception): pass
class ToyEtcd:
def __init__(self):
self.rev = 0
self.kv = {} # key -> (value, mod_revision)
self.history = [] # (revision, key, value)
self.compacted = 0
def put(self, key, value, expect_rv=None):
cur = self.kv.get(key)
if expect_rv is not None and cur and cur[1] != expect_rv:
raise Conflict(f"409 Conflict: {key} is at rv {cur[1]}, you sent {expect_rv}")
self.rev += 1 # a single txn: compare then put
self.kv[key] = (value, self.rev)
self.history.append((self.rev, key, value))
return self.rev
def get(self, key):
return self.kv[key]
def compact(self, upto):
self.history = [h for h in self.history if h[0] > upto]
self.compacted = upto
def watch_from(self, rv):
if rv < self.compacted:
raise Gone(f"410 Gone: rv {rv} older than compacted rv {self.compacted}")
return [h for h in self.history if h[0] > rv]
db = ToyEtcd()
db.put("deploy/checkout", {"replicas": 3, "image": "v1"})
# Two writers read the same version...
hpa_view, hpa_rv = db.get("deploy/checkout")
ci_view, ci_rv = db.get("deploy/checkout")
print("both read rv", hpa_rv)
# ...the HPA scales first
db.put("deploy/checkout", {**hpa_view, "replicas": 6}, expect_rv=hpa_rv)
print("HPA wrote replicas=6 -> rv", db.get("deploy/checkout")[1])
# CI tries a full replace built from its stale read (PUT with resourceVersion)
try:
db.put("deploy/checkout", {**ci_view, "image": "v2"}, expect_rv=ci_rv)
except Conflict as e:
print("CI:", e)
# RetryOnConflict pattern: re-read, re-apply the intent, write again
fresh, rv = db.get("deploy/checkout")
db.put("deploy/checkout", {**fresh, "image": "v2"}, expect_rv=rv)
print("CI retried on rv", rv, "->", db.get("deploy/checkout"))
# Without the precondition, last writer wins and the HPA's scale-up is lost
lost = ToyEtcd(); lost.put("d", {"replicas": 3, "image": "v1"})
stale, _ = lost.get("d")
lost.put("d", {**stale, "replicas": 6})
lost.put("d", {**stale, "image": "v2"}) # blind write from stale copy
print("blind writes result:", lost.get("d")[0], "<- replicas=6 silently lost")
# Watches: resume inside the window, or relist after compaction
for i in range(5):
db.put(f"pod/p{i}", {"phase": "Running"})
print("\ncurrent revision", db.rev)
print("watch from rv 4 ->", [r for r, _, _ in db.watch_from(4)])
db.compact(6) # kube-apiserver asks etcd to compact on an interval (default 5m)
try:
db.watch_from(4)
except Gone as e:
print("watch from rv 4 ->", e)
snapshot_rv = db.rev # relist: read everything at a current revision
print("relist at rv", snapshot_rv, "then watch from there ->", db.watch_from(snapshot_rv))Output:
both read rv 1
HPA wrote replicas=6 -> rv 2
CI: 409 Conflict: deploy/checkout is at rv 2, you sent 1
CI retried on rv 2 -> ({'replicas': 6, 'image': 'v2'}, 3)
blind writes result: {'replicas': 3, 'image': 'v2'} <- replicas=6 silently lost
current revision 8
watch from rv 4 -> [5, 6, 7, 8]
watch from rv 4 -> 410 Gone: rv 4 older than compacted rv 6
relist at rv 8 then watch from there -> []Two lessons from that run. First, the blind write lost the HPA's scale-up with no error at all; the 409 is a feature. Second, a watcher that falls behind the compaction point must relist, which is expensive on big clusters. BOOKMARK events exist so idle watchers can advance their resourceVersion and stay inside the window, and streaming lists (sendInitialEvents=true, beta and enabled by default in current releases) let a client get the initial state as a watch stream instead of one giant LIST response.
Server-side apply: optimistic concurrency per field
Whole-object CAS is coarse: the HPA changing replicas should not conflict with CI changing image. Server-side apply (SSA) tracks a field manager for every field in metadata.managedFields. Applying a field that another manager owns with a different value is a conflict that you either resolve or deliberately --force-conflicts.
| Update method | Conflict unit | Good for | Watch out for |
|---|---|---|---|
| PUT with resourceVersion | Whole object | Controllers that own the whole object | Retries under contention; dropping unknown fields |
| JSON merge patch / strategic merge patch | None by default | Quick edits | Lost updates unless you add a resourceVersion precondition |
JSON patch with test ops | The fields you test | Increment-style changes | Building patches carefully |
| Server-side apply | Per field, per manager | GitOps tools, Helm 4 new installs, multiple actors on one object | Ownership surprises; forcing takes fields away from others |
Informers: how controllers read without melting the API server
A controller built with client-go does not poll. It uses an informer: a Reflector does one LIST and then one long-running WATCH, pushes changes through a DeltaFIFO into a thread-safe Indexer (the local cache), and calls event handlers. Handlers do not do the work; they put the object's key (namespace/name) on a rate-limited work queue. Workers pop keys, read the latest object from the cache, and reconcile. Because the queue holds each key at most once, ten updates to one object while it waits produce one reconcile, against the newest state.
// Simulation: why controllers use informers (list+watch into a local cache)
// and a de-duplicating work queue instead of polling the API server.
type Pod = { name: string; rv: number; phase: string };
let apiReads = 0; // requests that hit the API server
const server = new Map<string, Pod>();
for (let i = 0; i < 200; i++) server.set(`p${i}`, { name: `p${i}`, rv: i + 1, phase: "Running" });
// Generate a burst: 1000 status updates spread over just 20 pods (hot keys)
const events: Pod[] = [];
let rv = 1000;
for (let i = 0; i < 1000; i++) {
const name = `p${i % 20}`;
const pod = { name, rv: ++rv, phase: i % 7 === 0 ? "Pending" : "Running" };
server.set(name, pod);
events.push(pod);
}
// Approach A: naive controller does a GET of every pod for every event it hears about
function naive(): number {
apiReads = 0;
for (const _ev of events) { for (const _k of server.keys()) apiReads++; } // "list everything to be safe"
return apiReads;
}
// Approach B: informer = 1 LIST + 1 WATCH stream feeding a cache; handlers enqueue KEYS
class WorkQueue {
private queue: string[] = [];
private dirty = new Set<string>(); // a key waits in the queue at most once
add(k: string) { if (!this.dirty.has(k)) { this.dirty.add(k); this.queue.push(k); } }
get(): string | undefined { const k = this.queue.shift(); if (k) this.dirty.delete(k); return k; }
get len() { return this.queue.length; }
}
function informer(): { reads: number; reconciles: number; staleSeen: number } {
apiReads = 1; // initial LIST (the watch is one long request)
const cache = new Map<string, Pod>(); // the "indexer"
const q = new WorkQueue();
// the burst arrives faster than workers drain the queue, so keys coalesce.
// The last 3 watch events are still in flight when the worker runs: the cache lags.
const delivered = events.slice(0, events.length - 3);
for (const ev of delivered) { cache.set(ev.name, ev); q.add(ev.name); }
let reconciles = 0, staleSeen = 0;
for (let k = q.get(); k; k = q.get()) {
const latest = cache.get(k)!; // read from cache, not the server
if (latest.rv !== server.get(k)!.rv) staleSeen++;
reconciles++; // level-triggered: act on current state of k
}
return { reads: apiReads, reconciles, staleSeen };
}
console.log("events in burst:", events.length, "on", new Set(events.map((e) => e.name)).size, "distinct pods");
console.log("naive polling API reads:", naive().toLocaleString("en-US"));
const r = informer();
console.log("informer API reads:", r.reads, "(1 LIST, then 1 open WATCH)");
console.log("reconciles after queue de-dup:", r.reconciles);
console.log("reconciles that saw a stale cache entry:", r.staleSeen, "(writes must carry resourceVersion; a 409 means requeue)");Output:
events in burst: 1000 on 20 distinct pods
naive polling API reads: 200,000
informer API reads: 1 (1 LIST, then 1 open WATCH)
reconciles after queue de-dup: 20
reconciles that saw a stale cache entry: 3 (writes must carry resourceVersion; a 409 means requeue)Expectedevents in burst: 1000 on 20 distinct pods naive polling API reads: 200,000 informer API reads: 1 (1 LIST, then 1 open WATCH) reconciles after queue de-dup: 20 reconciles that saw a stale cache entry: 3 (writes must carry resourceVersion; a 409 means requeue)
Press Run. Snippets must be self-contained — no network, files, or native modules.
The stale-cache line is the important caveat: informer caches lag the server slightly, so a reconcile can act on old data. That is safe only because writes carry resourceVersions (or use SSA): a stale write fails with a conflict and the key is requeued. Shared informer factories let many controllers in one process share a single watch per resource type.
CRDs and operators: extending the API
| Extension | What it is | Storage | When to use |
|---|---|---|---|
| CustomResourceDefinition (CRD) | Declare a new type; the API server serves it with the same machinery (RBAC, admission, watches, SSA) | etcd, like built-ins | Almost always: declarative config for your domain |
| Aggregated API server | Your own API server registered behind kube-apiserver | Anything you choose | Non-declarative APIs, custom storage, very high churn (metrics-server is the classic example) |
| Operator | A CRD plus a controller that encodes operational knowledge (create backups, fail over, upgrade) | The CRD in etcd plus real resources | Stateful software whose day-2 operations are automatable |
A good CRD has a structural schema with validation (OpenAPI plus CEL rules), a status subresource so controllers write status without touching spec, and versioning with conversion if the schema changes. A good operator is level-triggered, idempotent, uses finalizers for external cleanup, writes conditions to status, and never assumes it saw every event.
Decision chart: extending and talking to the control plane.
Decisions
- 1
A
- nextDirect GET or LIST, paginated with limit and continue
- nextStep 2: written in Go or a language with a mature client library?
- 2
Direct GET or LIST, paginated with limit and continue
- ?
Step 2: written in Go or a language with a mature client library?
- nextShared informer plus workqueue: one LIST then WATCH, local cache, retries with backoff
- nextRaw WATCH with resourceVersion and BOOKMARKs; relist on 410 Gone
- 4
Shared informer plus workqueue: one LIST then WATCH, local cache, retries with backoff
- nextStep 3: do you need new object types?
- 5
Raw WATCH with resourceVersion and BOOKMARKs; relist on 410 Gone
- nextStep 3: do you need new object types?
- ?
Step 3: do you need new object types?
- nextUse built-in kinds; write with server-side apply and a field manager
- nextCRD plus operator: schema validation, status subresource, finalizers
- nextKeep it out of etcd: own database or an aggregated API server
- 7
Use built-in kinds; write with server-side apply and a field manager
- nextStep 4: writes conflict with other actors?
- 8
CRD plus operator: schema validation, status subresource, finalizers
- nextStep 4: writes conflict with other actors?
- 9
Keep it out of etcd: own database or an aggregated API server
- ?
Step 4: writes conflict with other actors?
- nextRe-read and retry, or switch to server-side apply to own only your fields
- nextDone: watch metrics for list and watch load
- 11
Re-read and retry, or switch to server-side apply to own only your fields
- 12
Done: watch metrics for list and watch load
What happens if you choose an alternative
| Choice | Instead of | Consequence |
|---|---|---|
| Webhook for a simple rule | ValidatingAdmissionPolicy (CEL) | Extra latency and a new failure mode on every write |
failurePolicy: Ignore | Fail | Writes keep working during webhook outages, but policy is silently skipped |
| Polling with LIST | Informers | Linear growth in API load with cluster size; 429s; etcd pressure |
| Blind patches | resourceVersion preconditions or SSA | Lost updates between automations, with no error |
| Storing big blobs in ConfigMaps | Object storage plus a reference | etcd quota and latency problems for the whole cluster |
| One huge cluster | Several smaller clusters | Simpler fleet but a bigger blast radius and etcd scale ceiling |
Pitfalls
- Webhooks that match everything, including
kube-systemand their own namespace. When the webhook pod dies, it cannot be recreated because its own creation is blocked. Exclude system namespaces. - Watch history assumptions. Code that resumes a watch from an hour-old resourceVersion must handle
410 Goneby relisting. - Comparing resourceVersions across resource types. They are only orderable within one resource type, as decimal integers; treat them as opaque otherwise.
- etcd backups that were never restored. A snapshot you have not restored is a hope, not a backup; see Backups That Actually Restore.
- RBAC sprawl. Wildcard verbs on
*resources for a CI service account turn a leaked token into cluster admin; see RBAC and role explosion.
Interview Q&A
In what order does the API server process a write?
Answer
Authentication, then API Priority and Fairness classifies and queues the request using that identity, then authorization, mutating admission, schema validation, validating admission, and finally persistence to etcd, after which the watch cache fans the event out.
What is a resourceVersion and what is it used for?
Answer
An opaque string, in practice the etcd revision of the object's last modification. Clients use it to resume watches without missing events, to request read consistency levels, and as a precondition for optimistic concurrency so stale writes fail with 409.
A controller gets `410 Gone` on its watch. What happened and what should it do?
Answer
The resourceVersion it tried to resume from is older than the retained history (it was compacted). It must relist to rebuild its cache and start a new watch from the list's resourceVersion. Bookmarks reduce how often this happens.
Why is etcd often the scaling bottleneck?
Answer
All writes go through a single Raft leader and must be fsynced on a quorum, all objects share one keyspace with a size quota (2 GiB default, 8 GiB suggested max), and every watcher depends on its event stream. Large objects, high churn and heavy LISTs all land there.
Webhook or ValidatingAdmissionPolicy?
Answer
CEL-based ValidatingAdmissionPolicy for rules that can be expressed on the object itself: no network hop and no availability risk. Webhooks when you need external data or complex logic, with tight timeouts and scoped selectors.
Why do controllers use a work queue of keys instead of processing each event?
Answer
Coalescing and correctness. Many events for one object collapse into one reconcile of the newest state, retries get rate-limited backoff, and level-triggered logic means a missed event does not matter.
What does server-side apply add on top of resourceVersion checks?
Answer
Field-level ownership. Each field records its manager in managedFields, so two actors conflict only when they set the same field to different values.
What do BOOKMARK events do?
Answer
They let idle watchers advance their resourceVersion so they stay inside the retained history and avoid an expensive relist.
Check yourself
Run kubectl get deploy NAME -o yaml twice while scaling the Deployment in between. Note how metadata.resourceVersion changes, then try kubectl replace with the old file and read the 409.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Raft Consensus — Leader Election, Log Replication & Safety, Quorums & Majority — Why 2f+1, Read Quorums & Stale Reads, ZooKeeper / etcd Lock Recipes — Ephemeral Nodes, Sessions & Watches, Compare-And-Swap & Optimistic Coordination — When You Do Not Need a Lock, Load Shedding & Admission Control — Drop Early, Protect the Core, RBAC — Roles, Permissions & Role Explosion, Authorization — RBAC, ABAC, ReBAC & Policy Engines, Control Plane & Discovery — xDS-style Config, Endpoints & Convergence, MVCC, Snapshot Isolation & Write Skew.