Networking
Part 4 of 6 · Sidecar & Service MeshControl Plane & Discovery — xDS-style Config, Endpoints & Convergence
Sidecars are useless with stale config. The control plane discovers endpoints, compiles routes and listeners, and pushes them until the fleet converges. Envoy xDS is the best-known API shape — the ideas apply to any mesh.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
How proxies learn backends
Prefer
Versioned snapshots, ACK/NACK, ordered resources
The control plane watches EndpointSlices or Cloud Map, compiles listeners/routes/clusters/endpoints, and streams them. Proxies ACK applied config or NACK a bad push. ADS exists so a route never races a missing cluster.
- EDS churns as pods scale; CDS is the slower policy object.
- Last-good config keeps availability when the control plane is down.
- SDS outage is different: new connections may fail-closed without certs.
Alternative
DNS TTLs, or independent streams that race
DNS is coarse discovery. Caching fights rapid scale. Split LDS/RDS/CDS/EDS streams without ordering are why deploys 503.
- No ACK means the control plane cannot tell apply from reject.
- Flapping endpoints without debounce storm the fleet.
- Two control planes after a partition = split brain. Define stale SLA.
Convergence loop
Interviews start at the 503 after deploy, not at the acronym soup.
- 1
Watch sources
EndpointSlices, Cloud Map, Consul, or a static bootstrap to reach the control plane. - 2
Compile snapshots
Listeners, routes, clusters, endpoints, secrets. Version them. - 3
Push or stream
ADS-style ordered delivery, poll, or render files and reload. - 4
ACK or NACK
Proxy applied it, or rejected it. Do not thrash. - 5
Warmup weight
New endpoints get partial weight until healthy. Depth: health checks.
Overview
Sidecars are useless with stale config. The control plane discovers endpoints, compiles routes and listeners, and pushes them to data-plane proxies until the fleet converges. Envoy's xDS (CDS/EDS/LDS/RDS, often multiplexed as ADS) is the best-known API shape — but the ideas apply to any mesh: Linkerd's destination controller, nginx+Consul templates, AWS App Mesh, VPC Lattice associations.
Interview for: discovery sources, push vs pull, consistency, and failure under partition.
Abstract API families (xDS-shaped thinking)
| Family | Owns | Churn |
|---|---|---|
| LDS-like | Listeners: ports and protocols the proxy opens | Low |
| RDS-like | Routes: match → cluster / weighted dest | Medium |
| CDS-like | Clusters: logical upstreams, LB policy, outlier | Medium |
| EDS-like | Endpoints: concrete IPs/ports | High |
| SDS-like | Secrets: certs/keys for mTLS | Medium, high-stakes |
| ADS | One aggregated stream, ordered resources | Operational glue |
Is xDS Envoy-only? It originated for Envoy. Typed discovery resources are wider. Other meshes have analogues.
Discovery inputs
- Kubernetes Endpoints / EndpointSlice / Gateway API
- Cloud Map / Consul / Eureka / custom registries
- DNS (limited — TTLs fight rapid churn)
- Static bootstrap for the proxy's first hop to the control plane
Managed discovery plus policy association is the same job with a different API: App Mesh / Lattice.
Resource order
Routes reference clusters. Clusters need endpoints. Listeners reference routes. Independent streams race. ADS exists because a route pointing at a missing cluster is a 503 after deploy.
Decisions
- 1
1 EDS endpoints
- next2 CDS clusters
- 2
2 CDS clusters
- next3 RDS routes
- 3
3 RDS routes
- next4 LDS listeners
- 4
4 LDS listeners
- next5 Proxy ACK?
- ?
5 Proxy ACK?
- yes6 Converged
- no7 NACK revert alert
- 6
6 Converged
- 7
7 NACK revert alert
Lesson map
Control Plane & Discovery — xDS-style Config, Endpoints & Convergence
Sidecars are useless with stale config. The control plane discovers endpoints, compiles routes and listeners, and pushes them until the fleet converges. Envoy xDS is the best-known API shape — the ideas apply to any mesh.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB eds["1 EDS endpoints"] cds["2 CDS clusters"] rds["3 RDS routes"] lds["4 LDS listeners"] eds -->|1 EDS endpoints to 2 CDS clusters| cds cds -->|2 CDS clusters to 3 RDS routes| rds rds -->|3 RDS routes to 4 LDS listeners| lds
Convergence and consistency
- Eventual consistency is normal: a new pod may receive traffic before all sidecars learn it; an old IP remains until EDS deletes it.
- Versioning / ACK-NACK: the proxy confirms apply or rejects bad config. The control plane must not thrash.
- Warmup: delay full weight until health passes — health checks, slow start.
- Split brain: two control planes or cached snapshots after disconnect. Define stale-serve vs fail-closed.
Decisions
- 1
1 EndpointSlices
- next2 Config compiler
- 2
2 Config compiler
- next3 Versioned snapshot
- 3
3 Versioned snapshot
- next4 Push proxy A
- 4
4 Push proxy A
- next5 Then proxy B
- 5
5 Then proxy B
- next6 All ACK?
- ?
6 All ACK?
- yes7 Converged
- no7 Lagging proxies
- 7
7 Converged
- 8
7 Lagging proxies
Push vs pull vs template
| Mode | Wins | Loses |
|---|---|---|
| Streaming xDS / gRPC | Near-real-time | Connection health, gRPC in the path |
| Poll / long-poll | Simpler firewall story | Higher lag |
| Render files + reload | Operable, nginx-familiar | Reload cost, thundering herds |
Failure modes to name
- Control plane down — proxies should keep last-good config (stale-serve) for availability.
- Identity / SDS outage — new connections may fail-closed without certs.
- Endpoint flapping — churn storms; debounce and hysteresis with health/outlier.
- Over-fanout watches — control plane CPU; shard by namespace.
Sandbox: EDS convergence (Python)
Proxy b lags a push. The fleet is not converged until the second push.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same idea: activation order (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
New ReplicaSet. EndpointSlice adds a pod IP. Proxy A ACKs EDS+CDS+RDS. Proxy B still has the old snapshot. Who 503s, and for how long? Now the control plane dies. Do you stale-serve or fail-closed? Does SDS change that answer?
Interview Q&A
EDS vs CDS?
Answer
CDS is the logical cluster plus policy. EDS is the member IPs. EDS changes constantly as pods scale.
Why ADS?
Answer
One stream, ordered delivery, avoids route-to-missing-cluster races after deploy.
ACK / NACK?
Answer
The proxy confirms it applied config or rejects a bad push so the control plane can revert and alert instead of thrashing.
DNS instead of EDS?
Answer
Works for coarse discovery. TTL and caching fight rapid scale events.
Control plane partition?
Answer
Proxies usually serve last-good config. Define an SLA for how stale is acceptable. SDS/certs are the exception that may fail-closed.
Warmup?
Answer
New endpoints get partial weight until healthy. Ties to health checks and slow start.
Is xDS Envoy-only?
Answer
It originated for Envoy. The idea of typed discovery resources is wider. Other meshes have analogues.
Cloud Map / Lattice?
Answer
Managed discovery plus policy association — same job, different API. Depth: cloud equivalents.
What consumes this config?
Answer
The data plane: Envoy, nginx, Linkerd-proxy.
Endpoint flapping?
Answer
Debounce and hysteresis. Pair with outlier detection and health thresholds so one bad probe does not churn EDS for the fleet.
Watch fanout?
Answer
A control plane that watches every namespace from every proxy will melt. Shard by namespace or use a scalable watch cache.
Bootstrap?
Answer
The proxy still needs a static first hop to find the control plane. That file is not the fleet's EDS.