Networking
Part 5 of 6 · Load balancingGlobal Load Balancing — DNS, Anycast & Multi-Region Failover
DNS geo/latency, Anycast, health-aware DNS; active-active vs active-passive; split-brain; RTO vs RPO confusion.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Steering clients across regions
Prefer
Health-aware steering + an anycast or short-TTL DNS plan
Regional health is the input. How fast caches forget the dead region is the real RTO. Pair with retries and idempotent APIs.
- Anycast / Global Accelerator: sub-second BGP shift for a stable VIP.
- DNS failover is bounded by resolver TTLs you do not fully control.
- Active-active needs a write story or you split-brain checkouts.
Alternative
One regional ALB and hope
Multi-AZ is not multi-region. Geo-DNS without health pins users to a smoking crater until TTL.
- Corporate resolvers live in another country — geo is a lie.
- Sticky sessions trap users in the sick region. Depth: sticky lesson.
- RTO marketing that quotes DNS TTL and ignores data RPO.
Overview
Move traffic across regions and PoPs with acceptable RTO, avoid split-brain writes, and explain DNS geo/latency routing vs Anycast vs health-aware DNS (Route 53, Cloud DNS) at interview depth.
A regional ALB/NLB survives AZ loss (if multi-AZ). It does not survive a region outage or continent-scale RTT needs. GLB answers: which region/PoP should this client use?
Mechanisms compared
| Mechanism | Cutover speed | Health | Footgun |
|---|---|---|---|
| DNS geo / latency | TTL-bound | Optional | Resolver ≠ user; EDNS Client Subnet nuances |
| Health-aware DNS | Still TTL-bound | Yes | Cached answers are a hidden RTO gap |
| Anycast | BGP (fast) | Route withdraw | Mid-flow TCP may break; clients reconnect |
| Client-side steering | Precise | App-defined | Every client implements policy |
| Mesh / Traffic Director | Strong internals | Yes | Heavier ops |
DNS geo and latency routing
Geo DNS maps country → region. Latency-based uses resolver-to-region RTT.
Pitfalls: TTL tradeoff (failover speed vs DNS QPS); EDNS Client Subnet; corporate resolvers in another country.
Route 53 policies worth naming: simple, failover, geolocation, geoproximity, latency, weighted, multi-value.
Multi-value DNS is multiple IPs with crude health — not a serious LB.
Anycast
Multiple PoPs advertise the same IP via BGP. On PoP failure, routes withdraw and traffic shifts. TCP caveat: a mid-connection path move may lose state — clients reconnect. Great for DNS, CDN HTTP, DDoS absorption. Edge handshake version: TLS termination and Anycast.
Maglev-style packet LBs often sit behind anycast VIPs. The table is local; BGP is how you get there. Paper: Maglev. Do not re-teach vnode rings.
Multi-region topologies
Users hit DNS or an anycast GLB. Regions replicate asynchronously. Failover is a control-plane change, not a local RR tweak.
Flow
- 1
1 Users
- next2 DNS / Anycast GLB
- 2
2 DNS / Anycast GLB
- steer A3a Region A
- steer B3b Region B
- 3
3a Region A
- 4
3b Region B
Replication is the dashed data path between regions — keep it off the steering chart so Region A and B stay siblings. A write in A is not a GLB hop.
Diagrams - step by step
Three small diagrams. Step numbers in the labels give the animation order. The lesson map under Diagram 1 plays those steps.
Diagram 1 - Happy path: steer users to a healthy region
Decisions
- 1
Step 1 User resolves or connects
- nextStep 2 Steering mechanism?
- ?
Step 2 Steering mechanism?
- geo or latency DNSStep 3a DNS answers with the nearest healthy region
- anycastStep 3b BGP routes to the nearest PoP
- 3
Step 3a DNS answers with the nearest healthy region
- nextStep 4 Regional LB serves the request
- region failsFailure path - cached answers keep the old region until TTL
- 4
Step 3b BGP routes to the nearest PoP
- nextStep 4 Regional LB serves the request
- 5
Step 4 Regional LB serves the request
- nextStep 5 Data layer replicates across regions
- 6
Step 5 Data layer replicates across regions
- 7
Failure path - cached answers keep the old region until TTL
The user resolves or connects. Geo or latency DNS answers with the nearest healthy region, or anycast lets BGP route to the nearest point of presence. The regional load balancer serves the request, and the data layer replicates across regions. If the region fails, cached DNS answers keep the old region until the TTL expires.
Lesson map
Global Load Balancing — DNS, Anycast & Multi-Region Failover
Diagram 1 walks 6 steps from Step 1 User resolves or connects through Step 5 Data layer replicates across regions.
Architecture. Step 1 User resolves or connects Ready. Step 2 Steering mechanism? Ready. Step 3a DNS answers with the nearest healthy region Ready. Step 3b BGP routes to the nearest PoP Ready. Step 4 Regional LB serves the request Ready. Step 5 Data layer replicates across regions Ready. Failure path - cached answers keep the old region until TTL Ready
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB A["Step 1 User resolves or connects Ready"] B["Step 2 Steering mechanism? Ready"] C["Step 3a DNS answers with the nearest healthy region Ready"] D["Step 3b BGP routes to the nearest PoP Ready"] E["Step 4 Regional LB serves the request Ready"] F["Step 5 Data layer replicates across regions Ready"] X["Failure path - cached answers keep the old region until TTL Ready"] A -->|continues| B B -->|geo or latency DNS| C B -->|anycast| D C -->|continues| E D -->|continues| E E -->|continues| F C -->|region fails| X
Diagram 2 - Failure path: split-brain in active-active
Sequence
- 1
Region A → Region B
Step 1 replication link breaks
- 2
User on laptop → Region A
Step 2 checkout writes order 1
- 3
Same user on phone → Region B
Step 3 retry lands on B and writes order 1 again
- 4
Region A
Step 4 both actives accepted diverging writes - double charge
- 5
Region A
Fix - single writer per key, idempotency keys, or consensus
The replication link between the regions breaks. The laptop checks out in region A, and the same user's phone retry lands in region B and writes the order again. Both actives accepted diverging writes, so the customer is charged twice. A single writer per key, idempotency keys, or consensus stop that.
Diagram 3 - Decision: GLB mechanism and topology
Decisions
- ?
Step 1 Cutover speed needed?
- sub-secondAnycast or Global Accelerator style
- minutes are OKStep 2 Who runs the regions?
- 2
Anycast or Global Accelerator style
- ?
Step 2 Who runs the regions?
- managed edgeManaged anycast HTTP GLB
- BYO VMsRoute 53 latency plus regional ALBs
- 4
Managed anycast HTTP GLB
- 5
Route 53 latency plus regional ALBs
- nextStep 3 Write model?
- ?
Step 3 Write model?
- strict single regionActive-passive plus DNS failover
- read heavyActive-active reads, regional writes
- 7
Active-passive plus DNS failover
- 8
Active-active reads, regional writes
- Wrong pick - writes everywhere with no conflict planSplit-brain double writes
- 9
Split-brain double writes
Sub-second cutover wants anycast or a Global Accelerator style entry. If minutes are acceptable, a managed edge can use a managed anycast HTTP global load balancer, and bring-your-own VMs can use Route 53 latency routing plus regional ALBs. Strict single-region writes stay active-passive with DNS failover. Read-heavy systems can serve active-active reads with regional writes. Writes in every region with no conflict plan is split-brain.
Active-active vs active-passive
- Active-passive: clear primary; wasted capacity; longer warm-up. Simpler write story.
- Active-active: use all capacity; conflict resolution; sticky users; replication lag.
Split-brain: two regions accept diverging writes — double charge on checkout. Mitigate with single-writer, CRDTs, or consensus. Sticky affinity fights GLB by pinning users to a sick region. Depth: sticky vs stateless.
RTO vs RPO: GLB controls traffic RTO. Databases control data RPO. Separate them in interviews. Pre-warm standby regions or your RTO includes cache/JIT/pool ramp — slow start.
Health-aware DNS failover sequence
- Regional health check fails (use multi-probe so one path blip does not declare a region down).
- DNS provider marks the endpoint down.
- New answers omit or deprioritize the region.
- Cached answers continue until TTL — this is the hidden RTO gap.
Pair with app retries, short TTLs, or anycast HTTP GLB (stable VIP). Game-day: block region health and measure dig TTL behavior.
Decision guide
- Lowest ops complexity → managed anycast HTTP GLB
- BYO regions on VMs → Route 53 latency + regional ALBs
- Sub-second cutover → Anycast / Global Accelerator style
- Strict single-region writes → active-passive + DNS failover
- Read-heavy global → active-active reads, regional writes
Global Accelerator is anycast entry + backbone to regional stacks — not a data design.
Sandbox: DNS TTL lag (Python)
Authoritative flips to region-b at t=5. The resolver still answers region-a until expiry.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Sandbox: region picker (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Region A dies at T=0. DNS TTL is 60s. Anycast withdraws in 2s. Postgres async replica lag is 15s. What is traffic RTO on DNS vs anycast? What is RPO? Who owns each number?
Interview Q&A
Why is DNS failover slow?
Answer
Resolver TTLs and caches. Your 60s TTL is a lower bound, not a guarantee. Recursive resolvers and stub caches add delay.
Anycast vs DNS geo?
Answer
Anycast shifts with BGP on one VIP. DNS geo depends on TTL and resolver location, which is not the user.
RTO vs RPO?
Answer
RTO: how long until traffic hits a live region. RPO: how much data you may lose. GLB owns the first; replication owns the second.
Split-brain example?
Answer
Two actives accept checkout writes → double charge, divergent inventory. Single-writer or consensus, not "just GLB."
What does Global Accelerator actually do?
Answer
Anycast entry plus AWS backbone to regional stacks. It is not a database and not a conflict resolver.
Multi-value DNS?
Answer
Multiple IPs in one answer; crude client-side pick; limited health. Do not call it a load balancer in a senior interview.
Why multi-probe health for a region?
Answer
Avoid declaring a region down because one probe path failed (a single AZ, a single DNS checker, a single ISP).