Failover with DNS.
When Larkbook's Virginia site dies, nothing in DNS notices by itself. Health checkers in several regions have to agree it is down, the nameserver has to swap the answer, and every resolver still holding the old address has to let it expire. DNS failover works well for minutes-scale recovery of stateless front doors, and badly for anything that needs seconds.
Builds on Caching and TTLs.
The idea.
A name can point somewhere else when a site dies. The hard part is how long everyone keeps using the old answer.
What it does and what it doesn't
Larkbook's GeoDNS policy sends shoppers in the Americas to Virginia (192.0.2.30); Frankfurt (198.51.100.20) already serves Europe and is sized to take them too. If Virginia's front door goes dark, clients keep resolving the name to 192.0.2.30 and their connections time out. DNS has no idea: a nameserver hands out whatever records it holds.
DNS failover fixes that by attaching a health check to each answer. Probes from several regions hit each site's /healthz. When Virginia fails enough of them, the nameserver starts answering with Frankfurt instead. Nothing on the client changes; the next time it looks the name up, it gets the new address.
- Good at: a whole site or region going dark, in front of stateless web or API traffic. It is cheap, works for any protocol that starts with a name lookup, and spans clouds and providers.
- Not for: one bad server inside a pool (the load balancer's health checks handle that in seconds), anything that needs sub-second recovery, or moving a database primary (that needs fencing, not a new address).
The catch is time. The timeline below is Larkbook's worst case with a 10 s check interval, 3 failures to flip and a 60 s TTL: almost two minutes before the last well-behaved client moves.
Where a two-minute failover goes
Scenario 1 of 2: As described.
- 4 Everyone at first; a shrinking share after t=52.
- 11 An answer fetched just before t=52 lives until t=112.
Timeline as a list
Where a two-minute failover goes: 5 lanes, from 0 s to 140 s.
- 0–140 s · Virginia site · front door down
- 0–42 s · Health checkers · detecting
- 0–42 s · DNSCo nameservers · still answers Virginia
- 0–112 s · Clients · some connections fail — Everyone at first; a shrinking share after t=52.
- 0 s · Virginia site · Virginia goes dark (error)
- 14 s · Health checkers · fail 1 (error)
- 28 s · Health checkers · fail 2 (error)
- 42–52 s · DNSCo nameservers · publishing new status
- 42 s · Health checkers · fail 3 (error)
- 52–140 s · DNSCo nameservers · answers Frankfurt
- 52–112 s · Resolver caches · old answers drain (TTL 60) — An answer fetched just before t=52 lives until t=112.
- 112 s · Clients · last TTL-respecting client moves (ok, ok)
- 112–140 s · Clients · window: stragglers that ignore TTL
How it works.
Checkers vote, an aggregator decides, the nameserver answers from the decision, and caches hold the old answer until it expires.
Health-checked DNS for shop.example.com
One checker's vote on Virginia
States of5Health checkers
Route 53 uses one threshold in both directions: 3 results in a row to go down and 3 to come back. The aggregator counts these votes; Virginia leaves the answers once 18% or fewer of them say healthy.
| From → To | Event | Guard | Action |
|---|---|---|---|
| Healthy → Healthy, 1–2 failures seen | probe fails | count = 1 | |
| Healthy, 1–2 failures seen → Healthy, 1–2 failures seen | probe fails | count + 1 < 3 | count += 1 |
| Healthy, 1–2 failures seen → Unhealthy | probe fails | count + 1 = 3 | vote unhealthy |
| Healthy, 1–2 failures seen → Healthy | probe passes | reset count | |
| Unhealthy → Unhealthy, 1–2 passes seen | probe passes | count = 1 | |
| Unhealthy, 1–2 passes seen → Unhealthy, 1–2 passes seen | probe passes | count + 1 < 3 | count += 1 |
| Unhealthy, 1–2 passes seen → Healthy | probe passes | count + 1 = 3 | vote healthy |
| Unhealthy, 1–2 passes seen → Unhealthy | probe fails | reset count |
- Healthystart
- Healthy, 1–2 failures seen
- A single blip does not change the vote
- Unhealthyerror
A client caught mid-failover
- Note over Virginia site (primary): t=0 s: front door down
- Health checkers → Virginia site (primary): t=10 s: GET /healthz, 4 s timeout
- Health checkers → Virginia site (primary): t=24 s: fail 2
- Health checkers → Virginia site (primary): t=38 s: fail 3 in a row
- Health checkers → Health aggregator: all 5 regions: unhealthy
- Health aggregator → Authoritative nameservers: primary down, live by t=52 s
- Clients → Recursive resolvers: A? shop (t=55 s)
- Recursive resolvers → Clients (reply): 192.0.2.30 (55 s left)
- Clients → Virginia site (primary): connect: times out
- Note over Recursive resolvers: t=110 s: copy expires
- Clients → Recursive resolvers: A? shop (retry)
- Recursive resolvers → Authoritative nameservers: A? shop
- Authoritative nameservers → Recursive resolvers (reply): 198.51.100.20, TTL 60
- Recursive resolvers → Clients (reply): 198.51.100.20
- Clients → Frankfurt site (secondary): connect: OK
Ways to wire it
| Pattern | What DNS returns | Good for | Cost |
|---|---|---|---|
| Active–passive | The primary while it is healthy, else the secondary | Disaster-recovery sites and one-region products | The passive site idles, must stay warm and sized, and is only proven when you fail over |
| Active–active | Every healthy site, or the nearest or weighted pick among the healthy ones | Normal multi-region serving | Every site must absorb a failed peer's share on top of its own |
| Multi-address answers | Several healthy addresses at once; clients try the next on failure | Clients you control, or browsers that race addresses | Instant only for clients that actually retry; many libraries take the first address and stop |
| Calculated checks | A site counts as healthy only while at least x of its y child checks are | Sites with many front-end nodes; one bad node should not fail the whole site | More checks to configure and pay for; the x must be tuned |
The rules that surprise people
Managed DNS services lean hard towards always answering. In Route 53, which the others resemble:
- A record with no health check is always healthy. Forget to attach one and that site is never pulled.
- If every record in a group is unhealthy, all of them are treated as healthy and the normal policy (weights, latency) picks one.
- In a failover pair, if the primary and the secondary both fail their checks, the primary is returned. If the secondary has no check, it is returned whenever the primary fails, even if it is also down.
- Checks run on their own schedule. A query never triggers a probe; it reads the last status, so detection never speeds up because traffic is heavy.
The reasoning: an empty answer is a certain outage for everyone, while a possibly dead address is only a likely one. If the checkers themselves break, or a shared dependency makes every site fail its check, answering nothing would turn a monitoring fault into the whole domain disappearing.
A health endpoint worth failing over on
# Larkbook's /healthz: "can this site serve a shopper right now?"
# Check only what is LOCAL to the site. A shared global dependency
# (payments provider, central auth) would fail every site at once,
# and fail-open DNS would then send traffic everywhere anyway.
def healthz(request):
reasons = []
if draining_flag_set(): # operator is emptying the site
reasons.append("draining")
if not replica_ping(timeout_s=0.2): # this site's read replica
reasons.append("catalog replica unreachable")
if disk_free_fraction("/var/cache") < 0.05:
reasons.append("cache disk nearly full")
if healthy_app_servers() < 3: # of 12; below this we can't carry peak
reasons.append("too few app servers")
if reasons:
return Response(503, body="unhealthy: " + ", ".join(reasons))
return Response(200, body="ok larkbook") # string match looks for "ok larkbook"
In practice.
Put numbers on each phase, see which one dominates, and learn what goes wrong in real failovers.
How long Larkbook's users are stuck
- Check interval
- 10 sRoute 53 fast interval; measured from one result to the next probe
- Failures in a row to flip
- 3
- Probe timeout
- 4 sHTTP connect limit; a dead site times out rather than refusing
- Status reaches every nameserver
- ~10 sassumption; providers do not publish a figure
- Record TTL
- 60 s
- Whole-site failures per year
- 4assumption
- Detection, worst caseinterval (wait for first probe) + threshold × timeout + (threshold − 1) × interval = 10 + 3 × 4 + 2 × 1042 sfrom Check interval, Failures in a row to flip and Probe timeout
- Worst case for a TTL-respecting clientdetect + publish + ttl = 42 + 10 + 60112 s ≈ 2 minfrom Detection, worst case, Status reaches every nameserver and Record TTL
- Average(5 + 12 + 20) + 10 + 60 ÷ 2 = 37 + 10 + 3077 sfrom Check interval, Failures in a row to flip, Probe timeout, Status reaches every nameserver and Record TTL · Failure lands mid-interval on average; a cached answer has half its TTL left on average.
- Worst case with TTL 30042 + 10 + 300352 s ≈ 6 minfrom Detection, worst case and Status reaches every nameserver
- Failover lag per year, worst case4 × 112 s = 448 s ≈ 7.5 min; 7.5 ÷ 525,600 min (at the 77 s average: 308 s ≈ 5.1 min ≈ 0.001%)≈ 0.0014% (99.9986% for this cause alone)from Whole-site failures per year and Worst case for a TTL-respecting client
- Past detection, the TTL is the budget. 60 s is the usual floor worth paying for; below that, resolver load rises and some resolvers clamp very small TTLs up to a floor (see dns/caching-and-ttl).
- Anything that must recover in seconds needs another layer, such as anycast or a global L7 balancer (load-balancers/global-load-balancing).
- Clients that try the next address or re-resolve after a connect error beat any TTL.
Share of new connections still aimed at Virginia
- TTL 60 s
- TTL 300 s
Data
| Time since failure (s) | TTL 60 s (%) | TTL 300 s (%) |
|---|---|---|
| 0 | 100 | 100 |
| 52 | 100 | 100 |
| 112 | 3 | no value |
| 352 | no value | 3 |
| 400 | 3 | 3 |
- new answer published: Time since failure (s) = 52
- At 380: stragglers: caches that ignore TTL
Lessons from real failovers
Trade-offs.
Where failover should live, which record policy carries it, and what to do when every check is red.
- Pro:Cheap and protocol-agnostic
- Pro:Works across providers, clouds and on-premises sites
- Pro:No extra hop in the request path
- Con:Minutes rather than seconds
- Con:Stragglers that ignore TTL never move
Ties you to one provider's edge; Costs more; Its own outage is everyone's outage
Only works where you ship the client; Stale site lists live in old app versions for years
- Pro:Exactly one address per answer, so behaviour is easy to predict
- Pro:Fits Larkbook's Americas answer: Virginia while healthy, else Frankfurt
- Con:Frankfurt sees no Americas traffic until the switch, so drill it on a schedule (capacity planning for standby sites is in load-balancers/global-load-balancing)
- Con:Only two records per pair
Every site needs headroom for a failed peer: with three sites, at most two thirds each; If all checks fail, every record is returned again
Many client libraries take the first address and stop; AWS says it is not a substitute for a load balancer
- Pro:A checker fault or shared-dependency failure never removes the domain
- Con:Some users may be sent to a site that really is dead
Turns a monitoring bug into a global outage; Negative answers get cached too, which slows recovery
How DNS failover goes wrong
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| Flapping5Health checkers | Answers bounce between sites; users hit both halves of a sick site | Many status changes per hour on one check | Raise the threshold or interval; in your own checker, need more passes to recover than failures to fail | Both sites serve, some requests fail |
| Shallow health check6Virginia site (primary) | /healthz returns 200 while checkout fails, so no failover happens | User-facing error rate climbs while every check is green | Check what shoppers need from local dependencies; add a string match on the body | Site stays in DNS while broken |
| A shared dependency fails every site4Health aggregator | All checks go red at once; fail-open serves everything | Every record unhealthy at the same moment | Keep global dependencies out of /healthz; alert on them separately | DNS behaves as if checks did not exist |
| Cascading overload7Frankfurt site (secondary) | Frankfurt takes Virginia's traffic and falls over too | Frankfurt latency and errors spike right after the switch | Headroom, load shedding, or moving only a weighted share at first | Partial service from the survivor |
| Caches that outlive the TTL2Recursive resolvers | A tail of clients keeps dialling Virginia after the drain | Connections still arriving at the dead address minutes later | Clients re-resolve on connect errors; short JVM and pool TTLs | A small share of users stuck until restart |
| A stateful writer behind a DNS name | Two database primaries accept writes after a failover | Conflicting writes, split data | Use the database's own failover with fencing; keep DNS failover for stateless front doors | Data loss or manual repair |