DNSFailover with DNS

100%

Failover with DNS.

When Larkbook's Virginia site dies, nothing in DNS notices by itself. Health checkers in several regions have to agree it is down, the nameserver has to swap the answer, and every resolver still holding the old address has to let it expire. DNS failover works well for minutes-scale recovery of stateless front doors, and badly for anything that needs seconds.

Intermediate18 minUpdated 30 Sept 2026

Builds on Caching and TTLs.

The idea.

A name can point somewhere else when a site dies. The hard part is how long everyone keeps using the old answer.

What it does and what it doesn't

Larkbook's GeoDNS policy sends shoppers in the Americas to Virginia (192.0.2.30); Frankfurt (198.51.100.20) already serves Europe and is sized to take them too. If Virginia's front door goes dark, clients keep resolving the name to 192.0.2.30 and their connections time out. DNS has no idea: a nameserver hands out whatever records it holds.

DNS failover fixes that by attaching a health check to each answer. Probes from several regions hit each site's /healthz. When Virginia fails enough of them, the nameserver starts answering with Frankfurt instead. Nothing on the client changes; the next time it looks the name up, it gets the new address.

  • Good at: a whole site or region going dark, in front of stateless web or API traffic. It is cheap, works for any protocol that starts with a name lookup, and spans clouds and providers.
  • Not for: one bad server inside a pool (the load balancer's health checks handle that in seconds), anything that needs sub-second recovery, or moving a database primary (that needs fencing, not a new address).

The catch is time. The timeline below is Larkbook's worst case with a 10 s check interval, 3 failures to flip and a 60 s TTL: almost two minutes before the last well-behaved client moves.

Where a two-minute failover goes

Scenario 1 of 2: As described.

Notes
  • 4 Everyone at first; a shrinking share after t=52.
  • 11 An answer fetched just before t=52 lives until t=112.
Timeline as a list

Where a two-minute failover goes: 5 lanes, from 0 s to 140 s.

  1. 0–140 s · Virginia site · front door down
  2. 0–42 s · Health checkers · detecting
  3. 0–42 s · DNSCo nameservers · still answers Virginia
  4. 0–112 s · Clients · some connections fail — Everyone at first; a shrinking share after t=52.
  5. 0 s · Virginia site · Virginia goes dark (error)
  6. 14 s · Health checkers · fail 1 (error)
  7. 28 s · Health checkers · fail 2 (error)
  8. 42–52 s · DNSCo nameservers · publishing new status
  9. 42 s · Health checkers · fail 3 (error)
  10. 52–140 s · DNSCo nameservers · answers Frankfurt
  11. 52–112 s · Resolver caches · old answers drain (TTL 60) — An answer fetched just before t=52 lives until t=112.
  12. 112 s · Clients · last TTL-respecting client moves (ok, ok)
  13. 112–140 s · Clients · window: stragglers that ignore TTL
Worst case with assumed timings. Probes time out after 4 s; each checker waits 10 s after a result before probing again; publication to every nameserver is assumed to take 10 s. The second tab shows clients that hold both addresses and skip the wait.
Was this section helpful?

How it works.

Checkers vote, an aggregator decides, the nameserver answers from the decision, and caches hold the old answer until it expires.

Health-checked DNS for shop.example.com

Health-checked DNS for shop.example.com. The numbered component cards that follow describe each part.
Health-checked DNS for shop.example.comComponents: 1. Clients (Browsers, apps and their runtimes. Each asks a resolver for shop.example.com and may cache the answer itself.), 2. Recursive resolvers (ISP and public resolvers. They hold an answer until its TTL (60 s here) runs out and only then ask again.), 3. Authoritative nameservers (DNSCo's ns1 and ns2. They read the latest health status when choosing which record to return.), 4. Health aggregator (Combines checker votes. The site is healthy while more than 18% of checkers say so each checker needs 3 results in a row to change its vote.), 5. Health checkers (Probes from several regions. Each sends GET /healthz to every site 10 s after its last response.), 6. Virginia site (primary) (us-east front door at 192.0.2.30. Gets all of the Americas' traffic while healthy.), 7. Frankfurt site (secondary) (eu-central front door at 198.51.100.20. Already serves Europe, and is sized to take Virginia's shoppers on top.).

Sites

Health control

GET /healthz

GET /healthz

pass / fail

primary unhealthy

secondary, TTL 60

cached answer

connect

connect

5Health checkers
GET /healthz every 10 s
4 s connect timeout
TokyoSão PauloDublinOregonSydney

4Health aggregator
healthy while > 18% say so

6Virginia site (primary)
192.0.2.30

7Frankfurt site (secondary)
198.51.100.20

3Authoritative nameservers
primary 192.0.2.30
secondary 198.51.100.20
TTL 60

2Recursive resolvers
hold the answer until TTL ends

1Clients

One checker's vote on Virginia

States of5Health checkers

One checker's vote on Virginia. 4 states, 8 transitions. The table below lists them.
One checker's vote on VirginiaThe states of Health checkers. 4 states, 8 transitions. The table below lists them.

probe fails / count = 1

probe fails [count + 1 ‹ 3] / count += 1

probe fails [count + 1 = 3] / vote unhealthy

probe passes / reset count

probe passes / count = 1

probe passes [count + 1 ‹ 3] / count += 1

probe passes [count + 1 = 3] / vote healthy

probe fails / reset count

Healthy

Healthy, 1–2 failures seen

Unhealthy

Unhealthy, 1–2 passes seen

3 steps. The outage in the timeline, seen by one checker.

Route 53 uses one threshold in both directions: 3 results in a row to go down and 3 to come back. The aggregator counts these votes; Virginia leaves the answers once 18% or fewer of them say healthy.

Transitions of One checker's vote on Virginia
From → ToEventGuardAction
Healthy → Healthy, 1–2 failures seenprobe failscount = 1
Healthy, 1–2 failures seen → Healthy, 1–2 failures seenprobe failscount + 1 < 3count += 1
Healthy, 1–2 failures seen → Unhealthyprobe failscount + 1 = 3vote unhealthy
Healthy, 1–2 failures seen → Healthyprobe passesreset count
Unhealthy → Unhealthy, 1–2 passes seenprobe passescount = 1
Unhealthy, 1–2 passes seen → Unhealthy, 1–2 passes seenprobe passescount + 1 < 3count += 1
Unhealthy, 1–2 passes seen → Healthyprobe passescount + 1 = 3vote healthy
Unhealthy, 1–2 passes seen → Unhealthyprobe failsreset count
Healthystart
Healthy, 1–2 failures seen
A single blip does not change the vote
Unhealthyerror

A client caught mid-failover

A client caught mid-failover, as an ordered list of steps:
A client caught mid-failover15 steps between Clients, Recursive resolvers, Authoritative nameservers, Health aggregator, Health checkers, Virginia site (primary), Frankfurt site (secondary). The steps are listed as text after the diagram.Frankfurt site (secondary)Virginia site (primary)Health checkersHealth aggregatorAuthoritative nameserversRecursive resolversClientst=0 s: front door downt=110 s: copy expirest=10 s: GET /healthz, 4 s timeout1t=24 s: fail 22t=38 s: fail 3 in a row3all 5 regions: unhealthy4primary down, live by t=52 s5A? shop (t=55 s)6192.0.2.30 (55 s left)7connect: times out8A? shop (retry)9A? shop10198.51.100.20, TTL 6011198.51.100.2012connect: OK13
  1. Note over Virginia site (primary): t=0 s: front door down
  2. Health checkers → Virginia site (primary): t=10 s: GET /healthz, 4 s timeout
  3. Health checkers → Virginia site (primary): t=24 s: fail 2
  4. Health checkers → Virginia site (primary): t=38 s: fail 3 in a row
  5. Health checkers → Health aggregator: all 5 regions: unhealthy
  6. Health aggregator → Authoritative nameservers: primary down, live by t=52 s
  7. Clients → Recursive resolvers: A? shop (t=55 s)
  8. Recursive resolvers → Clients (reply): 192.0.2.30 (55 s left)
  9. Clients → Virginia site (primary): connect: times out
  10. Note over Recursive resolvers: t=110 s: copy expires
  11. Clients → Recursive resolvers: A? shop (retry)
  12. Recursive resolvers → Authoritative nameservers: A? shop
  13. Authoritative nameservers → Recursive resolvers (reply): 198.51.100.20, TTL 60
  14. Recursive resolvers → Clients (reply): 198.51.100.20
  15. Clients → Frankfurt site (secondary): connect: OK

Ways to wire it

PatternWhat DNS returnsGood forCost
Active–passiveThe primary while it is healthy, else the secondaryDisaster-recovery sites and one-region productsThe passive site idles, must stay warm and sized, and is only proven when you fail over
Active–activeEvery healthy site, or the nearest or weighted pick among the healthy onesNormal multi-region servingEvery site must absorb a failed peer's share on top of its own
Multi-address answersSeveral healthy addresses at once; clients try the next on failureClients you control, or browsers that race addressesInstant only for clients that actually retry; many libraries take the first address and stop
Calculated checksA site counts as healthy only while at least x of its y child checks areSites with many front-end nodes; one bad node should not fail the whole siteMore checks to configure and pay for; the x must be tuned

The rules that surprise people

Managed DNS services lean hard towards always answering. In Route 53, which the others resemble:

  • A record with no health check is always healthy. Forget to attach one and that site is never pulled.
  • If every record in a group is unhealthy, all of them are treated as healthy and the normal policy (weights, latency) picks one.
  • In a failover pair, if the primary and the secondary both fail their checks, the primary is returned. If the secondary has no check, it is returned whenever the primary fails, even if it is also down.
  • Checks run on their own schedule. A query never triggers a probe; it reads the last status, so detection never speeds up because traffic is heavy.

The reasoning: an empty answer is a certain outage for everyone, while a possibly dead address is only a likely one. If the checkers themselves break, or a shared dependency makes every site fail its check, answering nothing would turn a monitoring fault into the whole domain disappearing.

A health endpoint worth failing over on

# Larkbook's /healthz: "can this site serve a shopper right now?"
# Check only what is LOCAL to the site. A shared global dependency
# (payments provider, central auth) would fail every site at once,
# and fail-open DNS would then send traffic everywhere anyway.
def healthz(request):
    reasons = []
    if draining_flag_set():                    # operator is emptying the site
        reasons.append("draining")
    if not replica_ping(timeout_s=0.2):        # this site's read replica
        reasons.append("catalog replica unreachable")
    if disk_free_fraction("/var/cache") < 0.05:
        reasons.append("cache disk nearly full")
    if healthy_app_servers() < 3:              # of 12; below this we can't carry peak
        reasons.append("too few app servers")
    if reasons:
        return Response(503, body="unhealthy: " + ", ".join(reasons))
    return Response(200, body="ok larkbook")   # string match looks for "ok larkbook"
Was this section helpful?

In practice.

Put numbers on each phase, see which one dominates, and learn what goes wrong in real failovers.

How long Larkbook's users are stuck

Assumptions
Check interval
10 sRoute 53 fast interval; measured from one result to the next probe
Failures in a row to flip
3
Probe timeout
4 sHTTP connect limit; a dead site times out rather than refusing
Status reaches every nameserver
~10 sassumption; providers do not publish a figure
Record TTL
60 s
Whole-site failures per year
4assumption
Working
  1. Detection, worst caseinterval (wait for first probe) + threshold × timeout + (threshold − 1) × interval = 10 + 3 × 4 + 2 × 1042 sfrom Check interval, Failures in a row to flip and Probe timeout
  2. Worst case for a TTL-respecting clientdetect + publish + ttl = 42 + 10 + 60112 s ≈ 2 minfrom Detection, worst case, Status reaches every nameserver and Record TTL
  3. Average(5 + 12 + 20) + 10 + 60 ÷ 2 = 37 + 10 + 3077 sfrom Check interval, Failures in a row to flip, Probe timeout, Status reaches every nameserver and Record TTL · Failure lands mid-interval on average; a cached answer has half its TTL left on average.
  4. Worst case with TTL 30042 + 10 + 300352 s ≈ 6 minfrom Detection, worst case and Status reaches every nameserver
  5. Failover lag per year, worst case4 × 112 s = 448 s ≈ 7.5 min; 7.5 ÷ 525,600 min (at the 77 s average: 308 s ≈ 5.1 min ≈ 0.001%)≈ 0.0014% (99.9986% for this cause alone)from Whole-site failures per year and Worst case for a TTL-respecting client
What it means
  • Past detection, the TTL is the budget. 60 s is the usual floor worth paying for; below that, resolver load rises and some resolvers clamp very small TTLs up to a floor (see dns/caching-and-ttl).
  • Anything that must recover in seconds needs another layer, such as anycast or a global L7 balancer (load-balancers/global-load-balancing).
  • Clients that try the next address or re-resolve after a connect error beat any TTL.

Share of new connections still aimed at Virginia

  • TTL 60 s
  • TTL 300 s
Share of new connections still aimed at VirginiaWith a 60 s TTL the dead site stops getting new connections about 1 minute after the switch; with 300 s it takes 5 minutes.020%40%60%80%100%0100 s200 s300 s400 snew answer publishedstragglers: caches that ignore TTLTTL 60 sTTL 300 sNew connections to Virginia (%)Time since failure (s)Share of new connections still aimed at VirginiaWith a 60 s TTL the dead site stops getting new connections about 1 minute after the switch; with 300 s it takes 5 minutes.020%40%60%80%100%0100 s200 s300 s400 snew answer publishedstragglers: caches that ignore TTLTTL 60 sTTL 300 sNew connections to Virginia (%)Time since failure (s)
Assumes busy resolvers whose cached copies expire evenly over one TTL, and that 3% of traffic comes from runtimes that never re-resolve (illustrative). Status is live on the nameservers at t=52 s in both cases.
Data
Time since failure (s)TTL 60 s (%)TTL 300 s (%)
0100100
52100100
1123no value
352no value3
40033
  • new answer published: Time since failure (s) = 52
  • At 380: stragglers: caches that ignore TTL

Lessons from real failovers

Checkers must agree across regions
Route 53 keeps an endpoint healthy while more than 18% of its checkers say so, and asks for at least three checker regions. One isolated vantage point cannot declare you down.
Don't let DNS health-check itself into silence
A health signal that can pull the DNS layer itself off the internet must fail open. Meta learned this in 2021: its nameservers treated a lost backbone as their own sickness, took themselves out of routing, and every Meta name went dark with them once cached answers expired.
The secondary must be ready
A cold or undersized secondary turns one outage into two. Fail over on purpose, on a schedule, so the first real test is not during the incident.
Re-resolve on connection failure
Some JVM configurations cache a lookup until restart; AWS recommends a 5 s JVM TTL. Connection pools that keep sockets to a dead address, or never look the name up again, never fail over at all.
Was this section helpful?

Trade-offs.

Where failover should live, which record policy carries it, and what to do when every check is red.

01
Where failover between sites happens
Chosen:DNS failover with a 60 s TTL, plus clients that retry and re-resolve
  • Pro:Cheap and protocol-agnostic
  • Pro:Works across providers, clouds and on-premises sites
  • Pro:No extra hop in the request path
Downside we accept:
  • Con:Minutes rather than seconds
  • Con:Stragglers that ignore TTL never move
Ruled out:Anycast or a global L7 load balancer

Ties you to one provider's edge; Costs more; Its own outage is everyone's outage

Ruled out:Client-side failover (the app knows every site)

Only works where you ship the client; Stale site lists live in old app versions for years

02
Which record policy carries the failover
Chosen:A failover pair (primary and secondary records)
  • Pro:Exactly one address per answer, so behaviour is easy to predict
  • Pro:Fits Larkbook's Americas answer: Virginia while healthy, else Frankfurt
Downside we accept:
  • Con:Frankfurt sees no Americas traffic until the switch, so drill it on a schedule (capacity planning for standby sites is in load-balancers/global-load-balancing)
  • Con:Only two records per pair
Ruled out:Weighted or latency records, each with a health check

Every site needs headroom for a failed peer: with three sites, at most two thirds each; If all checks fail, every record is returned again

Ruled out:Multivalue answers

Many client libraries take the first address and stop; AWS says it is not a substitute for a load balancer

03
When every site looks unhealthy
Chosen:Fail open (keep answering with every site)
  • Pro:A checker fault or shared-dependency failure never removes the domain
Downside we accept:
  • Con:Some users may be sent to a site that really is dead
Ruled out:Fail closed (answer nothing, or only a static sorry page)

Turns a monitoring bug into a global outage; Negative answers get cached too, which slows recovery

How DNS failover goes wrong

FailureImpactDetectionMitigationMeanwhile
Flapping5Health checkersAnswers bounce between sites; users hit both halves of a sick siteMany status changes per hour on one checkRaise the threshold or interval; in your own checker, need more passes to recover than failures to failBoth sites serve, some requests fail
Shallow health check6Virginia site (primary)/healthz returns 200 while checkout fails, so no failover happensUser-facing error rate climbs while every check is greenCheck what shoppers need from local dependencies; add a string match on the bodySite stays in DNS while broken
A shared dependency fails every site4Health aggregatorAll checks go red at once; fail-open serves everythingEvery record unhealthy at the same momentKeep global dependencies out of /healthz; alert on them separatelyDNS behaves as if checks did not exist
Cascading overload7Frankfurt site (secondary)Frankfurt takes Virginia's traffic and falls over tooFrankfurt latency and errors spike right after the switchHeadroom, load shedding, or moving only a weighted share at firstPartial service from the survivor
Caches that outlive the TTL2Recursive resolversA tail of clients keeps dialling Virginia after the drainConnections still arriving at the dead address minutes laterClients re-resolve on connect errors; short JVM and pool TTLsA small share of users stuck until restart
A stateful writer behind a DNS nameTwo database primaries accept writes after a failoverConflicting writes, split dataUse the database's own failover with fencing; keep DNS failover for stateless front doorsData loss or manual repair

What this topic covers

This topic
Health-gated DNS answers between sites
The detection, publication and TTL timing budget
Fail-open rules and checker quorums
Elsewhere
Instance health inside one pool, drainingload-balancers/health-checks
Anycast and global load balancing, region evacuationload-balancers/global-load-balancing
How TTLs and caches workdns/caching-and-ttl
Choosing a site per userdns/geo-dns
Availability maths and SLOsnon-functional-requirements/availability
Client retries and backofffailure-models/timeouts-and-retries
Was this section helpful?
Related
GeoDNS and traffic steering
Read next