Caching and TTLs.
DNS is fast because almost nobody asks the authoritative server. Browsers, operating systems and resolvers all keep answers until their time-to-live runs out. That one number is a contract. A long TTL buys speed and resilience, a short one buys agility, and you cannot switch from one to the other at the moment you need to move.
Builds on How a name resolves.
The idea.
One number on every DNS answer decides how long the world may keep using it.
When Larkbook's zone says shop.example.com 300 IN A 203.0.113.10, the 300 is the time-to-live: the number of seconds any cache may reuse that answer without asking again. The zone owner sets it per record set. Every cache between the app and the authoritative server keeps its own copy and counts it down: the browser, the operating system, and the recursive resolver shared by thousands of users (often a home router's forwarder too).
So when Larkbook changes the address, nothing is pushed anywhere. A given user moves to the new address only when every copy on their path has run out. That makes the TTL four knobs at once: latency (a hit costs nothing, a miss costs a trip to the nameservers), load (each cache asks once per TTL), how fast a change spreads (up to one old TTL), and how long you survive a nameserver outage (cached users carry on until their copies expire).
Four caches, one old answer
Scenario 1 of 2: As described.
Timeline as a list
Four caches, one old answer: 4 lanes, from 0 s to 480 s.
- 0–300 s · Recursive resolver · old answer, TTL 300
- 0 s · DNSCo nameservers · query answered (tick)
- 0 s · DNSCo nameservers → Recursive resolver, arriving 4 s
- 90–300 s · Laptop OS cache · old answer, got 210 s left
- 90 s · Recursive resolver → Laptop OS cache, arriving 92 s: TTL 210
- 120 s · DNSCo nameservers · A → 203.0.113.20
- 120–300 s · Recursive resolver, Laptop OS cache, Browser cache · window: user still sent to the old address
- 240–300 s · Browser cache · old answer
- 240 s · Laptop OS cache → Browser cache, arriving 242 s: TTL 60
- 300 s · Recursive resolver · TTL runs out (deadline)
- 330–480 s · Recursive resolver · new answer, TTL 300
- 330–480 s · Laptop OS cache · new answer
- 330 s · DNSCo nameservers · query answered (tick)
- 330 s · DNSCo nameservers → Recursive resolver, arriving 334 s
- 335–480 s · Browser cache · new answer
How it works.
A shared resolver turns thousands of users into one query per TTL. Several layers cache, each with its own rules, and negative answers are cached too.
Two users, one resolver, one query upstream
- Alice's laptop → Recursive resolver: A? shop.example.com (t = 0 s)
- Note over Recursive resolver: miss
- Recursive resolver → Authoritative nameservers: A? shop.example.com
- Authoritative nameservers → Recursive resolver (reply): 203.0.113.10, TTL 300
- Recursive resolver → Alice's laptop (reply): 203.0.113.10, TTL 300
- Bob's phone → Recursive resolver: A? shop.example.com (t = 120 s)
- Note over Recursive resolver: hit, 180 s left
- Recursive resolver → Bob's phone (reply): 203.0.113.10, TTL 180
- Note over Recursive resolver: t = 300 s: expired; the next query goes upstream again
The life of one entry in a resolver's cache
States of3Recursive resolver
Prefetch keeps a popular entry from ever expiring under load; serve-stale keeps an expired one answering when the nameservers are unreachable. Both are resolver options, not DNS protocol.
| From → To | Event | Guard | Action |
|---|---|---|---|
| Not cached → Fresh | answer | store | |
| Fresh → Fresh | query | > 10% left | answer |
| Fresh → Refreshing in background | query | last 10% | answer + refetch |
| Refreshing in background → Fresh | refetch ok | reset TTL | |
| Fresh → Expired | TTL hits 0 | ||
| Expired → Fresh | query | upstream ok | store |
| Expired → Serving stale | query | silent 1.8 s | answer stale |
| Serving stale → Serving stale | query | rechecked < 30 s ago | answer stale |
| Serving stale → Fresh | refetch ok | store | |
| Serving stale → Evicted | stale 1-3 days |
- Not cachedstart
- Fresh
- Answers with the remaining TTL
- Refreshing in background
- Unbound prefetch, last 10% of the TTL
- Serving stale
- RFC 8767: answers carry TTL 30 s
- Evictedend
Who caches, and whose rules
| Layer | Typical behaviour | What to watch |
|---|---|---|
| Browser | Its own host cache in front of the OS. Chromium keeps answers it got through the OS for about a minute, because that OS call returns no TTL. | Can hold an answer up to a minute past its TTL. |
| OS stub resolver | systemd-resolved, mDNSResponder or the Windows DNS Client caches for every app on the machine and honours the TTL. | Flush it when testing a change locally. |
| Language runtime | The JVM caches in-process (networkaddress.cache.ttl); some configurations never refresh. Other runtimes and HTTP clients have their own caches. | Set it explicitly; AWS's guide suggests 5 s for the JVM. |
| Recursive resolver | Honours the TTL but may cap it: Unbound by default keeps at most 1 day, and negative answers at most 1 hour (BIND's default cap is a week). May prefetch and serve stale. | On Unbound a 7-day TTL is really 1 day; a large resolver keeps many separate caches. |
| Connection pools | Not DNS at all: a keep-alive connection opened to the old address lives until it is closed. | The most common reason "we changed DNS and nothing moved". |
One zone, several TTLs
$TTL 3600 ; default for records without their own TTL
@ 900 IN SOA ns1.dnsco.example. hostmaster.example.com. (
2026093001 ; serial
7200 ; refresh
900 ; retry
1209600 ; expire
3600 ) ; MINIMUM: negative TTL = min(900, 3600) = 900 s
; the storefront, as in the first figure (one site for now; once shop is
; steered across three sites it drops to 60 s, see GeoDNS)
shop 300 IN A 203.0.113.10
; steered names: short, because DNS failover may move them
api 300 IN CNAME api.lb.example.com.
api.lb 60 IN A 203.0.113.30
; stable names: long, because nobody plans to move them this week
static 86400 IN CNAME larkbook.cdnco.example.
mail IN MX 10 mx1.example.com. ; inherits $TTL 3600
"No such name" is cached too
An NXDOMAIN (the name does not exist) or NODATA (it exists, but not with that type) answer carries the zone's SOA record, and resolvers cache the negative answer for the smaller of the SOA record's own TTL and its MINIMUM field (RFC 2308 §5). Larkbook's SOA has TTL 900 and MINIMUM 3,600, so a negative answer lives 900 s.
The trap: a health checker starts probing beta.example.com at 09:58, two minutes before the record is created at 10:00. Its resolver cached NXDOMAIN at 09:58 and keeps saying "no such name" until 10:13, 13 minutes after the record exists. Create records before anyone looks them up, and keep the negative TTL modest: RFC 2308 suggests 1 to 3 hours and calls more than a day problematic; for names you create often, 5 to 15 minutes is kinder.
What resolvers do on top of the TTL
In practice.
Choosing a TTL means pricing it in queries, latency and minutes of lag, then planning moves around it.
How the TTL sets the load on Larkbook's nameservers
- Resolver caches with steady demand for shop.example.com
- 20,000Our assumption for a mid-size global site. A large public resolver counts once per cache shard, not once.
- TTL options compared
- 3,600 / 300 / 60 / 5 s
- TTL 1 hour20,000 ÷ 3,600 s≈ 6 queries/sfrom Resolver caches with steady demand for shop.example.com and TTL options compared
- TTL 5 minutes20,000 ÷ 300 s≈ 67 queries/sfrom Resolver caches with steady demand for shop.example.com and TTL options compared
- TTL 1 minute20,000 ÷ 60 s≈ 333 queries/sfrom Resolver caches with steady demand for shop.example.com and TTL options compared
- TTL 5 seconds20,000 ÷ 5 s4,000 queries/sfrom Resolver caches with steady demand for shop.example.com and TTL options compared · An upper bound. Each cache refetches once per TTL only while its users keep asking.
- From 300 s to 5 s4,000 ÷ 6760× the queriesfrom TTL 5 minutes and TTL 5 seconds
- Load on the authoritative side scales with 1 ÷ TTL and the number of caches, not with the number of users.
- Going from 300 s to 5 s is 60 times the upstream queries, and 60 times as many moments where a user pays for a miss.
Upstream queries per second by TTL
Data
| TTL | queries/s (q/s) |
|---|---|
| 5 s | 4,000 |
| 60 s | 333 |
| 300 s | 67 |
| 1 h | 5.6 |
| 1 day | 0.23 |
What measurements say
Moving shop to a new server without a long tail
Scenario 1 of 2: As described.
- 6 Pooled connections and runtimes that cache forever. Keep it running and watch its request rate fall to zero.
Timeline as a list
Moving shop to a new server without a long tail: 4 lanes, from T−26 h to T+26 h.
- T−26 h–T · Old server 203.0.113.10 · all traffic
- T−25–T+24 h · shop's record at DNSCo · TTL 60
- T−25–T−1 h · Resolver caches worldwide · 1-day copies still out there
- T−25 h · shop's record at DNSCo · lower TTL to 60
- T−1 h · Resolver caches worldwide · every copy now ≤ 60 s (ok, ok)
- T–T+12 h · Old server 203.0.113.10 · stragglers only — Pooled connections and runtimes that cache forever. Keep it running and watch its request rate fall to zero.
- T–T+26 h · New server 203.0.113.20 · all traffic within about a minute
- T · shop's record at DNSCo · A → 203.0.113.20
- T+12 h · Old server 203.0.113.10 · retire old server (deadline)
- T+24 h · shop's record at DNSCo · raise TTL to 3,600
Changing DNS providers is slower still
Moving the zone from DNSCo to another provider changes its NS records, and the copy that matters most lives in the parent: the .com servers hand out example.com's delegation with a 2-day TTL (172,800 s) that you do not control. Child-centric resolvers then follow the NS TTL in your own zone, parent-centric ones the parent's. So the NS TTL that bounds the move is 172,800 s or your own, whichever is longer, and the old provider must keep serving the identical zone until it has passed. The order of the change (child first, then parent) is in resolution.
Trade-offs.
Every TTL trades agility against latency, load and surviving an outage; serving stale trades correctness for availability.
- Pro:Low latency and load where agility does not matter
- Pro:Fast changes where it does
- Pro:A nameserver outage is invisible to most users for hours
- Con:Two classes of record to keep straight
- Con:Planned moves of long-TTL names need the lower-then-switch dance
More misses and worse tail latency; 60 times the queries of a 1-hour TTL; A nameserver outage reaches users after one minute
An emergency change takes up to a day; Some resolvers cap it (Unbound at 1 day by default)
- Pro:Users keep working through a nameserver outage
- Pro:Short 30 s stale TTL means clients recover soon after the zone does
- Con:May keep serving an address you removed on purpose
- Con:Hides the outage from users and from naive monitoring
An outage longer than the TTL becomes a total outage for every user whose copy expired
A 90-minute nameserver outage, three ways
Scenario 1 of 3: As described.
Timeline as a list
A 90-minute nameserver outage, three ways: 3 lanes, from 0 min to 120 min.
- 0–13 min · Users · resolving (ok)
- 8–13 min · Recursive resolver ·
- 10–100 min · DNSCo nameservers · unreachable
- 13–100 min · Users · SERVFAIL (error)
- 13 min · Recursive resolver · 5-min copy expires (deadline)
- 100–120 min · Users · resolving (ok)
How caching bites
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| Runtime caches forever | A service keeps using the old address until it restarts. | Traffic to a retired address that never falls to zero. | Set the runtime's DNS TTL explicitly (5 to 60 s) and test it. | Everything else follows the change; one service lags. |
| Pooled connections pin old addresses | Keep-alive connections opened before the change carry on to the old server. | Connection counts on the old server. | Give pooled connections a maximum age (minutes), and drain the old server rather than killing it. | |
| Negative cache hides a new record3Recursive resolver | "No such name" for up to the negative TTL after the record exists. | NXDOMAIN from some networks and success from others. | Create records before anything looks them up; keep the SOA's negative TTL small (5 to 15 min) for zones that change often. | |
| TTL of 0 or 1 s4Authoritative nameservers | Every lookup is a miss; nameserver load and user latency jump. | Query rate at the provider; resolver cache hit ratio. | Use 60 s as the floor for steered names. Some resolvers are configured with a TTL floor, but don't count on it. | |
| Cache poisoning3Recursive resolver | A forged answer is cached and served to every user of that resolver for its TTL. | Answers that disagree with the authoritative servers. | Source-port and query-ID randomisation (RFC 5452) and DNSSEC validation on the resolver. |