DNSCaching and TTLs

100%

Caching and TTLs.

DNS is fast because almost nobody asks the authoritative server. Browsers, operating systems and resolvers all keep answers until their time-to-live runs out. That one number is a contract. A long TTL buys speed and resilience, a short one buys agility, and you cannot switch from one to the other at the moment you need to move.

Intermediate19 minUpdated 30 Sept 2026

Builds on How a name resolves.

The idea.

One number on every DNS answer decides how long the world may keep using it.

When Larkbook's zone says shop.example.com 300 IN A 203.0.113.10, the 300 is the time-to-live: the number of seconds any cache may reuse that answer without asking again. The zone owner sets it per record set. Every cache between the app and the authoritative server keeps its own copy and counts it down: the browser, the operating system, and the recursive resolver shared by thousands of users (often a home router's forwarder too).

So when Larkbook changes the address, nothing is pushed anywhere. A given user moves to the new address only when every copy on their path has run out. That makes the TTL four knobs at once: latency (a hit costs nothing, a miss costs a trip to the nameservers), load (each cache asks once per TTL), how fast a change spreads (up to one old TTL), and how long you survive a nameserver outage (cached users carry on until their copies expire).

Four caches, one old answer

Scenario 1 of 2: As described.

Timeline as a list

Four caches, one old answer: 4 lanes, from 0 s to 480 s.

  1. 0–300 s · Recursive resolver · old answer, TTL 300
  2. 0 s · DNSCo nameservers · query answered (tick)
  3. 0 s · DNSCo nameservers → Recursive resolver, arriving 4 s
  4. 90–300 s · Laptop OS cache · old answer, got 210 s left
  5. 90 s · Recursive resolver → Laptop OS cache, arriving 92 s: TTL 210
  6. 120 s · DNSCo nameservers · A → 203.0.113.20
  7. 120–300 s · Recursive resolver, Laptop OS cache, Browser cache · window: user still sent to the old address
  8. 240–300 s · Browser cache · old answer
  9. 240 s · Laptop OS cache → Browser cache, arriving 242 s: TTL 60
  10. 300 s · Recursive resolver · TTL runs out (deadline)
  11. 330–480 s · Recursive resolver · new answer, TTL 300
  12. 330–480 s · Laptop OS cache · new answer
  13. 330 s · DNSCo nameservers · query answered (tick)
  14. 330 s · DNSCo nameservers → Recursive resolver, arriving 334 s
  15. 335–480 s · Browser cache · new answer
Larkbook moves shop from 203.0.113.10 to 203.0.113.20 at t = 120 s. The resolver fetched the old answer at t = 0 with TTL 300, and each cache below it got the remaining TTL, so every copy on this user's path expires together at t = 300. For three minutes the change is invisible to them. The second tab shows the same change with a 60 s TTL.
Was this section helpful?

How it works.

A shared resolver turns thousands of users into one query per TTL. Several layers cache, each with its own rules, and negative answers are cached too.

Two users, one resolver, one query upstream

Two users, one resolver, one query upstream, as an ordered list of steps:
Two users, one resolver, one query upstream9 steps between Alice's laptop, Bob's phone, Recursive resolver, Authoritative nameservers. The steps are listed as text after the diagram.Authoritative nameserversRecursive resolverBob's phoneAlice's laptopmisshit, 180 s leftt = 300 s: expired; the next query goes upstream againA? shop.example.com (t = 0 s)1A? shop.example.com2203.0.113.10, TTL 3003203.0.113.10, TTL 3004A? shop.example.com (t = 120 s)5203.0.113.10, TTL 1806
  1. Alice's laptop → Recursive resolver: A? shop.example.com (t = 0 s)
  2. Note over Recursive resolver: miss
  3. Recursive resolver → Authoritative nameservers: A? shop.example.com
  4. Authoritative nameservers → Recursive resolver (reply): 203.0.113.10, TTL 300
  5. Recursive resolver → Alice's laptop (reply): 203.0.113.10, TTL 300
  6. Bob's phone → Recursive resolver: A? shop.example.com (t = 120 s)
  7. Note over Recursive resolver: hit, 180 s left
  8. Recursive resolver → Bob's phone (reply): 203.0.113.10, TTL 180
  9. Note over Recursive resolver: t = 300 s: expired; the next query goes upstream again

The life of one entry in a resolver's cache

States of3Recursive resolver

The life of one entry in a resolver's cache. 6 states, 10 transitions. The table below lists them.
The life of one entry in a resolver's cacheThe states of Recursive resolver. 6 states, 10 transitions. The table below lists them.

answer / store

query [› 10% left] / answer

query [last 10%] / answer + refetch

refetch ok / reset TTL

TTL hits 0

query [upstream ok] / store

query [silent 1.8 s] / answer stale

query [rechecked ‹ 30 s ago] / answer stale

refetch ok / store

stale 1-3 days

Not cached

Fresh

Refreshing in background

Expired

Serving stale

Evicted

4 steps.

Prefetch keeps a popular entry from ever expiring under load; serve-stale keeps an expired one answering when the nameservers are unreachable. Both are resolver options, not DNS protocol.

Transitions of The life of one entry in a resolver's cache
From → ToEventGuardAction
Not cached → Freshanswerstore
Fresh → Freshquery> 10% leftanswer
Fresh → Refreshing in backgroundquerylast 10%answer + refetch
Refreshing in background → Freshrefetch okreset TTL
Fresh → ExpiredTTL hits 0
Expired → Freshqueryupstream okstore
Expired → Serving stalequerysilent 1.8 sanswer stale
Serving stale → Serving stalequeryrechecked < 30 s agoanswer stale
Serving stale → Freshrefetch okstore
Serving stale → Evictedstale 1-3 days
Not cachedstart
Fresh
Answers with the remaining TTL
Refreshing in background
Unbound prefetch, last 10% of the TTL
Serving stale
RFC 8767: answers carry TTL 30 s
Evictedend

Who caches, and whose rules

LayerTypical behaviourWhat to watch
BrowserIts own host cache in front of the OS. Chromium keeps answers it got through the OS for about a minute, because that OS call returns no TTL.Can hold an answer up to a minute past its TTL.
OS stub resolversystemd-resolved, mDNSResponder or the Windows DNS Client caches for every app on the machine and honours the TTL.Flush it when testing a change locally.
Language runtimeThe JVM caches in-process (networkaddress.cache.ttl); some configurations never refresh. Other runtimes and HTTP clients have their own caches.Set it explicitly; AWS's guide suggests 5 s for the JVM.
Recursive resolverHonours the TTL but may cap it: Unbound by default keeps at most 1 day, and negative answers at most 1 hour (BIND's default cap is a week). May prefetch and serve stale.On Unbound a 7-day TTL is really 1 day; a large resolver keeps many separate caches.
Connection poolsNot DNS at all: a keep-alive connection opened to the old address lives until it is closed.The most common reason "we changed DNS and nothing moved".

One zone, several TTLs

$TTL 3600                       ; default for records without their own TTL
@       900   IN SOA ns1.dnsco.example. hostmaster.example.com. (
                  2026093001    ; serial
                  7200          ; refresh
                  900           ; retry
                  1209600       ; expire
                  3600 )        ; MINIMUM: negative TTL = min(900, 3600) = 900 s
; the storefront, as in the first figure (one site for now; once shop is
; steered across three sites it drops to 60 s, see GeoDNS)
shop     300  IN A     203.0.113.10
; steered names: short, because DNS failover may move them
api      300  IN CNAME api.lb.example.com.
api.lb    60  IN A     203.0.113.30
; stable names: long, because nobody plans to move them this week
static 86400  IN CNAME larkbook.cdnco.example.
mail           IN MX    10 mx1.example.com.   ; inherits $TTL 3600

"No such name" is cached too

An NXDOMAIN (the name does not exist) or NODATA (it exists, but not with that type) answer carries the zone's SOA record, and resolvers cache the negative answer for the smaller of the SOA record's own TTL and its MINIMUM field (RFC 2308 §5). Larkbook's SOA has TTL 900 and MINIMUM 3,600, so a negative answer lives 900 s.

The trap: a health checker starts probing beta.example.com at 09:58, two minutes before the record is created at 10:00. Its resolver cached NXDOMAIN at 09:58 and keeps saying "no such name" until 10:13, 13 minutes after the record exists. Create records before anyone looks them up, and keep the negative TTL modest: RFC 2308 suggests 1 to 3 hours and calls more than a day problematic; for names you create often, 5 to 15 minutes is kinder.

What resolvers do on top of the TTL

TTL clamping
Resolvers impose a ceiling and sometimes a floor. Unbound's defaults are cache-max-ttl 86,400 and cache-min-ttl 0; operators who raise the floor (say to 60 s) make a TTL of 5 s behave like 60 s. Your TTL is a maximum, not a guarantee.
Sharded caches
Google Public DNS uses a small per-machine cache plus a pool partitioned by name to limit fragmentation, but a global resolver still keeps separate caches in many locations, so a short TTL costs more upstream queries and misses than one shared cache would.
Remaining TTL flows down
Every cache hands on what is left of the countdown, not the original TTL. That is why the resolver, the laptop's OS cache and the browser in the first figure all expire together, and why the old TTL bounds the whole chain (browser minute aside).
Was this section helpful?

In practice.

Choosing a TTL means pricing it in queries, latency and minutes of lag, then planning moves around it.

How the TTL sets the load on Larkbook's nameservers

Assumptions
Resolver caches with steady demand for shop.example.com
20,000Our assumption for a mid-size global site. A large public resolver counts once per cache shard, not once.
TTL options compared
3,600 / 300 / 60 / 5 s
Working
  1. TTL 1 hour20,000 ÷ 3,600 s≈ 6 queries/sfrom Resolver caches with steady demand for shop.example.com and TTL options compared
  2. TTL 5 minutes20,000 ÷ 300 s≈ 67 queries/sfrom Resolver caches with steady demand for shop.example.com and TTL options compared
  3. TTL 1 minute20,000 ÷ 60 s≈ 333 queries/sfrom Resolver caches with steady demand for shop.example.com and TTL options compared
  4. TTL 5 seconds20,000 ÷ 5 s4,000 queries/sfrom Resolver caches with steady demand for shop.example.com and TTL options compared · An upper bound. Each cache refetches once per TTL only while its users keep asking.
  5. From 300 s to 5 s4,000 ÷ 6760× the queriesfrom TTL 5 minutes and TTL 5 seconds
What it means
  • Load on the authoritative side scales with 1 ÷ TTL and the number of caches, not with the number of users.
  • Going from 300 s to 5 s is 60 times the upstream queries, and 60 times as many moments where a user pays for a miss.

Upstream queries per second by TTL

Upstream queries per second by TTLQueries per second rise in inverse proportion to the TTL: 0.23 at one day, 5.6 at one hour, 67 at 300 s, 333 at 60 s and 4,000 at 5 s.0.11101001k10k5 s60 s300 s1 h1 day4k q/s333 q/s67 q/s5.6 q/s0.23 q/sQueries per second (q/s)TTLUpstream queries per second by TTLQueries per second rise in inverse proportion to the TTL: 0.23 at one day, 5.6 at one hour, 67 at 300 s, 333 at 60 s and 4,000 at 5 s.0.11101001k10k5 s60 s300 s1 h1 day4k q/s333 q/s67 q/s5.6 q/s0.23 q/sQueries per second (q/s)TTL
20,000 caches with steady demand (our assumption), each refetching once per TTL. Log scale.
Data
TTLqueries/s (q/s)
5 s4,000
60 s333
300 s67
1 h5.6
1 day0.23

What measurements say

Median lookup, after .uy lengthened its TTLs
28.7 → 8 ms
75th percentile 183 → 21 ms (Moura et al., IMC'19, §5.3)
Recommended for most records
≥ 1 h
ideally 4 to 24 h (Moura et al.)
For DNS-based steering
5-15 min
the paper's range; managed failover often uses 60 s
Child-centric queries
52-90%
52-90% of queries follow the child zone's NS TTL (Moura et al.)

Moving shop to a new server without a long tail

Scenario 1 of 2: As described.

Notes
  • 6 Pooled connections and runtimes that cache forever. Keep it running and watch its request rate fall to zero.
Timeline as a list

Moving shop to a new server without a long tail: 4 lanes, from T−26 h to T+26 h.

  1. T−26 h–T · Old server 203.0.113.10 · all traffic
  2. T−25–T+24 h · shop's record at DNSCo · TTL 60
  3. T−25–T−1 h · Resolver caches worldwide · 1-day copies still out there
  4. T−25 h · shop's record at DNSCo · lower TTL to 60
  5. T−1 h · Resolver caches worldwide · every copy now ≤ 60 s (ok, ok)
  6. T–T+12 h · Old server 203.0.113.10 · stragglers only — Pooled connections and runtimes that cache forever. Keep it running and watch its request rate fall to zero.
  7. T–T+26 h · New server 203.0.113.20 · all traffic within about a minute
  8. T · shop's record at DNSCo · A → 203.0.113.20
  9. T+12 h · Old server 203.0.113.10 · retire old server (deadline)
  10. T+24 h · shop's record at DNSCo · raise TTL to 3,600
Suppose shop had been given a 1-day TTL instead of 300 s. Lowering it only takes effect once the old 1-day copies expire, so the lowering happens a day before the move: the step people forget. The second tab shows the move without it.

Changing DNS providers is slower still

Moving the zone from DNSCo to another provider changes its NS records, and the copy that matters most lives in the parent: the .com servers hand out example.com's delegation with a 2-day TTL (172,800 s) that you do not control. Child-centric resolvers then follow the NS TTL in your own zone, parent-centric ones the parent's. So the NS TTL that bounds the move is 172,800 s or your own, whichever is longer, and the old provider must keep serving the identical zone until it has passed. The order of the change (child first, then parent) is in resolution.

Was this section helpful?

Trade-offs.

Every TTL trades agility against latency, load and surviving an outage; serving stale trades correctness for availability.

01
What TTL to give a record
Chosen:Long (1 h to 1 day) for stable names; short (60 to 300 s) only for the names you steer
  • Pro:Low latency and load where agility does not matter
  • Pro:Fast changes where it does
  • Pro:A nameserver outage is invisible to most users for hours
Downside we accept:
  • Con:Two classes of record to keep straight
  • Con:Planned moves of long-TTL names need the lower-then-switch dance
Ruled out:Short everywhere (60 s)

More misses and worse tail latency; 60 times the queries of a 1-hour TTL; A nameserver outage reaches users after one minute

Ruled out:Very long everywhere (1 day or more)

An emergency change takes up to a day; Some resolvers cap it (Unbound at 1 day by default)

02
What a resolver does when the nameservers are down
Chosen:Serve stale (RFC 8767)
  • Pro:Users keep working through a nameserver outage
  • Pro:Short 30 s stale TTL means clients recover soon after the zone does
Downside we accept:
  • Con:May keep serving an address you removed on purpose
  • Con:Hides the outage from users and from naive monitoring
Ruled out:Strict expiry

An outage longer than the TTL becomes a total outage for every user whose copy expired

A 90-minute nameserver outage, three ways

Scenario 1 of 3: As described.

Timeline as a list

A 90-minute nameserver outage, three ways: 3 lanes, from 0 min to 120 min.

  1. 0–13 min · Users · resolving (ok)
  2. 8–13 min · Recursive resolver ·
  3. 10–100 min · DNSCo nameservers · unreachable
  4. 13–100 min · Users · SERVFAIL (error)
  5. 13 min · Recursive resolver · 5-min copy expires (deadline)
  6. 100–120 min · Users · resolving (ok)
DNSCo is unreachable from minute 10 to minute 100. The resolver last fetched shop at minute 8. What users see depends on the TTL and on whether the resolver serves stale. The first tab is TTL 300 with strict expiry; the other tabs add serve-stale or a 1-day TTL.

How caching bites

FailureImpactDetectionMitigationMeanwhile
Runtime caches foreverA service keeps using the old address until it restarts.Traffic to a retired address that never falls to zero.Set the runtime's DNS TTL explicitly (5 to 60 s) and test it.Everything else follows the change; one service lags.
Pooled connections pin old addressesKeep-alive connections opened before the change carry on to the old server.Connection counts on the old server.Give pooled connections a maximum age (minutes), and drain the old server rather than killing it.
Negative cache hides a new record3Recursive resolver"No such name" for up to the negative TTL after the record exists.NXDOMAIN from some networks and success from others.Create records before anything looks them up; keep the SOA's negative TTL small (5 to 15 min) for zones that change often.
TTL of 0 or 1 s4Authoritative nameserversEvery lookup is a miss; nameserver load and user latency jump.Query rate at the provider; resolver cache hit ratio.Use 60 s as the floor for steered names. Some resolvers are configured with a TTL floor, but don't count on it.
Cache poisoning3Recursive resolverA forged answer is cached and served to every user of that resolver for its TTL.Answers that disagree with the authoritative servers.Source-port and query-ID randomisation (RFC 5452) and DNSSEC validation on the resolver.

Where this topic stops

This topic
What the TTL means and who honours it
Negative caching, prefetch, serve-stale, clamping
Choosing TTLs and planning a move around them
Elsewhere
The walk from the root to the answerresolution
Different answers per user, and how they split cachesgeo-dns
Health checks that change records, and the detection-plus-TTL budgetdns-failover
Application caches, eviction and stampedesdistributed-cache
Was this section helpful?
Builds on this
GeoDNS and traffic steering
Read next