How a name resolves.
Nobody holds the whole map of names. The root knows who runs .com, .com knows who runs example.com, and only example.com's own servers know where shop lives. A resolver near you follows those referrals for you, caches every hop, and hands your app one address, usually in a few milliseconds and occasionally in a few hundred.
The idea.
Programs connect to addresses, people and configs use names. DNS is the lookup in between, and it works because no single server has to know everything.
A socket needs an IP address, yet nobody wants to hard-code 203.0.113.10 into a browser bookmark or a service config. Addresses change when a server moves to a new data centre or a new load balancer; the name shop.example.com should not. So something has to translate a name into an address at the moment a program connects, and it has to do it for billions of names that millions of owners edit every day.
One global table cannot work: no one organisation could accept every edit, and every lookup on Earth would hit it. DNS splits the name space by label instead. Read a name right to left: shop.example.com. is the root (the trailing dot), then com, then example, then shop. Each level hands the subtree below it to an owner. The root operators run only the list of top-level domains; the .com registry runs only the list of .com domains and their nameservers; Larkbook's DNS provider runs example.com and decides where shop points. The handover point is called a delegation, and the piece of the tree one owner runs is a zone.
Resolving a name means walking down those delegations. A resolver asks the root, which says 'I don't know shop, but here are the servers for com'. It asks those, which say 'here are the servers for example.com'. Those finally answer. Every step is cached, so the next person who asks for anything under .com skips the first two steps entirely.
The name space is a tree cut into zones
What each level can tell you about shop.example.com
| Server | Knows | Answers a question about shop with |
|---|---|---|
| Root | Which servers run com (and every other TLD) | referral: com NS a.gtld-servers.net … |
| .com TLD | Which servers run example.com | referral: example.com NS ns1.dnsco.example … |
| example.com (DNSCo) | Every record in the zone | answer: shop A 203.0.113.10 (authoritative) |
How it works.
There are two kinds of asking. The stub asks recursively: 'give me the final answer'. The resolver asks iteratively: 'give me the answer, or tell me who to ask next'.
Who asks whom
A cold lookup, every hop
- Application → Stub resolver: getaddrinfo(shop.example.com)
- Stub resolver → Recursive resolver: A? shop.example.com (RD=1)
- Note over Recursive resolver: miss: shop, example.com, com
- Recursive resolver → Root servers: A? com (QNAME minimised)
- Root servers → Recursive resolver (reply): referral: com NS + glue, TTL 2 d
- Recursive resolver → .com TLD servers: A? example.com
- .com TLD servers → Recursive resolver (reply): referral: ns1/ns2.dnsco.example, no glue
- Note over Recursive resolver: look up ns1 first (usually cached)
- Recursive resolver → Authoritative nameservers: A? shop.example.com (RD=0)
- Authoritative nameservers → Recursive resolver (reply): shop A 203.0.113.10, TTL 300, AA=1
- Recursive resolver → Stub resolver (reply): A 203.0.113.10, TTL 300 (cached)
- Stub resolver → Application (reply): 203.0.113.10
Where the 137 ms of a cold lookup go
Scenario 1 of 3: As described.
Timeline as a list
Where the 137 ms of a cold lookup go: 6 lanes, from 0 ms to 140 ms.
- 0–137 ms · App (getaddrinfo) · blocked
- 0 ms · Stub resolver → Recursive resolver, arriving 6 ms
- 6–131 ms · Recursive resolver · walking the hierarchy
- 6 ms · Recursive resolver → Root server, arriving 16 ms: A? com
- 16 ms · Root server → Recursive resolver, arriving 26 ms
- 26 ms · Recursive resolver → .com TLD server, arriving 38 ms: A? example.com
- 39 ms · .com TLD server → Recursive resolver, arriving 51 ms
- 51 ms · Recursive resolver → DNSCo nameserver, arriving 91 ms: A? shop
- 91 ms · DNSCo nameserver → Recursive resolver, arriving 131 ms: A, AA=1
- 131 ms · Recursive resolver → Stub resolver, arriving 137 ms: answer
- 137 ms · App (getaddrinfo) · connect to 203.0.113.10 (ok)
The flag bits that say what kind of message this is
- QR, 1 bit, Bit clear (0), 0 query, 1 response
- Opcode, 4 bits, Multi-bit code, value 0 QUERY
- AA, 1 bit, Bit clear (0), authoritative answer
- TC, 1 bit, Bit clear (0), truncated; retry over TCP
- RD, 1 bit, Bit set (1), recursion desired
- RA, 1 bit, Bit clear (0), recursion available
- Z, 1 bit, Reserved (always 0)
- AD, 1 bit, Bit clear (0), authentic data (DNSSEC)
- CD, 1 bit, Bit clear (0), checking disabled (DNSSEC)
- RCODE, 4 bits, Multi-bit code, value 0 NOERROR
- Bit set (1)
- Bit clear (0)
- Multi-bit code
- Reserved (always 0)
- QR
- 0 query, 1 response
- AA
- authoritative answer
- TC
- truncated; retry over TCP
- RD
- recursion desired
- RA
- recursion available
- AD
- authentic data (DNSSEC)
- CD
- checking disabled (DNSSEC)
As it starts. 6 steps follow.
Records you meet in a system design
| Type | Holds | Example | Used for |
|---|---|---|---|
| A | An IPv4 address | shop 300 A 203.0.113.10 | Where to connect |
| AAAA | An IPv6 address | shop 300 AAAA 2001:db8::10 | Where to connect over IPv6 |
| CNAME | An alias: "look up this other name instead". Not allowed at the zone apex or beside other records | www CNAME shop.example.com. | Pointing a name at a CDN or load balancer hostname |
| NS | The nameservers for a zone | example.com NS ns1.dnsco.example. | Delegation |
| SOA | Zone metadata: primary server, serial, and the TTL for negative answers | example.com SOA ns1.dnsco.example. … | Zone transfers and caching of "no such name" |
| MX | Mail servers with a preference number | example.com MX 10 mail.example.com. | Where email for the domain goes |
| TXT | Free text | example.com TXT "v=spf1 …" | SPF, domain ownership checks |
| CAA | Which certificate authorities may issue for the name | example.com CAA 0 issue "ca.example" | Limiting who can mint TLS certificates |
| HTTPS / SVCB | Connection hints: ALPN protocols, alternative endpoints, address hints (RFC 9460) | shop HTTPS 1 . alpn=h2,h3 | Letting a browser try HTTP/3 without a first round trip |
Glue, and why nameserver names matter
Suppose example.com used ns1.example.com as its nameserver. To find ns1.example.com's address the resolver would have to ask example.com's nameserver, which is ns1.example.com: a loop. The parent breaks it by serving glue, the A and AAAA records of an in-bailiwick nameserver, alongside the referral. The .com servers hand out ns1.example.com's address even though they are not authoritative for it.
Larkbook's nameservers live under a different domain, dnsco.example, so they are out of bailiwick and the .com referral carries no glue for them. The resolver must resolve ns1.dnsco.example first, a side walk of its own. Providers make that cheap by putting every customer's nameservers under one domain that every busy resolver already has cached. When the glue at the parent and the records in the child disagree (after a nameserver move, say), some resolvers follow stale addresses: a common source of intermittent failures.
Watching a walk with dig +trace
. 518400 IN NS a.root-servers.net.
. 518400 IN NS b.root-servers.net.
;; (11 more root NS records trimmed)
;; Received 239 bytes from 192.0.2.53#53(192.0.2.53) in 12 ms
com. 172800 IN NS a.gtld-servers.net.
com. 172800 IN NS b.gtld-servers.net.
;; (11 more .com NS records trimmed)
;; Received 1171 bytes from 198.41.0.4#53(a.root-servers.net) in 20 ms
example.com. 172800 IN NS ns1.dnsco.example.
example.com. 172800 IN NS ns2.dnsco.example.
;; Received 112 bytes from 192.5.6.30#53(a.gtld-servers.net) in 25 ms
shop.example.com. 300 IN A 203.0.113.10
;; Received 96 bytes from 198.51.100.53#53(ns1.dnsco.example) in 80 ms
On the wire
| Transport | Port | When it is used | Why it exists |
|---|---|---|---|
| UDP | 53 | The default: one datagram each way, no handshake | Cheapest possible round trip. Without EDNS(0) an answer is capped at 512 bytes; with it the requester advertises a bigger buffer, 1,232 bytes recommended (DNS Flag Day 2020) so answers fit in an unfragmented packet on practically every path |
| TCP | 53 | After a TC=1 answer, for zone transfers, and for large DNSSEC answers | Mandatory to support (RFC 7766). A firewall that blocks TCP 53 breaks exactly the big answers |
| DNS over TLS (DoT) | 853 | Stub or forwarder to resolver | Encrypts queries from on-path observers (RFC 7858) |
| DNS over HTTPS (DoH) | 443 | Browsers and apps to a resolver of their choice | Encrypted and rides on ordinary HTTPS, so it is hard to block or inspect separately (RFC 8484); it can bypass the resolver your network configured |
In practice.
The hierarchy looks slow on paper, but the top of it is almost always cached. What users feel is the miss to the zone's own servers, and the rare lookup that times out.
What a lookup costs Larkbook's users
- Stub to ISP resolver RTT
- 12 msassumption; same metro
- Resolver to nearest root instance
- 20 msassumption
- Resolver to nearest .com server
- 25 msassumption
- Resolver to DNSCo
- 80 msassumption: DNSCo has no site near this resolver
- Resolver to nearest DNSCo anycast site
- 20 msassumption: a provider with a site in this metro
- Resolver cache hit rate for shop
- 90%assumption; a busy ISP resolver
- Resolver has the answerstub-rtt12 msfrom Stub to ISP resolver RTT
- Usual miss (com delegation cached)stub-rtt + auth-rtt = 12 + 8092 msfrom Stub to ISP resolver RTT and Resolver to DNSCo
- Nothing cached12 + 20 + 25 + 80137 msfrom Stub to ISP resolver RTT, Resolver to nearest root instance, Resolver to nearest .com server and Resolver to DNSCo
- Average lookup0.9 × 12 + 0.1 × 92 = 10.8 + 9.220 msfrom Resolver cache hit rate for shop, Resolver has the answer and Usual miss (com delegation cached)
- Usual miss if DNSCo had an anycast site 20 ms awaystub-rtt + auth-anycast-rtt = 12 + 2032 msfrom Stub to ISP resolver RTT and Resolver to nearest DNSCo anycast site
- OS or browser cache hitno network~0 ms
- Misses, not the hierarchy, drive DNS latency: root and TLD delegations stay cached for days, so the cold 137 ms case is rare.
- Nearly half the average (9.2 of 20 ms) comes from the 10% of lookups that miss, which is why where the authoritative servers sit matters.
- Plan for the tail: Google Public DNS measures 130 ms on average for nameservers that respond, and 300-400 ms end to end once lost packets and timeouts are counted.
What a lookup costs, from cache hit to fully cold
Data
| Situation | ms |
|---|---|
| OS cache hit | 0 |
| Resolver hit | 12 |
| Average | 20 |
| Anycast miss | 32 |
| Usual miss | 92 |
| Fully cold | 137 |
Who runs each piece
Trade-offs.
Two choices you control, which resolver your servers use and where your zone is served from, and the ways resolution fails.
- Pro:Lowest round trip (often under 1 ms)
- Pro:One shared cache per node or VPC
- Pro:Resolves your private zones
- Con:Another component on every connection's critical path
- Con:Must be sized and monitored like any other service
An internet round trip on every miss; Rate limits at high query volumes; Cannot see private names
Every new connection may pay a full lookup; A connection storm becomes a query storm upstream
- Pro:Sites near most resolvers keep misses short
- Pro:Floods absorbed by the providers
- Pro:One provider's outage leaves the other answering
- Con:Two zones to keep in sync
- Con:Limited to features both providers support
If that provider is down your names stop resolving once caches expire
You own global reach and DDoS defence; Few sites means long misses for distant users
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| Every nameserver for the zone is unreachable6Authoritative nameservers | Every nameserver in the NS set is unreachable at once, so resolvers have nowhere to send misses and names stop resolving for everyone as cached answers expire (Meta, October 2021; see Failover with DNS). | External DNS probes from several networks, not only from inside your own | Spread the NS set over two providers or networks that fail independently | Users with a cached answer keep connecting until its TTL runs out |
| Resolver slow or down3Recursive resolver | Every new connection stalls, even to healthy services | Lookup latency and SERVFAIL rate from the clients' side | Configure two resolvers; run a node-local cache; keep lookups off hot paths by reusing connections | Stubs move to the next configured resolver after a timeout of seconds |
| Lame delegation or glue mismatch after a nameserver change5.com TLD servers | The parent names a server that does not answer for the zone; some lookups get SERVFAIL, depending on which server a resolver picks | A dig +trace check in CI and after every NS change, from outside your network | Change nameservers in the child first, then the parent, and keep the old ones serving until the parent's 2-day TTL has passed | Intermittent failures while resolvers retry other nameservers |
| Large answer dropped over UDP (many records, DNSSEC)2Stub resolver | Fragments are dropped by middleboxes, the query times out and the app sees a slow or failed lookup | Timeouts that affect only some names, often the ones with long answers | Keep the EDNS buffer at 1,232 bytes and allow TCP 53 through every firewall | Resolvers fall back to TCP, at the cost of an extra handshake |