DNSHow a name resolves

100%

How a name resolves.

Nobody holds the whole map of names. The root knows who runs .com, .com knows who runs example.com, and only example.com's own servers know where shop lives. A resolver near you follows those referrals for you, caches every hop, and hands your app one address, usually in a few milliseconds and occasionally in a few hundred.

Beginner20 minUpdated 30 Sept 2026

The idea.

Programs connect to addresses, people and configs use names. DNS is the lookup in between, and it works because no single server has to know everything.

A socket needs an IP address, yet nobody wants to hard-code 203.0.113.10 into a browser bookmark or a service config. Addresses change when a server moves to a new data centre or a new load balancer; the name shop.example.com should not. So something has to translate a name into an address at the moment a program connects, and it has to do it for billions of names that millions of owners edit every day.

One global table cannot work: no one organisation could accept every edit, and every lookup on Earth would hit it. DNS splits the name space by label instead. Read a name right to left: shop.example.com. is the root (the trailing dot), then com, then example, then shop. Each level hands the subtree below it to an owner. The root operators run only the list of top-level domains; the .com registry runs only the list of .com domains and their nameservers; Larkbook's DNS provider runs example.com and decides where shop points. The handover point is called a delegation, and the piece of the tree one owner runs is a zone.

Resolving a name means walking down those delegations. A resolver asks the root, which says 'I don't know shop, but here are the servers for com'. It asks those, which say 'here are the servers for example.com'. Those finally answer. Every step is cached, so the next person who asks for anything under .com skips the first two steps entirely.

The name space is a tree cut into zones

The name space is a tree cut into zones
The name space is a tree cut into zonesParts: "." (root), com, org, in, example.com, other.com, shop, api, mail.

example.com zonerun by DNSCo for Larkbook

.com zonerun by the .com registry

Root zonerun by 12 root operators

delegation: NS for com

NS for org

NS for in

delegation: NS for example.com

NS for other.com

"." (root)
13 names · ~2,000 anycast sites

com
NS records for every .com domain

example.com
SOA
NS ns1/ns2.dnsco.example

shop
A 203.0.113.10

api
CNAME api.lb.example.com

mail
MX target

org

in

other.com
run by a different owner

What each level can tell you about shop.example.com

ServerKnowsAnswers a question about shop with
RootWhich servers run com (and every other TLD)referral: com NS a.gtld-servers.net …
.com TLDWhich servers run example.comreferral: example.com NS ns1.dnsco.example …
example.com (DNSCo)Every record in the zoneanswer: shop A 203.0.113.10 (authoritative)
Was this section helpful?

How it works.

There are two kinds of asking. The stub asks recursively: 'give me the final answer'. The resolver asks iteratively: 'give me the answer, or tell me who to ask next'.

Who asks whom

Who asks whom. The numbered component cards that follow describe each part.
Who asks whomComponents: 1. Application (A browser or a service. Calls getaddrinfo() and blocks until it has an address to connect to.), 2. Stub resolver (The resolver library in the OS. A small cache and the address of one or two recursive resolvers, from DHCP or config.), 3. Recursive resolver (Run by your ISP, a public service (1.1.1.1, 8.8.8.8) or your platform. Walks the hierarchy on a miss and caches every answer it collects, for everyone it serves.), 4. Root servers (Know only which servers run each top-level domain. 13 identities, about 2,000 anycast instances.), 5. .com TLD servers (Know which nameservers each .com domain has delegated to. Nothing about the hosts inside it.), 6. Authoritative nameservers (The zone's own servers (DNSCo's ns1 and ns2 for example.com). The only source of shop's address.).

Authoritative hierarchy

Resolver operator

Your machine

getaddrinfo()

A? RD=1

who runs com?

who runs example.com?

A? shop

1Application

2Stub resolver
OS cache
1-2 resolver addresses

3Recursive resolver
one cache shared by every client

4Root servers

5.com TLD servers

6Authoritative nameservers
ns1.dnsco.examplens2.dnsco.example

A cold lookup, every hop

A cold lookup, every hop, as an ordered list of steps:
A cold lookup, every hop12 steps between Application, Stub resolver, Recursive resolver, Root servers, .com TLD servers, Authoritative nameservers. The steps are listed as text after the diagram.Authoritative nameservers.com TLD serversRoot serversRecursive resolverStub resolverApplicationmiss: shop, example.com, comlook up ns1 first (usually cached)getaddrinfo(shop.example.com)1A? shop.example.com (RD=1)2A? com (QNAME minimised)3referral: com NS + glue, TTL 2 d4A? example.com5referral: ns1/ns2.dnsco.example, no glue6A? shop.example.com (RD=0)7shop A 203.0.113.10, TTL 300, AA=18A 203.0.113.10, TTL 300 (cached)9203.0.113.1010
  1. Application → Stub resolver: getaddrinfo(shop.example.com)
  2. Stub resolver → Recursive resolver: A? shop.example.com (RD=1)
  3. Note over Recursive resolver: miss: shop, example.com, com
  4. Recursive resolver → Root servers: A? com (QNAME minimised)
  5. Root servers → Recursive resolver (reply): referral: com NS + glue, TTL 2 d
  6. Recursive resolver → .com TLD servers: A? example.com
  7. .com TLD servers → Recursive resolver (reply): referral: ns1/ns2.dnsco.example, no glue
  8. Note over Recursive resolver: look up ns1 first (usually cached)
  9. Recursive resolver → Authoritative nameservers: A? shop.example.com (RD=0)
  10. Authoritative nameservers → Recursive resolver (reply): shop A 203.0.113.10, TTL 300, AA=1
  11. Recursive resolver → Stub resolver (reply): A 203.0.113.10, TTL 300 (cached)
  12. Stub resolver → Application (reply): 203.0.113.10

Where the 137 ms of a cold lookup go

Scenario 1 of 3: As described.

Timeline as a list

Where the 137 ms of a cold lookup go: 6 lanes, from 0 ms to 140 ms.

  1. 0–137 ms · App (getaddrinfo) · blocked
  2. 0 ms · Stub resolver → Recursive resolver, arriving 6 ms
  3. 6–131 ms · Recursive resolver · walking the hierarchy
  4. 6 ms · Recursive resolver → Root server, arriving 16 ms: A? com
  5. 16 ms · Root server → Recursive resolver, arriving 26 ms
  6. 26 ms · Recursive resolver → .com TLD server, arriving 38 ms: A? example.com
  7. 39 ms · .com TLD server → Recursive resolver, arriving 51 ms
  8. 51 ms · Recursive resolver → DNSCo nameserver, arriving 91 ms: A? shop
  9. 91 ms · DNSCo nameserver → Recursive resolver, arriving 131 ms: A, AA=1
  10. 131 ms · Recursive resolver → Stub resolver, arriving 137 ms: answer
  11. 137 ms · App (getaddrinfo) · connect to 203.0.113.10 (ok)
Assumed round trips for one Larkbook user: 12 ms to the ISP resolver, 20 ms to the nearest root instance, 25 ms to a .com server and 80 ms to DNSCo, whose nearest site is far from this resolver. The app is blocked the whole time; the far authoritative alone is 58% of it.

The flag bits that say what kind of message this is

Flags word
  1. QR, 1 bit, Bit clear (0), 0 query, 1 response
  2. Opcode, 4 bits, Multi-bit code, value 0 QUERY
  3. AA, 1 bit, Bit clear (0), authoritative answer
  4. TC, 1 bit, Bit clear (0), truncated; retry over TCP
  5. RD, 1 bit, Bit set (1), recursion desired
  6. RA, 1 bit, Bit clear (0), recursion available
  7. Z, 1 bit, Reserved (always 0)
  8. AD, 1 bit, Bit clear (0), authentic data (DNSSEC)
  9. CD, 1 bit, Bit clear (0), checking disabled (DNSSEC)
  10. RCODE, 4 bits, Multi-bit code, value 0 NOERROR
  • Bit set (1)
  • Bit clear (0)
  • Multi-bit code
  • Reserved (always 0)
QR
0 query, 1 response
AA
authoritative answer
TC
truncated; retry over TCP
RD
recursion desired
RA
recursion available
AD
authentic data (DNSSEC)
CD
checking disabled (DNSSEC)
Start

As it starts. 6 steps follow.

The 16-bit flags word in every DNS message header (RFC 1035 section 4.1.1; AD and CD from RFC 4035). Step through the messages of the cold walk: the stub asks for recursion, authoritative servers never offer it, and only the zone's own server sets AA.

Records you meet in a system design

TypeHoldsExampleUsed for
AAn IPv4 addressshop 300 A 203.0.113.10Where to connect
AAAAAn IPv6 addressshop 300 AAAA 2001:db8::10Where to connect over IPv6
CNAMEAn alias: "look up this other name instead". Not allowed at the zone apex or beside other recordswww CNAME shop.example.com.Pointing a name at a CDN or load balancer hostname
NSThe nameservers for a zoneexample.com NS ns1.dnsco.example.Delegation
SOAZone metadata: primary server, serial, and the TTL for negative answersexample.com SOA ns1.dnsco.example. …Zone transfers and caching of "no such name"
MXMail servers with a preference numberexample.com MX 10 mail.example.com.Where email for the domain goes
TXTFree textexample.com TXT "v=spf1 …"SPF, domain ownership checks
CAAWhich certificate authorities may issue for the nameexample.com CAA 0 issue "ca.example"Limiting who can mint TLS certificates
HTTPS / SVCBConnection hints: ALPN protocols, alternative endpoints, address hints (RFC 9460)shop HTTPS 1 . alpn=h2,h3Letting a browser try HTTP/3 without a first round trip

Glue, and why nameserver names matter

Suppose example.com used ns1.example.com as its nameserver. To find ns1.example.com's address the resolver would have to ask example.com's nameserver, which is ns1.example.com: a loop. The parent breaks it by serving glue, the A and AAAA records of an in-bailiwick nameserver, alongside the referral. The .com servers hand out ns1.example.com's address even though they are not authoritative for it.

Larkbook's nameservers live under a different domain, dnsco.example, so they are out of bailiwick and the .com referral carries no glue for them. The resolver must resolve ns1.dnsco.example first, a side walk of its own. Providers make that cheap by putting every customer's nameservers under one domain that every busy resolver already has cached. When the glue at the parent and the records in the child disagree (after a nameserver move, say), some resolvers follow stale addresses: a common source of intermittent failures.

Watching a walk with dig +trace

.                    518400  IN  NS     a.root-servers.net.
.                    518400  IN  NS     b.root-servers.net.
;; (11 more root NS records trimmed)
;; Received 239 bytes from 192.0.2.53#53(192.0.2.53) in 12 ms

com.                 172800  IN  NS     a.gtld-servers.net.
com.                 172800  IN  NS     b.gtld-servers.net.
;; (11 more .com NS records trimmed)
;; Received 1171 bytes from 198.41.0.4#53(a.root-servers.net) in 20 ms

example.com.         172800  IN  NS     ns1.dnsco.example.
example.com.         172800  IN  NS     ns2.dnsco.example.
;; Received 112 bytes from 192.5.6.30#53(a.gtld-servers.net) in 25 ms

shop.example.com.    300     IN  A      203.0.113.10
;; Received 96 bytes from 198.51.100.53#53(ns1.dnsco.example) in 80 ms

On the wire

TransportPortWhen it is usedWhy it exists
UDP53The default: one datagram each way, no handshakeCheapest possible round trip. Without EDNS(0) an answer is capped at 512 bytes; with it the requester advertises a bigger buffer, 1,232 bytes recommended (DNS Flag Day 2020) so answers fit in an unfragmented packet on practically every path
TCP53After a TC=1 answer, for zone transfers, and for large DNSSEC answersMandatory to support (RFC 7766). A firewall that blocks TCP 53 breaks exactly the big answers
DNS over TLS (DoT)853Stub or forwarder to resolverEncrypts queries from on-path observers (RFC 7858)
DNS over HTTPS (DoH)443Browsers and apps to a resolver of their choiceEncrypted and rides on ordinary HTTPS, so it is hard to block or inspect separately (RFC 8484); it can bypass the resolver your network configured
Was this section helpful?

In practice.

The hierarchy looks slow on paper, but the top of it is almost always cached. What users feel is the miss to the zone's own servers, and the rare lookup that times out.

Root identities
13
a to m.root-servers.net, each with one IPv4 and one IPv6 anycast address
Root instances
2,045
anycast sites behind those 13 addresses, as of 30 Sep 2026 (root-servers.org)
Root operators
12
Verisign runs both A and J
Delegation TTL
2 days
172,800 s on TLD NS records in the root and on domain NS records in .com

What a lookup costs Larkbook's users

Assumptions
Stub to ISP resolver RTT
12 msassumption; same metro
Resolver to nearest root instance
20 msassumption
Resolver to nearest .com server
25 msassumption
Resolver to DNSCo
80 msassumption: DNSCo has no site near this resolver
Resolver to nearest DNSCo anycast site
20 msassumption: a provider with a site in this metro
Resolver cache hit rate for shop
90%assumption; a busy ISP resolver
Working
  1. Resolver has the answerstub-rtt12 msfrom Stub to ISP resolver RTT
  2. Usual miss (com delegation cached)stub-rtt + auth-rtt = 12 + 8092 msfrom Stub to ISP resolver RTT and Resolver to DNSCo
  3. Nothing cached12 + 20 + 25 + 80137 msfrom Stub to ISP resolver RTT, Resolver to nearest root instance, Resolver to nearest .com server and Resolver to DNSCo
  4. Average lookup0.9 × 12 + 0.1 × 92 = 10.8 + 9.220 msfrom Resolver cache hit rate for shop, Resolver has the answer and Usual miss (com delegation cached)
  5. Usual miss if DNSCo had an anycast site 20 ms awaystub-rtt + auth-anycast-rtt = 12 + 2032 msfrom Stub to ISP resolver RTT and Resolver to nearest DNSCo anycast site
  6. OS or browser cache hitno network~0 ms
What it means
  • Misses, not the hierarchy, drive DNS latency: root and TLD delegations stay cached for days, so the cold 137 ms case is rare.
  • Nearly half the average (9.2 of 20 ms) comes from the 10% of lookups that miss, which is why where the authoritative servers sit matters.
  • Plan for the tail: Google Public DNS measures 130 ms on average for nameservers that respond, and 300-400 ms end to end once lost packets and timeouts are counted.

What a lookup costs, from cache hit to fully cold

What a lookup costs, from cache hit to fully coldA resolver cache hit costs 12 ms, but a miss to a far authoritative costs 92 ms, so the rare misses set most of the 20 ms average.020 ms40 ms60 ms80 ms100 ms120 ms140 msOS cache hitResolver hitAverageAnycast missUsual missFully cold012 ms20 ms32 ms92 ms137 msSituationTime before the app can connect (ms)What a lookup costs, from cache hit to fully coldA resolver cache hit costs 12 ms, but a miss to a far authoritative costs 92 ms, so the rare misses set most of the 20 ms average.040 ms80 ms120 msOS cache hitResolver hitAverageAnycast missUsual missFully cold012 ms20 ms32 ms92 ms137 msSituationTime before the app can connect (ms)
Larkbook's assumed round trips from the estimate above.
Data
Situationms
OS cache hit0
Resolver hit12
Average20
Anycast miss32
Usual miss92
Fully cold137

Who runs each piece

ISP resolvers
Close to users and shared by a whole region's subscribers, so popular names are almost always cached. Quality varies; some are slow or filter answers.
Public resolvers (1.1.1.1, 8.8.8.8, 9.9.9.9)
Anycast, with very large shared caches. They can sit in a different city from the user's ISP, which matters when the answer depends on where the user is (see GeoDNS).
Inside your platform
Servers use a VPC resolver, systemd-resolved or a node-local cache; Kubernetes pods use CoreDNS. The pod default of ndots:5 with a search list means a name like api.example.com is first tried with each cluster suffix appended, several queries before the real one, unless you write it fully qualified with a trailing dot.
Authoritative providers
Managed anycast services (Route 53, Cloudflare, NS1 and others) answer from many sites near resolvers and absorb floods. Self-hosting BIND, Knot or PowerDNS gives full control, and makes global reach and DDoS your problem.
Was this section helpful?

Trade-offs.

Two choices you control, which resolver your servers use and where your zone is served from, and the ways resolution fails.

01
Which recursive resolver your servers use
Chosen:A caching resolver inside your platform (VPC resolver or node-local cache)
  • Pro:Lowest round trip (often under 1 ms)
  • Pro:One shared cache per node or VPC
  • Pro:Resolves your private zones
Downside we accept:
  • Con:Another component on every connection's critical path
  • Con:Must be sized and monitored like any other service
Ruled out:A public resolver over the internet

An internet round trip on every miss; Rate limits at high query volumes; Cannot see private names

Ruled out:Every process resolves with no cache in between

Every new connection may pay a full lookup; A connection storm becomes a query storm upstream

02
Where the authoritative zone is served from
Chosen:Managed anycast provider plus a second provider
  • Pro:Sites near most resolvers keep misses short
  • Pro:Floods absorbed by the providers
  • Pro:One provider's outage leaves the other answering
Downside we accept:
  • Con:Two zones to keep in sync
  • Con:Limited to features both providers support
Ruled out:One managed anycast provider

If that provider is down your names stop resolving once caches expire

Ruled out:Self-hosted primary and secondaries

You own global reach and DDoS defence; Few sites means long misses for distant users

FailureImpactDetectionMitigationMeanwhile
Every nameserver for the zone is unreachable6Authoritative nameserversEvery nameserver in the NS set is unreachable at once, so resolvers have nowhere to send misses and names stop resolving for everyone as cached answers expire (Meta, October 2021; see Failover with DNS).External DNS probes from several networks, not only from inside your ownSpread the NS set over two providers or networks that fail independentlyUsers with a cached answer keep connecting until its TTL runs out
Resolver slow or down3Recursive resolverEvery new connection stalls, even to healthy servicesLookup latency and SERVFAIL rate from the clients' sideConfigure two resolvers; run a node-local cache; keep lookups off hot paths by reusing connectionsStubs move to the next configured resolver after a timeout of seconds
Lame delegation or glue mismatch after a nameserver change5.com TLD serversThe parent names a server that does not answer for the zone; some lookups get SERVFAIL, depending on which server a resolver picksA dig +trace check in CI and after every NS change, from outside your networkChange nameservers in the child first, then the parent, and keep the old ones serving until the parent's 2-day TTL has passedIntermittent failures while resolvers retry other nameservers
Large answer dropped over UDP (many records, DNSSEC)2Stub resolverFragments are dropped by middleboxes, the query times out and the app sees a slow or failed lookupTimeouts that affect only some names, often the ones with long answersKeep the EDNS buffer at 1,232 bytes and allow TCP 53 through every firewallResolvers fall back to TCP, at the cost of an extra handshake
This topic
The walk from root to zone, referrals and glue
Record types and the header flags
Transport over UDP, TCP, DoT and DoH
Elsewhere
How long answers live, negative caching, serving staleCaching and TTLs
Giving each user a different answerGeoDNS and traffic steering
Pulling a dead address out of rotationFailover with DNS
DNSSEC signatures and validationout of scope for now
Was this section helpful?
Builds on this
Caching and TTLs
Read next