DNSGeoDNS and traffic steering

100%

GeoDNS and traffic steering.

Larkbook runs its shop in Mumbai, Frankfurt and Virginia. One name, three addresses, and the nameserver picks which one to hand out based on where the question came from. The catch is that the question comes from a resolver, not from the shopper, and the answer is cached for everyone behind that resolver.

Intermediate17 minUpdated 30 Sept 2026

Builds on Caching and TTLs.

The idea.

One name, several sites, and a nameserver that picks a site for each question it is asked.

Larkbook serves shop.example.com from three sites: Mumbai (203.0.113.10), Frankfurt (198.51.100.20) and Virginia (192.0.2.30). A shopper in Pune should reach Mumbai, about 120 km away, not Virginia, about 13,000 km away. Light in fibre covers roughly 200 km per millisecond, so that distance alone adds about 130 ms to every round trip to Virginia (13,000 km each way at 200 km/ms is 65 ms, twice), and a page load needs several round trips. Larkbook also wants to stop sending anyone to a site that is failing or being drained, and to send less to the smallest site.

There are two places to make that choice. Anycast makes it in the network: every site announces the same address and internet routing delivers each packet to a nearby one. GeoDNS makes it in the answer: every site keeps its own ordinary address, and the authoritative nameserver returns the address it wants this asker to use. Because the decision is a lookup in software, the policy can be anything: country, measured latency, fixed weights, cost, legal location of data, or site health.

Two limits come with it, and the rest of this topic is about living with them. The nameserver never talks to the shopper; it talks to the shopper's recursive resolver (the ISP or public service that looks names up on the shopper's behalf and caches the answers), so it steers by where the resolver is unless the resolver says more. And the answer is cached by that resolver for its TTL (see Caching and TTLs), so every user behind it gets the same answer until it expires, and no change of policy takes effect faster than that.

Who gets which site under a plain country policy (IN → Mumbai, EU → Frankfurt, US → Virginia, anything else → Virginia)

ShopperAsks throughNameserver seesAnswerRight site?
Ananya, Puneher ISP's resolver in Mumbaian Indian address203.0.113.10Yes — Mumbai
Lukas, Berlinhis ISP's resolver in Frankfurta German address198.51.100.20Yes — Frankfurt
Maya, Ohioa public resolver's Chicago sitea US address192.0.2.30Yes — Virginia
Priya, Guwahatia public resolver's Singapore site with no client subneta Singapore address192.0.2.30No — Singapore matched no rule, so the default sent her to Virginia
Tomás, Lisbonhis company's VPN resolver in Virginiaa US address192.0.2.30No — the nameserver never saw Lisbon
Was this section helpful?

How it works.

The nameserver joins three inputs (where the asker is, how fast each site is from there, and which sites are fit to serve) into one answer, and says how widely that answer may be reused.

GeoDNS for shop.example.com

GeoDNS for shop.example.com. The numbered component cards that follow describe each part.
GeoDNS for shop.example.comComponents: 1. Shoppers (Larkbook's shoppers in India, Europe and the Americas. Their devices ask a resolver, never the nameserver.), 2. Recursive resolvers (The shoppers' resolvers: an ISP's (usually near the user) or a public anycast service (near one of its own sites, which may not be near the user).), 3. Authoritative GeoDNS (DNSCo's nameservers (Larkbook's managed DNS provider) for shop.example.com. They evaluate Larkbook's steering policy on every query that reaches them.), 4. GeoIP and network map (Maps an address prefix to a country, city and network (ASN). Bought or built, refreshed weekly or so, and wrong for some networks.), 5. Latency map (Measured round-trip time from each client network to each site, refreshed from real-user and probe measurements.), 6. Health and capacity feed (Tells the nameserver which sites are up and how much traffic each should take. The checking itself is covered in dns-failover.), 7. Larkbook sites (Three full copies of the shop, each behind its own unicast address.).

Steering inputs

A? shop

A? shop + ECS /24

answer, scope /20, TTL 60

where is this prefix?

fastest site?

site up? weight?

HTTPS

4GeoIP and network map

5Latency map

6Health and capacity feed

1Shoppers

2Recursive resolvers
ISP or public anycast

3Authoritative GeoDNS
policy per name
answer per query

7Larkbook sites
Mumbai203.0.113.10Frankfurt198.51.100.20Virginia192.0.2.30

What happens on each query

  1. Work out who is asking: the client subnet if the query carries one, otherwise the resolver's own address.
  2. Look that prefix up in the network map: country, city, ISP. This is a database lookup, not a measurement, and it is only as good as the database.
  3. Walk the policy's rules in order and take the first pool that matches (India → ap-south, say).
  4. Drop sites the health feed marks as down; if the pool is empty, fall through to the next choice instead of answering with a dead address.
  5. Answer with the chosen address, a TTL, and (when the query had a client subnet) the widest prefix the answer is valid for.

Steering policies

PolicyDecides byGood forWatch out
GeolocationCountry or continent of the source, with a default for everything elseData residency, licensing, languageWrong for VPN and roaming users; with no default, unmatched places get no answer at all
Latency-basedMeasured round-trip time from the source's network to each siteRaw performanceNeeds a measurement system; can flap between two sites that are nearly equal
WeightedFixed proportions (weight ÷ sum of weights)Canaries, gradual migration between sitesThe split is per resolver per TTL window, not per user
Geoproximity with biasDistance to each site, with a bias that grows or shrinks a site's areaShifting load off a hot regionCoarse; one bias moves whole regions
IP / CIDR mapExplicit prefix → site listsLarge ISPs, partners, known officesYou maintain the list and it goes stale
MultivalueSeveral healthy addresses at random (up to 8 in Route 53)Letting the client pick and retryNot a load balancer: random per response and per resolver, with no notion of capacity
record: shop.example.com
type: A
ttl: 60                      # was 300 while shop had one site; steered names stay short
pools:
  ap-south:   { address: 203.0.113.10,  health_check: https-healthz }
  eu-central: { address: 198.51.100.20, health_check: https-healthz }
  us-east:    { address: 192.0.2.30,    health_check: https-healthz }
  us-east-canary: { address: 192.0.2.31, health_check: https-healthz }
rules:                       # first match wins
  - match: { country: [IN, LK, BD, NP] }
    pool: ap-south
  - match: { continent: EU }
    pool: eu-central           # also keeps EU orders in the EU
  - match: { continent: [NA, SA] }
    split:                     # weighted canary inside one region
      - { pool: us-east,        weight: 95 }
      - { pool: us-east-canary, weight: 5 }
  - default:                   # everyone else, e.g. a Singapore resolver
      by: latency              # lowest measured RTT among healthy pools
      fallback: [ap-south, eu-central, us-east]
on_unhealthy: next_rule        # never answer with a pool that is down

With EDNS Client Subnet

With EDNS Client Subnet, as an ordered list of steps:
With EDNS Client Subnet8 steps between Priya, Guwahati, Public resolver, Singapore site, Authoritative GeoDNS, GeoIP and network map. The steps are listed as text after the diagram.GeoIP and network mapAuthoritative GeoDNSPublic resolver, Singapore sitePriya, Guwahaticache for 100.64.0.0/20 onlystand-in for her ISP's prefixA? shop.example.com1A? shop, ECS 100.64.12.0/242100.64.12.0/24?3IN, Assam; whole /20 in Assam4203.0.113.10, scope /20, TTL 605203.0.113.106
  1. Priya, Guwahati → Public resolver, Singapore site: A? shop.example.com
  2. Public resolver, Singapore site → Authoritative GeoDNS: A? shop, ECS 100.64.12.0/24
  3. Authoritative GeoDNS → GeoIP and network map: 100.64.12.0/24?
  4. GeoIP and network map → Authoritative GeoDNS (reply): IN, Assam; whole /20 in Assam
  5. Authoritative GeoDNS → Public resolver, Singapore site (reply): 203.0.113.10, scope /20, TTL 60
  6. Note over Public resolver, Singapore site: cache for 100.64.0.0/20 only
  7. Public resolver, Singapore site → Priya, Guwahati (reply): 203.0.113.10
  8. Note over Priya, Guwahati: stand-in for her ISP's prefix

Priya's address, as the resolver sends it and as the answer covers it

IPv4 address
  1. 1st octet, 8 bits, Sent to the nameserver, value 100
  2. 2nd octet, 8 bits, Sent to the nameserver, value 64
  3. 3rd high, 4 bits, Sent to the nameserver, value 0
  4. 3rd low, 4 bits, Sent to the nameserver, value 12
  5. 4th octet, 8 bits, Zeroed; never sent, value 77
  • SOURCE /24, from 0 to 24
  • Sent to the nameserver
  • Covered by the answer's scope
  • Zeroed; never sent
Start

As it starts. 2 steps follow.

Her address is shown as 100.64.12.77, a stand-in. The resolver sends only the first 24 bits (RFC 7871 recommends /24 for IPv4 and /56 for IPv6). The nameserver replies with scope /20: the answer is good for any client whose first 20 bits match, which is 16 /24s, so the resolver can reuse it far more often.
Was this section helpful?

In practice.

Seeing the user rather than the resolver pays off, it costs cache hits, and weights behave differently than their numbers suggest.

When the nameserver sees the user, not the resolver (Akamai, SIGCOMM 2015)

Mapping distance
8× shorter
user to the chosen edge, for public-resolver users
RTT and download time
2× lower
same users, after end-user mapping via ECS
Time to first byte
30% better
same users
DNS queries
~8× more
the price, at Akamai's nameservers

What ECS does to one resolver site's cache

Assumptions
Queries for shop.example.com at one public-resolver site
500/sassumption
TTL of the answer
60 s
Distinct client /24s behind that site asking within a minute
5,000assumption
Distinct /20s those /24s fall into
600assumption; at least 5,000 ÷ 16 ≈ 313 if they were packed densely, so 600 means they are scattered
Working
  1. Misses without ECS1 answer per TTL = 1 ÷ 60 s≈ 0.017/s (hit rate ≈ 99.997%)from TTL of the answer and Queries for shop.example.com at one public-resolver site
  2. Misses with scope /24 (worst case)nets ÷ ttl = 5,000 ÷ 60 s; hit rate ≥ 1 − 83 ÷ 500≤ 83/s (hit rate ≥ 83%)from Distinct client /24s behind that site asking within a minute, TTL of the answer and Queries for shop.example.com at one public-resolver site
  3. Misses with scope /20wide ÷ ttl = 600 ÷ 60 s; hit rate ≥ 1 − 10 ÷ 500≤ 10/s (hit rate ≥ 98%)from Distinct /20s those /24s fall into, TTL of the answer and Queries for shop.example.com at one public-resolver site
  4. Authoritative load from this one site, /24 scopes vs none83 ÷ 0.017up to ~5,000× morefrom Misses without ECS and Misses with scope /24 (worst case)
What it means
  • Answer with the widest scope that is still accurate. If a whole /16 goes to Mumbai, say /16; it costs nothing in accuracy and multiplies cache reuse.
  • Use scope 0 for names you do not steer (static assets on one CDN, the API's docs host) so they stay one cache entry per resolver.
  • Size the authoritative tier for the ECS query rate, not the rate you saw before turning ECS on.
CDN request routing
The classic GeoDNS user. A CDN maps each client network to an edge with low measured latency and spare capacity, and re-decides every 20 to 60 seconds through short TTLs. Its latency map is built from probes and from real users' page loads, not from a GeoIP guess.
Managed DNS policies
Route 53 offers geolocation, latency, geoproximity, IP-based, weighted, multivalue and failover policies, and nests them as a tree of alias records: latency picks a region, weighted splits inside it, and a health check can prune any branch.
Resolvers that do not share
Cloudflare's 1.1.1.1 sends no client subnet, for privacy; Google Public DNS does, to nameservers that support it. Design so that a user whose resolver hides them still gets an acceptable, if not the best, site.

Share of Americas traffic on a 5% canary, minute by minute

Share of Americas traffic on a 5% canary, minute by minuteA 95/5 split averages 5% over twenty minutes, but the one minute a big resolver draws the canary sends 24% of the Americas to it.05%10%15%20%25%30%1234567891011121314151617181920weight 5%the big resolver drew the canaryTraffic on canary (%)MinuteShare of Americas traffic on a 5% canary, minute by minuteA 95/5 split averages 5% over twenty minutes, but the one minute a big resolver draws the canary sends 24% of the Americas to it.05%10%15%20%25%30%135791113151719weight 5%the big resolver drew thecanaryTraffic on canary (%)Minute
Illustrative. One resolver site carries 20% of Larkbook's Americas traffic (the NA/SA rule's 95/5 split); many small ones carry the other 80%. TTL is 60 s. Each resolver draws the canary for 5% of its windows, so the small ones add up to about 5% × 80% = 4% every minute, and the big one adds its whole 20% for one minute in twenty on average. The 20 bars average exactly 5%.
Data
MinuteShare on canary (%)
14.2
23.8
34.1
43.9
54.3
63.7
724
84
94.1
103.9
113.8
124.2
134
143.9
154.1
164
173.8
184.2
194.1
203.9
  • weight 5%: Traffic on canary (%) = 5
  • At 7: the big resolver drew the canary

Weights choose between resolvers, not between users. Each resolver gets one answer per TTL and hands it to its whole population, so a large resolver moves thousands of shoppers at once. Over an hour the split converges to 95/5; minute to minute it is lumpy, and a bad canary can hit 24% of a region for a minute. If the split must be exact per user (a 1% experiment, a sticky rollout), make it in an L7 balancer behind one address, where each request or cookie is decided on its own (see Load balancing basics).

Was this section helpful?

Trade-offs.

The chosen option is first; the others stay visible so the reasoning can be checked.

01
How to send users to a nearby site
Chosen:GeoDNS with a latency policy and health gating
  • Pro:Any policy you can write (latency, cost, data location, weights)
  • Pro:Plain unicast address per site; nothing special in the network
  • Pro:Drain a site by changing an answer
Downside we accept:
  • Con:Sees resolvers, not users, unless they send a client subnet
  • Con:Every change waits for cached answers to expire
  • Con:Needs a latency measurement system to be better than a GeoIP guess
Ruled out:Anycast one address from every site

Internet routing decides, not you; "nearby" in BGP is not always fast; Hard to move part of a region's load; A route change mid-connection can land a TCP flow on a site that does not know it (see Global load balancing)

Ruled out:One region plus a CDN for static files

Every dynamic request from India still crosses to Virginia and back; One region is one failure domain

02
Honour EDNS Client Subnet?
Chosen:Yes, answering with the widest accurate scope
  • Pro:Users of distant public resolvers get the right site
  • Pro:Wide scopes keep most of the cache hit rate
Downside we accept:
  • Con:Bigger resolver caches and more authoritative queries
  • Con:The client's /24 is revealed to every authoritative server the resolver sends ECS to (not to root or TLD servers)
Ruled out:No, steer by the resolver's address

Users of far-away public resolvers get far-away sites

What goes wrong

FailureImpactDetectionMitigationMeanwhile
GeoIP places the user in the wrong place4GeoIP and network mapVPN, mobile carrier or corporate egress users are sent to a distant siteReal-user latency per site and per country; a country whose p75 is far above its neighbours'Prefer latency-based rules over pure geography; keep a sensible default; add CIDR overrides for known networksThe shop works, slower for those users
Resolver far from the user and sends no client subnet2Recursive resolversA whole resolver site's users land on the site nearest the resolverQueries arriving without ECS from known public-resolver rangesAccept it, or put the sites behind an anycast front so the last hop is still nearCorrect answers, suboptimal site
No default rule3Authoritative GeoDNSUsers from unmatched countries get an empty answer and cannot reach the shopRising NODATA responses; support tickets from one countryAlways configure a default; test the policy with sample source addresses from every continent
Cached answers outlive a capacity change2Recursive resolversDraining a site keeps sending it traffic for up to a TTL (longer from resolvers that stretch TTLs)Traffic on the drained site after the changeKeep steering TTLs at 60 s or less and drain before maintenance, not during itThe site must keep serving until traffic falls away
Latency policy flaps between two nearly equal sites5Latency mapCaches on both sites stay cold and load swings every refreshThe chosen site for one network changes many times an hourSwitch only when the new site is better by a margin (say 20%) for several measurement rounds
This topic
Steering policies and how they combine
Client subnet (ECS), scopes and their cache cost
Why weights and changes are lumpy
Elsewhere
How long answers live and whyCaching and TTLs
Health checks and failover timingFailover with DNS
Anycast and evacuating a regionload-balancers / Global load balancing
What the edge does with the requestcdn / Edge caching
Was this section helpful?
Next in Core
Failover with DNS
Read next