GeoDNS and traffic steering.
Larkbook runs its shop in Mumbai, Frankfurt and Virginia. One name, three addresses, and the nameserver picks which one to hand out based on where the question came from. The catch is that the question comes from a resolver, not from the shopper, and the answer is cached for everyone behind that resolver.
Builds on Caching and TTLs.
The idea.
One name, several sites, and a nameserver that picks a site for each question it is asked.
Larkbook serves shop.example.com from three sites: Mumbai (203.0.113.10), Frankfurt (198.51.100.20) and Virginia (192.0.2.30). A shopper in Pune should reach Mumbai, about 120 km away, not Virginia, about 13,000 km away. Light in fibre covers roughly 200 km per millisecond, so that distance alone adds about 130 ms to every round trip to Virginia (13,000 km each way at 200 km/ms is 65 ms, twice), and a page load needs several round trips. Larkbook also wants to stop sending anyone to a site that is failing or being drained, and to send less to the smallest site.
There are two places to make that choice. Anycast makes it in the network: every site announces the same address and internet routing delivers each packet to a nearby one. GeoDNS makes it in the answer: every site keeps its own ordinary address, and the authoritative nameserver returns the address it wants this asker to use. Because the decision is a lookup in software, the policy can be anything: country, measured latency, fixed weights, cost, legal location of data, or site health.
Two limits come with it, and the rest of this topic is about living with them. The nameserver never talks to the shopper; it talks to the shopper's recursive resolver (the ISP or public service that looks names up on the shopper's behalf and caches the answers), so it steers by where the resolver is unless the resolver says more. And the answer is cached by that resolver for its TTL (see Caching and TTLs), so every user behind it gets the same answer until it expires, and no change of policy takes effect faster than that.
Who gets which site under a plain country policy (IN → Mumbai, EU → Frankfurt, US → Virginia, anything else → Virginia)
| Shopper | Asks through | Nameserver sees | Answer | Right site? |
|---|---|---|---|---|
| Ananya, Pune | her ISP's resolver in Mumbai | an Indian address | 203.0.113.10 | Yes — Mumbai |
| Lukas, Berlin | his ISP's resolver in Frankfurt | a German address | 198.51.100.20 | Yes — Frankfurt |
| Maya, Ohio | a public resolver's Chicago site | a US address | 192.0.2.30 | Yes — Virginia |
| Priya, Guwahati | a public resolver's Singapore site with no client subnet | a Singapore address | 192.0.2.30 | No — Singapore matched no rule, so the default sent her to Virginia |
| Tomás, Lisbon | his company's VPN resolver in Virginia | a US address | 192.0.2.30 | No — the nameserver never saw Lisbon |
How it works.
The nameserver joins three inputs (where the asker is, how fast each site is from there, and which sites are fit to serve) into one answer, and says how widely that answer may be reused.
GeoDNS for shop.example.com
What happens on each query
- Work out who is asking: the client subnet if the query carries one, otherwise the resolver's own address.
- Look that prefix up in the network map: country, city, ISP. This is a database lookup, not a measurement, and it is only as good as the database.
- Walk the policy's rules in order and take the first pool that matches (India → ap-south, say).
- Drop sites the health feed marks as down; if the pool is empty, fall through to the next choice instead of answering with a dead address.
- Answer with the chosen address, a TTL, and (when the query had a client subnet) the widest prefix the answer is valid for.
Steering policies
| Policy | Decides by | Good for | Watch out |
|---|---|---|---|
| Geolocation | Country or continent of the source, with a default for everything else | Data residency, licensing, language | Wrong for VPN and roaming users; with no default, unmatched places get no answer at all |
| Latency-based | Measured round-trip time from the source's network to each site | Raw performance | Needs a measurement system; can flap between two sites that are nearly equal |
| Weighted | Fixed proportions (weight ÷ sum of weights) | Canaries, gradual migration between sites | The split is per resolver per TTL window, not per user |
| Geoproximity with bias | Distance to each site, with a bias that grows or shrinks a site's area | Shifting load off a hot region | Coarse; one bias moves whole regions |
| IP / CIDR map | Explicit prefix → site lists | Large ISPs, partners, known offices | You maintain the list and it goes stale |
| Multivalue | Several healthy addresses at random (up to 8 in Route 53) | Letting the client pick and retry | Not a load balancer: random per response and per resolver, with no notion of capacity |
record: shop.example.com
type: A
ttl: 60 # was 300 while shop had one site; steered names stay short
pools:
ap-south: { address: 203.0.113.10, health_check: https-healthz }
eu-central: { address: 198.51.100.20, health_check: https-healthz }
us-east: { address: 192.0.2.30, health_check: https-healthz }
us-east-canary: { address: 192.0.2.31, health_check: https-healthz }
rules: # first match wins
- match: { country: [IN, LK, BD, NP] }
pool: ap-south
- match: { continent: EU }
pool: eu-central # also keeps EU orders in the EU
- match: { continent: [NA, SA] }
split: # weighted canary inside one region
- { pool: us-east, weight: 95 }
- { pool: us-east-canary, weight: 5 }
- default: # everyone else, e.g. a Singapore resolver
by: latency # lowest measured RTT among healthy pools
fallback: [ap-south, eu-central, us-east]
on_unhealthy: next_rule # never answer with a pool that is down
With EDNS Client Subnet
- Priya, Guwahati → Public resolver, Singapore site: A? shop.example.com
- Public resolver, Singapore site → Authoritative GeoDNS: A? shop, ECS 100.64.12.0/24
- Authoritative GeoDNS → GeoIP and network map: 100.64.12.0/24?
- GeoIP and network map → Authoritative GeoDNS (reply): IN, Assam; whole /20 in Assam
- Authoritative GeoDNS → Public resolver, Singapore site (reply): 203.0.113.10, scope /20, TTL 60
- Note over Public resolver, Singapore site: cache for 100.64.0.0/20 only
- Public resolver, Singapore site → Priya, Guwahati (reply): 203.0.113.10
- Note over Priya, Guwahati: stand-in for her ISP's prefix
Priya's address, as the resolver sends it and as the answer covers it
- 1st octet, 8 bits, Sent to the nameserver, value 100
- 2nd octet, 8 bits, Sent to the nameserver, value 64
- 3rd high, 4 bits, Sent to the nameserver, value 0
- 3rd low, 4 bits, Sent to the nameserver, value 12
- 4th octet, 8 bits, Zeroed; never sent, value 77
- SOURCE /24, from 0 to 24
- Sent to the nameserver
- Covered by the answer's scope
- Zeroed; never sent
As it starts. 2 steps follow.
In practice.
Seeing the user rather than the resolver pays off, it costs cache hits, and weights behave differently than their numbers suggest.
When the nameserver sees the user, not the resolver (Akamai, SIGCOMM 2015)
What ECS does to one resolver site's cache
- Queries for shop.example.com at one public-resolver site
- 500/sassumption
- TTL of the answer
- 60 s
- Distinct client /24s behind that site asking within a minute
- 5,000assumption
- Distinct /20s those /24s fall into
- 600assumption; at least 5,000 ÷ 16 ≈ 313 if they were packed densely, so 600 means they are scattered
- Misses without ECS1 answer per TTL = 1 ÷ 60 s≈ 0.017/s (hit rate ≈ 99.997%)from TTL of the answer and Queries for shop.example.com at one public-resolver site
- Misses with scope /24 (worst case)nets ÷ ttl = 5,000 ÷ 60 s; hit rate ≥ 1 − 83 ÷ 500≤ 83/s (hit rate ≥ 83%)from Distinct client /24s behind that site asking within a minute, TTL of the answer and Queries for shop.example.com at one public-resolver site
- Misses with scope /20wide ÷ ttl = 600 ÷ 60 s; hit rate ≥ 1 − 10 ÷ 500≤ 10/s (hit rate ≥ 98%)from Distinct /20s those /24s fall into, TTL of the answer and Queries for shop.example.com at one public-resolver site
- Authoritative load from this one site, /24 scopes vs none83 ÷ 0.017up to ~5,000× morefrom Misses without ECS and Misses with scope /24 (worst case)
- Answer with the widest scope that is still accurate. If a whole /16 goes to Mumbai, say /16; it costs nothing in accuracy and multiplies cache reuse.
- Use scope 0 for names you do not steer (static assets on one CDN, the API's docs host) so they stay one cache entry per resolver.
- Size the authoritative tier for the ECS query rate, not the rate you saw before turning ECS on.
Share of Americas traffic on a 5% canary, minute by minute
Data
| Minute | Share on canary (%) |
|---|---|
| 1 | 4.2 |
| 2 | 3.8 |
| 3 | 4.1 |
| 4 | 3.9 |
| 5 | 4.3 |
| 6 | 3.7 |
| 7 | 24 |
| 8 | 4 |
| 9 | 4.1 |
| 10 | 3.9 |
| 11 | 3.8 |
| 12 | 4.2 |
| 13 | 4 |
| 14 | 3.9 |
| 15 | 4.1 |
| 16 | 4 |
| 17 | 3.8 |
| 18 | 4.2 |
| 19 | 4.1 |
| 20 | 3.9 |
- weight 5%: Traffic on canary (%) = 5
- At 7: the big resolver drew the canary
Weights choose between resolvers, not between users. Each resolver gets one answer per TTL and hands it to its whole population, so a large resolver moves thousands of shoppers at once. Over an hour the split converges to 95/5; minute to minute it is lumpy, and a bad canary can hit 24% of a region for a minute. If the split must be exact per user (a 1% experiment, a sticky rollout), make it in an L7 balancer behind one address, where each request or cookie is decided on its own (see Load balancing basics).
Trade-offs.
The chosen option is first; the others stay visible so the reasoning can be checked.
- Pro:Any policy you can write (latency, cost, data location, weights)
- Pro:Plain unicast address per site; nothing special in the network
- Pro:Drain a site by changing an answer
- Con:Sees resolvers, not users, unless they send a client subnet
- Con:Every change waits for cached answers to expire
- Con:Needs a latency measurement system to be better than a GeoIP guess
Internet routing decides, not you; "nearby" in BGP is not always fast; Hard to move part of a region's load; A route change mid-connection can land a TCP flow on a site that does not know it (see Global load balancing)
Every dynamic request from India still crosses to Virginia and back; One region is one failure domain
- Pro:Users of distant public resolvers get the right site
- Pro:Wide scopes keep most of the cache hit rate
- Con:Bigger resolver caches and more authoritative queries
- Con:The client's /24 is revealed to every authoritative server the resolver sends ECS to (not to root or TLD servers)
Users of far-away public resolvers get far-away sites
What goes wrong
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| GeoIP places the user in the wrong place4GeoIP and network map | VPN, mobile carrier or corporate egress users are sent to a distant site | Real-user latency per site and per country; a country whose p75 is far above its neighbours' | Prefer latency-based rules over pure geography; keep a sensible default; add CIDR overrides for known networks | The shop works, slower for those users |
| Resolver far from the user and sends no client subnet2Recursive resolvers | A whole resolver site's users land on the site nearest the resolver | Queries arriving without ECS from known public-resolver ranges | Accept it, or put the sites behind an anycast front so the last hop is still near | Correct answers, suboptimal site |
| No default rule3Authoritative GeoDNS | Users from unmatched countries get an empty answer and cannot reach the shop | Rising NODATA responses; support tickets from one country | Always configure a default; test the policy with sample source addresses from every continent | |
| Cached answers outlive a capacity change2Recursive resolvers | Draining a site keeps sending it traffic for up to a TTL (longer from resolvers that stretch TTLs) | Traffic on the drained site after the change | Keep steering TTLs at 60 s or less and drain before maintenance, not during it | The site must keep serving until traffic falls away |
| Latency policy flaps between two nearly equal sites5Latency map | Caches on both sites stay cold and load swings every refresh | The chosen site for one network changes many times an hour | Switch only when the new site is better by a margin (say 20%) for several measurement rounds |