/Interview Study Guide/System design
ConceptsPart of Load balancers

DNS load balancing

SystemsMid priority~15 min

Balancing by handing out different answers — the decision is made before the connection exists, cached by resolvers you do not own, and every property follows from that.

Definition

DNS load balancing spreads traffic by answering one name with different addresses. A resolver asks for api.example.com, the authoritative server picks records out of a set, and the client connects to whatever it was handed. Nothing sits between client and server: the distribution happens during name resolution, before the connection exists.

That buys cross-region and cross-datacenter spread for the price of a zone file, with no new component on the request path to fail. It costs control. The answer is cached by a chain of resolvers you do not own, for a duration you can only ask for.

RFC 1035 defines the TTL as the interval a record may be cached before the source should again be consulted — a ceiling rather than a schedule, and §7.3 lets a resolver clamp one it judges too long. RFC 8767 tightens the second half to a MUST reconsult, then licenses one exception: a record whose refresh fails may be served as though it were unexpired.

When to use

The cue is a design that has to split traffic across regions or datacenters, or place a client before it can reach any load balancer at all. Name resolution is the first hop of every cold request, so no other layer can act earlier — and it is the cheapest distribution available: no capacity, no new box, no per-request cost.

Wrong page when the choice is between healthy instances inside one datacenter — that is the load balancer's job, decided per request with health information DNS never sees. Wrong page too when the requirement is stated in seconds of failover, because the cache chain below sets a floor you cannot bid under.

Techniques

Five ways to pick the answer

Every variant answers one question — which records go in this response — and they differ in what gets consulted first. AWS documents each of the five below as a Route 53 routing policy, which makes that a usable vocabulary.

Round-robin returns several records for one name and rotates or randomises among them; Route 53's multivalue-answer policy is the health-aware form, answering with up to eight healthy records chosen at random. Weighted attaches a share to each record, so a 90/10 pair sends a tenth of new resolutions at a new stack. Latency-based answers with the endpoint whose measured latency to the querying network is lowest. Geolocation maps the query's source to a region you enumerated and answers with that region's record. Failover pairs a primary record with a secondary and swaps them when the primary's health check goes unhealthy.

Needs to knowThe split is a share ofOperational cost
Round-robin / multivalueNothing beyond the record set — plus per-record health, in the multivalue formCache entries, not requestsZone edits
WeightedA share you assign per recordResolutions made during the window the weights are liveZone edits; every re-weighting is another one
Latency-basedLatency measured between networks and endpoints — AWS notes its data covers traffic between users and AWS data centers onlyQuerying networks, rankedThe provider's measurement network, which you rent rather than run
GeolocationThe source address of the query, mapped to a placeThe regions you enumeratedA region map plus a default record — without one, unmapped queries get no answer
FailoverOne health check's verdict on the primaryNothing — it is all-or-nothing for every clientA prober per endpoint, billed per check
The axis is what each policy has to know, and what the resulting split is actually a share of.

Related concepts

DNS owns the protocol itself — records, delegation, how a name resolves. This page owns one use of it and what caching does to that use.

DNS gets a client to a site; a load balancer then picks the machine, with the algorithm it picks by a page of its own. Anycast reaches the same goal in the routing layer, with no per-client answer to cache, and a CDN edge is often both in layers.

Worked examples

Withdrawing a datacenter

Two sites, api.example.com carrying one A record for each, a 60-second TTL, and a health check per endpoint. All three of those are chosen values, not defaults. The east site loses its database; follow what the withdrawal actually reaches.

The health checker notices first, and only after enough consecutive failures to be confident. The authoritative server then stops putting the east address in new answers. Every resolver already holding the old response keeps serving it to its own clients until its own copy expires — and so does every OS and every browser process below it. Traffic to the dead site does not stop. It decays.

BrowserOS resolverRecursive resolverAuthoritative NSEast sitecheck the in-process cacheA hit never leaves the process — your nameservers see no query at allresolve api.example.comcheck the OS cachequery A api.example.comcheck the cacheOn a hit it answers with whatever is left of the TTL and stops herequery A api.example.com — on a miss onlyA 203.0.113.10, TTL 60The routing policy runs here, once, for this resolverA 203.0.113.10, TTL counting down203.0.113.10TCP + TLS to 203.0.113.10The balancing decision is already spent; this hop consults nothinghealth check fails — pull the A recordTakes effect on the next cache miss, not the next requestreconnect — still 203.0.113.10Three caches upstream still hold it, each on its own clock
One lookup, three independent caches. The withdrawal happens at the right-hand edge and reaches the left-hand edge only as each cache runs its own copy down.
QuantityValueDerivation
Health checker declares east down90 s30 s request interval × 3 consecutive failures. AWS documents intervals of 10 s or 30 s and a consecutive-failure threshold; both values chosen here. Modelled as one prober — Route 53 in fact runs uncoordinated checkers and calls an endpoint unhealthy when 18% or fewer report it healthy
A TTL-honouring resolver stops handing out the old address≤ 150 s after the fault90 s detection + the record's 60 s TTL, assuming the resolver expires on schedule
OS and browser caches below it expirelater, and unmeasuredEach holds its own copy on its own clock and reports to nobody; your nameserver's query volume cannot see them
Expired answer retained when your nameservers are unreachable1–3 daysRFC 8767's suggested maximum stale timer. It applies to a failed refresh, not to a successful withdrawal — the tail case, not the normal one
The two clocks a DNS failover runs on, and the tail underneath them.

Two and a half minutes is the floor, and it holds only where every cache honours the number. The shape is the lesson: a DNS failover is a decay curve with a long tail.

Tradeoffs

What you payWhen it bites
Withdrawal is advisoryThe endpoint has to keep answering after you remove its record, so hardware you planned to switch off stays on and the decommission date is not yours to setEvery planned drain, and every incident where the failing site is still healthy enough to accept a connection
Short TTLs for agilityA full resolution round trip before the first byte for every cold client, and the query volume that buys the agility lands on your nameserversClients on high-latency links, and pages that resolve several names before anything renders
One resolver, one decisionA resolver serving a million clients takes one answer for all of them, so an even record set does not produce an even loadPopulations concentrated behind a few carrier, corporate or public resolvers
Geo and latency policiesThey read the query's source address, so you are routing the resolver — AWS documents falling back to the resolver's location whenever it does not send edns-client-subnetAny population on a centralised public resolver, or a corporate resolver in another continent
Health-checked recordsA verdict formed outside your network, by checkers that see only what the public internet seesAn endpoint reachable from your datacenter but not from the checkers' networks — and what a shallow probe proves at all is the balancer's subject

Which settles the decision: DNS is a distribution mechanism that can also fail over, not a failover mechanism. Use it to place clients near a site and to shift shares deliberately, and put something in the request path behind it — an in-datacenter load balancer, or anycast one layer down — to cover the seconds the cache chain owns.

Things to look out for

  • Planning a cutover around the TTL you set. RFC 1034 recommends reducing the TTL ahead of an anticipated change and restoring it afterwards — do that, and still read the number as the earliest an answer can change rather than the latest.
  • Treating your own TTL as the ceiling. It binds in one direction only: RFC 1035 §7.3 tells a resolver it may discard or clamp a TTL it judges excessive, and RFC 8767 recommends capping retention at a week. A long TTL is no more enforceable than a short one.
  • Expecting an even split out of a multi-record answer. RFC 1034 says record order within a set is not significant and need not be preserved, and that a resolver library may sort the addresses or return only the best one. Multiple records buy spread, not shares.
  • Measuring a withdrawal at the nameserver. Query volume there describes recursive resolvers and nothing nearer the user, so watch the dead endpoint's own connection count instead — it is the only place the long tail is visible.

In the interview

  • How would you fail traffic over to another region? Give two clocks with a number on each — detection at the prober, then expiry through the cache chain — and say which one you can shrink. Quoting a TTL as the recovery time is the answer that gets pushed on.
  • What TTL would you pick? Treat the number as a bid rather than a setting — 60 seconds for records you expect to move. Then name what a permanently low value costs, because that is the half most answers skip.
  • Why not just use DNS round-robin instead of a load balancer? Because the answer is chosen once per cache entry rather than once per request, and the response carries no live health signal.
  • You've put geolocation in front of three regions — what happens when one dies? A region map has no health input at all, so it keeps answering with the dead region's address. The depth answer composes: geolocation for placement, a failover pair or health-checked records underneath it for liveness.

Learning resources