Single point of failure (SPOF)
SystemsHigh priority~20 minA component whose fault becomes the whole system's failure — the three conditions that make one, the ones that never appear on the diagram, and what redundancy costs to remove them.
Definition
A single point of failure is a component whose fault becomes the system's failure — in Kleppmann's phrasing, a node or network link whose fault leads to failure.
Three conditions hold together, and separating them is what turns there is only one database into an argument. The component is on the critical path of something users need; there is no alternative while it is down; and the impact is unacceptable for as long as recovery takes. Drop one and it is not a SPOF — it is a component with a workaround, or one whose outage nobody notices.
Removing one buys a fault that stays a fault instead of spreading. It costs capacity you pay for and hope never to use, and it costs coordination — two of something must now agree which is authoritative.
So the question is never eliminate every SPOF. It is which to buy out, and which to keep with a recovery time you have measured.
When to use
The first pass is trivial: anywhere the diagram has exactly one box. The second pass is the one that pays — anywhere two boxes depend on the same third thing. That is shared fate, and it is why a pair of replicas fails on the same afternoon.
In a prompt, the cue is any availability target expressed in nines (Availability), any must survive a region going down, and any what happens if X is unreachable. It fires again the moment you add a second of something, because the second one is where coordination enters.
Techniques
The redundancy patterns
Removing a SPOF means standing up a second one and deciding what it does while the first is healthy. That decision is the taxonomy.
Active-active — every replica takes live traffic, so losing one removes capacity rather than function; it needs a tier holding no session state (stateless services). Active-passive (hot standby) — a fully provisioned replica stays in sync and is promoted when the primary is declared gone. Warm standby — provisioned small and lagging, needing a catch-up first. Cold standby — a template, a backup, and a restore.
Failure-domain isolation is the odd one out: rather than duplicate the component, you bound how much of the system one instance of it can take down — cells, zones, regions, per-tenant shards. When a failover fires is load balancers' subject; how the data got there is read replicas'.
| Cost while healthy | Recovery time | What it leaves unfixed | |
|---|---|---|---|
| Active-active | No dedicated idle node, but every node carries N/(N−1) headroom — at two nodes that is a full spare, at ten it is 11% | Seconds — drain the failed node, the rest absorb | Concurrent writes need a consistency story (Consistency models) |
| Active-passive (hot) | A full second copy earning no traffic | Seconds to minutes — promote, then re-point clients | Who is authoritative when the primary is slow rather than dead |
| Warm standby | Small instances plus the replication stream | Minutes to an hour — scale up, catch up, cut over | Writes accepted after the last successful sync |
| Cold standby | Backup storage and a template | Hours — provision, restore, verify | Any objective shorter than a shift |
| Failure-domain isolation | Routing, data placement, headroom held per domain | Unchanged inside the affected domain — only one is affected | The dependency itself; a control plane still spans every domain |
Related concepts
Availability owns the composition arithmetic — hard dependencies multiply, and replicas that fail independently multiply their failure probabilities instead. This page owns what that word costs: independence is a property of the topology, not of the formula, and it holds only as far as nothing is shared.
Reliability owns staying correct while a fault runs. Failure detection (heartbeats) and split brain are ch. 14's; shedding load off a failing dependency is ch. 13's (circuit breakers, bulkheads).
Implementation
Finding them — five passes, in this order
Each pass surfaces a class the one before it cannot see, so the order matters more than the thoroughness.
- 1 · Follow one critical user flow end to end. Sign-in, add-to-cart, checkout — a flow, not the picture. Record every hop, including the ones omitted as infrastructure: name resolution, TLS termination, the token issuer, the flag lookup.
- 2 · Add what each hop reads at boot. Config, service discovery, secrets, the image registry. None sit on the request path; a restart mid-incident puts all of them there.
- 3 · Walk the recovery path too. The pipeline that ships the fix, the VPN, the dashboards, the person who has done this failover before. Meta's October 2021 outage ran aground here: its primary and out-of-band network access went down together, so engineers were sent on site — and the security hardening that made the data centers hard to enter then slowed the recovery.
- 4 · Group the result by shared fate. Same host, rack, zone, region, config push, DNS record, certificate, upstream provider. Anything two components share is a candidate no box shows.
- 5 · Test the assumption rather than asserting it. Take the thing away — a game day, a drained zone — and watch. Meta credits its rehearsed storm drills for bringing traffic back without a second collapse; the practice belongs to Reliability.
Worked examples
Amazon S3, us-east-1, 28 February 2017
AWS's published summary is precise enough to read as an anatomy.
An engineer debugging the S3 billing system ran a playbook command to remove a small number of servers. One input was entered incorrectly and a larger set went — including servers supporting two other subsystems. The index subsystem holds the metadata and location of every object in the region and is required by every GET, LIST, PUT and DELETE. The placement subsystem allocates storage for new objects and needs index to function.
Losing that much capacity forced a full restart of both. AWS notes neither had been completely restarted in its larger regions for many years, and that the metadata-integrity safety checks took longer than expected. From the 9:37 AM PST command: GET, LIST and DELETE resumed at 12:26 PM (2 h 49 min), index fully recovered at 1:18 PM (3 h 41 min), placement at 1:54 PM (4 h 17 min).
The instructive part is not the storage-shaped SPOF; it is the two that no architecture picture carries.
First, the dependency chain: placement needed index, and EC2 instance launches, EBS volumes restoring from snapshot and Lambda needed S3. One subsystem's fault crossed boundaries that looked independent on paper.
Second, the recovery path was inside the failure. AWS could not update the Service Health Dashboard until 11:37 AM — two hours in — because that console itself depended on S3.
Both fixes AWS shipped are worth naming, because between them they are the two moves available. The tool now removes capacity more slowly and refuses a removal that would take a subsystem below its minimum — a floor, which removes the trigger. And partitioning index into cells was brought forward, so a restart is bounded to one cell.
The second move generalises further. Making a shared component genuinely redundant is often impossible or ruinous; making its failures small is usually neither.
Tradeoffs
| What you pay | When it bites | |
|---|---|---|
| Redundancy that shares fate | Full price for the copy, none of the independence the maths assumed — AWS derives the redundant-component formula for independent components only | The instant the shared thing fails, which is the instant you were counting on the copy |
| Two writers where there was one | A coordination problem you did not have, and a promotion decision that can be wrong | During an ambiguous failure, where a slow primary and a dead one look identical (split brain) |
| Failure-domain isolation | You stop being able to treat the fleet as one pool — every cross-domain operation becomes a design decision | When one request needs data from two domains — and when a control plane quietly spans all of them |
Which brings the decision this page is for: which SPOFs you keep. A single strongly-consistent primary, a single region, a single payment provider — each can be the right call, and the defence is never it is redundant. It is a stated recovery time, a rehearsed failover, and a bounded blast radius.
A SPOF with a measured recovery time is a risk someone accepted. One nobody has priced is a surprise waiting for a Tuesday.
Things to look out for
- Counting boxes instead of dependencies. Two app servers reading one config service are one component with a spare front end — and the second box is what makes the picture look safe.
- A failover that has never been fired. Standby capacity that never served production traffic is untested: the first promotion is where you find the full disk or the lapsed certificate.
- Failing over via DNS and forgetting the TTL. Resolvers cache past the value you set, so a record swap is a recovery time you do not control (DNS).
- A redundant data plane behind a singleton control plane. Meta's 2021 outage took down a globally redundant backbone with one command, because the audit tool meant to reject it had a bug.
- One person who has done the failover. A runbook that only works in its author's hands is a SPOF with a holiday calendar.
In the interview
- Where are the single points of failure in what you've drawn? Walk the request path, then the recovery path — the second walk is where the answer stops being the database. Name one thing that is not on the diagram: config, the deploy pipeline, a certificate, the status page.
- You've added a second load balancer. Are you done? The follow-up is shared fate: same zone, same config push, same DNS record, at what TTL? Then name the problem the second one introduced — deciding who is authoritative.
- What's your recovery time if that database dies? Two numbers, not one: time to detect, then time to promote. Thirty seconds plus four minutes is defensible; it fails over automatically is not.
- Price the removal before proposing it. Tie it to the target with the composition arithmetic on Availability, then say which SPOFs you keep and what recovery time you accept for each.