/Interview Study Guide/System design
ConceptsPart of Scalability

Consistency models

SystemsHigh priority~30 min

The contract for which orderings of reads and writes a store may expose — from linearizable at the top to eventual at the bottom, and the price of each rung.

Definition

A consistency model is a contract about which histories of reads and writes a storage system is allowed to expose. Jepsen frames it as a safety property — it says nothing about what the system will do, only about what it may never do.

What is system design? promised this page would settle whether a read is guaranteed to see the most recent write. That guarantee has a name — linearizability — and it is one end of a spectrum with a dozen named points on it, most of which are cheaper and none of which is wrong.

The cost runs the length of that spectrum. A stronger model forbids more anomalies, and it buys that by making nodes agree before answering: an extra round trip to a leader or a quorum on the hot path, and no answer at all while the nodes that must agree cannot reach each other. A weaker model answers from whichever replica is nearest and keeps answering when the network breaks — and charges you in application code that has to tolerate, or reconcile, the anomalies it now permits.

None of this is a property of a database. It is a property of an operation, and in real stores it is usually a flag with a price on it.

When to use

Reach for this vocabulary the moment a design has replicas and a read path that can land somewhere other than where the write did — a follower, a cache, a second region.

Then ask the question per data class — CAP theorem owns why the call is never per system — and ask it as: would a user notice the staleness, and would it be wrong? A like count three seconds behind is nobody's incident. A comment missing from the page immediately after you posted it is noticed within one second, by the one user who is certain they are right. Those two answers buy different machinery, at very different prices.

Techniques

The spectrum, strongest first

The named models form a lattice rather than a menu. From strict serializability down to PRAM each really does imply the one beneath, so a system that hands you the top has already handed you those. Below PRAM it branches: read-your-writes and monotonic reads are siblings, and a system can have either without the other.

Strict serializability sits at the top, and Jepsen's map draws it as the join of two otherwise disjoint families — multi-object transaction models on one side (serializable, snapshot isolation, read committed) and single-object consistency models on the other. That split is why serializable and linearizable are not synonyms and why a store can advertise one without the other. Isolation levels belong to ACID transactions; this page walks the single-object branch down.

Linearizable requires every operation to appear to take effect instantaneously, in an order consistent with real time. Sequential keeps the single agreed order but drops the real-time tie, so a process may sit arbitrarily far behind as long as nobody disagrees about the sequence. Causal keeps only the orderings that could be cause and effect, and lets independent operations be observed in different orders by different clients.

Below causal the guarantees stop describing the system and start describing one client's session. Two of the four get rows below. Monotonic writes means a client's own writes land in the order it issued them, and writes-follow-reads means a write you make after reading something lands after what you read. Jepsen bundles the first, second and third of those four as PRAM. At the floor, eventual consistency promises only convergence — replicas agree once the writes stop arriving, and says nothing about any read before then.

ModelWhat it forbidsAvailability under a partitionWhere you meet it
Strict serializableAny history not equivalent to running whole transactions one at a time, in real-time orderUnavailableSpanner, whose external consistency Google documents as its strictest guarantee
LinearizableReturning a value older than the last completed write to that objectUnavailableDynamoDB reads with ConsistentRead: true; an S3 GET after a PUT
SequentialTwo processes disagreeing about the order of the same operationsUnavailableChiefly a memory-model term — Lamport defined it for multiprocessors, not for stores
CausalObserving an effect before its cause — the reply before the questionSticky availableMongoDB causally consistent sessions
Read-your-writes (and PRAM above it)A client failing to observe its own completed writeSticky availableSession modes: MongoDB sessions, Cosmos DB's Session level
Monotonic readsOne client's successive reads moving backwards in timeTotally availableRarely sold alone; arrives bundled inside a session guarantee
EventualNothing about any individual read — only permanent divergence after writes stopTotally availableDynamoDB's default read mode; S3 bucket configuration
Availability classes for the named models are Jepsen's, for an asynchronous network: unavailable = some nodes cannot make progress; sticky available = every client attached to a healthy replica can, so long as it never moves; totally available = every client can, always. Eventual consistency is not on Jepsen's map — its class is the trivial one.

Related concepts

The models are this page's; the ground around them belongs to neighbours. Why a partition forces the choice at all — and why the latency half of the bill arrives on healthy days too — is CAP theorem. Once the vocabulary is in hand, the decision procedure is strong vs eventual consistency. When replicas do diverge, the machinery that reconciles them is vector clocks and CRDTs; how the data reached a second node at all is read replicas.

Worked examples

One shopping cart, three models

Amazon's cart is the canonical case because the Dynamo paper argues the requirement out loud: an add to cart must never be rejected, so the store is designed to be always writeable. A partition therefore produces two live carts, and the merge happens at read time (CAP theorem walks that choice and prices it). What matters here is the model it lands on: eventual, with the anomaly pushed into the merge rule. Nobody waits during a partition; someone writes and owns that rule, and it is lossy in one direction.

Now the same cart with only a session guarantee. The customer adds an item and immediately reloads the page; the reload is load-balanced onto a replica that hasn't received the write yet.

ClientPrimaryReplicaPOST /cart — add item200 OKreplicateasync — arrives after the read belowGET /cartbalancer picked the nearest replicacart without the itemGET /cart — session pinnedfix: route reads to the primary for a few seconds after a write
A read-your-writes violation, and the cheapest fix for it.

The fix is routing, not a stronger database — and the guarantee is configuration-shaped, so it is easy to believe you have it when you don't. MongoDB's docs are explicit: a causally consistent session delivers all four session guarantees with durability only with read concern majority and write concern majority; drop the read concern to local and read-own-writes goes with it.

Pinning is not free either. The pin has to outlast the replication lag to be worth anything, and every pinned read lands back on the primary — the node the replicas were added to spare. A session guarantee is cheap because it is narrow: Jepsen is explicit that it covers one client's own view and promises nothing about what a second client sees of that write.

Third version: pay for the strong read. AWS documents DynamoDB's eventually consistent reads as the default, and as half the cost of strongly consistent ones. So ConsistentRead: true is a 2× read bill — and it isn't offered everywhere: AWS supports strongly consistent reads on tables and local secondary indexes, but not on global secondary indexes or streams.

S3 makes the same choice per resource rather than per request. AWS documents strong read-after-write for PUT and DELETE of objects in all Regions, with concurrent writes to one key resolved by latest timestamp. Bucket configuration, though, stays eventually consistent — the docs suggest waiting 15 minutes after enabling versioning before writing.

Three models, one cart, and the model was never a property of the database. Azure Cosmos DB makes that literal: Microsoft documents five named levels, set as an account default and overridable on an individual request.

Tradeoffs

What you payWhen the bill arrives
Linearizable readsA round trip to the leader or a quorum on every read — and in DynamoDB, twice the read chargeOn read-heavy and cross-region paths, where that hop is the p99; and when the read path turns out to be a secondary index, where AWS doesn't offer the flag at all
Causal and session guaranteesStickiness — they hold only while the client keeps talking to the same replicaThe moment a load balancer re-routes a client mid-session, or the pinned replica restarts
Eventual consistencyA merge rule, written by you, in application codeThe first time two clients write one key concurrently. Last-write-wins is a merge rule, and under clock skew it silently discards the loser
Any weak modelNo staleness bound unless you measure one yourselfEventually carries no deadline: AWS describes DynamoDB global-table replication as typically within a second — a typical, not a guarantee. Alert on replication lag

One of these is routinely mispriced. The latency of a strong read is not a partition-time cost — it is charged on every healthy day, which is the argument CAP theorem carries and this page won't repeat.

Read the table as a bill rather than a menu. Every row is paid in a different currency — money on the first, routing discipline on the second, engineering time on the third, and a monitoring obligation on the fourth — which is why they are so rarely traded against each other honestly.

Things to look out for

  • Validating on an idle cluster. The staleness window is milliseconds when nothing is happening and unbounded under load, failover or partition — the anomaly is always available, only the window changes. Reproduce it with replication lag and with fault injection, not one or the other.
  • Believing you have a session guarantee because you opened a session. MongoDB's causal-consistency matrix is the cautionary table: read concern local with write concern majority gives you monotonic writes and none of the other three. The guarantee comes from the concern pair, not from the session object.
  • Assuming a session guarantee protects the second reader. Read-your-writes is scoped to the process that did the write; Jepsen is explicit that another client reading the same key gets no promise at all. Two users watching the same page are two sessions.

In the interview

  • Can this read be stale, and for how long? Answer per data class, not per system. The ledger balance is linearizable and pays the leader round trip; the timeline is eventual and a few seconds behind is fine; the user's own post has to be read-your-writes, so pin that session to the primary briefly. Naming read-your-writes is the signal — most answers stop at strong and eventual.
  • Users say the like count jumps around. What's happening? Monotonic reads violated by routing successive reads to replicas at different points in the log. Fixes in cost order: sticky routing, a session token carrying the last-read position, or serving the counter from one authority.
  • You've said the system is AP. What does the application now owe you? A merge rule, and a named one: last-write-wins, union-merge, or a CRDT whose merge is commutative so a retried delivery changes nothing. The follow-up is usually what happens to a delete, and Dynamo's resurfacing cart item is the answer worth having ready.

Learning resources