Consistency modelsStrong and eventual consistency

100%

Strong and eventual consistency.

Copy data to three regions and a question appears that one database never had to answer. When someone reads, which copy's answer counts? Strong consistency makes every copy act like one machine and pays for it in waiting. Eventual consistency answers from the nearest copy at once and lets the copies catch up afterwards. Most real systems use both, choosing per operation.

Beginner19 minUpdated 30 Sept 2026

The idea.

One database never has to decide which copy is right. Three copies do.

One rename, three copies

Tandoor, a recipe-sharing app, keeps every recipe on three replicas: Mumbai, Frankfurt and Virginia. Priya renames her recipe from "Dal Makhani" to "Dal Makhani (no cream)". Twenty milliseconds after her phone shows "Saved", a reader in Virginia opens the recipe. Which name should they see?

Under strong consistency the answer is fixed: the new name, exactly as if there were one copy. Under eventual consistency the reader may still see the old name for a short while, called the inconsistency window. The only promise is that once edits stop, every replica ends up holding the same value. Neither answer is wrong. They are two different contracts, and the difference is paid for in waiting on one side and in surprises on the other.

The same rename, under two promises

Scenario 1 of 2: As described.

Notes
  • 1 Returns at 122 ms, once a majority has entry 1042.
  • 7 Stored, but Virginia does not yet know it is committed.
  • 10 Started 20 ms after the save returned, so it must see the rename.
  • 13 Without a lease, Mumbai first confirms it is still leader with a heartbeat round to Frankfurt (~120 ms).
Timeline as a list

The same rename, under two promises: 5 lanes, from 0 ms to 360 ms.

  1. 0–122 ms · Priya (Mumbai) · rename (ok) — Returns at 122 ms, once a majority has entry 1042.
  2. 2–122 ms · Mumbai (leader) · waiting for a majority
  3. 2 ms · Mumbai (leader) → Frankfurt, arriving 62 ms: append 1042
  4. 2 ms · Mumbai (leader) → Virginia, arriving 97 ms
  5. 62 ms · Frankfurt · has 1042
  6. 62 ms · Frankfurt → Mumbai (leader), arriving 122 ms: ack
  7. 97 ms · Virginia · has 1042 — Stored, but Virginia does not yet know it is committed.
  8. 122 ms · Mumbai (leader) · commit 1042 (ok)
  9. 142–332 ms · Virginia · confirm commit point
  10. 142–332 ms · Reader (Virginia) · read → new name (ok) — Started 20 ms after the save returned, so it must see the rename.
  11. 142 ms · all lanes · Strong: read after save returns
  12. 142 ms · Virginia → Mumbai (leader), arriving 237 ms: committed index?
  13. 237 ms · Mumbai (leader) · answers under its lease — Without a lease, Mumbai first confirms it is still leader with a heartbeat round to Frankfurt (~120 ms).
  14. 237 ms · Mumbai (leader) → Virginia, arriving 332 ms: ≥ 1042
The first tab is the strong promise: Priya's save returns only once Frankfurt has the entry (a majority of 2 of 3), and the Virginia replica checks the commit point with Mumbai before it answers. Mumbai is assumed to hold a leader lease, so it replies at once; without one it would first confirm its leadership with a heartbeat round to Frankfurt, adding about 120 ms. The Eventual tab replays the same rename: the save returns after Mumbai's local write, and a read inside the shaded window gets the old name. Latencies are illustrative assumptions.

Four names you will meet

Linearizable
Forbids missing any write that finished before your read started. The system behaves like one copy.
Sequential
Forbids clients disagreeing on the order of writes, but a read may lag behind real time.
Causal
Forbids seeing an effect before its cause: a recipe step never shows without the ingredient it was written for.
Eventual
Forbids only permanent disagreement. Once writes stop, every replica converges.
Was this section helpful?

How it works.

The two promises come from two different replication paths. Everything between them is a question of which orderings a reader may observe.

How a strong write and read work

Strong rename, then a strong read from Virginia

Strong rename, then a strong read from Virginia, as an ordered list of steps:
Strong rename, then a strong read from Virginia12 steps between Priya's phone, Mumbai replica (leader), Frankfurt replica, Virginia replica, Reader in Virginia. The steps are listed as text after the diagram.Reader in VirginiaVirginia replicaFrankfurt replicaMumbai replica (leader)Priya's phone2 of 3 replicas have 1042: commitlease held: answers at once (no lease: +~120 ms)apply up to 1042, then answerrename(recipe 812, …)1append entry 10422append entry 10423ack 10424ok, version 10425get(recipe 812)6commit index? (read index)7commit index ≥ 10428"Dal Makhani (no cream)"9
  1. Priya's phone → Mumbai replica (leader): rename(recipe 812, …)
  2. Mumbai replica (leader) → Frankfurt replica: append entry 1042
  3. Mumbai replica (leader) → Virginia replica: append entry 1042
  4. Frankfurt replica → Mumbai replica (leader) (reply): ack 1042
  5. Note over Mumbai replica (leader): 2 of 3 replicas have 1042: commit
  6. Mumbai replica (leader) → Priya's phone (reply): ok, version 1042
  7. Reader in Virginia → Virginia replica: get(recipe 812)
  8. Virginia replica → Mumbai replica (leader): commit index? (read index)
  9. Note over Mumbai replica (leader): lease held: answers at once (no lease: +~120 ms)
  10. Mumbai replica (leader) → Virginia replica (reply): commit index ≥ 1042
  11. Note over Virginia replica: apply up to 1042, then answer
  12. Virginia replica → Reader in Virginia (reply): "Dal Makhani (no cream)"

The write waits for a majority, not for all three: two of three copies surviving any single failure is what makes the commit durable. The read is the subtle half. Virginia may already hold entry 1042 without knowing it is committed, or may be missing entries it has never heard of. Before it answers a linearizable read it asks the leader for the current commit index and waits until it has applied that far (Raft calls this a read index). The leader has its own doubt: a newer leader may have been elected without it hearing. So before it hands out the index it either confirms it is still leader with a heartbeat round to a majority (Mumbai to Frankfurt, another ~120 ms), or it holds a time-bounded leader lease that rules out a rival and answers at once. Tandoor's numbers assume the lease. Sending every strong read straight to the leader faces the same choice. Each option costs a cross-region round trip or a dependence on bounded clock drift; none is free.

How an eventual write and read work

Eventual rename, read from the nearest copy

Eventual rename, read from the nearest copy, as an ordered list of steps:
Eventual rename, read from the nearest copy7 steps between Priya's phone, Mumbai replica (leader), Frankfurt replica, Virginia replica, Reader in Virginia. The steps are listed as text after the diagram.Reader in VirginiaVirginia replicaFrankfurt replicaMumbai replica (leader)Priya's phoneall three replicas agree againrename(recipe 812, …)1ok (local commit only, ~2 ms)2ship entry 1042 (async)3get(recipe 812)4"Dal Makhani" (old)5ship entry 1042 (async, ~95 ms)6
  1. Priya's phone → Mumbai replica (leader): rename(recipe 812, …)
  2. Mumbai replica (leader) → Priya's phone (reply): ok (local commit only, ~2 ms)
  3. Mumbai replica (leader) → Frankfurt replica: ship entry 1042 (async)
  4. Reader in Virginia → Virginia replica: get(recipe 812)
  5. Virginia replica → Reader in Virginia (reply): "Dal Makhani" (old)
  6. Mumbai replica (leader) → Virginia replica: ship entry 1042 (async, ~95 ms)
  7. Note over Virginia replica: all three replicas agree again

The spectrum between them

Each step down the ladder allows one more kind of surprise, in exchange for answering from fewer machines.

ModelPromiseWhat it still allowsExample
LinearizableEach operation appears to take effect at one instant between its call and its return, in real-time order.Nothing a single copy would not also do.etcd; Spanner (external consistency, which extends it to transactions)
SequentialAll clients see one total order of operations that keeps each client's own order, but not wall-clock order.A read that starts after another client's write returned can still miss it.ZooKeeper: writes in one order; a server may serve stale reads unless the client calls sync
CausalOperations that could have influenced each other are seen in that order by everyone.Two co-authors' concurrent edits appearing in opposite orders to two readers.MongoDB causally consistent sessions
EventualIf writes stop, all replicas converge on the same value.A recipe step shown before the ingredient it uses; a value going backwards between two reads.Cassandra at consistency level ONE; DynamoDB default reads

Every row implies the rows below it: a linearizable system is also sequential, causal and eventually consistent (Jepsen's hierarchy). The price runs the other way. Jepsen classes linearizable and sequential as impossible to keep available on both sides of a network partition; causal can stay available to a client that keeps talking to the same replica. Session guarantees such as read-your-writes sit between causal and eventual and have their own topic.

Two histories, judged by each model

Scenario 1 of 4: As described.

Notes
  • 3 Linearizable forbids this: the write returned at 24 ms, before the read began at 34 ms.
Timeline as a list

Two histories, judged by each model: 3 lanes, from 0 ms to 120 ms.

  1. 0–24 ms · Priya · w(title = new) (ok)
  2. 24 ms · all lanes · Linearizable: write returned
  3. 34–62 ms · Meera · r(title) → old (violation) — Linearizable forbids this: the write returned at 24 ms, before the read began at 34 ms.
  4. 70–94 ms · Sam · r(title) → new (ok)
Three clients, times in milliseconds (illustrative). The first tab judges the history as linearizable; the Sequential tab judges the same history: Meera's read starts after Priya's write returned and still sees the old title. The Causal and Eventual tabs show a second history, where Sam sees Meera's new recipe step (step 6: crush the kasuri methi) without the ingredient it uses.

Converging is a choice, not a given

"Replicas converge" hides a decision. Tandoor's single leader orders every write, so its replicas only ever lag; they never disagree about which write came last. Conflicts appear when replicas accept writes independently, as in multi-leader setups and leaderless stores such as Dynamo and Cassandra. Picture a multi-leader Tandoor: Priya renames the recipe in Mumbai while a co-author renames it in Frankfurt during a partition, both replicas accept a write, and when they reconnect one value must win or both must be kept. Last-writer-wins compares timestamps and keeps the larger one: simple and always convergent, but it silently discards the other write, and with clock skew the discarded one may be the newer. Keeping both versions (siblings) and letting the application merge them loses nothing but pushes work onto the reader: Dynamo merged shopping carts this way, which is why a deleted item could reappear. How versions are tracked (vector clocks) is covered in the key-value store's versioning topic.

One key under eventual consistency

One key under eventual consistency. 3 states, 6 transitions. The table below lists them.
One key under eventual consistency3 states, 6 transitions. The table below lists them.

write at one replica

another write [same replica] / ship in order

last replica applies

write at another replica [multi-leader]

replicas sync [last-writer-wins] / drop older
replicas sync [siblings] / app merges

All replicas agree

Update spreading

Concurrent versions

2 steps.

Replicas of one key agree, drift apart while an update travels, and agree again. The conflict branch exists only where more than one replica accepts writes (multi-leader or leaderless); a single leader such as Tandoor's never takes it.

Transitions of One key under eventual consistency
From → ToEventGuardAction
All replicas agree → Update spreadingwrite at one replica
Update spreading → Update spreadinganother writesame replicaship in order
Update spreading → All replicas agreelast replica applies
Update spreading → Concurrent versionswrite at another replicamulti-leader
Concurrent versions → All replicas agreereplicas synclast-writer-winsdrop older
Concurrent versions → All replicas agreereplicas syncsiblingsapp merges
All replicas agreestart
Update spreading
Some replicas are stale; reads there return the old value.
Concurrent versions
Two replicas accepted different writes that neither saw.
Was this section helpful?

In practice.

How often eventual consistency actually bites, what strong costs per request, and the knobs real databases expose.

How often does a stale read actually happen?

Assumptions
Recipe reads at peak
40,000/sassumption
Recipe edits at peak
700/s2M daily users × 10 edits ≈ 230/s on average, × 3 at peak (derived in the read-your-writes topic)
Distinct recipes read in a day
8Massumption
Replication lag to a replica (p99)
200 msassumption
The editor's reload after saving
~100 ms after each editthe app reloads the recipe once the save returns (assumption)
Working
  1. Recipes mid-replication at any instantwrites × lag = 700 × 0.2 s140from Recipe edits at peak and Replication lag to a replica (p99)
  2. Chance a random read hits one of theminflight ÷ keys = 140 ÷ 8,000,000≈ 0.0018%from Recipes mid-replication at any instant and Distinct recipes read in a day · Assumes reads spread evenly. Freshly edited recipes are often hotter, so the real figure is higher, but still tiny.
  3. Editor reloads that can land inside the window (upper bound)writes × 1 reload = 700 × 1 (100 ms < 200 ms lag)≤ 700/sfrom Recipe edits at peak, The editor's reload after saving and Replication lag to a replica (p99) · At p99 lag every reload beats replication. At median lag (tens of ms) most reloads arrive after the window has closed, so the real count is lower.
  4. Share of all reads that are an editor's rereadown-stale ÷ reads = 700 ÷ 40,000≤ 1.75%from Editor reloads that can land inside the window (upper bound) and Recipe reads at peak
What it means
  • For strangers, eventual consistency is close to invisible.
  • For the person who just wrote, it is likely, and it happens exactly when the user is watching. That is a job for session guarantees (read-your-writes), not a reason to make all 40,000 reads a second strong.

What strong costs per request

Assumptions
Round trip Mumbai–Frankfurt
~120 msassumption
Round trip Mumbai–Virginia
~190 msassumption
Local durable append
2 msassumption
Local read from the nearest replica
1 msassumption
Working
  1. Strong write (leader plus nearest follower is a majority)disk + rtt-fra = 2 + 120~122 msfrom Local durable append and Round trip Mumbai–Frankfurt
  2. Eventual write (local commit, ship later)disk~2 msfrom Local durable append
  3. How much slower a strong write is122 ÷ 2~61×from Strong write (leader plus nearest follower is a majority) and Eventual write (local commit, ship later)
  4. Strong read in Virginia, leader answering under a leasertt-iad + local = 190 + 1~191 msfrom Round trip Mumbai–Virginia and Local read from the nearest replica
  5. Strong read in Virginia, no lease (leader confirms leadership with Frankfurt first)rtt-iad + rtt-fra + local = 190 + 120 + 1~311 msfrom Round trip Mumbai–Virginia, Round trip Mumbai–Frankfurt and Local read from the nearest replica
  6. Eventual read in Virginialocal~1 msfrom Local read from the nearest replica
What it means
  • Across regions, strong consistency costs a round trip on every write, and on every strong read not served locally by a leader holding a lease; without the lease, the read pays a second round trip.
  • The cost is set by geography, not by hardware. Faster disks do not shorten 120 ms of fibre.

Latency per operation, Tandoor's three regions

Latency per operation, Tandoor's three regionsA strong write costs about 60 times an eventual one, and a strong read in Virginia about 190 times a local read.100 µs1 ms10 ms100 ms1 sEventual read (Virginia)Eventual writeStrong writeStrong read (Virginia)1 ms2 ms122 ms191 msOperationLatency (ms)Latency per operation, Tandoor's three regionsA strong write costs about 60 times an eventual one, and a strong read in Virginia about 190 times a local read.100 µs10 ms1 sEventual read (Vi…Eventual read (Virginia)Eventual writeStrong writeStrong read (Virg…Strong read (Virginia)1 ms2 ms122 ms191 msOperationLatency (ms)
Log scale. Values come from the estimate above and are illustrative; the strong read assumes the leader holds a lease (about 311 ms without one).
Data
OperationLatency (ms)
Eventual read (Virginia)1
Eventual write2
Strong write122
Strong read (Virginia)191

The knob in real systems

Amazon DynamoDB
Reads are eventually consistent by default. ConsistentRead: true gives a strong read on a table or local index and uses twice the read capacity. Global secondary indexes only offer eventual reads. Global tables replicate across Regions asynchronously by default, with an opt-in multi-Region strong mode.
Amazon S3
Since December 2020, every PUT, DELETE and LIST is strongly read-after-write consistent, automatically and at no extra cost. Before that, overwrites, deletes and listings were eventually consistent.
Azure Cosmos DB
Five levels, set per account; a read can relax it per request, never strengthen it: strong, bounded staleness, session, consistent prefix and eventual. Strong and bounded-staleness reads use about twice the request units of the weaker levels.
Apache Cassandra
The level is chosen per request: ONE, QUORUM, ALL and others. A write is always sent to every replica; the level only sets how many acknowledgements the coordinator waits for. The quorums topic works through which levels overlap.
Google Spanner
Externally consistent by default, which is linearizability extended to transactions. Stale reads with a staleness bound can be served by the nearest replica without contacting the leader.
Dynamo (2007 paper)
Always writable; divergent versions are handed to the application to merge. Over 24 hours of shopping-cart traffic, 99.94% of requests saw exactly one version.
Was this section helpful?

Trade-offs.

The choice is made per operation, not per database. The question is what a stale or reordered read would break.

01
Tandoor: recipe edits and like counts
Chosen:Eventual, plus read-your-writes for the editor
  • Pro:Reads answer from the local region in about 1 ms.
  • Pro:Writes keep working in Mumbai even if Frankfurt and Virginia are unreachable.
  • Pro:The one visible anomaly (the editor seeing an old version) is fixed by a session guarantee.
Downside we accept:
  • Con:Other readers see an edit up to one replication window late.
  • Con:A like count can briefly go backwards if two reads hit different replicas.
Ruled out:Strong everywhere

About 122 ms on every write and about 191 ms on every strong read from Virginia (311 ms without a leader lease).; Mumbai cannot commit writes if it loses both followers; the eventual design keeps writing as long as Mumbai is up.

02
Tandoor Pro: prepaid cooking-class credits
Chosen:Strong (single leader, majority commit)
  • Pro:A credit can never be spent twice, and the balance never goes below zero.
  • Pro:Check-then-spend is safe because every read sees every earlier spend.
Downside we accept:
  • Con:Spends take a cross-region round trip.
  • Con:A region on the minority side of a partition cannot spend credits until it heals.
Ruled out:Eventual, reconciled nightly

Two regions can both approve the last credit; the overspend is found hours later.; Fixing it means refunds, clawbacks and apologies.

Where the promises break in practice

FailureImpactDetectionMitigationMeanwhile
Leader fails over while asynchronously shipped entries have not reached any follower3Mumbai replica (leader)Writes that were acknowledged are gone after the new leader takes over.A gap between the old leader's last log position and the new leader's.Commit on a majority before acknowledging, or state the recovery point you accept.
Follower falls minutes behind (slow disk, long GC, saturated link)5Virginia replicaReads there are far staler than the usual 200 ms at p99.Replica lag metric in seconds and in log positions.Take the replica out of read rotation above a lag bound.Virginia reads go to Frankfurt, slower but fresh.
Last-writer-wins with clock skewA newer edit loses to an older one stamped by a machine whose clock runs fast.Hard; the loss is silent.Conditional writes on a version number (compare-and-set), or keep siblings and merge. Hybrid logical clocks only stop a fast clock from beating a causally later write; concurrent writes still lose one.
A cache in front of a strong storeThe store is linearizable, but what users read is whatever the cache holds.Cache age of served objects versus the store's version.Invalidate on write and bound the TTL; see the distributed cache topics.

Reach for strong consistency when…

An invariant spans readers.
A read decides a write.
Uniqueness must hold across regions.
Coordination depends on it.
Otherwise, start eventual and add session guarantees where users notice.
Was this section helpful?
Builds on this
CAP and PACELC
Read next