Consistency modelsCAP and PACELC

100%

CAP and PACELC.

CAP is usually quoted as "pick two of three", which is the least useful way to read it. The actual result is narrow. During a network partition, a replica that can't reach the others must either answer with what it has or not answer at all. PACELC adds the half that matters every day. With no partition at all, you still trade consistency for latency on every request.

Intermediate16 minUpdated 30 Sept 2026

Builds on Strong and eventual consistency.

The idea.

One shop, two copies of its data, and a cable that can break.

A thought experiment you can't engineer around

Kirana Cloud, a fictional online shop, keeps a replica of its database in Mumbai and another in Singapore so each shopper gets a nearby copy. One pressure cooker is left. The undersea link between the regions goes quiet. A Mumbai shopper buys the cooker, and Mumbai writes stock = 0.

A moment later a Singapore shopper asks whether the cooker is in stock. Singapore has never heard of the sale. It has two moves. It can answer "yes, 1 left" from its own copy, which is available but wrong. Or it can wait for Mumbai and, when that never comes, return an error, which is correct but unavailable.

No clever protocol escapes this, because no message can cross the gap. That is the whole of the CAP theorem as Gilbert and Lynch proved it: while messages between replicas are being lost, you can't have every read see the latest write and also have every live replica answer.

The last cooker, three ways

Scenario 1 of 3: As described.

Timeline as a list

The last cooker, three ways: 4 lanes, from 0 ms to 2,800 ms.

  1. 0–500 ms · Mumbai shopper · buy last cooker (ok)
  2. 300 ms · Mumbai replica · stock = 0
  3. 300 ms · Mumbai replica → Singapore replica, arriving 330 ms
  4. 330 ms · Singapore replica · stock = 0 (replicated)
  5. 700–1,200 ms · Singapore shopper · in stock? → 0 left (ok)
As described: the link is healthy, the sale reaches Singapore in about 30 ms and the later read sees 0. The other tabs cut the link before the sale. Singapore either answers from its stale copy or waits 2 s for Mumbai and gives up. Times are illustrative.

The three letters, precisely

C, consistency
Linearizability: every read returns the latest completed write, as if there were one copy. It is not ACID's C, which is about constraints such as "balance ≥ 0".
A, availability
Every request that reaches a replica that hasn't crashed gets a non-error answer. It is not "99.99% uptime". A system can meet a strict SLA and still fail CAP's A.
P, partition tolerance
The network may drop any number of messages between replicas. It is not a feature you opt out of. Cables get cut and switches misroute whatever you choose.
Was this section helpful?

How it works.

Nobody "detects a partition". A timeout expires, and the code decides what to do next.

The moment of choice

The moment of choice, as an ordered list of steps:
The moment of choice7 steps between Mumbai shopper, Mumbai replica, Inter-region link, Singapore replica, Singapore shopper. The steps are listed as text after the diagram.Singapore shopperSingapore replicaInter-region linkMumbai replicaMumbai shopperlink down, packets droppedno word from Mumbai for 2 s. AP would answer "1 left" (stale); CP refusesbuy last cooker1replicate stock=0 ✕ lost2200 order placed (Mumbai leads this item)3in stock?4CP choice: 503, try again later5
  1. Note over Inter-region link: link down, packets dropped
  2. Mumbai shopper → Mumbai replica: buy last cooker
  3. Mumbai replica → Singapore replica: replicate stock=0 ✕ lost
  4. Mumbai replica → Mumbai shopper (reply): 200 order placed (Mumbai leads this item)
  5. Singapore shopper → Singapore replica: in stock?
  6. Note over Singapore replica: no word from Mumbai for 2 s. AP would answer "1 left" (stale); CP refuses
  7. Singapore replica → Singapore shopper (reply): CP choice: 503, try again later

One replica's view of the link

States of3Mumbai replica

One replica's view of the link. 4 states, 8 transitions. The table below lists them.
One replica's view of the linkThe states of Mumbai replica. 4 states, 8 transitions. The table below lists them.

no ack for 500 ms

ack arrives

no ack for 2 s

add to cart / log locally
checkout [stock led remotely] / reply 503

link back / swap logs

merge done / union carts

link drops again

Normal, replicating

Suspect, acks overdue

Partition mode

Recovering

2 steps.

Brewer's three steps as states: detect (a timeout), run in partition mode with some operations limited, then recover by merging and compensating. A short blip never leaves the first two states. The 500 ms and 2 s thresholds are illustrative.

Transitions of One replica's view of the link
From → ToEventGuardActionActor
Normal, replicating → Suspect, acks overdueno ack for 500 ms
Suspect, acks overdue → Normal, replicatingack arrives
Suspect, acks overdue → Partition modeno ack for 2 sreplication monitor
Partition mode → Partition modeadd to cartlog locally
Partition mode → Partition modecheckoutstock led remotelyreply 503
Partition mode → Recoveringlink backswap logs
Recovering → Normal, replicatingmerge doneunion carts
Recovering → Partition modelink drops again
Normal, replicatingstart
Partition mode
Carts accepted on both sides; checkout only where the item's stock leader lives
Recovering
Logs exchanged, carts unioned; no oversell to compensate

Brewer's three steps, in general

StepWhat any system must decide
DetectWhich ack or heartbeat deadline counts as a partition, and how long it is.
Partition modePer operation: which keep going on both sides, which are refused and which queue for later.
RecoverHow to exchange logs, merge divergent state and compensate for mistakes made while apart (a refund for an oversold item is Brewer's kind of example).

The timeout is the choice

A dead link and a slow one look the same from inside a replica: no reply yet. Brewer calls a partition a time bound on communication. If the answer doesn't come within the bound, you treat the other side as unreachable. So the number you type into a config file is where you choose between C and A.

Wait 30 s and you almost never give up consistency, but a shopper stares at a spinner for 30 s. Give up after 200 ms and you answer quickly with whatever you have, even on a merely congested link. That tension, waiting for the other side against answering now, doesn't go away when the network is healthy. It is present on every request, and that is where PACELC starts.

PACELC: the else branch

Abadi's extension: if there is a Partition, choose Availability or Consistency; Else, choose Latency or Consistency. The else branch applies only to replicated data, and only because replicas must talk before a write counts.

PACELC as a decision

PACELC as a decision
PACELC as a decisionParts: Is the network partitioned right now?, Partitioned: answer, or stay correct?, Healthy: answer fast, or wait for replicas?, PA, answer from the local copy, PC, refuse or wait on the cut-off side, EL, reply before remote replicas confirm, EC, wait for the leader or a quorum.

yes (P)

no (E)

A

C

L

C

Is the network partitioned right now?

Partitioned: answer, or stay correct?

Healthy: answer fast, or wait for replicas?

PA, answer from the local copy
Reads may be stale, and writes may conflict
Needs a merge or compensation plan

PC, refuse or wait on the cut-off side
Minority side returns errors
No stale reads, no conflicting writes

EL, reply before remote replicas confirm
Local-replica latency
A window where replicas disagree

EC, wait for the leader or a quorum
At least one inter-replica round trip per write
Reads see the latest write

A system gets one label from each branch, for four combinations: PA/EL (fast and loose all the time), PC/EC (correct all the time, and you pay for it), PA/EC (correct while healthy, answers anyway when cut off) and PC/EL (fast while healthy, but doesn't get weaker when cut off). The labels describe defaults. Most real stores let you choose per request, so the useful unit is the operation, not the product.

Was this section helpful?

In practice.

Partitions are rare, but latency applies to every request. What the numbers say, and where real stores sit.

Why the else branch shapes the design

Assumptions
Partitions per year
3Illustrative assumption
Length of each partition
10 minIllustrative assumption
Writes per second, steady
1,000
Extra wait for a synchronous ack from the other region
60 msOne Mumbai–Singapore round trip, an illustrative figure
Working
  1. Partitioned time per year3 × 10 min30 min/yearfrom Partitions per year and Length of each partition
  2. Share of the year spent partitioned30 ÷ 525,600 min0.0057%from Partitioned time per year
  3. Writes that face the A-or-C choice1,000/s × 1,800 s1.8M/yearfrom Writes per second, steady and Partitioned time per year
  4. Writes that face the L-or-C choice1,000/s × 86,400 s86.4M/day, every one of themfrom Writes per second, steady
  5. Total added waiting if every write is EC86.4M × 0.06 s = 5.18M s ÷ 86,400≈ 60 days of waiting per dayfrom Writes that face the L-or-C choice and Extra wait for a synchronous ack from the other region
What it means
  • A whole year's partition decisions cover 1.8M writes, about 2% of a single day's 86.4M (1.8 ÷ 86.4 ≈ 2.1%). Every one of those 86.4M faces the latency choice.
  • Choose PA or PC for correctness during rare incidents. Choose EL or EC with latency budgets, because it applies every second.

Write p99 by how long the write waits

Write p99 by how long the write waitsEach extra region a strong write must hear from adds a round trip, so p99 climbs from single digits to hundreds of milliseconds.050 ms100 ms150 ms200 ms250 ms300 ms350 msLocal replica only (EL)Ack from SingaporeStrong across Mumbai + SingaporeStrong with Frankfurt added5 ms65 ms130 ms330 msWhat the write waits forp99 write latency (ms)Write p99 by how long the write waitsEach extra region a strong write must hear from adds a round trip, so p99 climbs from single digits to hundreds of milliseconds.0200 msLocal replica onl…Local replica only (EL)Ack from SingaporeStrong across Mum…Strong across Mumbai + SingaporeStrong with Frank…Strong with Frankfurt added5 ms65 ms130 ms330 msWhat the write waits forp99 write latency (ms)
Illustrative. Local write 5 ms; an ack from Singapore adds one 60 ms round trip (5 + 60 = 65 ms). The two strong bars are derived in the estimate below.
Data
What the write waits forp99 write (ms)
Local replica only (EL)5
Ack from Singapore65
Strong across Mumbai + Singapore130
Strong with Frankfurt added330

Where the two strong bars come from

Assumptions
RTT Mumbai–Singapore (about 3,900 km)
60 msIllustrative assumption
RTT Singapore–Frankfurt (about 10,300 km)
160 msIllustrative assumption
Working
  1. Strong write p99, Mumbai + Singapore2 × 60 ms + 10 ms130 msfrom RTT Mumbai–Singapore (about 3,900 km) · Cosmos DB documents 2 × RTT between the farthest regions + 10 ms at p99 for strong multi-region writes.
  2. Strong write p99, adding Frankfurt2 × 160 ms + 10 ms330 ms, and blocked by defaultfrom RTT Singapore–Frankfurt (about 10,300 km) · Cosmos DB blocks strong consistency by default when regions are more than 8,000 km apart. Singapore–Frankfurt is about 10,300 km.
  3. Session or eventual, in-regionserved by the nearest replicaunder 10 ms p99Cosmos DB's documented in-region figure for the weaker levels.
What it means
  • Distance sets the price of EC. Two nearby regions cost about 26× the chart's 5 ms local write (130 ÷ 5), and three spread-out regions about 66× (330 ÷ 5).
  • Strong consistency isn't offered at all with several write regions. Multi-master and linearizable don't go together there.

Where real stores sit, as classified by their sources

SystemIf partitionedElseNote
Dynamo, Cassandra, Riak (defaults)PAELAbadi 2012. Tunable: raise R and W and they move toward C on both branches.
VoltDB/H-Store, MegastorePCECAbadi 2012. Fully ACID; they pay latency always and refuse on the minority side.
MongoDB (as of 2012)PAECAbadi 2012. Reads and writes go to the primary. After failover, the old primary's unreplicated writes are rolled back.
PNUTS (Yahoo)PCELAbadi 2012. "PC" means it doesn't get weaker during a partition, not that it is linearizable.
SpannerPCECBrewer 2017: technically CP. Called "effectively CA" only because partitions on Google's network are very rare. EC is our reading: commits wait for a Paxos majority.
Cosmos DBper levelper levelOur reading of the docs: strong behaves PC/EC; session and eventual behave PA/EL. The account chooses.

Stores that are neither CP nor AP

Single leader, async followers
Reads from a follower can be stale, so it isn't linearizable (not C). A replica cut off from the leader can't take writes, so it isn't CAP-available (not A). This is the most common database set-up, and neither label fits it.
ZooKeeper
Writes need a majority, so the minority side refuses them (not A). Reads are served by whichever server you are connected to and can lag, so they aren't linearizable unless you call sync first (not quite C).
The lesson
Replace the two-letter label with two sentences: which consistency model each operation gets, and what each side does when cut off.
What to say in an interview
"Stock decrements are linearizable through the item's leader. During a partition the other region refuses them. Carts stay available and merge as a union."
Was this section helpful?

Trade-offs.

Choose per operation, for both branches, and write down what happens when you guess wrong.

01
During a partition, per operation
Chosen:Mixed: AP for browsing and carts, CP for stock decrements and payment capture
  • Pro:Most of the site keeps working on both sides of the cut
  • Pro:No item is sold twice, because only its stock leader sells it, so there is nothing to refund
Downside we accept:
  • Con:Two code paths, and partition mode needs its own tests
  • Con:A merge step for carts, and a log replay after the cut
  • Con:Shoppers far from an item's leader see errors for that item
Ruled out:CP everywhere

The cut-off region can't even browse without errors; Every partition becomes a regional outage

Ruled out:AP everywhere

Oversells the last unit whenever both sides sell it; Refunds, vouchers and apology emails become a recovery process you must build

02
Normal operation, per operation
Chosen:EL for catalogue and cart reads (local replica, bounded staleness); EC for stock writes (item's leader region)
  • Pro:Pages load at local-replica speed (about 5 ms)
  • Pro:Only checkout pays the round trip, about 65 ms from the far region
Downside we accept:
  • Con:A product page may show "in stock" for up to the replication lag after the last unit sold
  • Con:The UI must handle "sold out at checkout"
Ruled out:EC everywhere with synchronous cross-region replication

+60 ms on every write between two regions, and 330 ms p99 with three spread-out regions; Strong across more than 8,000 km is blocked by default on some managed stores

Ruled out:EL everywhere, async replication only

Two regions can each sell the last unit even without a partition

Readings of CAP that lead to bad designs

FailureImpactDetectionMitigationMeanwhile
"We're a CA system"No plan for the day the link drops, so behaviour during a partition is whatever the code happens to doCA only truly describes a single node; a one-rack cluster is CA only by assuming its network never fails. Across a network, decide A or C per operation in advance.
"Pick 2 of 3, once, for the whole system"The whole product is forced into the strictest or the loosest behaviourChoose per operation and per moment, as Brewer argues. Carts and stock differ.
"AP means it will be consistent enough"Divergent writes pile up with no rule to merge them. Oversells are found by customers.Every AP operation needs a named merge (union, last-writer-wins, CRDT) or a compensation step.
"CAP-available means highly available"A CP store is rejected for an HA service it could have served wellCAP's A requires every live replica to answer. A majority-quorum CP store can still meet a 99.99% SLA.
"CAP covers latency"A design is labelled CP and still ships 300 ms writes nobody budgeted forUse PACELC. The else branch sets everyday latency, as in the Cosmos formula above.
A short timeout on the replication path5Inter-region linkA congested but working link is treated as a partition, and the store flips into partition mode many times a dayPartition-mode entries per day compared with actual link incidentsSet the timeout from measured RTT percentiles (well above p99.9) and alert on flapping
Was this section helpful?
Next in Core
Session guarantees
Read next