CAP and PACELC.
CAP is usually quoted as "pick two of three", which is the least useful way to read it. The actual result is narrow. During a network partition, a replica that can't reach the others must either answer with what it has or not answer at all. PACELC adds the half that matters every day. With no partition at all, you still trade consistency for latency on every request.
Builds on Strong and eventual consistency.
The idea.
One shop, two copies of its data, and a cable that can break.
A thought experiment you can't engineer around
Kirana Cloud, a fictional online shop, keeps a replica of its database in Mumbai and another in Singapore so each shopper gets a nearby copy. One pressure cooker is left. The undersea link between the regions goes quiet. A Mumbai shopper buys the cooker, and Mumbai writes stock = 0.
A moment later a Singapore shopper asks whether the cooker is in stock. Singapore has never heard of the sale. It has two moves. It can answer "yes, 1 left" from its own copy, which is available but wrong. Or it can wait for Mumbai and, when that never comes, return an error, which is correct but unavailable.
No clever protocol escapes this, because no message can cross the gap. That is the whole of the CAP theorem as Gilbert and Lynch proved it: while messages between replicas are being lost, you can't have every read see the latest write and also have every live replica answer.
The last cooker, three ways
Scenario 1 of 3: As described.
Timeline as a list
The last cooker, three ways: 4 lanes, from 0 ms to 2,800 ms.
- 0–500 ms · Mumbai shopper · buy last cooker (ok)
- 300 ms · Mumbai replica · stock = 0
- 300 ms · Mumbai replica → Singapore replica, arriving 330 ms
- 330 ms · Singapore replica · stock = 0 (replicated)
- 700–1,200 ms · Singapore shopper · in stock? → 0 left (ok)
The three letters, precisely
How it works.
Nobody "detects a partition". A timeout expires, and the code decides what to do next.
The moment of choice
- Note over Inter-region link: link down, packets dropped
- Mumbai shopper → Mumbai replica: buy last cooker
- Mumbai replica → Singapore replica: replicate stock=0 ✕ lost
- Mumbai replica → Mumbai shopper (reply): 200 order placed (Mumbai leads this item)
- Singapore shopper → Singapore replica: in stock?
- Note over Singapore replica: no word from Mumbai for 2 s. AP would answer "1 left" (stale); CP refuses
- Singapore replica → Singapore shopper (reply): CP choice: 503, try again later
One replica's view of the link
States of3Mumbai replica
Brewer's three steps as states: detect (a timeout), run in partition mode with some operations limited, then recover by merging and compensating. A short blip never leaves the first two states. The 500 ms and 2 s thresholds are illustrative.
| From → To | Event | Guard | Action | Actor |
|---|---|---|---|---|
| Normal, replicating → Suspect, acks overdue | no ack for 500 ms | |||
| Suspect, acks overdue → Normal, replicating | ack arrives | |||
| Suspect, acks overdue → Partition mode | no ack for 2 s | replication monitor | ||
| Partition mode → Partition mode | add to cart | log locally | ||
| Partition mode → Partition mode | checkout | stock led remotely | reply 503 | |
| Partition mode → Recovering | link back | swap logs | ||
| Recovering → Normal, replicating | merge done | union carts | ||
| Recovering → Partition mode | link drops again |
- Normal, replicatingstart
- Partition mode
- Carts accepted on both sides; checkout only where the item's stock leader lives
- Recovering
- Logs exchanged, carts unioned; no oversell to compensate
Brewer's three steps, in general
| Step | What any system must decide |
|---|---|
| Detect | Which ack or heartbeat deadline counts as a partition, and how long it is. |
| Partition mode | Per operation: which keep going on both sides, which are refused and which queue for later. |
| Recover | How to exchange logs, merge divergent state and compensate for mistakes made while apart (a refund for an oversold item is Brewer's kind of example). |
The timeout is the choice
A dead link and a slow one look the same from inside a replica: no reply yet. Brewer calls a partition a time bound on communication. If the answer doesn't come within the bound, you treat the other side as unreachable. So the number you type into a config file is where you choose between C and A.
Wait 30 s and you almost never give up consistency, but a shopper stares at a spinner for 30 s. Give up after 200 ms and you answer quickly with whatever you have, even on a merely congested link. That tension, waiting for the other side against answering now, doesn't go away when the network is healthy. It is present on every request, and that is where PACELC starts.
PACELC: the else branch
Abadi's extension: if there is a Partition, choose Availability or Consistency; Else, choose Latency or Consistency. The else branch applies only to replicated data, and only because replicas must talk before a write counts.
PACELC as a decision
A system gets one label from each branch, for four combinations: PA/EL (fast and loose all the time), PC/EC (correct all the time, and you pay for it), PA/EC (correct while healthy, answers anyway when cut off) and PC/EL (fast while healthy, but doesn't get weaker when cut off). The labels describe defaults. Most real stores let you choose per request, so the useful unit is the operation, not the product.
In practice.
Partitions are rare, but latency applies to every request. What the numbers say, and where real stores sit.
Why the else branch shapes the design
- Partitions per year
- 3Illustrative assumption
- Length of each partition
- 10 minIllustrative assumption
- Writes per second, steady
- 1,000
- Extra wait for a synchronous ack from the other region
- 60 msOne Mumbai–Singapore round trip, an illustrative figure
- Partitioned time per year3 × 10 min30 min/yearfrom Partitions per year and Length of each partition
- Share of the year spent partitioned30 ÷ 525,600 min0.0057%from Partitioned time per year
- Writes that face the A-or-C choice1,000/s × 1,800 s1.8M/yearfrom Writes per second, steady and Partitioned time per year
- Writes that face the L-or-C choice1,000/s × 86,400 s86.4M/day, every one of themfrom Writes per second, steady
- Total added waiting if every write is EC86.4M × 0.06 s = 5.18M s ÷ 86,400≈ 60 days of waiting per dayfrom Writes that face the L-or-C choice and Extra wait for a synchronous ack from the other region
- A whole year's partition decisions cover 1.8M writes, about 2% of a single day's 86.4M (1.8 ÷ 86.4 ≈ 2.1%). Every one of those 86.4M faces the latency choice.
- Choose PA or PC for correctness during rare incidents. Choose EL or EC with latency budgets, because it applies every second.
Write p99 by how long the write waits
Data
| What the write waits for | p99 write (ms) |
|---|---|
| Local replica only (EL) | 5 |
| Ack from Singapore | 65 |
| Strong across Mumbai + Singapore | 130 |
| Strong with Frankfurt added | 330 |
Where the two strong bars come from
- RTT Mumbai–Singapore (about 3,900 km)
- 60 msIllustrative assumption
- RTT Singapore–Frankfurt (about 10,300 km)
- 160 msIllustrative assumption
- Strong write p99, Mumbai + Singapore2 × 60 ms + 10 ms130 msfrom RTT Mumbai–Singapore (about 3,900 km) · Cosmos DB documents 2 × RTT between the farthest regions + 10 ms at p99 for strong multi-region writes.
- Strong write p99, adding Frankfurt2 × 160 ms + 10 ms330 ms, and blocked by defaultfrom RTT Singapore–Frankfurt (about 10,300 km) · Cosmos DB blocks strong consistency by default when regions are more than 8,000 km apart. Singapore–Frankfurt is about 10,300 km.
- Session or eventual, in-regionserved by the nearest replicaunder 10 ms p99Cosmos DB's documented in-region figure for the weaker levels.
- Distance sets the price of EC. Two nearby regions cost about 26× the chart's 5 ms local write (130 ÷ 5), and three spread-out regions about 66× (330 ÷ 5).
- Strong consistency isn't offered at all with several write regions. Multi-master and linearizable don't go together there.
Where real stores sit, as classified by their sources
| System | If partitioned | Else | Note |
|---|---|---|---|
| Dynamo, Cassandra, Riak (defaults) | PA | EL | Abadi 2012. Tunable: raise R and W and they move toward C on both branches. |
| VoltDB/H-Store, Megastore | PC | EC | Abadi 2012. Fully ACID; they pay latency always and refuse on the minority side. |
| MongoDB (as of 2012) | PA | EC | Abadi 2012. Reads and writes go to the primary. After failover, the old primary's unreplicated writes are rolled back. |
| PNUTS (Yahoo) | PC | EL | Abadi 2012. "PC" means it doesn't get weaker during a partition, not that it is linearizable. |
| Spanner | PC | EC | Brewer 2017: technically CP. Called "effectively CA" only because partitions on Google's network are very rare. EC is our reading: commits wait for a Paxos majority. |
| Cosmos DB | per level | per level | Our reading of the docs: strong behaves PC/EC; session and eventual behave PA/EL. The account chooses. |
Stores that are neither CP nor AP
Trade-offs.
Choose per operation, for both branches, and write down what happens when you guess wrong.
- Pro:Most of the site keeps working on both sides of the cut
- Pro:No item is sold twice, because only its stock leader sells it, so there is nothing to refund
- Con:Two code paths, and partition mode needs its own tests
- Con:A merge step for carts, and a log replay after the cut
- Con:Shoppers far from an item's leader see errors for that item
The cut-off region can't even browse without errors; Every partition becomes a regional outage
Oversells the last unit whenever both sides sell it; Refunds, vouchers and apology emails become a recovery process you must build
- Pro:Pages load at local-replica speed (about 5 ms)
- Pro:Only checkout pays the round trip, about 65 ms from the far region
- Con:A product page may show "in stock" for up to the replication lag after the last unit sold
- Con:The UI must handle "sold out at checkout"
+60 ms on every write between two regions, and 330 ms p99 with three spread-out regions; Strong across more than 8,000 km is blocked by default on some managed stores
Two regions can each sell the last unit even without a partition
Readings of CAP that lead to bad designs
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| "We're a CA system" | No plan for the day the link drops, so behaviour during a partition is whatever the code happens to do | CA only truly describes a single node; a one-rack cluster is CA only by assuming its network never fails. Across a network, decide A or C per operation in advance. | ||
| "Pick 2 of 3, once, for the whole system" | The whole product is forced into the strictest or the loosest behaviour | Choose per operation and per moment, as Brewer argues. Carts and stock differ. | ||
| "AP means it will be consistent enough" | Divergent writes pile up with no rule to merge them. Oversells are found by customers. | Every AP operation needs a named merge (union, last-writer-wins, CRDT) or a compensation step. | ||
| "CAP-available means highly available" | A CP store is rejected for an HA service it could have served well | CAP's A requires every live replica to answer. A majority-quorum CP store can still meet a 99.99% SLA. | ||
| "CAP covers latency" | A design is labelled CP and still ships 300 ms writes nobody budgeted for | Use PACELC. The else branch sets everyday latency, as in the Cosmos formula above. | ||
| A short timeout on the replication path5Inter-region link | A congested but working link is treated as a partition, and the store flips into partition mode many times a day | Partition-mode entries per day compared with actual link incidents | Set the timeout from measured RTT percentiles (well above p99.9) and alert on flapping |