Rate limiterLimits across many servers

100%

Limits across many servers.

A limit of 1,000 requests a second means nothing if each of 50 gateway pods counts on its own. Keeping one count means shared state, and shared state brings races, hot keys, outages and cross-region latency. This design keeps the count in a sharded in-memory store, makes each check atomic, answers most checks locally, and decides ahead of time what happens when the counter cannot be reached.

Intermediate26 minUpdated 2 Oct 2026

Builds on Rate-limiting algorithms.

Requirements.

150 gateway pods in 3 regions, and one limit per merchant. Which algorithm to run is settled in Algorithms; here the question is where the count lives and how every pod agrees on it. Shiplane and every figure about it are our own assumptions.

Functional requirements

#1A merchant's limit holds in total, whichever pods and regions its requests land on.
#2The limit check adds at most 1 ms at p99 inside a region.
#3Two pods never both spend the same last token (no lost updates).
#4If the counter store cannot be reached, the API keeps serving; abuse-sensitive rules such as login attempts fail closed instead.
#5New rules and per-account overrides reach every pod within 30 s, without a deploy.
#6Overshoot, the requests admitted above the limit, stays under 5% for keys running near their limit.

Capacity estimates

Sizing the busiest region

Assumptions
Peak requests (all regions)
300K/sassumption
us-east share of traffic
60%split 60/25/15 across us-east, eu-west, ap-south
Gateway pods per region
50
Buckets checked per request
2the account-wide limit plus the route's limit
Bucket updates per Redis primary
50K/sa conservative planning figure for a short script; benchmark your own
Memory per bucket
~120 Bkey, two fields, hash and TTL overhead; a login log of up to 5 entries is a little larger, which this round figure absorbs
Live buckets
3.24M40,000 accounts × 6 route classes = 240K, plus ~3M login logs: the 2M anonymous IPs an hour that Algorithms assumes plus ~1M usernames tried (assumption; an upper bound, since per-IP logs expire 60 s after the newest attempt)
Working
  1. Peak requests in us-east300K × 0.6180K/sfrom Peak requests (all regions) and us-east share of traffic
  2. Bucket updates in us-east180K × 2360K/sfrom Peak requests in us-east and Buckets checked per request · Both buckets share a hash tag, so they go in one script call; the round trips are 180K/s.
  3. Primaries at 50% headroom360K ÷ (50K × 0.5) = 14.416 primariesfrom Bucket updates in us-east and Bucket updates per Redis primary · Rounded up to 16, so each owns an even 1,024 of the 16,384 hash slots.
  4. Requests per gateway pod180K ÷ 503.6K/s (7.2K bucket updates/s)from Peak requests in us-east and Gateway pods per region
  5. Memory for all buckets3.24M × 120 B~390 MBfrom Live buckets and Memory per bucket · Fits in any one node. Memory is not what sizes this store.
What it means
  • Round trips per second, not memory, size the counter store.
  • Every one of those round trips is on a request's critical path, so each must stay inside the region and well under 1 ms.
Was this section helpful?

High-level design.

The data path (gateway to counter store) runs on every request; the control path (rules) does not, so a rules outage never blocks traffic. Each region is a full copy of the picture below.

Distributed rate limiter (one region)

Distributed rate limiter (one region). The numbered component cards that follow describe each part.
Distributed rate limiter (one region)Components: 1. API clients (Merchants' integrations. One may hold a few long-lived connections, or thousands of short ones spread over many pods.), 2. Gateway pods (50 per region. Authenticate, match rules from a local copy, check the limit, then forward or answer 429.), 3. Counter store (A Redis Cluster per region with 16 primaries, each with a replica. Holds bucket state, sharded by the hash slot of the limit key.), 4. Rules API (Where operators and billing set rules and per-account overrides. Validates, versions and publishes them.), 5. Rules store (Versioned rules, overrides and region budget shares. Gateways watch it and apply changes.), 6. Budget rebalancer (Every 10 s, reads per-region usage of multi-region keys and re-splits their limits between regions.), 7. Shiplane API (Quotes and label purchases. Only admitted traffic reaches it.), 8. Metrics (Decisions per rule, shadow denies, fail-open counts, store latency and per-key usage for the top keys.).

Data plane (every request)

Control plane (off the request path)

HTTPS, any pod

EVALSHA on the key's shard

admitted requests

new version

watch, apply ≤ 30 s

per-region usage

region shares

4Rules API

5Rules store
versioned

6Budget rebalancer
every 10 s

2Gateway pods
rule cache
leases and deny cache
decisions and fail-open counts → Metrics
pod-01pod-02… pod-50

3Counter store
Redis Cluster
16,384 slots
shard 1shard 2… shard 16

7Shiplane API

1API clients
API key or client IP

8Metrics

Why each pod can't just count alone

The cheapest design gives each pod its slice of the limit and no shared state. For an Enterprise key at 1,000 req/s, each of 50 pods allows 1,000 ÷ 50 = 20 req/s. It is right only when traffic happens to spread evenly.

CaseWhat happensRate the client gets
Pinned connectionsThe client keeps 2 long-lived HTTP/2 connections, which the load balancer pins to 2 pods. Only those 2 pods see it.2 × 20 = 40/s (4% of 1,000)
Spread evenlyThousands of short connections land on all 50 pods about equally.~1,000/s (right by luck)
AutoscalingThe fleet grows to 75 pods but each still allows 20/s.75 × 20 = 1,500/s (150%)
Pod restartA restarted pod starts with a full bucket, and its clients briefly get an extra burst.up to +40 at once

Sharing the count safely

Moving the bucket into a shared store fixes the arithmetic, but a naive read-then-write lets two pods spend the same token.

The lost update

The lost update, as an ordered list of steps:
The lost update8 steps between pod-07, pod-31, Counter store. The steps are listed as text after the diagram.Counter storepod-31pod-07both see 1 ≥ 1, both allowtwo requests admitted on one tokenHGET rl:{acct_7Q2}:global tokens1HGET rl:{acct_7Q2}:global tokens2"1"3"1"4HSET … tokens 05HSET … tokens 06
  1. pod-07 → Counter store: HGET rl:{acct_7Q2}:global tokens
  2. pod-31 → Counter store: HGET rl:{acct_7Q2}:global tokens
  3. Counter store → pod-07 (reply): "1"
  4. Counter store → pod-31 (reply): "1"
  5. Note over pod-07 and pod-31: both see 1 ≥ 1, both allow
  6. pod-07 → Counter store: HSET … tokens 0
  7. pod-31 → Counter store: HSET … tokens 0
  8. Note over pod-07 and Counter store: two requests admitted on one token

One atomic script

One atomic script, as an ordered list of steps:
One atomic script5 steps between pod-07, pod-31, Counter store. The steps are listed as text after the diagram.Counter storepod-31pod-07reads TIME, refills, checks every key, spends, sets PEXPIRE: one stepEVALSHA bucket 2 rl:{acct_7Q2}:global rl:{acct_7Q2}:labels …1grant 1, retry 02EVALSHA bucket 2 … (queued behind pod-07's script)3grant 0, retry 10 ms4
  1. pod-07 → Counter store: EVALSHA bucket 2 rl:{acct_7Q2}:global rl:{acct_7Q2}:labels …
  2. Note over Counter store: reads TIME, refills, checks every key, spends, sets PEXPIRE: one step
  3. Counter store → pod-07 (reply): grant 1, retry 0
  4. pod-31 → Counter store: EVALSHA bucket 2 … (queued behind pod-07's script)
  5. Counter store → pod-31 (reply): grant 0, retry 10 ms

Redis runs one script at a time and nothing else runs while it does, so the read, the refill and the spend cannot interleave with another pod's. The same trap exists for fixed-window counters: the Redis docs' INCR pattern warns that a client which dies between INCR and EXPIRE leaves a key that never expires, and fixes it with a script. Both of a request's buckets use the hash tag {acct_7Q2}, so Redis Cluster hashes only that part and puts them in the same slot, which a multi-key script requires. How slots map to primaries is covered in Sharding a cache.

The check, as one script (ours)

-- KEYS: the request's buckets, all tagged {acct} so they share one slot
-- ARGV: want, pods, then rate (tokens/s), burst and cost for each key in order
local want, pods = tonumber(ARGV[1]), tonumber(ARGV[2])
local t = redis.call('TIME')              -- the store's clock, never a pod's
local now = t[1] * 1000 + math.floor(t[2] / 1000)
local tok, grant, retry = {}, want, 0
for i, key in ipairs(KEYS) do
  local rate  = tonumber(ARGV[3 * i])     -- tokens/s = milli-tokens per ms
  local burst = tonumber(ARGV[3 * i + 1]) * 1000
  local cost  = tonumber(ARGV[3 * i + 2]) -- tokens one request spends here
  local s = redis.call('HMGET', key, 'tokens_milli', 'ts_ms')
  local v = (tonumber(s[1]) or burst) + (now - (tonumber(s[2]) or now)) * rate
  tok[i] = math.min(burst, v)
  local whole = math.floor(tok[i] / 1000)
  if whole < cost then                    -- too few tokens here: deny all
    grant = 0
    retry = math.max(retry, math.ceil((cost * 1000 - tok[i]) / rate))
  else                                    -- lease: at most half of what is left
    grant = math.min(grant, math.max(1, math.floor(whole / (2 * pods * cost))))
  end
end
for i, key in ipairs(KEYS) do
  local rate  = tonumber(ARGV[3 * i])
  local burst = tonumber(ARGV[3 * i + 1]) * 1000
  local cost  = tonumber(ARGV[3 * i + 2])
  redis.call('HSET', key, 'tokens_milli', tok[i] - grant * cost * 1000, 'ts_ms', now)
  redis.call('PEXPIRE', key, math.ceil(burst / rate) * 2)
end
return { grant, retry }                   -- grant (requests) 0 means deny
LineWhy
TIMEPods' clocks drift by milliseconds; one clock per bucket keeps refill arithmetic consistent. Scripts may call TIME because Redis replicates a script's writes, not the script (the default since Redis 5).
tokens_milliThousandths of a token. For per-second plans it stays an integer (a Free key refills 10 milli-tokens per ms), so there is no rounding drift. A per-minute or per-hour bucket would pass a fractional rate (5/min = 0.083 milli-tokens per ms); Lua numbers are doubles, and the rounding error over a bucket's life is far below one token.
all-or-nothingThe first loop decides and the second spends, so a request never uses the account's token when its route bucket is empty.
PEXPIREA bucket untouched for 2 × B ÷ R is full anyway, so deleting it loses nothing and frees memory. That is 4 s on every plan (Free 2 × 20 ÷ 10, Growth 2 × 200 ÷ 100, Enterprise 2 × 2,000 ÷ 1,000). Login rules are not buckets but sliding logs (see Algorithms), checked by a similar script on a sorted set and expired one window after the newest entry.
want and podsWith want 1 it is an exact per-request check. With want 50 it grants a lease, sized in Optimizations.
costA bulk call that spends 10 tokens is denied unless every bucket holds 10, and a lease of g requests takes g × 10. Rules reject a burst smaller than the largest cost (422).
Was this section helpful?

Data model.

Three places hold state. Bucket state is hot and disposable; rules are small, audited and versioned; each pod keeps a working copy of both in memory.

Redis Cluster (per region)Owned by 3Counter store

Sub-millisecond atomic scripts, and the state is disposable.

bucketTTL 2 × B ÷ R (4 s on every plan)Key pattern rl:{acct_7Q2}:labels, a Redis hash
ColumnTypeKeyNote
keytextprimaryhash-tagged so an account's buckets share a slot
tokens_milliintinteger milli-tokens, no float drift
ts_msintlast refill time from Redis TIME
login_logTTL one window after the newest entry (60 s per IP, 1 h per username)Sliding log for login rules, a sorted set; key rl:login:ip:203.0.113.9
ColumnTypeKeyNote
keytextprimary
membertextone per accepted attempt, at most 5 per IP or 20 per username
scoreintattempt time in ms from Redis TIME; older than the window is trimmed first
window_counterTTL 25 hDaily quotas, key q:{acct_7Q2}:2026-09-30 (Free: 10,000 a day)
ColumnTypeKeyNote
keytextprimary
countint
PostgreSQLOwned by 5Rules store

Small, relational and audited; every change is a new version.

rules
ColumnTypeKeyNote
rule_idtextprimary
match_plantext
match_routetext
identityenumapi_key | ip | user
algorithmenumtoken_bucket | sliding_log | fixed_window
rate_per_sint nullabletoken_bucket
burstint nullabletoken_bucket
limitint nullablecount per window: sliding_log, fixed_window
window_sint nullable
modeenumenforce | shadow
fail_modeenumopen | closed
versionint
overrides
ColumnTypeKeyNote
account_idtextprimary
rule_idtextprimary
rate_per_sint
burstint
expires_attimestamptz nullable
reasontext
region_budgets
ColumnTypeKeyNote
rule_idtextprimary
limit_keytextprimary
regiontextprimary
sharefloatshares of one key sum to 1.0
updated_attimestamptz
In memory, per gateway podOwned by 2Gateway pods

Read on every request; rebuilt from the rules store and the counter store within seconds after a restart.

rule_cache
ColumnTypeKeyNote
rule_idtextprimary
versionint
leasesTTL 1 s
ColumnTypeKeyNote
limit_keytextprimary
tokens_leftint
expires_at_msint
deny_cache
ColumnTypeKeyNote
limit_keytextprimary
denied_until_msintnow + retry_ms from the script
fallback_buckets
ColumnTypeKeyNote
limit_keytextprimary
tokensintlimit ÷ 50 × 2, used only while the shard is unreachable
ts_msint
Access patterns
QueryUsesHow
Is this request allowed?bucketone EVALSHA on the account's shard, unless a lease or the deny cache answers first
Which rules apply to plan X, route Y?rule_cachein memory, per request
What changed since version v?ruleswatch from the pod's last version
What share of this key's limit does my region get?region_budgetswatched with the rules, applied to rate and burst
Was this section helpful?

Interface.

Most deployments run the check as a library inside the gateway. When it runs as a sidecar or a separate service, the call below is the contract. Its shape follows Envoy's rate limit service: a domain plus a list of entries, and the overall verdict is deny if any entry is over. The 429 and its headers are covered in Headers, 429s and backoff.

1POST/v1/ratelimit:check

Checks and spends every entry for one request in one atomic step. The verdict is the strictest entry's.

Request
{
"domain": "public-api",
"entries": [
{
"rule": "growth-global",
"key": "acct_7Q2",
"cost": 1
},
{
"rule": "labels-route",
"key": "acct_7Q2",
"cost": 1
}
]
}
Response
{
"verdict": "allow",
"entries": [
{
"rule": "growth-global",
"remaining": 141,
"reset_ms": 590
},
{
"rule": "labels-route",
"remaining": 12,
"reset_ms": 880
}
]
}

Admin endpoints

MethodPathDoes
PUT/v1/rules/{id}Replace a rule. Needs If-Match with the current version.
POST/v1/overridesRaise or lower one account's limit, usually with expires_at (a merchant's launch week).
POST/v1/rules/{id}:shadowPut a rule in shadow mode: evaluated and counted, never enforced.
GET/v1/keys/{key}/usageTokens left, recent rate and region shares for one key, for support.

Errors

HTTPTypeBody codeClient behaviour
409retry
version_conflict
The rule changed since the caller read it. Re-read and re-apply.
422error
invalid_rule
Rejected before it can hurt: rate ≤ 0, burst below the largest cost, or a match that covers every route.
504network
store_timeout
The shard took over 5 ms. Apply each rule's fail_mode: open allows and counts fail_open; closed denies with 503.
503network
store_unavailable
The shard's breaker is open. Open-mode rules use the pod's fallback bucket at limit ÷ 50 × 2; closed-mode rules (login) deny.
Was this section helpful?

Optimizations.

The design so far costs one store round trip per request. At Shiplane's traffic mix these changes cut that by about 4×, protect the store from floods and hot keys, and extend the limit across regions.

Local first, global second
Each pod runs a coarse token bucket per key at 2× its fair share before any global check. A bot sending 100K req/s is cut down on the pods and never reaches Redis. Envoy's docs describe this pairing: a local token bucket absorbs bursts that could overwhelm the global service, and the global limit finishes the job.
Token leases
A pod asks for up to 50 tokens at once and spends them locally for up to 1 s. The script grants min(50, max(1, tokens ÷ (2 × pods))), so leases never hold more than half of what is left, and near the limit the lease shrinks to 1: exact per-request checks again. With 50 pods a full Free bucket (20) gives leases of 1, Growth (200) 2 and Enterprise (2,000) 20, so leases pay off for big keys. Leases are pre-paid, so they cannot overshoot; their error is tokens stranded on idle pods.
Remember the no
When the script denies with retry 80 ms, the pod caches that and denies the key locally until then. A client hammering past its limit costs one store call per retry interval per pod, not one per request.
Count after, at the edge
Where a few extra requests do not matter (DDoS-scale per-IP rules), admit first and count asynchronously. Cloudflare keeps an isolated memcached cluster per PoP and runs increments off the request path, checking only a cached 'mitigated' flag inline.

Leases on a hot key

Assumptions
A marketplace partner's traffic
20K req/son a 25,000/s override, burst 50,000 (assumption)
Pods it lands on
50
Lease size granted
50min(50, 50,000 ÷ 100) = 50 while the bucket is full
Working
  1. Rate per pod20K ÷ 50400 req/sfrom A marketplace partner's traffic and Pods it lands on
  2. Store calls per pod400 ÷ 508/sfrom Rate per pod and Lease size granted
  3. Store calls for this key8 × 50400/s (was 20,000/s)from Store calls per pod and Pods it lands on
  4. Most tokens out on leases50 pods × 502,500 (5% of the burst)from Pods it lands on and Lease size granted · Never more than half of what the bucket holds, by construction of the grant.
What it means
  • One hot key is one slot on one primary; leases turn 20,000 calls a second there into 400.
  • Twenty-five pods holding leases of 20 can strand at most 500 tokens, which is why leases shrink as the bucket empties.
  • Hot counters (sharded-counters) covers contended counters in general; its sharded, slightly stale sums do not suit a limiter, which must decide on a current count, so the limiter leases instead.

Bucket updates per second in us-east

Bucket updates per second in us-eastCaching denials and leasing tokens cut the counter store's load from 360K to about 81K bucket updates a second, about 23%.050k100k150k200k250k300k350k400kPer-request checks+ deny cache+ leases360k updates/s · 100%324k updates/s · 90%81k updates/s · 23%StageBucket updates per second (updates/s)Bucket updates per second in us-eastCaching denials and leasing tokens cut the counter store's load from 360K to about 81K bucket updates a second, about 23%.0200k400kPer-request checks+ deny cache+ leases360k updates/s · 100%324k updates/s · 90%81k updates/s · 23%StageBucket updates per second (updates/s)
Assumptions: at peak, 10% of checks come from clients already over their limit and the deny cache answers them: 360K × 0.9 = 324K. Of the rest, 75% is Enterprise and override keys whose tightest bucket holds at least 1,000 tokens, so leases are at least 1,000 ÷ 100 = 10 and cut calls by 90%; 15% is Growth (lease 2, cut 50%); 10% is Free (lease 1, no cut). Left: 0.75 × 0.1 + 0.15 × 0.5 + 0.10 × 1 = 0.25, and 324K × 0.25 = 81K. Four primaries could carry that at 50% headroom; Shiplane keeps 16, because the store must still take per-request checks if the mix shifts toward small plans or leases are turned off.
Data
Stageupdates/s
Per-request checks360,000 (100%)
+ deny cache324,000 (90%)
+ leases81,000 (23%)

One limit, three regions

A round trip from us-east to ap-south takes on the order of 200 ms, so a single global counter breaks requirement #2 by two orders of magnitude. Instead, each region gets a share of each key's limit, set by where the key's traffic actually is.

Stepus-easteu-westap-south
Default split (60/25/15)600/s250/s150/s
Observed usage, last 10 s70%30%0%
Raise any region under 10% to 10%––10%
Split the other 90% as observed (70:30)63%27%10%
New budgets for a 1,000/s key630/s270/s100/s
  1. 01One gateway
    Load
    2K req/s
    Bottleneck
    None yet
    Change
    In-process token buckets on the one gateway
    Adds
    1API clients2Gateway pods7Shiplane API
  2. 0210 pods
    Load
    20K req/s
    Bottleneck
    Limit ÷ 10 per pod; pinned connections get a fraction of their limit
    Change
    One shared Redis with an atomic script; rules move out of config files
    Adds
    3Counter store4Rules API5Rules store
  3. 0350 pods
    Load
    180K req/s per region
    Bottleneck
    One Redis at 360K bucket updates/s
    Change
    Redis Cluster with 16 primaries, keys hash-tagged per account
    Adds
    8Metrics
  4. 04Hot keys
    Load
    One key at 20K req/s
    Bottleneck
    One slot, one primary; store cost per request
    Change
    Local buckets, leases and the deny cache in every pod
  5. 053 regions
    Load
    300K req/s
    Bottleneck
    Cross-region round trips of 70 to 200 ms per check
    Change
    A counter store per region, each key's limit split by traffic share
    Adds
    6Budget rebalancer
Was this section helpful?

Trade-offs.

The chosen option is first; the others stay visible so the reasoning can be checked.

01
Where the count lives
Chosen:A sharded in-memory store per region
  • Pro:One count per key, whichever pod the request hits
  • Pro:Sub-millisecond atomic scripts
  • Pro:Easy to reason about and to size
Downside we accept:
  • Con:A network hop on the request path
  • Con:A store to run, and to plan for when it fails
Ruled out:Each pod counts alone

Wrong whenever traffic is uneven (40/s instead of 1,000/s); Limits drift with autoscaling

Ruled out:Sticky routing by key to one pod

A hot key pins one pod; Scaling or a crash moves keys and loses their state

Ruled out:Gossip or CRDT counters between pods

Convergence takes tens to hundreds of ms and bursts slip through meanwhile; Traffic between pods grows with pods × keys

02
When the store cannot be reached
Chosen:Fail open, with a local fallback bucket (API rules)
  • Pro:An outage of the limiter never becomes an outage of the API
  • Pro:Still bounded: at most 2× the limit while the store is gone
Downside we accept:
  • Con:A client can briefly exceed its paid limit
  • Con:Needs alerting on the fail_open rate so it is not silent
Ruled out:Fail closed (per rule; login attempts)

Every covered request fails while the store is down

Real systems default to open. Stripe's limiters catch any error in the limiter code or Redis and let the request through. Envoy's rate limit filter allows the request when the rate limit service fails, and counts it as failure_mode_allowed; setting failure_mode_deny returns a 500 instead. Shiplane chooses per rule, with fail_mode on each.

A pod's view of one counter shard

States of2Gateway pods

A pod's view of one counter shard. 3 states, 5 transitions. The table below lists them.
A pod's view of one counter shardThe states of Gateway pods. 3 states, 5 transitions. The table below lists them.

reply under 5 ms / reset timeout count

3rd timeout in a row / start fallback, fail_open += 1

1 s timer

probe replies / drop fallback buckets

probe times out

Global checks

Fallback

Probing

5 steps.

Each pod keeps one of these per shard. It stops paying the 5 ms timeout on every request once a shard is clearly gone, and lets one probe a second find out when it is back.

Transitions of A pod's view of one counter shard
From → ToEventActionActor
Global checks → Global checksreply under 5 msreset timeout count
Global checks → Fallback3rd timeout in a rowstart fallback, fail_open += 1
Fallback → Probing1 s timerpod timer
Probing → Global checksprobe repliesdrop fallback buckets
Probing → Fallbackprobe times out
Global checksstart
every check goes to the shard
Fallbackerror
open rules use pod buckets; closed rules deny
Probing
one real check goes to the shard

Shard 7 fails over

Timeline as a list

Shard 7 fails over: 4 lanes, from 0 s to 24 s.

  1. 0–3 s · Shard 7 primary · serving
  2. 0–19 s · Shard 7 replica · async copy
  3. 0–3 s · Gateway pods · global checks
  4. 3–24 s · Shard 7 primary · down
  5. 3–20 s · Gateway pods · fallback: limit ÷ 50 × 2 (allowed)
  6. 3–20 s · Login rule on shard 7 · fail closed (503) (denied)
  7. 3 s · all lanes · primary dies (error)
  8. 3 s · Gateway pods · breaker opens
  9. 10 s · Gateway pods · probe times out (tick, denied)
  10. 18 s · all lanes · marked FAIL (15 s) (deadline)
  11. 19–24 s · Shard 7 replica · primary
  12. 19 s · Shard 7 replica · promoted (ok)
  13. 20–24 s · Gateway pods · global checks
  14. 20 s · Gateway pods · probe ok (ok, ok)
With cluster-node-timeout at 15 s (the sample redis.conf value), the replica is promoted about 16 s after the primary dies. API rules run on fallback buckets meanwhile; the login rule denies. Replication is asynchronous, so the new primary may be a few ms behind: some buckets come back slightly fuller than they were. Set cluster-require-full-coverage no: with the default (yes), from FAIL at 18 s until promotion the whole cluster refuses queries, and all 16 shards' keys fall back, not just shard 7's.
03
One limit, three regions
Chosen:Split each key's limit by traffic share, rebalanced every 10 s
  • Pro:Every check stays inside its region
  • Pro:Budgets sum to the limit so there is no overshoot
Downside we accept:
  • Con:A sudden shift of traffic to another region is under-served for up to 10 s
  • Con:A rebalancer and its inputs to run
Ruled out:The full limit in every region

A client spread over 3 regions gets up to 3× its limit

Ruled out:One global counter store

70 to 200 ms added to requests far from it; One region's outage breaks limiting everywhere

04
Check before, or count after
Chosen:Check before admitting (paid plans, login)
  • Pro:The limit holds to within the lease error
  • Pro:Denials are immediate
Downside we accept:
  • Con:A store call (or lease) on the request path
Ruled out:Admit now, count after (DDoS-scale per-IP rules)

A flood gets through until the count catches up; Not acceptable for quotas a customer pays for

FailureImpactDetectionMitigationMeanwhile
A shard's primary dies3Counter storeChecks on its 1/16 of keys time outPod breakers open; fail_open rate alertReplica promoted after the node timeout; pods probe once a second; cluster-require-full-coverage no, so the other 15 shards keep servingAPI rules run on pod buckets (up to 2× limit); login denies
Network partition between pods and the store2Gateway podsEvery check waits for the 5 ms timeout until breakers openStore latency and timeout rate per podBreaker per shard stops paying the timeoutEach rule's fail_mode applies
Pod clocks drift2Gateway podsNone on bucketsNTP offset metricThe script reads Redis TIME; pod clocks never enter the refillLease expiry (1 s, pod-local) is off by the drift at most
A bad rule is pushed (rate 0 on every route)4Rules APIEveryone is deniedShadow-mode deny rate before enforcing; deny rate by ruleValidation (422), then shadow, then canary pods first; pods keep the last good versionPods ignore a version that fails their own checks
One hot key saturates its primary3Counter storeLatency for every key on that primaryPer-slot ops and the top keys by callsLeases and deny cache; if still hot, migrate its slot to a dedicated primaryOther keys on the shard see higher p99 until moved
Was this section helpful?
Next in Core
Headers, 429s and backoff
Read next