Rate limiterHeaders, 429s and backoff

100%

Headers, 429s and backoff.

A rate limiter is half an API contract. The other half is what the client sees, which is headers that show the quota running down, a 429 that says which limit was hit and exactly when to come back, and an SDK that backs off with jitter and throttles itself before the server has to.

Beginner22 minUpdated 2 Oct 2026

Builds on Rate-limiting algorithms.

Requirements.

The server-side limiter already decides allow or deny. This topic designs everything a client sees and does with that decision: the headers, the rejection, and the retries that follow.

Functional requirements

#1Every response tells the client its policy and how much is left, so it can slow down before it is rejected.
#2A rejection names the limit that was hit and when to retry, in a machine-readable form.
#3"You are over your quota" (429) is distinguishable from "we are overloaded" (503).
#4Retried writes never create a second label.
#5During an incident the official SDK's retries stay at or below 10% of its normal traffic.
#6A rejection costs the server almost nothing. It is decided at the edge, before the body is read.

Why the client half matters

A limiter that only says no protects the server once. Then comes the second wave: thousands of rejected clients retrying at the same moment, retries stacked on retries at every layer, and merchants who cannot tell a plan limit from an outage. The server's answer has to carry enough information for a well-behaved client to do the right thing, and the SDK Shiplane ships has to actually do it.

Capacity estimates

Assumptions for Shiplane: 300K req/s at peak across three regions; after a 5-second blip, 10,000 SDK instances have been told to retry.

What one rejection turns into

Assumptions
SDK instances rejected during a 5 s blip
10,000assumption
Retry-After sent with each rejection
5 s
Extra random wait on top of Retry-After
uniform over 0–5 sJ = min(20 s, max(1 s, Retry-After))
Working
  1. Retries at t = 5 s without jitterevery client sleeps exactly Retry-After10,000 in the same instantfrom SDK instances rejected during a 5 s blip and Retry-After sent with each rejection
  2. Retry rate with jitter10,000 ÷ 5 s window, over t = 5 to 10 s2,000/sfrom SDK instances rejected during a 5 s blip, Retry-After sent with each rejection and Extra random wait on top of Retry-After
  3. Rate headers on every response, uncompressed300K req/s × ~120 B (RateLimit + RateLimit-Policy)≈ 36 MB/s (≈ 290 Mbit/s) across all regionsWith HTTP/2 header compression the unchanging Policy line shrinks to an index of a byte or two after the first response.
What it means
  • Retry-After says when; jitter says not everyone at once.
  • Retry at one layer only, and cap retries at 10% of traffic (about 1.1× load at worst). How retries multiply across layers, and the budget arithmetic, are worked in failure-models/timeouts-and-retries.

Retries per second after the blip

  • No jitter
  • Retry-After + random(0, 5 s)
Retries per second after the blipWithout jitter all 10,000 retries land in one second; spreading them over five seconds caps the wave at 2,000 per second.02k4k6k8k10k5–6 s6–7 s7–8 s8–9 s9–10 s10k req/s0 req/s0 req/s0 req/s0 req/s2k req/s2k req/s2k req/s2k req/s2k req/sRetries arriving (req/s)Seconds after the rejectionRetries per second after the blipWithout jitter all 10,000 retries land in one second; spreading them over five seconds caps the wave at 2,000 per second.02k4k6k8k10k5–6 s6–7 s7–8 s8–9 s9–10 s10k req/s0 req/s0 req/s0 req/s0 req/s2k req/s2k req/s2k req/s2k req/s2k req/sRetries arriving (req/s)Seconds after the rejection
Same 10,000 clients, same Retry-After: 5. Only the jitter differs.
Data
Seconds after the rejectionNo jitter (req/s)Retry-After + random(0, 5 s) (req/s)
5–6 s10,0002,000
6–7 s02,000
7–8 s02,000
8–9 s02,000
9–10 s02,000
Was this section helpful?

High-level design.

The gateway turns the limiter's verdict into headers and status codes; the SDK turns those back into pacing and retries. Everything else is the normal request path.

Rate limiting, client to edge

Rate limiting, client to edge. The numbered component cards that follow describe each part.
Rate limiting, client to edgeComponents: 1. Merchant app (The merchant's checkout backend. It calls Shiplane through the SDK and decides what to tell its own users.), 2. Shiplane SDK (The official client library. Reads rate headers, paces requests, retries with jitter under a retry budget and adds idempotency keys.), 3. Gateway (Authenticates the key, asks the limiter and the shedder, writes RateLimit headers on every response and returns 429 or 503.), 4. Limiter (Returns allow or deny, what is left, when it resets and how long to wait (the algorithms and distributed-limits topics).), 5. Load shedder (Rejects by priority when the API itself is overloaded, whatever each client's quota says.), 6. Shiplane API (Rate quotes and label purchases. It only ever sees admitted requests.), 7. Usage API (`GET /v1/usage`: the key's plan, its limits and what is left, for dashboards and the SDK's start-up.).

Shiplane edge

Merchant side

createLabel()

HTTPS + Idempotency-Key; back: RateLimit headers, 429 / 503

check key + route

priority, current load

admitted only

at start-up

read counters

1Merchant app

2Shiplane SDK
pacing from RateLimit headers
jittered retries under a budget
Idempotency-Key on every POST

3Gateway
decides before reading the body
429 / 503 + Retry-After

4Limiter
allow · remaining · reset

5Load shedder
criticality
fleet load

6Shiplane API

7Usage API
GET /v1/usage

One allowed call, one rejected call

The burst’s last token

The burst’s last token, as an ordered list of steps:
The burst’s last token7 steps between Merchant app, Shiplane SDK, Gateway, Limiter, Shiplane API. The steps are listed as text after the diagram.Shiplane APILimiterGatewayShiplane SDKMerchant appcreateLabel(order8812)1POST /v1/labels ·Idempotency-Key:k-88122check key sk_live_…,route labels3allow, r=0, t=14forward5201 label6201 ·RateLimit-Policy:"burst";q=200;w=2 ·RateLimit:"burst";r=0;t=17
  1. Merchant app → Shiplane SDK: createLabel(order 8812)
  2. Shiplane SDK → Gateway: POST /v1/labels · Idempotency-Key: k-8812
  3. Gateway → Limiter: check key sk_live_…, route labels
  4. Limiter → Gateway (reply): allow, r=0, t=1
  5. Gateway → Shiplane API: forward
  6. Shiplane API → Gateway (reply): 201 label
  7. Gateway → Shiplane SDK (reply): 201 · RateLimit-Policy: "burst";q=200;w=2 · RateLimit: "burst";r=0;t=1

The next request: 429, then a retry

The next request: 429, then a retry, as an ordered list of steps:
The next request: 429, then a retry9 steps between Merchant app, Shiplane SDK, Gateway, Limiter. The steps are listed as text after the diagram.LimiterGatewayShiplane SDKMerchant appr=0: the last token is gone. A second worker on this key hasn't seen that and sends 2 ms later.Waits 1 s + random(0, 1 s) = 1.6 s (8 ms rounds up to 1 s). Refill 1.6 s × 100/s = 160; other workers spend 10.POST /v1/labels · Idempotency-Key: k-88131check2deny, next token in 8 ms3429 · Retry-After: 1 · problem+json limit=burst4retry POST /v1/labels · same Idempotency-Key: k-88135201 · RateLimit: "burst";r=149;t=1 (160 − 10 − 1)6label for order 88137
  1. Note over Shiplane SDK: r=0: the last token is gone. A second worker on this key hasn't seen that and sends 2 ms later.
  2. Shiplane SDK → Gateway: POST /v1/labels · Idempotency-Key: k-8813
  3. Gateway → Limiter: check
  4. Limiter → Gateway (reply): deny, next token in 8 ms
  5. Gateway → Shiplane SDK (reply): 429 · Retry-After: 1 · problem+json limit=burst
  6. Note over Shiplane SDK: Waits 1 s + random(0, 1 s) = 1.6 s (8 ms rounds up to 1 s). Refill 1.6 s × 100/s = 160; other workers spend 10.
  7. Shiplane SDK → Gateway: retry POST /v1/labels · same Idempotency-Key: k-8813
  8. Gateway → Shiplane SDK (reply): 201 · RateLimit: "burst";r=149;t=1 (160 − 10 − 1)
  9. Shiplane SDK → Merchant app (reply): label for order 8813

What a rejection costs the edge

Rejected at the edge vs admitted

Scenario 1 of 2: As described.

Timeline as a list

Rejected at the edge vs admitted: 6 lanes, from 0 ms to 2.5 ms.

  1. 0–0.8 ms · Request at the gateway · 429 in 0.8 ms
  2. 0–0.2 ms · Parse headers + look up key · key cache hit
  3. 0.2–0.7 ms · Limiter check · check (denied)
  4. 0.7–0.8 ms · Write response · 429 + headers
  5. 0.8 ms · Read request body · body never parsed (ok)
  6. 1 ms · all lanes · 1 ms limiter budget (p99) (deadline)
Illustrative timings (our assumptions) for one request on a warm keep-alive connection. A 429 is written before the body is parsed, so it finishes in about 0.8 ms; an admitted label purchase is still at the API when the axis is cut, around 120 ms later. On HTTP/1.1 the gateway still drains or closes the connection for a body already in flight; on HTTP/2 it resets the stream.

Headers in the wild

HeaderWhere it appearsMeaning
Retry-AfterRFC 9110. On 429 and 503 (and 3xx).Wait at least this long. Either delay-seconds ("120") or an HTTP-date. The only one that is a finished standard for this purpose.
RateLimit-PolicyIETF httpapi draft (-11). Any response; Shiplane sends it on every one.The policies that apply: q (quota), w (window in seconds), optional qu (unit: requests, content-bytes, concurrent-requests) and pk (partition key).
RateLimitIETF httpapi draft (-11). Any response; Shiplane sends it on every one.What is left, per named policy: r (remaining) and t (seconds in which the remaining quota r applies; in effect, until it resets). Clients must cope with responses that omit it.
x-ratelimit-limit · -remaining · -used · -resetGitHub REST API, every response.The widespread pre-draft convention. GitHub's reset is an epoch timestamp in seconds, not a delay.
x-envoy-ratelimitedEnvoy on the 429s it generates itself.Marks the 429 as the proxy's, so a client or a retry policy can tell it from a 429 the upstream sent.
Was this section helpful?

Data model.

The contract is small. The server sends policies, what is left and a problem body; the SDK keeps four little pieces of state to act on them.

What crosses the wire, and what the SDK keeps

The server sends three shapes (policies, what is left, a problem body on rejection); the SDK folds them into pacing, per-call retry state, a 2-minute throttle window and a 1-minute retry budget.keywordmodelfieldprimitivevalue
// 01 · What the server sends
// one item of RateLimit-Policy
type RateLimitPolicy = {
name: string // "burst", "daily"
q: int // quota
w_s: int // window in seconds
qu?: QuotaUnit
}
// one item of RateLimit
type RateLimit = {
name: string
r: int // remaining
t_s: int // seconds in which r applies (≈ until reset)
}
// RFC 9457, application/problem+json
type Problem = {
type: uri
title: string
status: int // 429
detail: string
limit: string // which policy was hit
retry_after_s: int
}
type QuotaUnit =
| "requests" | "content-bytes"
| "concurrent-requests"
// 02 · What the SDK keeps
type PacingState = {
remaining: int
reset_at: timestamp
policy: RateLimitPolicy[]
}
// one per call in flight
type RetryState = {
attempt: int // 1 to 3
idempotency_key: string
next_at: timestamp
}
// last 2 minutes
type ThrottleWindow = {
requests: int
accepts: int
}
// last 1 minute (our window; the SRE book sets none)
type RetryBudget = {
requests: int
retries: int // ≤ 10% of requests
}

The 429 body when a Free plan's daily quota runs out at 18:40 UTC

{
  "type": "https://docs.shiplane.example/errors/quota_exhausted",
  "title": "Daily quota used up",
  "status": 429,
  "detail": "10,000 of 10,000 requests used today (Free plan).",
  "limit": "daily",
  "retry_after_s": 19200
}

The daily quota resets at 00:00 UTC. From 18:40 that is 5 h 20 min = 5 × 3,600 + 20 × 60 = 19,200 s, the same number as the Retry-After header. The limit field is what lets the SDK treat this 429 differently from a burst 429 that clears in a second.

One SDK call, from send to answer

States of2Shiplane SDK

One SDK call, from send to answer. 5 states, 6 transitions. The table below lists them.
One SDK call, from send to answerThe states of Shiplane SDK. 5 states, 6 transitions. The table below lists them.

2xx / update PacingState

rate_limited or 503 [attempt ‹ 3, budget left]

quota_exhausted / raise QuotaExceeded
rate_limited or 503 [retries used up]

timer fires / attempt++, same key

local throttle check [random ‹ p_drop]

Sending

Waiting to retry

Succeeded

Error returned to app

Dropped locally

3 steps.

RetryState moves through these states. Only a burst 429 or a 503 leads back to sending, and only while the attempt count and the retry budget allow it.

Transitions of One SDK call, from send to answer
From → ToEventGuardAction
Sending → Succeeded2xxupdate PacingState
Sending → Waiting to retryrate_limited or 503attempt < 3, budget left
Sending → Error returned to appquota_exhaustedraise QuotaExceeded
Sending → Error returned to apprate_limited or 503retries used up
Waiting to retry → Sendingtimer firesattempt++, same key
Waiting to retry → Dropped locallylocal throttle checkrandom < p_drop
Sendingstart
Waiting to retry
Retry-After + random(0, J)
Succeededend
Error returned to apperror
quota used up, budget spent or 3 attempts
Dropped locallyerror
adaptive throttle said no
Was this section helpful?

Interface.

Every endpoint returns the rate headers; the label endpoint shows every way a call can end. Multiple response headers are shown separated by · .

1POST/v1/labels

Buys a shipping label. Shown for a Free-plan key (10 req/s, burst 20, 10,000 a day) at about 18:00 UTC.

Request
Headers
Authorization: Bearer sk_live_…Idempotency-Key: 5f1c2b7e-…
{
"shipment_id": "shp_4410",
"rate_id": "rate_ups_ground"
}
Response
{
"label_id": "lbl_77120",
"tracking_number": "1Z999AA10123456784"
}
Also setsRateLimit-Policy: "burst";q=20;w=2, "daily";q=10000;w=86400 · RateLimit: "burst";r=7;t=1, "daily";r=2310;t=21600
7 of the burst and 2,310 of the day are left; the daily reset is 6 h (21,600 s) away.
2GET/v1/usage

The key's plan, limits and what is left. The SDK calls it once at start-up to seed its pacing before the first response arrives.

Request
Headers
Authorization: Bearer sk_live_…
Response200
{
"plan": "free",
"limits": [
{ "name": "burst", "q": 20, "w": 2, "r": 7 },
{
"name": "daily",
"q": 10000,
"w": 86400,
"r": 2310,
"reset_at": "2026-10-01T00:00:00Z"
}
]
}
Also setsCache-Control: no-store

Errors

{ "type": "https://docs.shiplane.example/errors/<code>", "title": "Short summary", "status": 429, "detail": "What happened, in words", "limit": "burst | daily", "retry_after_s": 1 }
HTTPTypeBody codeClient behaviour
429retry
rate_limited
Honour Retry-After, add jitter, count the retry against the budget.
429error
quota_exhausted
Do not auto-retry. Surface it to the merchant; the next attempt is after the reset.
503retry
overloaded
Back off with jitter; the SDK's adaptive throttle engages. Not charged to the quota.
409error
idempotency_conflict
The same Idempotency-Key with a different body is a client bug. Do not retry.
400error
cost_exceeds_burst
The call costs more than the key's burst, so it can never fit: e.g. a 50-label batch if batches were priced per label (today POST /v1/labels:batch is a flat 10, which always fits a Free burst of 20). Split the request.
Was this section helpful?

Optimizations.

Six things that keep rejections cheap and keep retries from becoming the outage.

Adaptive client-side throttling
Over the last 2 minutes the SDK counts requests and accepts, and drops a new request locally with probability max(0, (requests − K·accepts) / (requests + 1)), K = 2 (Google SRE book). Worked: 1,000 requests and 300 accepts give (1,000 − 600) / 1,001 ≈ 40% that never leave the client. A lower K (1.1) starts dropping sooner.
Retry budgets, one layer only
At most 3 attempts per request, and retries at most 10% of a client's requests (SRE book); the Shiplane SDK measures that ratio over the last minute (our choice). Only the SDK retries; the app and internal proxies pass errors up. Why one layer: failure-models/timeouts-and-retries.
Retry-After plus jitter
wait = Retry-After + random(0, J), J = min(20 s, max(1 s, Retry-After)). Without Retry-After (a timeout, a reset connection) fall back to full jitter, random(0, min(cap, base·2^attempt)), which did the least total work in AWS's comparison. Backoff theory lives in failure-models/timeouts-and-retries.
Cheap rejection
Decide from the API key and route in the headers, before reading the body or touching a database. Caveat from the SRE book: rejecting is not free, and for requests about as cheap as the rejection itself it saves little. That is why client-side throttling matters: the cheapest rejection never reaches the server.
Soft limits before hard ones
At 80% of the daily quota (8,000 of 10,000 on Free) add a warning to the response and email the merchant; stop at 100%. Many quota 429s catch the merchant by surprise (our expectation); an early warning gives them hours to upgrade or slow down.
Criticality for the shedder
Tag each route: a label purchase is critical, a rate quote is sheddable. Under overload the shedder drops quotes first. Stripe reserves 20% of its fleet for critical requests; the SRE book's levels run from CRITICAL_PLUS down to SHEDDABLE.

When the SDK starts dropping locally

  • K = 2
  • K = 1.1
When the SDK starts dropping locallyWith K = 2 the SDK drops nothing until fewer than half its requests are accepted; K = 1.1 starts dropping once acceptance falls below about 91%.00.20.40.60.8100.20.40.60.811,000 sent, 300 accepted → 40%K = 2K = 1.1Share dropped by the SDKShare of requests the server acceptedWhen the SDK starts dropping locallyWith K = 2 the SDK drops nothing until fewer than half its requests are accepted; K = 1.1 starts dropping once acceptance falls below about 91%.00.20.40.60.8100.20.40.60.811,000 sent, 300 accepted → 40%K = 2K = 1.1Share dropped by the SDKShare of requests the server accepted
p_drop ≈ max(0, 1 − K × accept ratio) for large request counts. At a 30% accept ratio, K = 2 drops 40%.
Data
Share of requests the server acceptedK = 2K = 1.1
011
0.10.80.89
0.20.60.78
0.30.40.67
0.40.20.56
0.500.45
0.600.34
0.700.23
0.800.12
0.900.01
100
  • At 0.3: 1,000 sent, 300 accepted → 40%

The SDK's retry decision

function nextWait(res: Response, attempt: number, budget: RetryBudget): number | null {
  const problem = parseProblem(res);
  if (res.status === 429 && problem?.limit === "daily") return null; // surface, don't retry
  if (res.status !== 429 && res.status !== 503) return null;
  if (attempt >= 3 || budget.retries >= 0.1 * budget.requests) return null;
  const ra = parseRetryAfterSeconds(res.headers.get("Retry-After")); // seconds or HTTP-date
  if (ra !== null) {
    const j = Math.min(20, Math.max(1, ra));
    return (ra + Math.random() * j) * 1000;
  }
  const cap = 20, base = 0.5;
  return Math.random() * Math.min(cap, base * 2 ** attempt) * 1000; // full jitter
}
What the SDK does with each answer
CodeOutcomeKindWhat happensReacts
429rate_limitederrorWait Retry-After plus jitter and retry with the same Idempotency-Key, if the budget allows.2Shiplane SDK
429quota_exhaustederrorRaise QuotaExceeded with the reset time; the merchant's app decides whether to queue orders or upgrade.1Merchant app
503overloadederrorRetry with jitter under the budget; the adaptive throttle starts dropping locally as accepts fall.2Shiplane SDK
201r below 10%successSuccess, but little is left. Space the next calls out until the reset instead of spending the rest at once.2Shiplane SDK
Was this section helpful?

Trade-offs.

The chosen option is first; the others stay visible so the reasoning can be checked.

01
Which status for over-quota
Chosen:429 plus Retry-After
  • Pro:The standard meaning (RFC 6585) that clients and proxies understand
  • Pro:Caches must not store it
  • Pro:Retry-After gives an exact wait
Downside we accept:
  • Con:Burst and quota rejections share a code; the body has to tell them apart
Ruled out:403

Clients read it as a permission error and stop, or never retry at all

Ruled out:503

Means "the server is overloaded", which blames Shiplane; Fires the wrong alerts; Clients cannot tell a plan limit from an outage

02
How much to reveal
Chosen:Exact remaining for authenticated keys, nothing for anonymous login limits
  • Pro:Merchants can pace precisely
  • Pro:Attackers on /v1/login get no counter to pace against
Downside we accept:
  • Con:Two behaviours to document and test
Ruled out:Headers everywhere

Tells a credential-stuffing script exactly how to stay under 5 attempts a minute per IP

Ruled out:No headers, only 429s

Clients learn the limit by hitting it, so every client produces rejections

03
Reject or hold the excess
Chosen:Reject fast
  • Pro:Constant, tiny cost per excess request
  • Pro:The client sees the problem and can react
Downside we accept:
  • Con:Well-behaved clients must implement retries
Ruled out:Delay at the server (NGINX burst with delay)

Holds a connection and memory per waiting request; Hides overload until timeouts appear

04
Who retries
Chosen:The SDK, under a budget
  • Pro:One place with Retry-After, jitter and idempotency keys done right
  • Pro:Bounded extra load
Downside we accept:
  • Con:Merchants on raw HTTP must copy the behaviour from the docs
Ruled out:Every layer retries

Retries multiply layer by layer (see failure-models/timeouts-and-retries)

Ruled out:Nobody retries

Every merchant writes their own loop, usually without jitter or budgets

What goes wrong on the client side

FailureImpactDetectionMitigationMeanwhile
A client ignores Retry-After and retries in a tight loop2Shiplane SDKIts own traffic is mostly rejections; the gateway spends CPU on themShare of a key's requests arriving inside its own Retry-After windowEscalate for that key (a longer Retry-After, then a temporary block); contact the merchantOther keys are unaffected; their limits are separate
A CDN or proxy caches a 4293GatewayEvery client behind it is rejected long after the limit cleared429s with an Age header, or 429s for keys with quota leftRFC 6585 forbids caching 429; also send Cache-Control: no-store and check the CDN configDirect clients work
Clock skew with an HTTP-date Retry-After2Shiplane SDKA client clock 30 s fast retries 30 s early; one 30 s slow waits 30 s too longRetries arriving before the stated dateSend delay-seconds, never a dateOnly skewed clients misbehave
A retried POST without an idempotency key1Merchant appTwo labels bought and charged for one orderDuplicate labels per order in reconciliationThe SDK always sets Idempotency-Key; the API rejects label POSTs without oneNothing; the duplicate is prevented, not repaired
RateLimit headers read from a cached response2Shiplane SDKThe SDK paces on stale numbersRateLimit on responses carrying AgeIgnore the fields on cached responses, as the draft saysPacing falls back to reacting to 429s
Shedder mislabels label purchases as sheddable5Load shedderPaid work is dropped first during overload503 share by route during load testsCriticality set per route in reviewed config, with a test that critical routes are shed lastQuotes still work
Was this section helpful?
Related
Limits across many servers
Read next