Headers, 429s and backoff.
A rate limiter is half an API contract. The other half is what the client sees, which is headers that show the quota running down, a 429 that says which limit was hit and exactly when to come back, and an SDK that backs off with jitter and throttles itself before the server has to.
Builds on Rate-limiting algorithms.
Requirements.
The server-side limiter already decides allow or deny. This topic designs everything a client sees and does with that decision: the headers, the rejection, and the retries that follow.
Functional requirements
Why the client half matters
A limiter that only says no protects the server once. Then comes the second wave: thousands of rejected clients retrying at the same moment, retries stacked on retries at every layer, and merchants who cannot tell a plan limit from an outage. The server's answer has to carry enough information for a well-behaved client to do the right thing, and the SDK Shiplane ships has to actually do it.
Capacity estimates
Assumptions for Shiplane: 300K req/s at peak across three regions; after a 5-second blip, 10,000 SDK instances have been told to retry.
What one rejection turns into
- SDK instances rejected during a 5 s blip
- 10,000assumption
- Retry-After sent with each rejection
- 5 s
- Extra random wait on top of Retry-After
- uniform over 0–5 sJ = min(20 s, max(1 s, Retry-After))
- Retries at t = 5 s without jitterevery client sleeps exactly Retry-After10,000 in the same instantfrom SDK instances rejected during a 5 s blip and Retry-After sent with each rejection
- Retry rate with jitter10,000 ÷ 5 s window, over t = 5 to 10 s2,000/sfrom SDK instances rejected during a 5 s blip, Retry-After sent with each rejection and Extra random wait on top of Retry-After
- Rate headers on every response, uncompressed300K req/s × ~120 B (RateLimit + RateLimit-Policy)≈ 36 MB/s (≈ 290 Mbit/s) across all regionsWith HTTP/2 header compression the unchanging Policy line shrinks to an index of a byte or two after the first response.
- Retry-After says when; jitter says not everyone at once.
- Retry at one layer only, and cap retries at 10% of traffic (about 1.1× load at worst). How retries multiply across layers, and the budget arithmetic, are worked in failure-models/timeouts-and-retries.
Retries per second after the blip
- No jitter
- Retry-After + random(0, 5 s)
Data
| Seconds after the rejection | No jitter (req/s) | Retry-After + random(0, 5 s) (req/s) |
|---|---|---|
| 5–6 s | 10,000 | 2,000 |
| 6–7 s | 0 | 2,000 |
| 7–8 s | 0 | 2,000 |
| 8–9 s | 0 | 2,000 |
| 9–10 s | 0 | 2,000 |
High-level design.
The gateway turns the limiter's verdict into headers and status codes; the SDK turns those back into pacing and retries. Everything else is the normal request path.
Rate limiting, client to edge
One allowed call, one rejected call
The burst’s last token
- Merchant app → Shiplane SDK: createLabel(order 8812)
- Shiplane SDK → Gateway: POST /v1/labels · Idempotency-Key: k-8812
- Gateway → Limiter: check key sk_live_…, route labels
- Limiter → Gateway (reply): allow, r=0, t=1
- Gateway → Shiplane API: forward
- Shiplane API → Gateway (reply): 201 label
- Gateway → Shiplane SDK (reply): 201 · RateLimit-Policy: "burst";q=200;w=2 · RateLimit: "burst";r=0;t=1
The next request: 429, then a retry
- Note over Shiplane SDK: r=0: the last token is gone. A second worker on this key hasn't seen that and sends 2 ms later.
- Shiplane SDK → Gateway: POST /v1/labels · Idempotency-Key: k-8813
- Gateway → Limiter: check
- Limiter → Gateway (reply): deny, next token in 8 ms
- Gateway → Shiplane SDK (reply): 429 · Retry-After: 1 · problem+json limit=burst
- Note over Shiplane SDK: Waits 1 s + random(0, 1 s) = 1.6 s (8 ms rounds up to 1 s). Refill 1.6 s × 100/s = 160; other workers spend 10.
- Shiplane SDK → Gateway: retry POST /v1/labels · same Idempotency-Key: k-8813
- Gateway → Shiplane SDK (reply): 201 · RateLimit: "burst";r=149;t=1 (160 − 10 − 1)
- Shiplane SDK → Merchant app (reply): label for order 8813
What a rejection costs the edge
Rejected at the edge vs admitted
Scenario 1 of 2: As described.
Timeline as a list
Rejected at the edge vs admitted: 6 lanes, from 0 ms to 2.5 ms.
- 0–0.8 ms · Request at the gateway · 429 in 0.8 ms
- 0–0.2 ms · Parse headers + look up key · key cache hit
- 0.2–0.7 ms · Limiter check · check (denied)
- 0.7–0.8 ms · Write response · 429 + headers
- 0.8 ms · Read request body · body never parsed (ok)
- 1 ms · all lanes · 1 ms limiter budget (p99) (deadline)
Headers in the wild
| Header | Where it appears | Meaning |
|---|---|---|
| Retry-After | RFC 9110. On 429 and 503 (and 3xx). | Wait at least this long. Either delay-seconds ("120") or an HTTP-date. The only one that is a finished standard for this purpose. |
| RateLimit-Policy | IETF httpapi draft (-11). Any response; Shiplane sends it on every one. | The policies that apply: q (quota), w (window in seconds), optional qu (unit: requests, content-bytes, concurrent-requests) and pk (partition key). |
| RateLimit | IETF httpapi draft (-11). Any response; Shiplane sends it on every one. | What is left, per named policy: r (remaining) and t (seconds in which the remaining quota r applies; in effect, until it resets). Clients must cope with responses that omit it. |
| x-ratelimit-limit · -remaining · -used · -reset | GitHub REST API, every response. | The widespread pre-draft convention. GitHub's reset is an epoch timestamp in seconds, not a delay. |
| x-envoy-ratelimited | Envoy on the 429s it generates itself. | Marks the 429 as the proxy's, so a client or a retry policy can tell it from a 429 the upstream sent. |
Data model.
The contract is small. The server sends policies, what is left and a problem body; the SDK keeps four little pieces of state to act on them.
What crosses the wire, and what the SDK keeps
The 429 body when a Free plan's daily quota runs out at 18:40 UTC
{
"type": "https://docs.shiplane.example/errors/quota_exhausted",
"title": "Daily quota used up",
"status": 429,
"detail": "10,000 of 10,000 requests used today (Free plan).",
"limit": "daily",
"retry_after_s": 19200
}The daily quota resets at 00:00 UTC. From 18:40 that is 5 h 20 min = 5 × 3,600 + 20 × 60 = 19,200 s, the same number as the Retry-After header. The limit field is what lets the SDK treat this 429 differently from a burst 429 that clears in a second.
One SDK call, from send to answer
States of2Shiplane SDK
RetryState moves through these states. Only a burst 429 or a 503 leads back to sending, and only while the attempt count and the retry budget allow it.
| From → To | Event | Guard | Action |
|---|---|---|---|
| Sending → Succeeded | 2xx | update PacingState | |
| Sending → Waiting to retry | rate_limited or 503 | attempt < 3, budget left | |
| Sending → Error returned to app | quota_exhausted | raise QuotaExceeded | |
| Sending → Error returned to app | rate_limited or 503 | retries used up | |
| Waiting to retry → Sending | timer fires | attempt++, same key | |
| Waiting to retry → Dropped locally | local throttle check | random < p_drop |
- Sendingstart
- Waiting to retry
- Retry-After + random(0, J)
- Succeededend
- Error returned to apperror
- quota used up, budget spent or 3 attempts
- Dropped locallyerror
- adaptive throttle said no
Interface.
Every endpoint returns the rate headers; the label endpoint shows every way a call can end. Multiple response headers are shown separated by · .
Buys a shipping label. Shown for a Free-plan key (10 req/s, burst 20, 10,000 a day) at about 18:00 UTC.
The key's plan, limits and what is left. The SDK calls it once at start-up to seed its pacing before the first response arrives.
Errors
| HTTP | Type | Body code | Client behaviour |
|---|---|---|---|
| 429 | retry | rate_limited | Honour Retry-After, add jitter, count the retry against the budget. |
| 429 | error | quota_exhausted | Do not auto-retry. Surface it to the merchant; the next attempt is after the reset. |
| 503 | retry | overloaded | Back off with jitter; the SDK's adaptive throttle engages. Not charged to the quota. |
| 409 | error | idempotency_conflict | The same Idempotency-Key with a different body is a client bug. Do not retry. |
| 400 | error | cost_exceeds_burst | The call costs more than the key's burst, so it can never fit: e.g. a 50-label batch if batches were priced per label (today POST /v1/labels:batch is a flat 10, which always fits a Free burst of 20). Split the request. |
Optimizations.
Six things that keep rejections cheap and keep retries from becoming the outage.
When the SDK starts dropping locally
- K = 2
- K = 1.1
Data
| Share of requests the server accepted | K = 2 | K = 1.1 |
|---|---|---|
| 0 | 1 | 1 |
| 0.1 | 0.8 | 0.89 |
| 0.2 | 0.6 | 0.78 |
| 0.3 | 0.4 | 0.67 |
| 0.4 | 0.2 | 0.56 |
| 0.5 | 0 | 0.45 |
| 0.6 | 0 | 0.34 |
| 0.7 | 0 | 0.23 |
| 0.8 | 0 | 0.12 |
| 0.9 | 0 | 0.01 |
| 1 | 0 | 0 |
- At 0.3: 1,000 sent, 300 accepted → 40%
The SDK's retry decision
function nextWait(res: Response, attempt: number, budget: RetryBudget): number | null {
const problem = parseProblem(res);
if (res.status === 429 && problem?.limit === "daily") return null; // surface, don't retry
if (res.status !== 429 && res.status !== 503) return null;
if (attempt >= 3 || budget.retries >= 0.1 * budget.requests) return null;
const ra = parseRetryAfterSeconds(res.headers.get("Retry-After")); // seconds or HTTP-date
if (ra !== null) {
const j = Math.min(20, Math.max(1, ra));
return (ra + Math.random() * j) * 1000;
}
const cap = 20, base = 0.5;
return Math.random() * Math.min(cap, base * 2 ** attempt) * 1000; // full jitter
}| Code | Outcome | Kind | What happens | Reacts |
|---|---|---|---|---|
| 429 | rate_limited | error | Wait Retry-After plus jitter and retry with the same Idempotency-Key, if the budget allows. | 2Shiplane SDK |
| 429 | quota_exhausted | error | Raise QuotaExceeded with the reset time; the merchant's app decides whether to queue orders or upgrade. | 1Merchant app |
| 503 | overloaded | error | Retry with jitter under the budget; the adaptive throttle starts dropping locally as accepts fall. | 2Shiplane SDK |
| 201 | r below 10% | success | Success, but little is left. Space the next calls out until the reset instead of spending the rest at once. | 2Shiplane SDK |
Trade-offs.
The chosen option is first; the others stay visible so the reasoning can be checked.
- Pro:The standard meaning (RFC 6585) that clients and proxies understand
- Pro:Caches must not store it
- Pro:Retry-After gives an exact wait
- Con:Burst and quota rejections share a code; the body has to tell them apart
Clients read it as a permission error and stop, or never retry at all
Means "the server is overloaded", which blames Shiplane; Fires the wrong alerts; Clients cannot tell a plan limit from an outage
- Pro:Merchants can pace precisely
- Pro:Attackers on /v1/login get no counter to pace against
- Con:Two behaviours to document and test
Tells a credential-stuffing script exactly how to stay under 5 attempts a minute per IP
Clients learn the limit by hitting it, so every client produces rejections
- Pro:Constant, tiny cost per excess request
- Pro:The client sees the problem and can react
- Con:Well-behaved clients must implement retries
Holds a connection and memory per waiting request; Hides overload until timeouts appear
- Pro:One place with Retry-After, jitter and idempotency keys done right
- Pro:Bounded extra load
- Con:Merchants on raw HTTP must copy the behaviour from the docs
Retries multiply layer by layer (see failure-models/timeouts-and-retries)
Every merchant writes their own loop, usually without jitter or budgets
What goes wrong on the client side
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| A client ignores Retry-After and retries in a tight loop2Shiplane SDK | Its own traffic is mostly rejections; the gateway spends CPU on them | Share of a key's requests arriving inside its own Retry-After window | Escalate for that key (a longer Retry-After, then a temporary block); contact the merchant | Other keys are unaffected; their limits are separate |
| A CDN or proxy caches a 4293Gateway | Every client behind it is rejected long after the limit cleared | 429s with an Age header, or 429s for keys with quota left | RFC 6585 forbids caching 429; also send Cache-Control: no-store and check the CDN config | Direct clients work |
| Clock skew with an HTTP-date Retry-After2Shiplane SDK | A client clock 30 s fast retries 30 s early; one 30 s slow waits 30 s too long | Retries arriving before the stated date | Send delay-seconds, never a date | Only skewed clients misbehave |
| A retried POST without an idempotency key1Merchant app | Two labels bought and charged for one order | Duplicate labels per order in reconciliation | The SDK always sets Idempotency-Key; the API rejects label POSTs without one | Nothing; the duplicate is prevented, not repaired |
| RateLimit headers read from a cached response2Shiplane SDK | The SDK paces on stale numbers | RateLimit on responses carrying Age | Ignore the fields on cached responses, as the draft says | Pacing falls back to reacting to 429s |
| Shedder mislabels label purchases as sheddable5Load shedder | Paid work is dropped first during overload | 503 share by route during load tests | Criticality set per route in reviewed config, with a test that critical routes are shed last | Quotes still work |