Abstractions and RPCIdempotency and retries

100%

Idempotency and retries.

The label call timed out. Retry, and you may pay the carrier twice; don't, and the parcel may never ship. An idempotent API removes the dilemma: the client sends the same key with every attempt, and the server does the work once and replays the answer.

Intermediate22 minUpdated 30 Sept 2026

Builds on RPC and what it hides.

The idea.

A retry is the only way to learn what happened to a call that timed out. Idempotency is what makes that retry harmless.

The Fulfilment service asks the Label service for a shipping label for order 918273, a 1,450 g parcel going by ground. The call's 250 ms deadline passes with no answer. Either the request never arrived, or the label was bought and the reply got lost on the way back. From the caller's side the two look identical. If it gives up, a packed parcel may sit without a label. If it retries, the carrier may bill for two labels and two tracking numbers may end up on two boxes.

Retrying is the only way to get an answer, so the operation has to tolerate being repeated. Some operations already do: setting a value, deleting by id, replacing a whole resource with PUT. Running them twice ends in the same state as running them once. Others do not: creating something, charging, sending a message. Each run makes a new thing. Those can be made safe by having the server remember what it has already done, and recognise a repeat by a key the client sends.

HTTP writes this down. RFC 9110 calls PUT, DELETE and the read-only methods idempotent, and says a client should not automatically retry a non-idempotent request unless it knows the request is idempotent anyway. An idempotency key is how it knows.

Is it safe to run twice?

OperationIdempotent?Why
SET hold = trueYesThe second run writes the value that is already there.
DELETE /labels/L-5521YesThe second run finds it gone and may return 404, but the state is the same: no label L-5521.
PUT /labels/L-5521/hold {"on":true}YesA full replacement: the end state does not depend on how many times it ran.
weight_g = weight_g + 50NoA relative update: each run adds another 50 g.
CreateLabel(order 918273)No, unless keyedEach run buys a new label from the carrier.
send "your parcel shipped" emailNo, unless keyedEach run sends another email.

The reply is lost; the retry gets the same label

The reply is lost; the retry gets the same label, as an ordered list of steps:
The reply is lost the retry gets the same label10 steps between Fulfilment service, Label service, Idempotency records. The steps are listed as text after the diagram.Idempotency recordsLabel serviceFulfilment servicebuys label L-5521, saves the response with the keyno reply by the 250 ms deadline; retry with the same keyCreateLabel(order 918273), key label:918273:1:v11claim label:918273:1:v12new key, lease taken3201 Label L-5521 (lost on the way back ✕)4CreateLabel(order 918273), key label:918273:1:v15claim label:918273:1:v16already completed, saved response7201 Label L-5521 (replayed, no second purchase)8
  1. Fulfilment service → Label service: CreateLabel(order 918273), key label:918273:1:v1
  2. Label service → Idempotency records: claim label:918273:1:v1
  3. Idempotency records → Label service (reply): new key, lease taken
  4. Note over Label service: buys label L-5521, saves the response with the key
  5. Label service → Fulfilment service (reply): 201 Label L-5521 (lost on the way back ✕)
  6. Note over Fulfilment service: no reply by the 250 ms deadline; retry with the same key
  7. Fulfilment service → Label service: CreateLabel(order 918273), key label:918273:1:v1
  8. Label service → Idempotency records: claim label:918273:1:v1
  9. Idempotency records → Label service (reply): already completed, saved response
  10. Label service → Fulfilment service (reply): 201 Label L-5521 (replayed, no second purchase)
Was this section helpful?

How it works.

A key made by the client, a record kept by the server, and two short transactions around the side effect, the second tying the record to the result.

1

The client makes a key per intent, not per attempt

for (let attempt = 1; attempt <= 4; attempt++) {
  const key = crypto.randomUUID();          // new key each time
  try {
    return await labels.createLabel(parcel, { idempotencyKey: key });
  } catch (e) {
    if (!isRetryable(e)) throw e;
    await backoff(attempt);
  }
}
// Every retry looks like a new request: up to 4 labels.
// Derived and stable: the same parcel always gets the same key,
// even after Fulfilment restarts mid-retry. Nothing to save.
const key = `label:${parcel.orderId}:${parcel.parcelNo}:v${parcel.labelVersion}`;
// e.g. label:918273:1:v1

for (let attempt = 1; attempt <= 4; attempt++) {
  try {
    return await labels.createLabel(parcel, { idempotencyKey: key });
  } catch (e) {
    if (!isRetryable(e)) throw e;
    await backoff(attempt);
  }
}

The key names the intent: "a label for parcel 1 of order 918273". There are two sound ways to get one. A key derived from the parcel's own identity, as in the snippet (label:918273:1:v1), needs no storage: a restarted caller derives the same key again. A random UUIDv4 also works, but only if it is saved with the parcel before the first attempt; otherwise a restart mints a fresh one and the retry looks new. The labelVersion suffix is there for the day a label genuinely has to be reprinted: that is a new intent and gets a new key. Keep personal data (email addresses, names) out of keys; they are logged and stored for a day or more. How to generate unique ids at scale is its own topic (Unique ID generator).

2

The server keeps a record per key

PostgreSQLOwned by 3Idempotency records

The key records live in the same database as the labels, so a key and the label it produced commit together or not at all.

idempotency_keysTTL 24 h after completion
ColumnTypeKeyNote
caller_idtextprimarykeys are scoped per caller
idem_keytextprimarysecond half of the primary key
request_hashbyteaSHA-256 of method + path + canonical body
stateenumstarted | completed
locked_untiltimestamptz nullablelease for the attempt in flight
lease_tokenuuid nullablewhich attempt holds the lease; set on claim and on takeover
response_statusint nullable
response_bodyjsonb nullable
created_attimestamptz
expires_attimestamptzindex
labels
ColumnTypeKeyNote
label_idtextprimary
order_idbigintuniqueunique together with parcel_no
parcel_nointunique
tracking_notext
cost_centsint
created_by_keytextthe idempotency key that made it
Access patterns
QueryUsesHow
Have I seen this key from this caller?idempotency_keysprimary key (caller_id, idem_key)
Purge expired recordsidempotency_keysrange scan on expires_at, in batches
At most one label per parcellabelsunique (order_id, parcel_no), a second line of defence
3

The life of a key

One idempotency key, from first request to purge

States of3Idempotency records

One idempotency key, from first request to purge. 4 states, 9 transitions. The table below lists them.
One idempotency key, from first request to purgeThe states of Idempotency records. 4 states, 9 transitions. The table below lists them.

first request / insert, lease 30 s

fails validation / delete, 400

repeat [same body, lease live] / 409
repeat [same body, lease expired] / take over
repeat [new body] / 422

handler returns [lease still ours] / save response

repeat [same body] / replay
repeat [new body] / 422

24 h later / delete row

No record

Started, lease held

Completed, response saved

Purged

3 steps.

Every repeat is a self-loop: what it gets back depends on the lease and on whether the body matches.

Transitions of One idempotency key, from first request to purge
From → ToEventGuardActionActor
No record → Started, lease heldfirst requestinsert, lease 30 sLabel service
Started, lease held → No recordfails validationdelete, 400Label service
Started, lease held → Started, lease heldrepeatsame body, lease live409Label service
Started, lease held → Started, lease heldrepeatsame body, lease expiredtake overLabel service
Started, lease held → Started, lease heldrepeatnew body422Label service
Started, lease held → Completed, response savedhandler returnslease still ourssave responseLabel service
Completed, response saved → Completed, response savedrepeatsame bodyreplayLabel service
Completed, response saved → Completed, response savedrepeatnew body422Label service
Completed, response saved → Purged24 h laterdelete rowsweeper
No recordstart
The server has never seen this (caller, key).
Started, lease held
One attempt owns the key until locked_until.
Completed, response saved
Status and body stored; every repeat gets them back.
Purgedend
Row deleted. A repeat after this is treated as brand new, so retention must outlast every retry.
4

The server algorithm

-- Txn 1: claim the key (or inspect it) and commit at once.
BEGIN;
INSERT INTO idempotency_keys
       (caller_id, idem_key, request_hash, state, locked_until, lease_token, created_at, expires_at)
VALUES ($caller, $key, $hash, 'started', now() + interval '30 s', $my_token, now(), now() + interval '24 h')
ON CONFLICT (caller_id, idem_key) DO NOTHING;
-- 1 row inserted → we hold the lease: COMMIT and go to the carrier.
-- 0 rows → the key already exists as a committed row:
SELECT request_hash, state, locked_until, response_status, response_body
  FROM idempotency_keys
 WHERE caller_id = $caller AND idem_key = $key
   FOR UPDATE;
--   request_hash <> $hash   → ROLLBACK; 422 idempotency_key_reused
--   state = 'completed'     → ROLLBACK; replay response_status + response_body
--   locked_until > now()    → ROLLBACK; 409 request_in_progress
--   lease expired           → take over with a fresh token:
UPDATE idempotency_keys
   SET locked_until = now() + interval '30 s', lease_token = $my_token
 WHERE caller_id = $caller AND idem_key = $key;
COMMIT;

-- Between the transactions: no transaction open, no row lock held.
-- POST carrier /labels, Idempotency-Key 'carrier:' || $key, timeout 20 s (< 30 s lease)

-- Txn 2: the effect and the outcome commit together, only if the lease is still ours.
BEGIN;
UPDATE idempotency_keys
   SET state = 'completed', locked_until = NULL, lease_token = NULL,
       response_status = 201, response_body = $response_json,
       expires_at = now() + interval '24 h'
 WHERE caller_id = $caller AND idem_key = $key
   AND state = 'started' AND lease_token = $my_token;
-- 0 rows updated → another attempt took over: ROLLBACK and write nothing.
INSERT INTO labels (label_id, order_id, parcel_no, tracking_no, cost_cents, created_by_key)
VALUES ($label_id, 918273, 1, $tracking_no, $cost_cents, $key);
COMMIT;

Txn 1 commits straight away so that other attempts can see the claim: an uncommitted row is invisible to them, and a crash would roll it back, so there would be no lease to answer 409 with or to take over. The carrier call then runs between the two transactions, because holding a database transaction open across a network call ties up a connection and row locks for as long as the carrier takes. In txn 2 the label row and the completed key commit together. If they didn't, a crash between them would leave a label with no record, so the retry buys a second one, or a record saying "completed" with no label behind it. The Amazon Builders' Library article gives the same rule for AWS APIs: record the client token and the mutation in one ACID transaction. The lease_token check makes a slow attempt that lost its lease roll back instead of overwriting the result of the attempt that took over.

This is Brandur Leach's Postgres design in miniature: split the handler into atomic phases around each foreign call, and save a recovery point on the key row after each phase when there is more than one. A retry that takes over an expired lease reads the recovery point and resumes from there instead of starting again; here the only foreign call is the carrier, and its own key makes redoing it safe.

5

Side effects outside your database

Pass a derived key downstream

Pass a derived key downstream, as an ordered list of steps:
Pass a derived key downstream5 steps between Label service, Idempotency records, Carrier API, Labels table. The steps are listed as text after the diagram.Labels tableCarrier APIIdempotency recordsLabel serviceCrash before this commit? The retry resends the carrier key and gets GRD000918273 back, not a second label.claim the key, lease 30 s, commit1POST /labels, key carrier:label:918273:1:v12label, tracking no. GRD0009182733insert label + complete key, one txn4
  1. Label service → Idempotency records: claim the key, lease 30 s, commit
  2. Label service → Carrier API: POST /labels, key carrier:label:918273:1:v1
  3. Carrier API → Label service (reply): label, tracking no. GRD000918273
  4. Label service → Labels table: insert label + complete key, one txn
  5. Note over Label service and Carrier API: Crash before this commit? The retry resends the carrier key and gets GRD000918273 back, not a second label.

The carrier call cannot join our database transaction, so a crash between buying the label and committing is always possible. The fix is to make the downstream call idempotent too, with a key derived from ours (carrier: plus our key). Each hop dedupes its own repeats, and idempotency composes along the chain. One exception to the usual rule from RPC basics (stop working once the caller's deadline passes and the call is cancelled): the Label service does not abandon a purchase it has started. Once money may be moving at the carrier, it finishes the call (carrier timeout 20 s, inside the 30 s lease) and records the outcome, so the caller's next retry is answered from the key. If a downstream offers no key, write the intent to an outbox table in the same transaction, have one worker call out with a unique job id, and reconcile against the provider's records afterwards (Payment system: reconciliation covers the money version).

The lease decides who may run

Scenario 1 of 2: As described.

Timeline as a list

The lease decides who may run: 4 lanes, from -4 s to 40 s.

  1. 0–10 s · Attempt 1 · attempt 1
  2. 0–30 s · Attempt 1, Retries · window: lease held by attempt 1
  3. 0 s · Caller → Attempt 1, arriving 0.2 s: CreateLabel
  4. 0.5–9.5 s · Carrier · buy label (carrier key)
  5. 4.8 s · Caller → Retries, arriving 5 s: retry
  6. 5 s · Retries · 409 (delayed)
  7. 10 s · Attempt 1 · process killed (error)
  8. 17.8 s · Caller → Retries, arriving 18 s: retry
  9. 18 s · Retries · 409 (delayed)
  10. 30 s · all lanes · lease expires (deadline)
  11. 31.8 s · Caller → Retries, arriving 32 s: retry
  12. 32–38 s · Retries · attempt 2 takes over
  13. 34 s · Carrier · same label replayed (ok, ok)
  14. 38 s · Retries → Caller, arriving 38.2 s: 201
  15. 38.2 s · Caller · 201 GRD000918273 (ok, ok)
Attempt 1 buys the label and dies before committing. Retries bounce off its lease with 409 until it expires; then attempt 2 takes over and the carrier's own key hands back the same label.
What the Label service answersStatus mapping follows the IETF Idempotency-Key draft (expired, but widely copied); the replay header is Stripe's convention.
CodeOutcomeKindWhat happensReacts
201First runsuccessThe label is created and the response saved under the key.2Label service
201ReplaysuccessSame key, same body, already completed: the saved response, with Idempotent-Replayed: true. Nothing runs.2Label service
409request_in_progresschallengeAn attempt with this key holds the lease. Wait briefly and retry with the same key.1Fulfilment service
422idempotency_key_reusederrorSame key, different body. A client bug; never retry it blindly.1Fulfilment service
400idempotency_key_missingerrorCreateLabel requires a key. Fix the client.1Fulfilment service
Was this section helpful?

In practice.

What the key store costs, what it saves, how long to keep keys, and how real APIs do it.

How big is the key store?

Assumptions
Average label creates
400/sour assumption for a large warehouse network
Peak label creates
2,000/sthe pre-holiday packing rush
Retention after completion
24 h (86,400 s)
Bytes per record
~512 Ba key of up to ~36 B, 32 B hash, ~300 B saved response, the rest row and index overhead
Working
  1. Keys per day400/s × 86,400 s34.56Mfrom Average label creates and Retention after completion
  2. Store size at average load34.56M × 512 B≈ 17.7 GBfrom Keys per day and Bytes per record
  3. Keys if the peak lasted all day2,000/s × 86,400 s172.8Mfrom Peak label creates and Retention after completion
  4. Worst-case store size172.8M × 512 B≈ 88.5 GBfrom Keys if the peak lasted all day and Bytes per record
What it means
  • One indexed table and a batch sweeper on expires_at are enough; no special store is needed.
  • Every create costs one extra insert and one update on the primary, which is the real price.

What the keys save

Assumptions
Label creates per day
34.56Mfrom the estimate above
Calls that time out
0.1%illustrative
Of those, already done server-side
60%illustrative: the reply was slow or lost, not the request
Average label price
$6.00illustrative
Working
  1. Timed-out calls per day34.56M × 0.00134,560from Label creates per day and Calls that time out
  2. Duplicate labels if retried without keys34,560 × 0.620,736/dayfrom Timed-out calls per day and Of those, already done server-side
  3. Money spent on duplicates20,736 × $6.00$124,416/dayfrom Duplicate labels if retried without keys and Average label price
  4. Parcels left unlabelled if never retried34,560 × 0.413,824/dayfrom Timed-out calls per day and Of those, already done server-side
What it means
  • Without keys the choice is between 20,736 duplicate labels and 13,824 stuck parcels a day. With keys, retrying is free.

How late do retries arrive?

How late do retries arrive?A 24-hour window still misses 0.8% of retries, the ones replayed from queues and by hand, and each of those can buy a second label.50%60%70%80%90%100%10s1m10m1h6h24h3d24 h purgeArrivedShare of retries that have arrived (%)Time since the first attempt (m = minutes)How late do retries arrive?A 24-hour window still misses 0.8% of retries, the ones replayed from queues and by hand, and each of those can buy a second label.50%60%70%80%90%100%10s1m10m1h6h24h3d24 h purgeArrivedShare of retries that have arrived (%)Time since the first attempt (m = minutes)
Illustrative distribution, not measured: in-process retries land within seconds, queued redeliveries within hours, manual re-runs after an incident within days. Of 34,560 retried calls a day, 0.8% is 276 that arrive after a 24 h purge; a 3-day window cuts that to 35.
Data
Time since the first attempt (m = minutes)Arrived (%)
10s62
1m85
10m91
1h94
6h97.5
24h99.2
3d99.9
  • 24 h purge: Time since the first attempt (m = minutes) = 24h

Set retention from the longest retry horizon, not the typical one: the oldest message a queue can redeliver, the longest a caller keeps a pending parcel before retrying, the day an operator re-runs yesterday's failed batch. Alert on requests that arrive with a key the server has already purged; each one is a possible duplicate. At 512 B a record, going from 24 h to 3 days triples the store to about 53 GB, which is cheap next to one duplicate label per few minutes.

How real APIs recognise a repeat

SystemMechanismWindowSame key, different request
Stripe APIIdempotency-Key header on POST, up to 255 characters. The first status and body are saved once execution begins, 500s included; validation failures and concurrent conflicts are not saved.Keys may be pruned once at least 24 h oldError: parameters are compared with the original request
Amazon EC2ClientToken, a case-sensitive string of up to 64 ASCII characters, on RunInstances and about 70 other actions.Retention not stated; idempotent per Region, or per Availability Zone when RunInstances names a zone or subnetIdempotentParameterMismatch error
Google APIs (AIP-155)Optional request_id field, a UUID4, on mutating methods.Left to each APIThe duplicate gets the earlier successful response
Protocol Buffers / gRPCMethod option idempotency_level = IDEMPOTENT or NO_SIDE_EFFECTS. A declaration for tools and clients, not a dedupe mechanism.n/an/a
Kafka producerenable.idempotence: a producer id plus a sequence number per partition; the broker drops a batch it has already written.The producer sessionn/a (see Queue semantics)
Birrell & Nelson RPC (1984)Each call carries a call identifier with a per-caller sequence number; the callee discards calls it has already seen.The last call per callern/a

Idempotent without a key store

Unique constraint on a natural key
UNIQUE(order_id, parcel_no) on labels. A second insert fails, and the handler returns the label that already exists. Needs no extra table and never expires, but only works when the domain has such a key.
Conditional writes
PATCH /labels/L-5521 with If-Match: "v7" (RFC 9110 §13.1.1), or a version column checked in the UPDATE. The first write moves the label to v8; a replay against v7 fails cleanly with 412 instead of applying twice.
Absolute, not relative, updates
Send weight_g = 1500, not weight_g + 50. The final state no longer depends on how many copies of the request arrive.
Was this section helpful?

Trade-offs.

The chosen option is first; the others stay visible so the reasoning can be checked.

01
How the server recognises a repeat
Chosen:Client-supplied idempotency key
  • Pro:Works for any operation, including ones with no natural key
  • Pro:The client decides what counts as the same intent
Downside we accept:
  • Con:A key store with a TTL to run
  • Con:Clients must make the key once and reuse it; a bug there silently disables the protection
Ruled out:Natural unique key, e.g. (order_id, parcel_no)

Only where the domain has one; A genuine second intent (a reprint) becomes impossible without adding a version to the key

Ruled out:Server-side hash of the request body alone

Two genuinely separate requests with identical bodies (two 1,450 g parcels) are merged into one label

02
Replay failures too, or only successes?
Chosen:Replay every final outcome once execution began, errors included
  • Pro:The caller always sees one consistent answer per key
  • Pro:A half-finished run is never started a second time
Downside we accept:
  • Con:A transient 500 sticks to the key; the client needs a new key to try again
Ruled out:Save only successes

A failure part-way through can run the first half twice unless the handler is resumable

03
Where the key records live
Chosen:Same database as the effect
  • Pro:Key and effect commit atomically
  • Pro:Nothing new to operate
Downside we accept:
  • Con:Adds an insert and an update per create to the primary
Ruled out:Separate cache (Redis SET NX with a TTL)

Not atomic with the effect: a crash in between still duplicates; Eviction under memory pressure silently forgets keys

How idempotency quietly fails

FailureImpactDetectionMitigationMeanwhile
A new key on every retry attempt1Fulfilment serviceDuplicate labels despite an "idempotent" APIMore than one label per (order_id, parcel_no)One key per intent, derived or saved before the first attemptParcels ship, but some carry a second paid label
Retention shorter than the retry horizon3Idempotency recordsA redelivery from a 2-day-old queue buys a second labelRequests carrying keys that were already purgedRetention at least the maximum queue age and manual re-run windowOnly late retries duplicate; fresh ones still replay
Keys scoped globally, not per caller2Label serviceOne caller's key collides with another's and gets their saved responseReplays whose caller differs from the record's creatorPrimary key (caller_id, idem_key)Most calls work; a colliding one gets the wrong label back
Effect and record in different stores3Idempotency recordsA crash between the two leaves a label with no record, and the retry buys anotherReconciliation of labels against completed keysOne transaction, or recovery points on the key rowRare doubles after crashes; normal traffic is fine
Downstream call sent without a key5Carrier APIOur retry buys a second label at the carrierCarrier invoice lines with no matching labelPass a derived key on every downstream callOur records look right; the carrier bills twice
Lease shorter than the slowest downstream call2Label serviceTwo attempts run at onceTakeovers while the original attempt is still loggingLease longer than the downstream timeout; complete only WHERE lease_token matchesThe derived carrier key still prevents a second label
Was this section helpful?
Builds on this
Delivery semantics
Read next