RPC and what it hides.
One line of code, labels.createLabel(order), turns into marshalling, a network round trip, a thread on another machine and a reply that may never arrive. RPC hides the first four. It cannot hide the last, and every design that uses it has to decide what to do when the answer is "I don't know".
The idea.
A function call between machines, and the one promise it cannot fully keep.
Why RPC exists
A warehouse's Fulfilment service packs orders. For each parcel it needs a shipping label, which a separate Label service buys from a carrier. Without RPC, the Fulfilment team would open a socket, invent a byte layout for the order ID, carrier and weight, write it, wait, parse whatever comes back, and handle every error by hand. The Label team would write the mirror image. Every new call between every pair of services would repeat that work.
Remote procedure call, as Birrell and Nelson built it at Xerox PARC in 1984, removes that work. Both teams agree on an interface. A tool generates code for each side, and the Fulfilment programmer writes one ordinary-looking line: labels.createLabel(order). Their design goal was to make a remote call behave as much like a local one as possible.
Ten years later, Waldo and colleagues at Sun pointed out four things no generated code can hide: latency (a remote call is thousands of times slower), memory (the other machine cannot follow your pointers), concurrency (other callers hit the same server at the same moment) and partial failure (one machine can fail while the other carries on, and the survivor may not know what happened). The last one shapes everything else in this topic.
What one call becomes
The same call, local and remote
| Aspect | Local call | Remote call |
|---|---|---|
| Time | Nanoseconds: a few for a plain call, tens for a small lookup | About 0.5 ms inside one datacenter; up to ~150 ms across continents (Dean, 2009). About 10,000× a 50 ns local lookup (assumption). |
| Arguments | Passed by reference; both sides share memory | Copied into bytes. A pointer means nothing on the other machine. |
| Failure | The call returns, or the whole process dies with it | Can fail halfway: request lost, server crashed mid-call, or reply lost |
| Concurrency | Runs on your thread | Runs on another machine's thread, next to other callers' requests |
How it works.
Five pieces, three jobs: agree on the interface, find the server, and move one call there and back. Then decide what "there and back" means when it fails.
1. Agree on an interface
package labels.v1;
service Labels {
rpc CreateLabel(CreateLabelRequest) returns (Label);
}
message CreateLabelRequest {
int64 order_id = 1;
string carrier = 2; // "ground" or "express"
int32 weight_g = 3;
}
message Label {
string label_id = 1;
string tracking_no = 2;
int64 cost_cents = 3;
}
ctx, cancel := context.WithTimeout(ctx, 250*time.Millisecond)
defer cancel()
label, err := labels.CreateLabel(ctx, &pb.CreateLabelRequest{
OrderId: 918273,
Carrier: "ground",
WeightG: 1450,
})
if err != nil {
// lost request? crashed server? lost reply?
}
The interface definition is the contract. A compiler turns it into a client stub and a server stub for each language: Birrell and Nelson's Lupine did this from Mesa interface modules, protoc does it from .proto files today. The numbers after each field (= 1, = 2, = 3) matter more than the names, as the next step shows.
2. Marshal the arguments
CreateLabel(918273, ground, 1450) on the wire
- 1 byte, JSON punctuation, value {
- order_id, 11 bytes, Field name (JSON only), value "order_id":
- order, 6 bytes, Value, value 918273
- 1 byte, JSON punctuation, value ,
- carrier, 10 bytes, Field name (JSON only), value "carrier":
- carrier, 8 bytes, Value, value "ground"
- 1 byte, JSON punctuation, value ,
- weight_g, 11 bytes, Field name (JSON only), value "weight_g":
- weight, 4 bytes, Value, value 1450
- 1 byte, JSON punctuation, value }
- 54 bytes of text, from 0 to 54
- field 1, 1 byte, Protobuf tag (field + wire type), value 08, 1 << 3 | 0 (varint)
- order, 3 bytes, Value, value 81 86 38, 918273 in three 7-bit groups
- field 2, 1 byte, Protobuf tag (field + wire type), value 12, 2 << 3 | 2 (length-delimited)
- len, 1 byte, Length prefix, value 06
- carrier, 6 bytes, Value, value ground
- field 3, 1 byte, Protobuf tag (field + wire type), value 18, 3 << 3 | 0 (varint)
- weight, 2 bytes, Value, value aa 0b, 1450 = 42 + 11 × 128
- 15 bytes, from 0 to 15
- compressed?, 1 byte, gRPC frame header, value 00
- length, 4 bytes, gRPC frame header, value 00 00 00 0f, 15 in 4 big-endian bytes
- the 15 protobuf bytes, 15 bytes, Value
- 5-byte prefix, from 0 to 5
- JSON punctuation
- Field name (JSON only)
- Protobuf tag (field + wire type)
- Length prefix
- Value
- gRPC frame header
- field 1
- 1 << 3 | 0 (varint)
- order
- 918273 in three 7-bit groups
- field 2
- 2 << 3 | 2 (length-delimited)
- field 3
- 3 << 3 | 0 (varint)
- weight
- 1450 = 42 + 11 × 128
- length
- 15 in 4 big-endian bytes
3. Find the server (binding)
The stub knows a service name, Labels, not an address. Binding turns the name into live instances: in 1984 the Grapevine database mapped an interface's type and instance names to machines; today it is DNS, Consul, a Kubernetes Service or an xDS control plane. The runtime then picks one instance, often through a load balancer, and keeps the connection open for later calls. If the registry is stale, the call goes to a machine that is gone, which the runtime sees as one more failed attempt.
4. One call, there and back
CreateLabel, there and back
- Fulfilment code → Client stub: createLabel(918273, ground, 1450)
- Client stub → RPC runtime (caller): 15 bytes + method name
- Note over RPC runtime (caller): registry: Labels → 10.2.7.14, 10.2.7.15
- Note over RPC runtime (caller): call ID 7f3a, deadline 250 ms
- RPC runtime (caller) → RPC runtime (callee): request 7f3a + bytes
- Note over RPC runtime (callee): Birrell & Nelson: seen 7f3a? no → dispatch (gRPC skips this)
- RPC runtime (callee) → Server stub: dispatch
- Server stub → Label handler: createLabel(order)
- Label handler → Server stub (reply): Label{L-5521, cost 612}
- Server stub → RPC runtime (callee) (reply): marshal the Label
- RPC runtime (callee) → RPC runtime (caller) (reply): reply 7f3a + bytes
- RPC runtime (caller) → Client stub (reply): match 7f3a, hand over bytes
- Client stub → Fulfilment code (reply): unmarshalled Label
5. Three failures that look the same
What the caller sees in 250 ms
Scenario 1 of 4: As described.
Timeline as a list
What the caller sees in 250 ms: 3 lanes, from -40 ms to 300 ms.
- 0 ms · Caller's runtime → Callee's runtime, arriving 3 ms: request 7f3a
- 5–180 ms · Label handler · buy label from carrier
- 182 ms · Callee's runtime → Caller's runtime, arriving 185 ms: reply
- 185 ms · Caller's runtime · Label L-5521 arrives (ok, ok)
- 250 ms · all lanes · deadline (deadline)
What the runtime promises about a call
| Semantics | Runtime does | Duplicates | Lost calls | Fits |
|---|---|---|---|---|
| Maybe | Sends once, never retries | Never | Possible: a lost request is simply gone | Fire-and-forget hints; UI actions a person can repeat |
| At least once | Retries until a reply or the deadline | Possible: a lost reply means the work runs twice | Never, while retries last | Reads such as GetLabel, and operations that are idempotent |
| At most once | Retries, and the server filters duplicate call IDs and resends its cached reply (Birrell and Nelson, Java RMI) | Never, unless the server crashes and forgets the IDs it saw | Possible if the server crashes: the caller gets an error and cannot tell whether the call ran | Calls where a double is worse than a miss |
| "Exactly once" (effectively once) | At most once, with the seen IDs and stored replies kept in durable storage, e.g. idempotency keys in a database | Never, even across server restarts | Never, while the caller keeps retrying with the same key | Anything with side effects, like CreateLabel. Costs a durable dedupe store on the server. |
Birrell and Nelson's runtime is the textbook at-most-once design. Every call carried an identifier made of the calling machine, the calling process and a sequence number; the callee remembered the last sequence number it had seen from each caller and discarded retransmissions. Their promise: if the call returns, the procedure ran precisely once; if it raises an exception, it ran once or not at all, and the caller is not told which. The table lived in memory, so a server crash erased it. Keeping that table in durable storage, keyed by a request ID the client chooses, is what idempotency keys do, and the idempotency topic builds it out for CreateLabel.
In practice.
What remote calls cost, how gRPC carries a deadline through a chain of them, and which failures a runtime may retry on its own.
What a remote call costs
A loop of remote calls
- Round trip inside one datacenter
- 0.5 msOrder of magnitude from Dean (2009), the same figure numbers to know uses; a modern same-zone hop is often lower.
- Order lines to price, one call each
- 40Assumption for a large order.
- The same lookup as an in-process call
- 50 nsAssumption; tens of nanoseconds.
- Extra encode, transfer and handler work per item in a batch
- 12.5 µsAssumption; puts one batched call at about 1 ms, the same as the cache multi-get in numbers to know.
- Chance any one backend call is slow
- 1%Assumption, for the fan-out line.
- 40 calls, one after anotherlines × rtt = 40 × 0.5 ms20 msfrom Order lines to price, one call each and Round trip inside one datacenter
- One call carrying all 40 itemsrtt + lines × per-item = 0.5 ms + 0.5 ms~1 msfrom Round trip inside one datacenter, Order lines to price, one call each and Extra encode, transfer and handler work per item in a batch
- Batching savessequential ÷ batched = 20 ÷ 1~20×from 40 calls, one after another and One call carrying all 40 items
- The same loop in-processlines × local = 40 × 50 ns2 µsfrom Order lines to price, one call each and The same lookup as an in-process call
- 40 calls in parallel; share of requests that wait on at least one slow call1 − (1 − slow)^lines = 1 − 0.99^40 = 1 − 0.669~33%from Chance any one backend call is slow and Order lines to price, one call each
- A remote call inside a loop is the classic RPC mistake. Offer a batch method (PriceLines, CreateLabels) instead. The same 20 ms against ~1 ms arithmetic for a cache multi-get is in numbers to know.
- Going parallel fixes the 20 ms but makes rare slowness common. Tail latency is covered in latency percentiles.
Pricing 40 order lines three ways
Data
| Approach | Time (ms) |
|---|---|
| 40 calls in a row | 20 |
| 1 batched call | 1 |
| In-process loop | 0.002 |
One deadline, passed down
An 800 ms budget, passed down as grpc-timeout
| Hop | Starts at | What it sends on |
|---|---|---|
| API gateway → Fulfilment (PackOrder) | 20 ms | grpc-timeout: 780m |
| Fulfilment → Label service (CreateLabel) | 60 ms | grpc-timeout: 740m |
| Label service → Carrier API (POST /labels) | 60 ms | HTTP client timeout 700 ms (HTTP has no deadline header) |
gRPC sets no deadline by default, so a call to a hung server can wait forever. A deadline travels as the grpc-timeout header in relative units (780m is 780 milliseconds; H, M, S, m, u and n are allowed), so the machines' clocks need not agree; each hop sends what is left of the user's 800 ms, as the table shows. Java and Go servers pass the incoming deadline on to their own outgoing calls automatically; C++ needs propagation enabled, and elsewhere the handler has to pass it on itself. When the deadline passes, the client gets DEADLINE_EXCEEDED and the server sees the call CANCELLED, and it should stop working on it. Birrell and Nelson deliberately had no call timeout: while waiting, their runtime probed the server and raised an exception only if the server stopped answering, just as a local call waits for as long as the procedure runs. Choosing the numbers, and what happens down a chain that has no deadline (work nobody will read), are worked through in timeouts and retries.
Which failures the runtime may retry
| gRPC status | What it usually means | Retry? |
|---|---|---|
| UNAVAILABLE (14) | Server unreachable, restarting or shedding load | Yes, with backoff, if the method is idempotent |
| DEADLINE_EXCEEDED (4) | Time ran out. The work may or may not have happened. | Only if idempotent and the caller still has budget |
| RESOURCE_EXHAUSTED (8) | A quota or rate limit was hit | Later, after backing off; not in a tight loop |
| INVALID_ARGUMENT (3) | The request is wrong, e.g. weight_g below zero | Never: it will fail the same way |
| NOT_FOUND (5) / ALREADY_EXISTS (6) | State on the server, not a transient fault | Never |
| INTERNAL (13) | A bug or broken invariant on the server | Usually not |
One gRPC call, from the caller's runtime
States of3RPC runtime (caller)
Retries happen only while an attempt is uncommitted. Once response headers arrive the call is committed and never retried, and a passed deadline leaves the outcome unknown whether or not headers came.
| From → To | Event | Guard | Action | Actor |
|---|---|---|---|---|
| Call created → Attempt in flight | stub sends | start deadline | client stub | |
| Attempt in flight → Attempt in flight | never reached the app | first time | resend | runtime |
| Attempt in flight → Failed with a status | UNAVAILABLE | attempts = max | retry policy | |
| Attempt in flight → Attempt in flight | UNAVAILABLE | attempts < max | back off | retry policy |
| Attempt in flight → Committed | headers arrive | server | ||
| Attempt in flight → Failed with a status | non-retryable status | server | ||
| Committed → OK | trailers: OK | server | ||
| Committed → Failed with a status | trailers: error | server | ||
| Attempt in flight → Deadline passed: outcome unknown | deadline passes | cancel | runtime | |
| Committed → Deadline passed: outcome unknown | deadline passes | cancel | runtime |
- Call createdstart
- Attempt in flight
- No response headers yet; retries allowed
- Committed
- Response headers received
- OKend
- Failed with a statuserror
- Deadline passed: outcome unknownerror
gRPC retries a call on its own in two cases. A transparent retry happens when the request never left the client, or never reached the server's application, which is always safe. A configured retry policy names the status codes to retry, and gRFC A6 treats any maxAttempts above 5 as 5. Neither applies once response headers have arrived. Retrying CreateLabel after DEADLINE_EXCEEDED is only safe once the Label service deduplicates requests, which is what the idempotency topic adds.
RPC systems you will meet
Trade-offs.
The chosen option is first; the others stay visible so the reasoning can be checked.
- Pro:A deadline and an error are part of every call's signature (context, futures, RemoteException)
- Pro:Callers plan for latency and failure; reviewers can see each network hop
- Con:More ceremony at every call site
- Con:Moving code between local and remote means changing its callers
Hides a thousands-fold latency jump, so chatty object graphs end up on the wire; Partial failure surfaces as odd exceptions from what looked like a getter (Waldo et al.)
- Pro:One round trip per task (about 1 ms instead of 20 ms for 40 items)
- Pro:One place to set the deadline and handle failure
- Con:Bigger messages
- Con:Partial success has to be reported item by item
N round trips for N items; Each call is another chance to fail halfway
- Pro:Simple, with an immediate error; right for CreateLabel, because the packer prints the label on the spot
- Con:The caller is only up when the callee is: two services at 99.9% each give 99.9% × 99.9% ≈ 99.8% for the chain
No immediate answer; needs status polling or a callback; Delivery semantics move to the broker (see message queues)
Common RPC mistakes
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| No deadline set on calls3RPC runtime (caller) | Threads pile up waiting on a hung Label service until Fulfilment stops answering too | Saturated thread or connection pools; calls far older than any normal latency | Set a deadline on every call and pass the remaining time downstream | Packers see a fast error instead of a frozen screen |
| Retrying a non-idempotent call3RPC runtime (caller) | A lost reply plus a retry buys two labels | The carrier invoice shows two labels for one parcel | Send an idempotency key and deduplicate on the server (see idempotency) | Parcels still ship; the extra label is a cost to void later |
| A remote call inside a loop1Fulfilment code | Latency grows with order size; 40 lines cost 20 ms | Traces show N sequential spans to the same method | Add a batch method | Large orders pack slowly; small ones feel fine |
| A schema change breaks old clients2Client stub | Unmarshal errors or silently wrong values after a deploy | Error spike on the first calls from the old version | Only add fields with new numbers; never reuse or renumber a field | Old clients fail or print wrong labels until they upgrade |
| Retry storm while the callee is slow3RPC runtime (caller) | Every layer retries, multiplying load on a service that is already struggling | Request rate climbs while successful calls fall | Retry budgets and backoff at one layer only (see timeouts and retries) | Most calls time out until load drops or retries stop |