Abstractions and RPCRPC and what it hides

100%

RPC and what it hides.

One line of code, labels.createLabel(order), turns into marshalling, a network round trip, a thread on another machine and a reply that may never arrive. RPC hides the first four. It cannot hide the last, and every design that uses it has to decide what to do when the answer is "I don't know".

Beginner22 minUpdated 30 Sept 2026

The idea.

A function call between machines, and the one promise it cannot fully keep.

Why RPC exists

A warehouse's Fulfilment service packs orders. For each parcel it needs a shipping label, which a separate Label service buys from a carrier. Without RPC, the Fulfilment team would open a socket, invent a byte layout for the order ID, carrier and weight, write it, wait, parse whatever comes back, and handle every error by hand. The Label team would write the mirror image. Every new call between every pair of services would repeat that work.

Remote procedure call, as Birrell and Nelson built it at Xerox PARC in 1984, removes that work. Both teams agree on an interface. A tool generates code for each side, and the Fulfilment programmer writes one ordinary-looking line: labels.createLabel(order). Their design goal was to make a remote call behave as much like a local one as possible.

Ten years later, Waldo and colleagues at Sun pointed out four things no generated code can hide: latency (a remote call is thousands of times slower), memory (the other machine cannot follow your pointers), concurrency (other callers hit the same server at the same moment) and partial failure (one machine can fail while the other carries on, and the survivor may not know what happened). The last one shapes everything else in this topic.

What one call becomes

What one call becomes. The numbered component cards that follow describe each part.
What one call becomesComponents: 1. Fulfilment code (Business logic that wants a label for a parcel calls what looks like an ordinary method.), 2. Client stub (Generated from the interface definition. Marshals the arguments, names the method, unmarshals the reply.), 3. RPC runtime (caller) (Picks a server address, owns connections, sets the deadline, retransmits, and matches replies to calls by call ID.), 4. RPC runtime (callee) (Accepts connections, dispatches to the right handler and sends the reply. Birrell and Nelson's runtime also dropped duplicate call IDs here gRPC's does not.), 5. Server stub (Unmarshals the arguments, calls the real handler, marshals the result or the error.), 6. Label handler (The actual CreateLabel implementation. Buys a label from the carrier and stores it.), 7. Service registry (Maps a service name to the addresses of its live instances (binding).).

Label host

Fulfilment host

createLabel(order)

bytes + method name

where is Labels?

request over the network

dispatch

createLabel(order)

reply (may be lost)

1Fulfilment code
labels.createLabel(order)

2Client stub
marshal arguments

3RPC runtime (caller)
call ID
deadline

4RPC runtime (callee)
dispatch by method

5Server stub
unmarshal arguments

6Label handler
buy label from carrier

7Service registry
Labels → 10.2.7.14, 10.2.7.15

The same call, local and remote

AspectLocal callRemote call
TimeNanoseconds: a few for a plain call, tens for a small lookupAbout 0.5 ms inside one datacenter; up to ~150 ms across continents (Dean, 2009). About 10,000× a 50 ns local lookup (assumption).
ArgumentsPassed by reference; both sides share memoryCopied into bytes. A pointer means nothing on the other machine.
FailureThe call returns, or the whole process dies with itCan fail halfway: request lost, server crashed mid-call, or reply lost
ConcurrencyRuns on your threadRuns on another machine's thread, next to other callers' requests
Was this section helpful?

How it works.

Five pieces, three jobs: agree on the interface, find the server, and move one call there and back. Then decide what "there and back" means when it fails.

1. Agree on an interface

package labels.v1;

service Labels {
  rpc CreateLabel(CreateLabelRequest) returns (Label);
}

message CreateLabelRequest {
  int64  order_id = 1;
  string carrier  = 2;   // "ground" or "express"
  int32  weight_g = 3;
}

message Label {
  string label_id    = 1;
  string tracking_no = 2;
  int64  cost_cents  = 3;
}
ctx, cancel := context.WithTimeout(ctx, 250*time.Millisecond)
defer cancel()

label, err := labels.CreateLabel(ctx, &pb.CreateLabelRequest{
    OrderId: 918273,
    Carrier: "ground",
    WeightG: 1450,
})
if err != nil {
    // lost request? crashed server? lost reply?
}

The interface definition is the contract. A compiler turns it into a client stub and a server stub for each language: Birrell and Nelson's Lupine did this from Mesa interface modules, protoc does it from .proto files today. The numbers after each field (= 1, = 2, = 3) matter more than the names, as the next step shows.

2. Marshal the arguments

CreateLabel(918273, ground, 1450) on the wire

JSON (54 bytes)
  1. 1 byte, JSON punctuation, value {
  2. order_id, 11 bytes, Field name (JSON only), value "order_id":
  3. order, 6 bytes, Value, value 918273
  4. 1 byte, JSON punctuation, value ,
  5. carrier, 10 bytes, Field name (JSON only), value "carrier":
  6. carrier, 8 bytes, Value, value "ground"
  7. 1 byte, JSON punctuation, value ,
  8. weight_g, 11 bytes, Field name (JSON only), value "weight_g":
  9. weight, 4 bytes, Value, value 1450
  10. 1 byte, JSON punctuation, value }
  • 54 bytes of text, from 0 to 54
Protobuf (15 bytes)
  1. field 1, 1 byte, Protobuf tag (field + wire type), value 08, 1 << 3 | 0 (varint)
  2. order, 3 bytes, Value, value 81 86 38, 918273 in three 7-bit groups
  3. field 2, 1 byte, Protobuf tag (field + wire type), value 12, 2 << 3 | 2 (length-delimited)
  4. len, 1 byte, Length prefix, value 06
  5. carrier, 6 bytes, Value, value ground
  6. field 3, 1 byte, Protobuf tag (field + wire type), value 18, 3 << 3 | 0 (varint)
  7. weight, 2 bytes, Value, value aa 0b, 1450 = 42 + 11 × 128
  • 15 bytes, from 0 to 15
gRPC message (20 bytes)
  1. compressed?, 1 byte, gRPC frame header, value 00
  2. length, 4 bytes, gRPC frame header, value 00 00 00 0f, 15 in 4 big-endian bytes
  3. the 15 protobuf bytes, 15 bytes, Value
  • 5-byte prefix, from 0 to 5
  • JSON punctuation
  • Field name (JSON only)
  • Protobuf tag (field + wire type)
  • Length prefix
  • Value
  • gRPC frame header
field 1
1 << 3 | 0 (varint)
order
918273 in three 7-bit groups
field 2
2 << 3 | 2 (length-delimited)
field 3
3 << 3 | 0 (varint)
weight
1450 = 42 + 11 × 128
length
15 in 4 big-endian bytes
Only 16 of JSON's 54 bytes are values; the other 38 are quoted field names (32) and punctuation. Protobuf replaces each name with a 1-byte tag (field number << 3 | wire type) and writes numbers as varints, 7 bits per byte: 918273 fits in 3 bytes, 1450 in 2. gRPC then adds a 5-byte prefix, so 20 bytes travel. Both sides need the schema to read protobuf.

3. Find the server (binding)

The stub knows a service name, Labels, not an address. Binding turns the name into live instances: in 1984 the Grapevine database mapped an interface's type and instance names to machines; today it is DNS, Consul, a Kubernetes Service or an xDS control plane. The runtime then picks one instance, often through a load balancer, and keeps the connection open for later calls. If the registry is stale, the call goes to a machine that is gone, which the runtime sees as one more failed attempt.

4. One call, there and back

CreateLabel, there and back

CreateLabel, there and back, as an ordered list of steps:
CreateLabel, there and back13 steps between Fulfilment code, Client stub, RPC runtime (caller), RPC runtime (callee), Server stub, Label handler. The steps are listed as text after the diagram.Label handlerServer stubRPC runtime (callee)RPC runtime (caller)Client stubFulfilment coderegistry: Labels → 10.2.7.14, 10.2.7.15call ID 7f3a, deadline 250 msBirrell & Nelson: seen 7f3a? no → dispatch (gRPC skips this)createLabel(918273, ground, 1450)115 bytes + method name2request 7f3a + bytes3dispatch4createLabel(order)5Label{L-5521, cost 612}6marshal the Label7reply 7f3a + bytes8match 7f3a, hand over bytes9unmarshalled Label10
  1. Fulfilment code → Client stub: createLabel(918273, ground, 1450)
  2. Client stub → RPC runtime (caller): 15 bytes + method name
  3. Note over RPC runtime (caller): registry: Labels → 10.2.7.14, 10.2.7.15
  4. Note over RPC runtime (caller): call ID 7f3a, deadline 250 ms
  5. RPC runtime (caller) → RPC runtime (callee): request 7f3a + bytes
  6. Note over RPC runtime (callee): Birrell & Nelson: seen 7f3a? no → dispatch (gRPC skips this)
  7. RPC runtime (callee) → Server stub: dispatch
  8. Server stub → Label handler: createLabel(order)
  9. Label handler → Server stub (reply): Label{L-5521, cost 612}
  10. Server stub → RPC runtime (callee) (reply): marshal the Label
  11. RPC runtime (callee) → RPC runtime (caller) (reply): reply 7f3a + bytes
  12. RPC runtime (caller) → Client stub (reply): match 7f3a, hand over bytes
  13. Client stub → Fulfilment code (reply): unmarshalled Label

5. Three failures that look the same

What the caller sees in 250 ms

Scenario 1 of 4: As described.

Timeline as a list

What the caller sees in 250 ms: 3 lanes, from -40 ms to 300 ms.

  1. 0 ms · Caller's runtime → Callee's runtime, arriving 3 ms: request 7f3a
  2. 5–180 ms · Label handler · buy label from carrier
  3. 182 ms · Callee's runtime → Caller's runtime, arriving 185 ms: reply
  4. 185 ms · Caller's runtime · Label L-5521 arrives (ok, ok)
  5. 250 ms · all lanes · deadline (deadline)
Switch tabs. In every failure the caller's runtime sees the same thing: silence until the 250 ms deadline. Only the first case is safe to resend blindly, and the caller cannot tell it apart from the other two.

What the runtime promises about a call

SemanticsRuntime doesDuplicatesLost callsFits
MaybeSends once, never retriesNeverPossible: a lost request is simply goneFire-and-forget hints; UI actions a person can repeat
At least onceRetries until a reply or the deadlinePossible: a lost reply means the work runs twiceNever, while retries lastReads such as GetLabel, and operations that are idempotent
At most onceRetries, and the server filters duplicate call IDs and resends its cached reply (Birrell and Nelson, Java RMI)Never, unless the server crashes and forgets the IDs it sawPossible if the server crashes: the caller gets an error and cannot tell whether the call ranCalls where a double is worse than a miss
"Exactly once" (effectively once)At most once, with the seen IDs and stored replies kept in durable storage, e.g. idempotency keys in a databaseNever, even across server restartsNever, while the caller keeps retrying with the same keyAnything with side effects, like CreateLabel. Costs a durable dedupe store on the server.

Birrell and Nelson's runtime is the textbook at-most-once design. Every call carried an identifier made of the calling machine, the calling process and a sequence number; the callee remembered the last sequence number it had seen from each caller and discarded retransmissions. Their promise: if the call returns, the procedure ran precisely once; if it raises an exception, it ran once or not at all, and the caller is not told which. The table lived in memory, so a server crash erased it. Keeping that table in durable storage, keyed by a request ID the client chooses, is what idempotency keys do, and the idempotency topic builds it out for CreateLabel.

Was this section helpful?

In practice.

What remote calls cost, how gRPC carries a deadline through a chain of them, and which failures a runtime may retry on its own.

What a remote call costs

A loop of remote calls

Assumptions
Round trip inside one datacenter
0.5 msOrder of magnitude from Dean (2009), the same figure numbers to know uses; a modern same-zone hop is often lower.
Order lines to price, one call each
40Assumption for a large order.
The same lookup as an in-process call
50 nsAssumption; tens of nanoseconds.
Extra encode, transfer and handler work per item in a batch
12.5 µsAssumption; puts one batched call at about 1 ms, the same as the cache multi-get in numbers to know.
Chance any one backend call is slow
1%Assumption, for the fan-out line.
Working
  1. 40 calls, one after anotherlines × rtt = 40 × 0.5 ms20 msfrom Order lines to price, one call each and Round trip inside one datacenter
  2. One call carrying all 40 itemsrtt + lines × per-item = 0.5 ms + 0.5 ms~1 msfrom Round trip inside one datacenter, Order lines to price, one call each and Extra encode, transfer and handler work per item in a batch
  3. Batching savessequential ÷ batched = 20 ÷ 1~20×from 40 calls, one after another and One call carrying all 40 items
  4. The same loop in-processlines × local = 40 × 50 ns2 µsfrom Order lines to price, one call each and The same lookup as an in-process call
  5. 40 calls in parallel; share of requests that wait on at least one slow call1 − (1 − slow)^lines = 1 − 0.99^40 = 1 − 0.669~33%from Chance any one backend call is slow and Order lines to price, one call each
What it means
  • A remote call inside a loop is the classic RPC mistake. Offer a batch method (PriceLines, CreateLabels) instead. The same 20 ms against ~1 ms arithmetic for a cache multi-get is in numbers to know.
  • Going parallel fixes the 20 ms but makes rare slowness common. Tail latency is covered in latency percentiles.

Pricing 40 order lines three ways

Pricing 40 order lines three waysForty sequential remote calls take 20 ms, about 20 times one batched call and 10,000 times the same loop in-process at an assumed 50 ns per lookup.1 µs10 µs100 µs1 ms10 ms100 ms40 calls in a row1 batched callIn-process loop20 ms1 ms2 µsApproachTime (ms)Pricing 40 order lines three waysForty sequential remote calls take 20 ms, about 20 times one batched call and 10,000 times the same loop in-process at an assumed 50 ns per lookup.1 µs100 µs10 ms40 calls in a row1 batched callIn-process loop20 ms1 ms2 µsApproachTime (ms)
Log scale; each gridline is 10×. Numbers from the estimate above.
Data
ApproachTime (ms)
40 calls in a row20
1 batched call1
In-process loop0.002

One deadline, passed down

An 800 ms budget, passed down as grpc-timeout

HopStarts atWhat it sends on
API gateway → Fulfilment (PackOrder)20 msgrpc-timeout: 780m
Fulfilment → Label service (CreateLabel)60 msgrpc-timeout: 740m
Label service → Carrier API (POST /labels)60 msHTTP client timeout 700 ms (HTTP has no deadline header)

gRPC sets no deadline by default, so a call to a hung server can wait forever. A deadline travels as the grpc-timeout header in relative units (780m is 780 milliseconds; H, M, S, m, u and n are allowed), so the machines' clocks need not agree; each hop sends what is left of the user's 800 ms, as the table shows. Java and Go servers pass the incoming deadline on to their own outgoing calls automatically; C++ needs propagation enabled, and elsewhere the handler has to pass it on itself. When the deadline passes, the client gets DEADLINE_EXCEEDED and the server sees the call CANCELLED, and it should stop working on it. Birrell and Nelson deliberately had no call timeout: while waiting, their runtime probed the server and raised an exception only if the server stopped answering, just as a local call waits for as long as the procedure runs. Choosing the numbers, and what happens down a chain that has no deadline (work nobody will read), are worked through in timeouts and retries.

Which failures the runtime may retry

gRPC statusWhat it usually meansRetry?
UNAVAILABLE (14)Server unreachable, restarting or shedding loadYes, with backoff, if the method is idempotent
DEADLINE_EXCEEDED (4)Time ran out. The work may or may not have happened.Only if idempotent and the caller still has budget
RESOURCE_EXHAUSTED (8)A quota or rate limit was hitLater, after backing off; not in a tight loop
INVALID_ARGUMENT (3)The request is wrong, e.g. weight_g below zeroNever: it will fail the same way
NOT_FOUND (5) / ALREADY_EXISTS (6)State on the server, not a transient faultNever
INTERNAL (13)A bug or broken invariant on the serverUsually not

One gRPC call, from the caller's runtime

States of3RPC runtime (caller)

One gRPC call, from the caller's runtime. 6 states, 10 transitions. The table below lists them.
One gRPC call, from the caller's runtimeThe states of RPC runtime (caller). 6 states, 10 transitions. The table below lists them.

stub sends / start deadline

never reached the app [first time] / resend
UNAVAILABLE [attempts ‹ max] / back off

UNAVAILABLE [attempts = max]
non-retryable status

headers arrive

trailers: OK

trailers: error

deadline passes / cancel

deadline passes / cancel

Call created

Attempt in flight

Committed

OK

Failed with a status

Deadline passed: outcome unknown

4 steps.

Retries happen only while an attempt is uncommitted. Once response headers arrive the call is committed and never retried, and a passed deadline leaves the outcome unknown whether or not headers came.

Transitions of One gRPC call, from the caller's runtime
From → ToEventGuardActionActor
Call created → Attempt in flightstub sendsstart deadlineclient stub
Attempt in flight → Attempt in flightnever reached the appfirst timeresendruntime
Attempt in flight → Failed with a statusUNAVAILABLEattempts = maxretry policy
Attempt in flight → Attempt in flightUNAVAILABLEattempts < maxback offretry policy
Attempt in flight → Committedheaders arriveserver
Attempt in flight → Failed with a statusnon-retryable statusserver
Committed → OKtrailers: OKserver
Committed → Failed with a statustrailers: errorserver
Attempt in flight → Deadline passed: outcome unknowndeadline passescancelruntime
Committed → Deadline passed: outcome unknowndeadline passescancelruntime
Call createdstart
Attempt in flight
No response headers yet; retries allowed
Committed
Response headers received
OKend
Failed with a statuserror
Deadline passed: outcome unknownerror

gRPC retries a call on its own in two cases. A transparent retry happens when the request never left the client, or never reached the server's application, which is always safe. A configured retry policy names the status codes to retry, and gRFC A6 treats any maxAttempts above 5 as 5. Neither applies once response headers have arrived. Retrying CreateLabel after DEADLINE_EXCEEDED is only safe once the Label service deduplicates requests, which is what the idempotency topic adds.

RPC systems you will meet

gRPC
Protobuf over HTTP/2 with generated stubs, deadlines and status codes. The wire details and call kinds are in REST, gRPC and GraphQL.
Apache Thrift
An interface language and code generator from Facebook, now an Apache project. The protocol (binary, compact or JSON) and the transport (plain or framed sockets, HTTP) are chosen separately, so the same service definition can run over different stacks.
CORBA and Java RMI
Distributed-object systems. CORBA aimed for location transparency, the design Waldo et al. argued against. Java RMI made the network explicit: remote interfaces extend java.rmi.Remote, and every remote method must declare RemoteException, so callers cannot forget it can fail.
JSON-RPC 2.0
A two-page specification: a JSON object with method, params and id, over any transport. The id matches a response to its request, like Birrell and Nelson's call identifier; a request without an id is a notification and gets no response at all.
Was this section helpful?

Trade-offs.

The chosen option is first; the others stay visible so the reasoning can be checked.

01
How much should a remote call look local?
Chosen:Explicit remote calls
  • Pro:A deadline and an error are part of every call's signature (context, futures, RemoteException)
  • Pro:Callers plan for latency and failure; reviewers can see each network hop
Downside we accept:
  • Con:More ceremony at every call site
  • Con:Moving code between local and remote means changing its callers
Ruled out:Fully transparent remote objects

Hides a thousands-fold latency jump, so chatty object graphs end up on the wire; Partial failure surfaces as odd exceptions from what looked like a getter (Waldo et al.)

02
Fine- or coarse-grained interface
Chosen:Coarse, batch-friendly methods, e.g. CreateLabels(repeated parcel)
  • Pro:One round trip per task (about 1 ms instead of 20 ms for 40 items)
  • Pro:One place to set the deadline and handle failure
Downside we accept:
  • Con:Bigger messages
  • Con:Partial success has to be reported item by item
Ruled out:Fine-grained getters and setters

N round trips for N items; Each call is another chance to fail halfway

03
Wait for the answer, or hand the work off
Chosen:Synchronous RPC when the caller needs the result now
  • Pro:Simple, with an immediate error; right for CreateLabel, because the packer prints the label on the spot
Downside we accept:
  • Con:The caller is only up when the callee is: two services at 99.9% each give 99.9% × 99.9% ≈ 99.8% for the chain
Ruled out:Put the work on a queue

No immediate answer; needs status polling or a callback; Delivery semantics move to the broker (see message queues)

Common RPC mistakes

FailureImpactDetectionMitigationMeanwhile
No deadline set on calls3RPC runtime (caller)Threads pile up waiting on a hung Label service until Fulfilment stops answering tooSaturated thread or connection pools; calls far older than any normal latencySet a deadline on every call and pass the remaining time downstreamPackers see a fast error instead of a frozen screen
Retrying a non-idempotent call3RPC runtime (caller)A lost reply plus a retry buys two labelsThe carrier invoice shows two labels for one parcelSend an idempotency key and deduplicate on the server (see idempotency)Parcels still ship; the extra label is a cost to void later
A remote call inside a loop1Fulfilment codeLatency grows with order size; 40 lines cost 20 msTraces show N sequential spans to the same methodAdd a batch methodLarge orders pack slowly; small ones feel fine
A schema change breaks old clients2Client stubUnmarshal errors or silently wrong values after a deployError spike on the first calls from the old versionOnly add fields with new numbers; never reuse or renumber a fieldOld clients fail or print wrong labels until they upgrade
Retry storm while the callee is slow3RPC runtime (caller)Every layer retries, multiplying load on a service that is already strugglingRequest rate climbs while successful calls fallRetry budgets and backoff at one layer only (see timeouts and retries)Most calls time out until load drops or retries stop
Was this section helpful?
Builds on this
REST, gRPC and GraphQL
Read next