News feedFan-out on write

100%

Fan-out on write.

A feed read that has to ask 380 accounts "what's new?" is slow, and it repeats the same work every time someone scrolls. Fan-out on write moves that work to posting time: a queue hands each new post to workers that append its ID to the feed of every active follower. Posting stays fast, feeds are ready before anyone asks, and a handful of huge accounts are left out of the push on purpose.

Intermediate29 minUpdated 2 Oct 2026

Builds on Delivery semantics.

Requirements.

The write path of Chorus's home feed: from tapping Post to the post sitting in followers' feeds. Reading, merging and ranking those feeds are the next two topics.

Functional requirements

#1Publishing a post returns in well under a second, whatever the author's follower count.
#2The post appears in the Following feed of 95% of the author's active followers within 10 seconds.
#3Following someone brings their recent posts into your feed; unfollowing, muting or blocking removes them.
#4A deleted post disappears from every feed and is never served again.
#5A post from an account with millions of followers reaches its readers without slowing everyone else's posts.

Non-functional requirements

NFR-01Creating a post takes under 300 ms at p99, measured at the gateway.
NFR-02Fan-out lag (post created → in a follower's feed) is under 10 s at p95 and under 60 s at p99.
NFR-03Delivery into feeds is at-least-once, and appends are idempotent, so a post never shows twice.
NFR-04Losing the feed store loses no posts; every feed can be rebuilt from the post store.

Capacity estimates

New posts per second, peak
1,750
50M a day, ×3 in the evening (assumption)
Feed appends per second, peak
260K
each post to 250 × 60% = 150 active followers
Feed store
60 TB
400M feeds × 500 entries × ~100 B in memory, three copies

What the push costs

Assumptions
New posts per day
50MAssumption for Chorus (400M monthly and 200M daily users).
Mean followers of a pushed post
250Median 90. Excludes pull-mode accounts over 100K followers.
Followers active in the last 30 days
60%Assumption. Active means opened the feed.
Evening peak over the daily mean
×3Assumption.
Logical bytes per feed entry
20 Bpost_id 8 B + author_id 8 B + type and flags 4 B: what the entry means, not what Redis spends.
In-memory bytes per entry
100 BAssumption. Over 128 members a sorted set leaves listpack for a skiplist plus a hash table: a text member of up to ~40 B, an 8 B score, a skiplist node and a hash entry.
Entries kept per feed
500About 25 pages of 20; older posts come from author timelines.
Feeds kept
400MEveryone active in the last 30 days.
Copies of each feed
3
Data per Redis node
64 GBAssumption.
Kai Ren's followers
30MA footballer; our example of a very large account.
Working
  1. Appends per postfollowers × active = 250 × 0.6150from Mean followers of a pushed post and Followers active in the last 30 days
  2. Appends per dayposts × per-post = 50M × 1507.5B/dayfrom New posts per day and Appends per post
  3. Average append ratedaily ÷ 86,400 s~86,800/sfrom Appends per day
  4. Peak append rateavg × peak = 86,806 × 3~260K/sfrom Average append rate and Evening peak over the daily mean
  5. One feedkeep × entry = 500 × 20 B10 KBfrom Entries kept per feed and Logical bytes per feed entry · Logical size. All feeds would be 400M × 10 KB = 4 TB of raw references.
  6. One feed in Rediskeep × mem = 500 × 100 B50 KBfrom Entries kept per feed and In-memory bytes per entry
  7. All feeds, one copyfeeds × per-feed-mem = 400M × 50 KB20 TBfrom Feeds kept and One feed in Redis
  8. Redis shards (primaries)store ÷ node = 20,000 GB ÷ 64 GB = 312.5~313from All feeds, one copy and Data per Redis node
  9. All feeds, three copiesstore × replicas = 20 TB × 3; 313 shards × 360 TB on ~939 nodesfrom All feeds, one copy, Copies of each feed and Redis shards (primaries) · Raising zset-max-listpack-entries above 500 keeps each feed in the compact encoding at a fraction of this, at the cost of O(n) inserts; benchmark before counting on it.
  10. Kai Ren's post, if pushedkai-followers × active = 30M × 0.6 = 18M appends; 18M ÷ 260K/s69 s of the whole fleetfrom Kai Ren's followers, Followers active in the last 30 days and Peak append rate
What it means
  • Pushing is cheap for the typical author: 150 small appends.
  • Keep references, not posts: about 100 B per entry in Redis, against ~1 KB for a post, is what lets 400M feeds live in memory.
  • One very large account would take over the whole fleet for a minute, so it can't be pushed like everyone else.

Appends needed for one post

Appends needed for one postKai Ren's one post would need 18M appends, 120,000 times a typical post's 150, which is why the largest accounts are not pushed.101001k10k100k1M10M100MMedian author (90)Mean author (250)Dara (1,840)Cut-off (100K)Kai Ren (30M)541501.1k60k18MAuthorFeed appendsAppends needed for one postKai Ren's one post would need 18M appends, 120,000 times a typical post's 150, which is why the largest accounts are not pushed.101k100k10MMedian author (90)Mean author (250)Dara (1,840)Cut-off (100K)Kai Ren (30M)541501.1k60k18MAuthorFeed appends
Active followers only (60% of followers). The mean is over pushed posts; the cut-off is the largest pushed account. Log scale: each gridline is ten times the one before.
Data
Authorappends per post
Median author (90)54
Mean author (250)150
Dara (1,840)1,104
Cut-off (100K)60,000
Kai Ren (30M)18,000,000
Was this section helpful?

High-level design.

Posting and fan-out are split by a queue. The author waits for one database write; the copies are made afterwards, in parallel, by workers that can fall behind without anyone noticing.

Fan-out on write

Fan-out on write. The numbered component cards that follow describe each part.
Fan-out on writeComponents: 1. Chorus apps (iOS, Android and web. Post, scroll, pull to refresh.), 2. API gateway (Authenticates the session, rate-limits, routes posting calls to the post service.), 3. Post service (Validates and stores a post, assigns its time-ordered ID, emits post events.), 4. Post store (Posts by ID and each author's posts newest first (the author timeline). Source of truth.), 5. Post events (Kafka topics of post and follow events, partitioned so each author's events stay in order.), 6. Fan-out workers (Turn one post into appends on followers' feeds.), 7. Social graph (Followers and followees of a user, paged backed by its own sharded store.), 8. Activity store (Last feed open per user filters fan-out before any Redis call.), 9. Feed store (Redis cluster holding one capped, ordered set of post references per active user.).

Stores

Asynchronous fan-out

Write API

POST /v1/posts

authenticated

post + timeline row

post.created

consume

followers

active?

append if key exists

2API gateway
auth
posting limits

3Post service
time-ordered post_id
writes, then emits

5Post events
post.created
post.deleted
follow
p0p1… p127

6Fan-out workers
followers, 1,000 a page
skip inactive
batch per shard
worker-1worker-2worker-3

4Post store
posts by id
author timeline

8Activity store
last_active_at

9Feed store
feed:{user} → 500 newest refs
shard 0shard 1… shard 312

1Chorus apps
Post button

7Social graph
followers of X, paged

Components

ComponentResponsibilityOwns
API gatewayAuthenticates, applies per-user posting limits, routes.rate-limit counters
Post serviceValidates the post, assigns a time-ordered ID, writes it, publishes post.created.idempotency records
Post storeEvery post once, plus each author's posts newest first.posts, author_posts
Post eventsBuffers post and follow events; keeps each author's events in order.post.events, graph.events
Fan-out workersPage followers, drop inactive ones, append the reference to each feed.nothing (stateless)
Social graphAnswers who follows whom in pages of 1,000.followers, following
Activity storeLast feed open per user, read in batches of 1,000.last_active
Feed storeOne capped, ordered set of post references per active user.feed:{user_id}

Dara posts a new loaf

Dara posts a new loaf, as an ordered list of steps:
Dara posts a new loaf14 steps between Dara's app, API gateway, Post service, Post store, Post events, Fan-out workers, Social graph, Feed store. The steps are listed as text after the diagram.Feed storeSocial graphFan-out workersPost eventsPost storePost serviceAPI gatewayDara's appDara is done: about 60ms1,104 of 1,840 active:1,104 appendsCommit offset after allbatches succeedPOST /v1/posts(Idempotency-Key 7c1e…)1forward, authenticated2insert post +author_posts row3ok4post.created, key =Dara's author_id5201 {post_id}6deliver post.created7followers of Dara, page181,000 IDs; page 2 has8409313 shard pipelines inparallel, 3-4 each10ok11
  1. Dara's app → API gateway: POST /v1/posts (Idempotency-Key 7c1e…)
  2. API gateway → Post service: forward, authenticated
  3. Post service → Post store: insert post + author_posts row
  4. Post store → Post service (reply): ok
  5. Post service → Post events: post.created, key = Dara's author_id
  6. Post service → Dara's app (reply): 201 {post_id}
  7. Note over Dara's app: Dara is done: about 60 ms
  8. Post events → Fan-out workers: deliver post.created
  9. Fan-out workers → Social graph: followers of Dara, page 1
  10. Social graph → Fan-out workers (reply): 1,000 IDs; page 2 has 840
  11. Note over Fan-out workers: 1,104 of 1,840 active: 1,104 appends
  12. Fan-out workers → Feed store: 313 shard pipelines in parallel, 3-4 each
  13. Feed store → Fan-out workers (reply): ok
  14. Note over Fan-out workers: Commit offset after all batches succeed

What one append does

# One pipeline per shard; followers are
# already grouped by shard.
pipe = shard.pipeline(transaction=False)
for uid in followers_on_this_shard:
    key = f"feed:{uid}"
    # LPUSHX pushes only if the feed exists:
    # never create a stub feed.
    pipe.lpushx(key, f"{post_id}:{author_id}")
    pipe.ltrim(key, 0, 499)  # newest 500
pipe.execute()
# LTRIM right after LPUSH removes one
# element: O(1) in practice. But a
# redelivered batch pushes the post twice.
# Plain ZADD on a missing key creates a
# one-entry feed the read path would trust.
# So append only if the key exists.
APPEND = shard.register_script("""
  if redis.call('EXISTS', KEYS[1]) == 0
    then return 0 end
  -- same member twice = one entry
  redis.call('ZADD', KEYS[1],
             ARGV[1], ARGV[2])
  -- keep the newest 500
  redis.call('ZREMRANGEBYRANK',
             KEYS[1], 0, -501)
  return 1""")

# The ID's 41-bit ms timestamp + epoch.
score = created_ms(post_id)
member = f"{post_id}:{author_id}"
pipe = shard.pipeline(transaction=False)
for uid in followers_on_this_shard:
    APPEND(keys=[f"feed:{uid}"],
           args=[score, member], client=pipe)
pipe.execute()
# Scores are doubles, exact only to 2^53:
# a full 64-bit post_id would lose its low
# bits, so score by the timestamp inside it.

The list is the smallest and fastest structure, and Twitter's home timeline used one. It has two weaknesses in a system with at-least-once delivery: a redelivered batch adds the post again, and entries sit in arrival order rather than post order, so a post from a lagging partition lands above newer ones. A sorted set keyed by the post fixes both at the cost of O(log n) inserts and more memory per entry. Either way the append must not create a missing feed: a key holding one entry looks like a live feed to the read path, which then skips the rebuild. So the list uses LPUSHX and the sorted set a four-line script that checks EXISTS first. Step through the strip to see the difference.

Ines's feed as a list and as a sorted set

List (LPUSH + LTRIM)9Feed store
  1. Tomas, 1 slot, Existing entry, value 06:01:40
  2. Dara, 1 slot, Existing entry, value 05:58:02
  3. Rowing club, 1 slot, Existing entry, value 05:51:30
  4. Tomas, 1 slot, Existing entry, value 05:40:15
  5. 496 more, 3 slots, Older entries (to 500)
Sorted set (ZADD)9Feed store
  1. Tomas, 1 slot, Existing entry, value 06:01:40
  2. Dara, 1 slot, Existing entry, value 05:58:02
  3. Rowing club, 1 slot, Existing entry, value 05:51:30
  4. Tomas, 1 slot, Existing entry, value 05:40:15
  5. 496 more, 3 slots, Older entries (to 500)
  • Just appended
  • Existing entry
  • Duplicate
  • Older entries (to 500)
Start

As it starts. 4 steps follow.

The newest five entries of each structure; the grey block stands for the rest of the 500. Entries are labelled by author and post time; the stored member is the full post_id:author_id.
Was this section helpful?

Data model.

Posts live once, in the post store. Feeds hold only references, and can be thrown away and rebuilt.

Cassandra (or sharded MySQL)Owned by 4Post store

Written once, read by ID from everywhere; the author timeline is the one range query the system needs.

posts
ColumnTypeKeyNote
post_idbigintprimaryTime-ordered: 41-bit ms timestamp, then shard and sequence bits, so ID order is time order.
author_idbigint
bodytext
mediajsonBlob keys only; bytes live in the blob store.
visibilityenumpublic | followers | close_friends
created_attimestamp
deleted_attimestamp nullable
author_postsServes pull-mode authors, new-follow backfill and cold rebuilds.
ColumnTypeKeyNote
author_idbigintpartition
post_idbigintclustering desc
Graph service over sharded MySQL (TAO-style)Owned by 7Social graph

Two directions of the same edge, each read as a paged list.

followers
ColumnTypeKeyNote
user_idbigintpartition
follower_idbigintclustering
created_attimestamp
following
ColumnTypeKeyNote
user_idbigintpartition
followee_idbigintclustering
mutedbool
created_attimestamp
Key-value storeOwned by 8Activity store

One small value per user, read 1,000 at a time by fan-out.

last_active
ColumnTypeKeyNote
user_idbigintprimary
last_active_attimestampLast feed open, the same event that refreshes the feed key's TTL. Written at most once an hour, so it can lag the TTL by an hour; the conditional append covers that gap.
Redis clusterOwned by 9Feed store

Every feed read starts here; it fits in RAM (20 TB a copy) because entries are references, not posts.

feed:{user_id}TTL 30 days without a readSorted set capped at 500. Only the read path's rebuild creates the key; fan-out appends only to a key that exists, so a missing key always means rebuild, never a feed that silently missed posts.
ColumnTypeKeyNote
membertextpost_id:author_id
scoredoubleCreated time in ms, from the ID's top 41 bits.
Access patterns
QueryUsesHow
Who follows Dara?followerspartition user_id, paged by follower_id, 1,000 per page
Append to Ines's feedfeed:{user_id}script: if EXISTS, ZADD then ZREMRANGEBYRANK 0 -501
Dara's last 20 postsauthor_postspartition author_id, first 20 rows
Is this follower active?last_activemulti-get of 1,000 user IDs
Broker: Kafka
post.eventskeyed by author_id · 128 partitions · kept 3 days · at-least-once
Produced by3Post serviceConsumed by6Fan-out workers
{
"type": "created | deleted",
"post_id": "int64",
"author_id": "int64",
"visibility": "string",
"at": "timestamp"
}

Keyed by author so a delete can never overtake its create. Every feed-service node also tails this topic for large-account posts and deletes (feed-reads).

graph.eventskeyed by follower_id · 64 partitions · kept 3 days · at-least-once
Produced by7Social graphConsumed by6Fan-out workers
{
"type": "follow | unfollow | mute | block",
"follower_id": "int64",
"followee_id": "int64",
"at": "timestamp"
}

follow backfills the followee's last 20 posts into the follower's feed; unfollow, mute and block scan the follower's 500 entries and remove that author's, then repeat 60 s later. This topic is keyed by follower and post.events by author, so an append in flight can land after the cleanup: the read path also filters muted and blocked authors (feed-reads), and a stray post from someone just unfollowed can survive until the repeat pass or the trim.

fanout.batcheskeyed by post_id + batch_no · 256 partitions · kept 1 day · at-least-once
Produced by6Fan-out workersConsumed by6Fan-out workers
{
"post_id": "int64",
"author_id": "int64",
"batch_no": "int32",
"follower_ids": "int64[] (up to 1,000)"
}

A big author's post split into 1,000-follower batches so many workers share it (see Optimizations).

The life of one reader's feed key

States of9Feed store

The life of one reader's feed key. 3 states, 8 transitions. The table below lists them.
The life of one reader's feed keyThe states of Feed store. 3 states, 8 transitions. The table below lists them.

post.created / nothing (filtered, or no key)

app opened / key + placeholder, TTL 60 s

post.created / ZADD lands

rebuild done / 500 refs in, placeholder out

post.created [key exists] / ZADD + trim
feed read / refresh TTL to 30 d

30 d unread / key expires
shard lost

No feed key

Rebuilding

Live and pushed to

5 steps.

Fan-out never creates a feed key; only the rebuild does. So a reader whose key is missing, for any reason, has no feed at all and gets a fresh one on her next open, never a feed that silently missed posts.

Transitions of The life of one reader's feed key
From → ToEventGuardActionActor
No feed key → No feed keypost.creatednothing (filtered, or no key)Fan-out workers
No feed key → Rebuildingapp openedkey + placeholder, TTL 60 sFeed service (feed-reads)
Rebuilding → Rebuildingpost.createdZADD landsFan-out workers
Rebuilding → Live and pushed torebuild done500 refs in, placeholder outFeed service (feed-reads)
Live and pushed to → Live and pushed topost.createdkey existsZADD + trimFan-out workers
Live and pushed to → Live and pushed tofeed readrefresh TTL to 30 dFeed service (feed-reads)
Live and pushed to → No feed key30 d unreadkey expiresRedis
Live and pushed to → No feed keyshard lostRedis cluster
No feed keystart
New user, 30 days without a feed read, or shard lost
Rebuilding
The read path refills it from author timelines (feed-reads)
Was this section helpful?

Interface.

One public call to post, one to delete, two for following. Everything else is internal: events and a paged follower list.

1POST/v1/posts

Stores the post and publishes post.created. Returns before any follower's feed is touched.

Request
Headers
Authorization: Bearer …Idempotency-Key: 7c1e…
{
"text": "Rye came out of the oven at 6",
"media": ["blob:7f3a…"],
"visibility": "public"
}
Response
{
"post_id": "2105176361031421961",
"created_at": "2026-09-30T06:02:11Z"
}

Endpoints

MethodPathDoesReturns
POST/v1/postsCreate a post; fan-out follows asynchronously.201 { post_id, created_at }
DELETE/v1/posts/{id}Mark the post deleted and emit post.deleted.204
POST/v1/followsFollow { followee_id }; emits follow, which backfills recent posts.201
DELETE/v1/follows/{followee_id}Unfollow; emits unfollow.204
GET/internal/users/{id}/followers?cursor&limit=1000Internal; the page fan-out reads.{ ids[], next_cursor }

Errors

{ "error": { "code": "invalid_post", "message": "A post can have at most 4 photos." } }
HTTPTypeBody codeClient behaviour
400error
invalid_post
Empty, over 2,000 characters, or more than 4 photos. Show the message; keep the draft.
403error
not_author
Deleting someone else's post. Nothing to retry.
409error
idempotency_mismatch
The key was used with a different body. Generate a new key only if this is really a new post.
413error
media_too_large
Reject before upload; tell the author the limit.
429retry
rate_limited
Wait for Retry-After. Posting limits stop a spam flood from turning into a fan-out flood.
Was this section helpful?

Optimizations.

What keeps fan-out lag under ten seconds when posting spikes.

Push only to people who will read
Followers who have not opened their feed for 30 days are filtered out before any Redis call; their key has expired too, and the conditional append would skip them anyway. They get a rebuild if they return (feed-reads). Without the filter the peak would be 50M × 250 ÷ 86,400 × 3 ≈ 434K appends/s; with it, 260K/s. The filter removes about 174K/s.
Leave large accounts out of the push
Accounts above 100K followers (about 80K accounts, 0.02% of users) are pull-mode: their posts go to the post store and the queue only, and readers merge them in at read time. The worst pushed post becomes 60K appends (0.23 s of fleet capacity) instead of Kai Ren's 18M (69 s). The cut-off is a dial: tune it against fan-out lag and read-time merge cost. The full story is twitter/celebrity-fan-out.
Split big fan-outs into batches
A worker that pages 90K followers alone holds its partition for seconds while small posts queue behind it. Stage 1 lists followers and emits one fanout.batches message per 1,000 active followers; stage 2 workers append in parallel. Posts under 10K followers skip the split.
Group appends by shard
Dara's 1,104 appends spread over all 313 shards, 3 or 4 each (1,104 ÷ 313 ≈ 3.5). The worker sends each shard's appends as one pipeline and all 313 pipelines in parallel: one round of round trips instead of 1,104 in turn. The gain grows with the post: Harbour FC's 54K appends are about 170 per shard.

A big post on Dara's partition

Scenario 1 of 2: As described.

Timeline as a list

A big post on Dara's partition: 3 lanes, from 0 s to 4.5 s.

  1. 0–3.6 s · Partition 17 · Harbour FC: 90 pages + 54K appends
  2. 0.3 s · Partition 17 · Dara posts
  3. 0.3–3.68 s · all lanes · window: Dara waits 3.4 s
  4. 3.6–3.68 s · Partition 17 · Dara: 1,104
  5. 3.68 s · Feeds · Dara delivered (delayed)
Harbour FC (90K followers, 54K active; under the cut-off, so pushed) posts on the same partition as Dara, 0.3 s before her. With one worker, Dara's 1,104 appends wait behind 54K of someone else's; three such posts in a row would push her past the 10 s target. Assumptions: a worker appends about 20K entries/s pipelined, and a follower page of 1,000 takes 10 ms.
  1. 01Launch
    Load
    50 posts/s
    Bottleneck
    None
    Change
    Feeds computed on read with one SQL query over following JOIN posts; no feed store.
    Adds
    4Post store7Social graph
  2. 0210× launch
    Load
    500 posts/s
    Bottleneck
    Reads slow down as follow counts grow
    Change
    Add the feed store and push synchronously from the post service.
    Adds
    9Feed store
  3. 03Today's peak
    Load
    ~1,750 posts/s
    Bottleneck
    Posting latency tied to the author's follower count
    Change
    Move the push behind a queue with workers; skip inactive followers.
    Adds
    5Post events6Fan-out workers8Activity store
  4. 04Today, with very large accounts
    Load
    ~1,750 posts/s, 260K appends/s
    Bottleneck
    One post stalls the fleet or its partition
    Change
    Pull-mode accounts over 100K followers; two-stage batched fan-out.
Was this section helpful?

Trade-offs.

The chosen option is first; the others stay visible so the reasoning can be checked.

01
When the copies are made
Chosen:Hybrid: push for most authors, pull for large ones
  • Pro:A read is one lookup plus a few merges
  • Pro:The cost of any single post is bounded (60K appends at most)
Downside we accept:
  • Con:Two code paths to build, test and keep consistent
  • Con:Readers who follow many large accounts pay extra merge latency
Ruled out:Push to everyone

One large account's post is millions of writes and minutes of lag for everyone else; Writes to feeds nobody opens

Ruled out:Pull only (fan-out on read)

Every page asks hundreds of author timelines and merges them; reads outnumber posts about 50:1 here (2.4B page loads vs 50M posts a day); Works at scale only with heavy in-memory infrastructure, as Facebook's Multifeed does: aggregators query leaf servers holding recent actions at read time

Decide per pair, not per author

The cost of pushing Tomas's posts to Ines is Tomas's posting rate: one write per post. The cost of pulling them is Ines's reading rate: one timeline query per feed read. Push the pair when the author posts less often than the reader reads, weighted by the relative cost of a write and a read-time query. Silberstein et al. (Feeding Frenzy, SIGMOD 2010) show that making this choice per producer-consumer pair from the ratio of rates minimises total cost.

Worked, with a write and a query costing the same: Ines opens her feed 8 times a day. Tomas posts once a day, so push: one write serves 8 reads. @metro_alerts posts 150 times a day, so pull: 150 writes a day, most scrolled past, which would also push Tomas's posts out of her 500 entries. Follower count, which Chorus uses, is only the easy proxy for this rule.

One author and one reader, per day

  • Push
  • Pull
One author and one reader, per dayPushing costs one write per post and pulling one query per read, so for a reader who opens the feed 8 times a day the cheaper choice flips at 8 posts a day.0.11101001k0.11101001k8 posts/dayTomas: push@metro_alerts: pullPushPullOperations per day for this pairAuthor's posts per dayOne author and one reader, per dayPushing costs one write per post and pulling one query per read, so for a reader who opens the feed 8 times a day the cheaper choice flips at 8 posts a day.0.11101001k0.11101001k8 posts/dayTomas: push@metro_alerts: pullPushPullOperations per day for this pairAuthor's posts per day
Push costs one write per post; pull costs one query per feed read, and Ines reads 8 times a day. Assumes one write costs the same as one read-time query. Both axes are logarithmic.
Data
Author's posts per dayPushPull
0.10.18
11no value
88no value
150150no value
1,0001,0008
  • 8 posts/day: Author's posts per day = 8
  • At 1: Tomas: push
  • At 150: @metro_alerts: pull
02
What a feed entry is and where it lives
Chosen:Sorted set of post references in Redis, capped at 500
  • Pro:Idempotent; a redelivered append overwrites itself
  • Pro:Ordered by post time even when partitions lag
  • Pro:Small entries (a reference, ~100 B in Redis)
Downside we accept:
  • Con:Needs hydration on read
  • Con:O(log n) inserts and more memory per entry than a list
  • Con:History beyond 500 comes from author timelines
Ruled out:Capped list (LPUSH + LTRIM)

Duplicates on redelivery; Arrival order, not post order; Removing one entry is O(n)

Ruled out:Full post copies in each feed

About 50× the raw bytes (1 KB post vs 20 B reference); Edits and deletes must touch every copy

03
Removing a deleted post
Chosen:Filter at read time, clean up lazily
  • Pro:Instant for readers: hydration sees deleted_at, and a small set of recently deleted IDs is checked
  • Pro:No second fan-out storm
Downside we accept:
  • Con:Dead references occupy slots until trimmed
Ruled out:Fan out the delete to every feed

As expensive as the post itself; Still races with appends in flight, so the read filter is needed anyway

FailureImpactDetectionMitigationMeanwhile
Workers fall behind (lag 2 min)6Fan-out workersPosts appear late in feedsConsumer lag per partition against the 10 s p95 targetAutoscale workers; split big posts into batches soonerPosts are still on the author's profile at once
A shard and its replica lost9Feed storeThose users' feeds are emptyShard health checksAppends skip the missing keys; rebuild on read from author timelines (feed-reads)A slower first page for about 0.3% of users (one shard of 313)
Duplicate delivery after a worker crash5Post eventsNone visibleNot neededZADD is idempotent per postNo change
Follower lookups slow7Social graphFan-out lag grows for every authorp99 of follower-page latencyRetry with backoff; posts wait safely in the queue for up to 3 daysFeeds lag; posting is unaffected
Activity store unavailable8Activity storeFan-out can't tell active from inactive followersError rate on the batch lookupFail open and push to all followers of small accounts; delay large onesAbout 1.7× the append load until it recovers
Was this section helpful?
Builds on this
Reading a feed
Read next