News feedRanking the feed

100%

Ranking the feed.

A ranked feed is a pipeline, not a sort. For each session the feed service gathers candidates from the reader's feed list, the large accounts she follows and recommendations, fetches their features, scores them in two passes so the expensive model only sees the best 200, reorders them for variety, and stores the result so the next pages cost almost nothing. This topic is the serving system around the models; what the models predict is the ML track's job.

Advanced24 minUpdated 2 Oct 2026

Builds on Reading a feed.

Requirements.

The For you tab: the same posts as Following plus recommendations, ordered by what each reader is likely to value. The models are a black box here; the pipeline that runs them is not.

Functional requirements

#1Opening or refreshing For you returns a ranked first page; later pages continue the same ranking with no repeats.
#2Candidates come from accounts the reader follows and from recommendations, in a controlled mix.
#3The page is varied: no run of one author, no wall of videos.
#4Posts that fail integrity checks, or that the reader has hidden, are never ranked.
#5If ranking is slow or down, the reader still gets a feed (newest first), never an error.
#6Every impression and reaction is logged with its position, for training and measurement.

Non-functional requirements

NFR-01First page p99 under 400 ms server-side.
NFR-02Next pages p99 under 150 ms.
NFR-03A fallback path that depends on no model and no recommendation service.
NFR-04Logging never slows the response.

Capacity estimates

Ranking sessions per second, peak
28K
200M daily users × 4 opens or refreshes, ×3 evening peak
Light-model scores per second
23M
820 candidates per session after filters
Heavy-model scores per second
5.6M
the best 200 per session

Why rank once per session

Assumptions
Ranking sessions per day
800M200M DAU × 4 opens or refreshes (assumption)
Pages per session
32.4B page loads / 800M sessions
Peak-to-average ratio
3
Candidates left after filters
820900 gathered (500 feed list + ~100 large accounts + 300 recommended), minus seen, hidden and flagged
Shortlist for the heavy model
200
Heavy scores per second per model server
25Kassumption for one GPU host with batching
Headroom
30%
Session lifetime
30 min (1,800 s)
One ranked ref
20 Bpost_id 8 B + author_id 8 B + score and source 4 B
Working
  1. Sessions per second800M / 86,400 = 9.26K/s; × 3~27.8K/s at peakfrom Ranking sessions per day and Peak-to-average ratio
  2. Light-model scores27.8K × 820~22.8M scores/sfrom Sessions per second and Candidates left after filters
  3. Heavy-model scores27.8K × 200~5.56M scores/sfrom Sessions per second and Shortlist for the heavy model
  4. Heavy-model servers5.56M / 25K = 222; × 1.3~290 serversfrom Heavy-model scores, Heavy scores per second per model server and Headroom
  5. If every page re-ranked instead290 × 3 pages~870 serversfrom Heavy-model servers and Pages per session · What the session store saves.
  6. Session store at peak27.8K/s × 1,800 s = 50M live sessions; × 200 refs × 20 B (4 KB)~200 GBfrom Sessions per second, Session lifetime, Shortlist for the heavy model and One ranked ref · About 67 GB at the daily average rate; a small Redis cluster either way.
What it means
  • The heavy model's cost sets the shortlist size; the light model exists so the heavy one sees 200 posts, not 820.
  • Ranking once per session, not once per page, cuts the model fleet by two-thirds and costs ~200 GB of RAM.

From 900 candidates to one page

From 900 candidates to one pageEach stage throws most of its input away, so the costly heavy model runs on under a quarter of what was gathered.02004006008001kGatheredAfter filtersHeavy shortlistShown (3 pages)Page 1900 · 100%820 · 91%200 · 22%60 · 6.7%20 · 2.2%StagePostsFrom 900 candidates to one pageEach stage throws most of its input away, so the costly heavy model runs on under a quarter of what was gathered.02004006008001kGatheredAfter filtersHeavy shortlistShown (3 pages)Page 1900 · 100%820 · 91%200 · 22%60 · 6.7%20 · 2.2%StagePosts
One session for Ines. Counts shrink at each stage; the heavy model, the costliest per post, sees 200 of the ~900 gathered (22%). The multiplication is in the estimate above.
Data
StagePosts
Gathered900 (100%)
After filters820 (91%)
Heavy shortlist200 (22%)
Shown (3 pages)60 (6.7%)
Page 120 (2.2%)
Was this section helpful?

High-level design.

The mixer is the conductor. It calls the candidate sources in parallel, the feature store once, and the model servers twice, then applies rules it owns.

Ranked feed pipeline

Ranked feed pipeline. The numbered component cards that follow describe each part.
Ranked feed pipelineComponents: 1. Chorus apps (iOS, Android and web. Open the For you tab, scroll, pull to refresh.), 2. API gateway (Authenticates the session, rate-limits refreshes, routes to the feed service, and forwards reader actions to the impression log.), 3. Feed service (mixer) (Runs the ranking pipeline for a session, gather → filter → score twice → rerank → page, and owns the rules and fallbacks.), 4. Feed store (Redis cluster with one capped sorted set of post refs per active user, written by fan-out the in-network candidates.), 5. Recommendations (Out-of-network candidate source (posts from accounts the reader does not follow). A black box here.), 6. Feature store (Online copy of user and post features for ranking, read by user ID and post ID.), 7. Model servers (Light and heavy scoring models behind one RPC requests are batched, the heavy model on GPUs.), 8. Session store (Redis holding each feed session's ranked list and each reader's recently seen posts.), 9. Post cache (Post bodies and author profiles by ID used to hydrate the page, exactly as in feed-reads.), 10. Impression log (Kafka topics of what was shown, where, and what the reader did feeds training and measurement.).

Serving

Scoring

Candidate sources

GET /v1/feed

500 refs

~300 refs, 60 ms

features

820, then 200

ranked list

hydrate

shown (async)

reactions

4Feed store
500 in-network refs

5Recommendations
~300 recommended refs

6Feature store
user once
posts by ID

7Model servers
lightCPUheavyGPU

3Feed service (mixer)
gather
filter
score twice
rerank
page
gatherfilterrerank

8Session store
200 ranked refs
30 min

9Post cache
hydrate 20

1Chorus apps

2API gateway

10Impression log
shown
position
action

Components

ComponentResponsibilityOwns
Feed service (mixer)Runs the pipeline for a session: gathers, filters, calls the models, reranks, pages. Owns every deadline and fallback.rerank rules
Feed storeThe reader's precomputed list of 500 refs, written by fan-out; large accounts are merged in as in feed-reads.feed:{user_id}
RecommendationsPosts from accounts the reader does not follow. A black box with a deadline.retrieval indexes
Feature storeNumbers the models read about the reader and each post; refreshed by offline and streaming jobs.user_features, post_features
Model serversScore batches of (reader, post) pairs; light on CPU, heavy on GPU.model versions
Session storeThe ranked list per session and the posts each reader saw recently.rank_session:{id}, seen:{user_id}
Impression logWhat was shown, at which position, and what happened next.feed.impressions

Ines pulls to refresh For you

Ines pulls to refresh For you, as an ordered list of steps:
Ines pulls to refresh For you18 steps between Chorus apps, Feed service (mixer), Feed store, Recommendations, Feature store, Model servers, Session store. The steps are listed as text after the diagram.Session storeModel serversFeature storeRecommendationsFeed storeFeed service (mixer)Chorus appsIn parallel; recommendations get 60 ms+ ~100 large-account posts from memoryFilter seen (48 h), hidden, flagged: 820 leftPost features: ~90% from an in-process cacheOne score, then rerank for varietyGET /v1/feed?tab=for_you (new session)1in-network refs (all 500)2recommended refs (300)3500 refs4~300 refs5user + 820 post features6feature vectors7light model, 820 candidates8scores; keep top 2009heavy model, 200 candidates10P(like), P(comment), P(share), P(hide)11session s-81c3: 200 refs, TTL 30 min12page 1 (20 posts) + cursor {s-81c3, 20}13
  1. Chorus apps → Feed service (mixer): GET /v1/feed?tab=for_you (new session)
  2. Feed service (mixer) → Feed store: in-network refs (all 500)
  3. Feed service (mixer) → Recommendations: recommended refs (300)
  4. Note over Feed service (mixer) and Recommendations: In parallel; recommendations get 60 ms
  5. Feed store → Feed service (mixer) (reply): 500 refs
  6. Note over Feed service (mixer): + ~100 large-account posts from memory
  7. Recommendations → Feed service (mixer) (reply): ~300 refs
  8. Note over Feed service (mixer): Filter seen (48 h), hidden, flagged: 820 left
  9. Feed service (mixer) → Feature store: user + 820 post features
  10. Feature store → Feed service (mixer) (reply): feature vectors
  11. Note over Feed service (mixer) and Feature store: Post features: ~90% from an in-process cache
  12. Feed service (mixer) → Model servers: light model, 820 candidates
  13. Model servers → Feed service (mixer) (reply): scores; keep top 200
  14. Feed service (mixer) → Model servers: heavy model, 200 candidates
  15. Model servers → Feed service (mixer) (reply): P(like), P(comment), P(share), P(hide)
  16. Note over Feed service (mixer): One score, then rerank for variety
  17. Feed service (mixer) → Session store: session s-81c3: 200 refs, TTL 30 min
  18. Feed service (mixer) → Chorus apps (reply): page 1 (20 posts) + cursor {s-81c3, 20}
Was this section helpful?

Data model.

Three kinds of state. Features are read on every request, the ranked session lives for 30 minutes, and the log of what was shown lives in the warehouse.

Key-value feature store (online copy of the offline tables)Owned by 6Feature store

One multi-get by ID per request; the offline tables are too slow and too large to read on the request path.

user_features
ColumnTypeKeyNote
user_idbigintprimary
embeddingvector(64)
follow_countsmap
recent_affinitymap(author_id → score)How much the reader engaged with each author lately
updated_attimestamp
post_featuresRefreshed every few minutes while a post is young, when its engagement moves fastest.
ColumnTypeKeyNote
post_idbigintprimary
author_idbigint
typeenum(text, photo, video)
created_attimestamp
engagement_1hmap(action → count)
integrity_scorefloat
embeddingvector(64)
updated_attimestamp
RedisOwned by 8Session store

Small, short-lived, read by key and range; losing one costs a re-rank, not data.

rank_session:{session_id}TTL 30 min
ColumnTypeKeyNote
user_idbigint
created_attimestamp
model_versiontext
refslist(post_id, author_id, score, source)200 entries of 20 B
seen:{user_id}TTL 48 h (rolling)About 480 IDs (2 days × 12 pages × 20). A Bloom filter cannot expire single entries, so keep two 24 h filters (today, yesterday): check both, and each day drop the older one and start a new one. At 1% false positives that is ~600 B per reader in total.
ColumnTypeKeyNote
post_idstwo rotating 24 h Bloom filters
Kafka topic → warehouseOwned by 10Impression log

Append-only and high volume; joined with actions offline.

impressions
ColumnTypeKeyNote
impression_iduuidprimary
user_idbigintpartition
post_idbigint
session_idtext
positionint
sourceenum(in_network, large_account, recommended)
model_versiontext
scoresmap(label → p)
shown_attimestamp
Access patterns
QueryUsesHow
Features for one request*_featuresOne multi-get, user_id plus up to 820 post_ids
Next page of this sessionrank_session:{session_id}Range read [offset, offset + 20)
Has Ines seen this post in the last 48 h?seen:{user_id}Membership test during filtering
What did we show and what happened next?impressionsJoined with feed.actions offline by impression_id
Broker: Kafka
feed.impressionskeyed by user_id · 256 partitions · kept 7 days · at-least-once
Produced by3Feed service (mixer)Consumed byNo consumer
{
"impression_id": "uuid",
"user_id": "bigint",
"post_id": "bigint",
"session_id": "string",
"position": "int",
"source": "in_network | large_account | recommended",
"model_version": "string",
"scores": "map",
"shown_at": "timestamp"
}

Read by the training pipeline and the experiment dashboards (outside this design). Positions must be logged; without them a model learns that top slots hold good posts simply because they were on top.

feed.actionskeyed by user_id · 256 partitions · kept 7 days · at-least-once
Produced by2API gatewayConsumed byNo consumer
{
"impression_id": "uuid",
"user_id": "bigint",
"post_id": "bigint",
"action": "like | comment | share | hide | dwell",
"value": "number",
"acted_at": "timestamp"
}

Dwell carries milliseconds; the rest carry 1. Duplicates are removed by impression_id and action at join time.

Was this section helpful?

Interface.

One public call with a session cursor, and six internal calls, each with a deadline and a plan for when it misses.

1GET/v1/feed?tab=for_you&limit=20

Without a cursor, starts a ranking session and returns page 1. With a cursor, reads the next 20 from that session's ranked list. ranked says which path served it.

Request
Headers
Authorization: Bearer <session token>
Response
{
"posts": ["… 20 hydrated posts …"],
"session_id": "s-81c3",
"next_cursor": "eyJzIjoiczgxYzMiLCJvIjoyMH0",
"ranked": true
}
The cursor encodes {session s-81c3, offset 20} and is signed like the Following cursor (feed-reads), so a hand-edited offset is a 400.

Internal calls

CallDeadlineIf it misses
Mixer → feed store (in-network refs)30 msSkip in-network refs; rank recommendations and large-account posts, and trigger a background rebuild (feed-reads).
Mixer → recommendations60 msSkip them; rank in-network candidates only.
Mixer → feature store50 msDefault features; run the light model only.
Mixer → light model20 msTake the newest 200 candidates as the shortlist.
Mixer → heavy model120 msOrder by light-model scores.
Mixer → session store10 msServe page 1 anyway; do not cache, so the next page re-ranks.

A ranking session's life

States of8Session store

A ranking session's life. 5 states, 8 transitions. The table below lists them.
A ranking session's lifeThe states of Session store. 5 states, 8 transitions. The table below lists them.

open or refresh

pipeline done / store, page 1

no usable scores / serve Following

next page [offset ‹ 200] / range read

refresh or list used up

30 min or evicted / 410 on next page

next page / chronological

refresh

No session

Ranking page 1

List cached

Fallback

Expired

4 steps.

The cursor points into a session. When the session is gone, the app gets 410 and starts a new one from the top, keeping what is already on screen.

Transitions of A ranking session's life
From → ToEventGuardActionActor
No session → Ranking page 1open or refreshapp
Ranking page 1 → List cachedpipeline donestore, page 1mixer
Ranking page 1 → Fallbackno usable scoresserve Followingmixer
List cached → List cachednext pageoffset < 200range readapp
List cached → Ranking page 1refresh or list used upapp
List cached → Expired30 min or evicted410 on next pagesession store
Fallback → Fallbacknext pagechronologicalapp
Fallback → Ranking page 1refreshapp
No sessionstart
Ranking page 1
The full pipeline runs once
List cached
200 ranked refs, 30 min TTL
Fallback
Newest first; ranked: false and a Following cursor
Expiredend
30 min old, or evicted

Errors

{ "error": { "code": "cursor_expired", "message": "This feed session has ended." } }
HTTPTypeBody codeClient behaviour
410error
cursor_expired
The session was created more than 30 min ago or was evicted. Start a new session from the top and keep what is on screen.
429retry
rate_limited
A refresh storm (repeated pull to refresh). Retry after the given delay; the app also debounces pull to refresh.
Was this section helpful?

Optimizations.

Where the 400 ms goes, how scores become an order, and the four changes that keep the fleet small and the feed up.

Where 400 ms goes

Scenario 1 of 3: As described.

Notes
  • 12 Fire and forget; the response does not wait.
Timeline as a list

Where 400 ms goes: 12 lanes, from 0 ms to 400 ms.

  1. 0–340 ms · Mixer · 340 ms
  2. 0–30 ms · In-network refs · 30
  3. 0–60 ms · Recommended refs · 60 (deadline)
  4. 60–70 ms · Filter · 10
  5. 70–120 ms · Features · 50
  6. 120–140 ms · Light model · 20
  7. 140–260 ms · Heavy model · 120 (GPU)
  8. 260–270 ms · Combine + rerank · 10
  9. 270–280 ms · Store session · 10
  10. 270–310 ms · Hydrate 20 posts · 40
  11. 310–340 ms · Serialise + network · 30
  12. 310–325 ms · Log impressions · async — Fire and forget; the response does not wait.
  13. 340–400 ms · all lanes · window: headroom
  14. 400 ms · all lanes · p99 budget 400 ms (deadline)
Page 1 of a new session. Candidate calls overlap; everything after them is sequential. Budgets are our targets, not measurements: 340 ms of work leaves 60 ms of headroom.

From predictions to an order

One score from four predictions

PostP(like)P(comment)P(share)P(hide)Score (weights 1, 4, 6, -20)
A: Tomas's photo0.300.050.010.0050.30 + 0.20 + 0.06 - 0.10 = 0.46
B: a page's video0.120.010.040.020.12 + 0.04 + 0.24 - 0.40 = 0.00
C: a baking club's question0.080.120.000.0020.08 + 0.48 + 0.00 - 0.04 = 0.52

The order is C, A, B. C wins on comments, which carry four times the weight of a like; B's likely shares are cancelled by its hide risk. The weights are illustrative and a product decision, tuned by experiments; how they are chosen is social-feed-ranking/value-model. Meta describes the same shape: many predictions per post, combined with weights into one value.

Rules after scoring

Scored order7Model servers
  1. V1, 1 item, Video, value Kofi .71
  2. V2, 1 item, Video, value Kofi .69
  3. P3, 1 item, Photo, value Lena .66
  4. V4, 1 item, Video, value Arun .60
  5. T5, 1 item, Text, value Mei .55
  6. V6, 1 item, Video, value Sam .52
  7. P7, 1 item, Photo, value Ola .50
Page slots3Feed service (mixer)
  1. slot 1, 1 item, Open slot
  2. slot 2, 1 item, Open slot
  3. slot 3, 1 item, Open slot
  4. slot 4, 1 item, Open slot
  5. slot 5, 1 item, Open slot
  6. slot 6, 1 item, Open slot
  • Video
  • Photo
  • Text
  • Open slot
  • Already placed
  • Fails a rule for this slot
Start

As it starts. 6 steps follow.

A greedy walk down the scored list. Rules: never the same author twice in a row; at most 2 videos in any 4 consecutive slots. A post that fails waits for a later slot, so the top of the list still leads.

Four changes that matter most

Two passes, sized by cost
The light model scores all 820 in ~20 ms on CPU; only 200 reach the GPU model. Real systems do the same: Meta narrows the eligible posts to about 500 with a lightweight pass before the main one, and X sources about 1,500 candidates for a ~48M-parameter heavy ranker. Outcome: a GPU fleet sized for 200 posts per session, not 820.
Rank once per session
Store the 200 ranked refs; pages 2 and 3 are range reads. Outcome: ~290 model servers instead of ~870, and next pages in ~75 ms instead of ~340 ms.
Degrade by stage, never to an error
Every call has a deadline and a fallback (see Internal calls). The last resort is the Following feed, newest first, flagged ranked: false. Outcome: a slightly worse feed instead of a spinner.
Cache post features by ID
Post features are the same for every reader, and the same popular posts are candidates in millions of sessions. An in-process LRU keyed by post_id answers ~90%. Outcome: feature reads drop from ~5.7 GB/s (27.8K sessions × 820 posts × ~250 B ≈ 200 KB each) to ~0.57 GB/s.
Was this section helpful?

Trade-offs.

The chosen option is first; the others stay visible so the reasoning can be checked.

01
When ranking runs
Chosen:At request time, once per session
  • Pro:Uses the freshest posts and engagement signals
  • Pro:No work for people who do not open the app
Downside we accept:
  • Con:A 400 ms budget on the request path
  • Con:GPU fleet sized for the evening peak
Ruled out:Precompute ranked feeds in the background

Ranks for the 200M monthly users who will not open the app today (half of 400M); Stale by the time the reader arrives

Ruled out:Score at fan-out time, into per-user pools

Cannot use the reader's context at read time; Misses engagement that arrives after the post is scored

02
How recommendations are mixed in
Chosen:A target share per session, enforced in rerank
  • Pro:Predictable
  • Pro:Protects the followed-accounts experience
Downside we accept:
  • Con:A hand-tuned knob (say, at most 30% recommended; an assumption to test)
Ruled out:Let the scores decide

Recommendations with high predicted engagement can crowd out friends; X reports an average near 50% out-of-network, which shows how far this can go

03
What to serve when ranking fails
Chosen:Following feed, newest first
  • Pro:Depends only on the feed-reads path
  • Pro:Every post is from someone the reader chose
Downside we accept:
  • Con:Noticeably worse for readers who follow many busy accounts
Ruled out:The last ranked session, if under an hour old

Repeats what the reader just saw; Nothing new since

Ruled out:An error

Breaks requirement #5; readers retry and add load to a failing system

FailureImpactDetectionMitigationMeanwhile
GPU pool degraded7Model serversHeavy stage misses its 120 ms deadlineDeadline-exceeded rate per stageOrder by light-model scores; shed load before adding retriesA slightly worse order, on time
Feature store stale or slow6Feature storeScores drift, or the features call misses 50 msFeature freshness lag and call latencyAlert; default features and the light model onlyA plainer order
Session store node lost or evicting8Session storecursor_expired on page 2410 rateThe app starts a new session; more nodes or a longer TTLThe list jumps back to the top
Recommendations down5RecommendationsNo out-of-network candidatesSkipped-source rateRank in-network candidates onlyA feed of followed accounts only
Impression log backlog10Impression logTraining and dashboards run lateConsumer lagBuffered by the 7-day retention; the request path never waitsNothing visible to readers
Was this section helpful?
Related
Fan-out on write
Read next