Ranking the feed.
A ranked feed is a pipeline, not a sort. For each session the feed service gathers candidates from the reader's feed list, the large accounts she follows and recommendations, fetches their features, scores them in two passes so the expensive model only sees the best 200, reorders them for variety, and stores the result so the next pages cost almost nothing. This topic is the serving system around the models; what the models predict is the ML track's job.
Builds on Reading a feed.
Requirements.
The For you tab: the same posts as Following plus recommendations, ordered by what each reader is likely to value. The models are a black box here; the pipeline that runs them is not.
Functional requirements
Non-functional requirements
Capacity estimates
Why rank once per session
- Ranking sessions per day
- 800M200M DAU × 4 opens or refreshes (assumption)
- Pages per session
- 32.4B page loads / 800M sessions
- Peak-to-average ratio
- 3
- Candidates left after filters
- 820900 gathered (500 feed list + ~100 large accounts + 300 recommended), minus seen, hidden and flagged
- Shortlist for the heavy model
- 200
- Heavy scores per second per model server
- 25Kassumption for one GPU host with batching
- Headroom
- 30%
- Session lifetime
- 30 min (1,800 s)
- One ranked ref
- 20 Bpost_id 8 B + author_id 8 B + score and source 4 B
- Sessions per second800M / 86,400 = 9.26K/s; × 3~27.8K/s at peakfrom Ranking sessions per day and Peak-to-average ratio
- Light-model scores27.8K × 820~22.8M scores/sfrom Sessions per second and Candidates left after filters
- Heavy-model scores27.8K × 200~5.56M scores/sfrom Sessions per second and Shortlist for the heavy model
- Heavy-model servers5.56M / 25K = 222; × 1.3~290 serversfrom Heavy-model scores, Heavy scores per second per model server and Headroom
- If every page re-ranked instead290 × 3 pages~870 serversfrom Heavy-model servers and Pages per session · What the session store saves.
- Session store at peak27.8K/s × 1,800 s = 50M live sessions; × 200 refs × 20 B (4 KB)~200 GBfrom Sessions per second, Session lifetime, Shortlist for the heavy model and One ranked ref · About 67 GB at the daily average rate; a small Redis cluster either way.
- The heavy model's cost sets the shortlist size; the light model exists so the heavy one sees 200 posts, not 820.
- Ranking once per session, not once per page, cuts the model fleet by two-thirds and costs ~200 GB of RAM.
From 900 candidates to one page
Data
| Stage | Posts |
|---|---|
| Gathered | 900 (100%) |
| After filters | 820 (91%) |
| Heavy shortlist | 200 (22%) |
| Shown (3 pages) | 60 (6.7%) |
| Page 1 | 20 (2.2%) |
High-level design.
The mixer is the conductor. It calls the candidate sources in parallel, the feature store once, and the model servers twice, then applies rules it owns.
Ranked feed pipeline
Components
| Component | Responsibility | Owns |
|---|---|---|
| Feed service (mixer) | Runs the pipeline for a session: gathers, filters, calls the models, reranks, pages. Owns every deadline and fallback. | rerank rules |
| Feed store | The reader's precomputed list of 500 refs, written by fan-out; large accounts are merged in as in feed-reads. | feed:{user_id} |
| Recommendations | Posts from accounts the reader does not follow. A black box with a deadline. | retrieval indexes |
| Feature store | Numbers the models read about the reader and each post; refreshed by offline and streaming jobs. | user_features, post_features |
| Model servers | Score batches of (reader, post) pairs; light on CPU, heavy on GPU. | model versions |
| Session store | The ranked list per session and the posts each reader saw recently. | rank_session:{id}, seen:{user_id} |
| Impression log | What was shown, at which position, and what happened next. | feed.impressions |
Ines pulls to refresh For you
- Chorus apps → Feed service (mixer): GET /v1/feed?tab=for_you (new session)
- Feed service (mixer) → Feed store: in-network refs (all 500)
- Feed service (mixer) → Recommendations: recommended refs (300)
- Note over Feed service (mixer) and Recommendations: In parallel; recommendations get 60 ms
- Feed store → Feed service (mixer) (reply): 500 refs
- Note over Feed service (mixer): + ~100 large-account posts from memory
- Recommendations → Feed service (mixer) (reply): ~300 refs
- Note over Feed service (mixer): Filter seen (48 h), hidden, flagged: 820 left
- Feed service (mixer) → Feature store: user + 820 post features
- Feature store → Feed service (mixer) (reply): feature vectors
- Note over Feed service (mixer) and Feature store: Post features: ~90% from an in-process cache
- Feed service (mixer) → Model servers: light model, 820 candidates
- Model servers → Feed service (mixer) (reply): scores; keep top 200
- Feed service (mixer) → Model servers: heavy model, 200 candidates
- Model servers → Feed service (mixer) (reply): P(like), P(comment), P(share), P(hide)
- Note over Feed service (mixer): One score, then rerank for variety
- Feed service (mixer) → Session store: session s-81c3: 200 refs, TTL 30 min
- Feed service (mixer) → Chorus apps (reply): page 1 (20 posts) + cursor {s-81c3, 20}
Data model.
Three kinds of state. Features are read on every request, the ranked session lives for 30 minutes, and the log of what was shown lives in the warehouse.
One multi-get by ID per request; the offline tables are too slow and too large to read on the request path.
| Column | Type | Key | Note |
|---|---|---|---|
| user_id | bigint | primary | |
| embedding | vector(64) | ||
| follow_counts | map | ||
| recent_affinity | map(author_id → score) | How much the reader engaged with each author lately | |
| updated_at | timestamp |
| Column | Type | Key | Note |
|---|---|---|---|
| post_id | bigint | primary | |
| author_id | bigint | ||
| type | enum(text, photo, video) | ||
| created_at | timestamp | ||
| engagement_1h | map(action → count) | ||
| integrity_score | float | ||
| embedding | vector(64) | ||
| updated_at | timestamp |
Small, short-lived, read by key and range; losing one costs a re-rank, not data.
| Column | Type | Key | Note |
|---|---|---|---|
| user_id | bigint | ||
| created_at | timestamp | ||
| model_version | text | ||
| refs | list(post_id, author_id, score, source) | 200 entries of 20 B |
| Column | Type | Key | Note |
|---|---|---|---|
| post_ids | two rotating 24 h Bloom filters |
Append-only and high volume; joined with actions offline.
| Column | Type | Key | Note |
|---|---|---|---|
| impression_id | uuid | primary | |
| user_id | bigint | partition | |
| post_id | bigint | ||
| session_id | text | ||
| position | int | ||
| source | enum(in_network, large_account, recommended) | ||
| model_version | text | ||
| scores | map(label → p) | ||
| shown_at | timestamp |
| Query | Uses | How |
|---|---|---|
| Features for one request | *_features | One multi-get, user_id plus up to 820 post_ids |
| Next page of this session | rank_session:{session_id} | Range read [offset, offset + 20) |
| Has Ines seen this post in the last 48 h? | seen:{user_id} | Membership test during filtering |
| What did we show and what happened next? | impressions | Joined with feed.actions offline by impression_id |
Read by the training pipeline and the experiment dashboards (outside this design). Positions must be logged; without them a model learns that top slots hold good posts simply because they were on top.
Dwell carries milliseconds; the rest carry 1. Duplicates are removed by impression_id and action at join time.
Interface.
One public call with a session cursor, and six internal calls, each with a deadline and a plan for when it misses.
Without a cursor, starts a ranking session and returns page 1. With a cursor, reads the next 20 from that session's ranked list. ranked says which path served it.
Internal calls
| Call | Deadline | If it misses |
|---|---|---|
| Mixer → feed store (in-network refs) | 30 ms | Skip in-network refs; rank recommendations and large-account posts, and trigger a background rebuild (feed-reads). |
| Mixer → recommendations | 60 ms | Skip them; rank in-network candidates only. |
| Mixer → feature store | 50 ms | Default features; run the light model only. |
| Mixer → light model | 20 ms | Take the newest 200 candidates as the shortlist. |
| Mixer → heavy model | 120 ms | Order by light-model scores. |
| Mixer → session store | 10 ms | Serve page 1 anyway; do not cache, so the next page re-ranks. |
A ranking session's life
States of8Session store
The cursor points into a session. When the session is gone, the app gets 410 and starts a new one from the top, keeping what is already on screen.
| From → To | Event | Guard | Action | Actor |
|---|---|---|---|---|
| No session → Ranking page 1 | open or refresh | app | ||
| Ranking page 1 → List cached | pipeline done | store, page 1 | mixer | |
| Ranking page 1 → Fallback | no usable scores | serve Following | mixer | |
| List cached → List cached | next page | offset < 200 | range read | app |
| List cached → Ranking page 1 | refresh or list used up | app | ||
| List cached → Expired | 30 min or evicted | 410 on next page | session store | |
| Fallback → Fallback | next page | chronological | app | |
| Fallback → Ranking page 1 | refresh | app |
- No sessionstart
- Ranking page 1
- The full pipeline runs once
- List cached
- 200 ranked refs, 30 min TTL
- Fallback
- Newest first; ranked: false and a Following cursor
- Expiredend
- 30 min old, or evicted
Errors
| HTTP | Type | Body code | Client behaviour |
|---|---|---|---|
| 410 | error | cursor_expired | The session was created more than 30 min ago or was evicted. Start a new session from the top and keep what is on screen. |
| 429 | retry | rate_limited | A refresh storm (repeated pull to refresh). Retry after the given delay; the app also debounces pull to refresh. |
Optimizations.
Where the 400 ms goes, how scores become an order, and the four changes that keep the fleet small and the feed up.
Where 400 ms goes
Scenario 1 of 3: As described.
- 12 Fire and forget; the response does not wait.
Timeline as a list
Where 400 ms goes: 12 lanes, from 0 ms to 400 ms.
- 0–340 ms · Mixer · 340 ms
- 0–30 ms · In-network refs · 30
- 0–60 ms · Recommended refs · 60 (deadline)
- 60–70 ms · Filter · 10
- 70–120 ms · Features · 50
- 120–140 ms · Light model · 20
- 140–260 ms · Heavy model · 120 (GPU)
- 260–270 ms · Combine + rerank · 10
- 270–280 ms · Store session · 10
- 270–310 ms · Hydrate 20 posts · 40
- 310–340 ms · Serialise + network · 30
- 310–325 ms · Log impressions · async — Fire and forget; the response does not wait.
- 340–400 ms · all lanes · window: headroom
- 400 ms · all lanes · p99 budget 400 ms (deadline)
From predictions to an order
One score from four predictions
| Post | P(like) | P(comment) | P(share) | P(hide) | Score (weights 1, 4, 6, -20) |
|---|---|---|---|---|---|
| A: Tomas's photo | 0.30 | 0.05 | 0.01 | 0.005 | 0.30 + 0.20 + 0.06 - 0.10 = 0.46 |
| B: a page's video | 0.12 | 0.01 | 0.04 | 0.02 | 0.12 + 0.04 + 0.24 - 0.40 = 0.00 |
| C: a baking club's question | 0.08 | 0.12 | 0.00 | 0.002 | 0.08 + 0.48 + 0.00 - 0.04 = 0.52 |
The order is C, A, B. C wins on comments, which carry four times the weight of a like; B's likely shares are cancelled by its hide risk. The weights are illustrative and a product decision, tuned by experiments; how they are chosen is social-feed-ranking/value-model. Meta describes the same shape: many predictions per post, combined with weights into one value.
Rules after scoring
- V1, 1 item, Video, value Kofi .71
- V2, 1 item, Video, value Kofi .69
- P3, 1 item, Photo, value Lena .66
- V4, 1 item, Video, value Arun .60
- T5, 1 item, Text, value Mei .55
- V6, 1 item, Video, value Sam .52
- P7, 1 item, Photo, value Ola .50
- slot 1, 1 item, Open slot
- slot 2, 1 item, Open slot
- slot 3, 1 item, Open slot
- slot 4, 1 item, Open slot
- slot 5, 1 item, Open slot
- slot 6, 1 item, Open slot
- Video
- Photo
- Text
- Open slot
- Already placed
- Fails a rule for this slot
As it starts. 6 steps follow.
Four changes that matter most
Trade-offs.
The chosen option is first; the others stay visible so the reasoning can be checked.
- Pro:Uses the freshest posts and engagement signals
- Pro:No work for people who do not open the app
- Con:A 400 ms budget on the request path
- Con:GPU fleet sized for the evening peak
Ranks for the 200M monthly users who will not open the app today (half of 400M); Stale by the time the reader arrives
Cannot use the reader's context at read time; Misses engagement that arrives after the post is scored
- Pro:Predictable
- Pro:Protects the followed-accounts experience
- Con:A hand-tuned knob (say, at most 30% recommended; an assumption to test)
Recommendations with high predicted engagement can crowd out friends; X reports an average near 50% out-of-network, which shows how far this can go
- Pro:Depends only on the feed-reads path
- Pro:Every post is from someone the reader chose
- Con:Noticeably worse for readers who follow many busy accounts
Repeats what the reader just saw; Nothing new since
Breaks requirement #5; readers retry and add load to a failing system
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| GPU pool degraded7Model servers | Heavy stage misses its 120 ms deadline | Deadline-exceeded rate per stage | Order by light-model scores; shed load before adding retries | A slightly worse order, on time |
| Feature store stale or slow6Feature store | Scores drift, or the features call misses 50 ms | Feature freshness lag and call latency | Alert; default features and the light model only | A plainer order |
| Session store node lost or evicting8Session store | cursor_expired on page 2 | 410 rate | The app starts a new session; more nodes or a longer TTL | The list jumps back to the top |
| Recommendations down5Recommendations | No out-of-network candidates | Skipped-source rate | Rank in-network candidates only | A feed of followed accounts only |
| Impression log backlog10Impression log | Training and dashboards run late | Consumer lag | Buffered by the 7-day retention; the request path never waits | Nothing visible to readers |