Social feed rankingCandidate sources

100%

Candidate sources.

When Maya opens Stoop, about 940 unseen posts from her friends, groups and pages are eligible, and millions more from outside her network could be. The ranker can afford 600. This topic decides which 600 get through. It covers where candidates come from, how each source cheaply ranks its own pile, how the slots are split between sources, and how to tell when a source has stopped earning its share.

Intermediate18 minUpdated 1 Oct 2026

Builds on Multi-stage funnels and Two towers and a nearest-neighbour index.

Framing.

Candidate generation decides what the ranker is allowed to consider. A post that never becomes a candidate can't be ranked, however good the ranker is.

From business goal to ML task

Stoop is a fictional social app for friends, neighbourhood groups and local pages such as a bakery, the library or the council. Its Home feed mixes posts from people you know, groups you joined, pages you follow, and a small share from outside your network. Every Stoop number here is illustrative. Each Home request starts with one question: out of everything this person could see, which few hundred posts are worth scoring at all?

LayerDefinition
Business goalViewers come back because their feed shows what their people and places are up to (measured as 28-day return).
ML objectiveOf the posts a viewer would engage with in this session, maximise the share that are among the 600 candidates (engagement recall@600).
ML taskSeveral retrievers (inbox reads, index reads, a graph walk, an embedding search), each trimmed by a cheap per-source model, merged under quotas.
Out of scopeScoring and ordering (engagement model, value model), never-show and demote decisions (integrity), serving plumbing such as fan-out and cursors (news-feed).

What candidate generation must do

#1Return 600 candidates for every Home request, each tagged with every source that found it.
#2Never return a post the viewer has already seen, hidden, or is not allowed to see (blocked authors, private groups she left).
#3Keep every unseen friend post from the last 72 h when there are fewer of them than the friends quota.
#4Reach posts from outside the viewer's network, while keeping them at most 25% of impressions.
#5Finish all sources within 50 ms. A source that misses its deadline loses its slots for that request and the page does not wait.
#6Report each source's share of candidates and its marginal recall every day.

Scale and budget

Feed requests a day
160M
40M viewers × 4 sessions (illustrative)
Eligible connected posts
~940
per request, unseen, last 72 h
Candidates to the ranker
600
all sources merged
Deadline, all sources
50 ms
of a ~350 ms page budget

Stoop's Home feed, from sources to screen

Stoop's Home feed, from sources to screen. The numbered component cards that follow describe each part.
Stoop's Home feed, from sources to screenComponents: 1. Friend inbox (References to friends' posts from the last 72 h, pushed in at write time (see news-feed/fan-out-on-write).), 2. Group and page index (Recent posts per group and per page, read at request time.), 3. Graph walker (Finds posts that the viewer's closest friends liked or commented on.), 4. Interest retriever (A two-tower user vector queried against an ANN index of recent public posts.), 5. Source mixer (Calls every source with a deadline, filters, dedupes, applies per-source quotas and sends 600 on.), 6. Ranking stack (The integrity check, the engagement model and the value scorer (the next three topics).).

Candidate sources

600 + source tags

ranked

1Friend inbox
~200 unseen
quota 240

2Group and page index
groups ~500, pages ~240
quotas 165 + 60

3Graph walker
friends engaged with
quota 60

4Interest retriever
two-tower + ANN
quota 65

Fresh local
< 2 h, within 5 km
quota 10

5Source mixer
dedupe · seen filter · quotas
600 out

6Ranking stack
1. integrity: drop or multiplier m
2. engagement model: 5 probabilities
3. value scorer: V × m

Feed page
top 25, after variety rules

Read the diagram from right to left to see why this stage matters. Everything after the mixer can only reorder or drop posts: the integrity check drops what must not be shown and passes on a multiplier, the engagement model predicts, the value scorer turns predictions into one number. None can add a post the sources missed. So the ceiling on feed quality is set here: if a close friend's news never reaches the 600, the best ranker in the world can't show it. The later stages exist to pick well among the candidates. This one exists to make sure the right ones are in the pile.

Was this section helpful?

Data.

Every candidate carries a tag saying which source found it. Without that tag you can't tell which source earns its slots.

SourceUsed asNote
Impression log with source tagsLabels for per-source light rankersOne row per shown post: viewer, post, the source or sources that found it, position, and what the viewer did. A post found by two sources keeps both tags.
Viewer-to-viewer interactions (likes, comments, messages, profile visits; 90 days, decayed)Edge weights for the graph walk and the friends rankerTurned into a predicted chance that one person interacts with another, the same idea as the real-graph model in X's open-sourced recommender.
Engagements that started outside the feed (search, notifications, a friend's profile)Recall ground truthPosts the viewer clearly wanted but the feed never offered. This is the only direct evidence of what the sources missed.
Random-inventory sample (0.1% of requests: the heavy ranker scores all ~940 connected posts, logged, not shown)Offline recall baselineShows how much the light rankers throw away that the heavy ranker would have liked.

What is eligible for one request (Maya, a typical viewer)

Assumptions
Friends who post
1901.5 posts each per 3 days
Groups joined
1445 posts each per 3 days
Pages followed
704 posts each per 3 days
Already seen (friends · groups · pages)
30% · 20% · 15%
Working
  1. Unseen friend posts190 × 1.5 = 285; × 0.7~200from Friends who post and Already seen (friends · groups · pages)
  2. Unseen group posts14 × 45 = 630; × 0.8504from Groups joined and Already seen (friends · groups · pages)
  3. Unseen page posts70 × 4 = 280; × 0.85238from Pages followed and Already seen (friends · groups · pages)
  4. Eligible connected posts200 + 504 + 238942from Unseen friend posts, Unseen group posts and Unseen page posts
  5. Eligible out-of-network postspublic posts from the last 72 hmillionsOnly a graph walk or an embedding search can reach these; no filter over the inbox will.
What it means
  • Friends' posts (~200) fit under their quota of 240, so they are never trimmed. The unused slots spill over to groups first, then to interest.
  • Groups and pages need a light ranker to cut them down (504 → 165 and 238 → 60).
  • Out-of-network posts need a retriever that goes looking for them, not a filter over what already arrived.

The impression log has a blind spot built in: a viewer can only engage in the feed with posts that some source already offered. Train a light ranker on feed engagements alone and it learns to agree with the sources it already has. The off-feed engagements and the random-inventory sample exist to break that loop. They are small, so weight them up when you measure recall, and keep them out of the rows that train the models (see exploration-vs-exploitation/feedback-loops).

Was this section helpful?

Features.

Each source's light ranker sees only cheap features: a few dozen per post, mostly precomputed.

FeatureUsed byType
Tie strength (predicted chance the viewer interacts with the author in the next 7 days)friends · graph walkdaily batch score
Days since last interaction with the authorfriendsnumeric, bucketed
Viewer's engagement rate in this group or on this page (28 days)groups · pagesnumeric
Group activity (posts per day, members)groupsnumeric, log-scaled
Post age · type (text, photo, video)allbucketed · categorical
Early engagement (likes and comments in the first hour per 100 impressions)allstreaming counter
Close friends who engaged, weighted by tie strengthgraph walkcomputed during the walk
Distance from the viewer's home areafresh localnumeric, km

Tie strength, the feature everything reuses

Tie strength is a small model of its own. For each pair of viewer and friend it predicts whether the viewer will like, comment on or message that friend in the next 7 days, from interaction counts with time decay, mutual friends and shared groups. X's open-sourced code has a component with the same job, real-graph, which predicts how likely one account is to interact with another. Stoop scores Maya's ~190 friend edges every night, and three consumers read the result: the friends light ranker, the graph walker, and the engagement model in the next topic.

A friend added yesterday has no interaction history, so for the first few weeks her score leans on mutual-friend count and shared groups. Without that fallback, every new friendship would start at zero and never get the exposure that creates the history the model needs.

01
The model inside each light ranker
Chosen:Logistic regression on ~30 cheap features
  • Pro:Scores ~900 posts in well under a millisecond on one core
  • Pro:Easy to read which feature moved a post
  • Pro:Retrains in minutes per source
Downside we accept:
  • Con:No feature interactions unless you add crosses by hand
  • Con:Its scores are not comparable across sources
Ruled out:Small distilled network trained to mimic the heavy ranker

Needs the heavy ranker's scores on the dropped posts too (the random-inventory sample); Harder to debug when one source's share shifts

Was this section helpful?

Model.

Six sources, each cheap and each biased in its own way, then one decision about how many slots each gets.

Stoop's sources

SourceHow it finds postsLight rankerBlind spot
FriendsReads the fan-out inbox (news-feed/fan-out-on-write).None needed; all ~200 pass.Friends who post rarely can still be buried by busy ones, but that happens in ranking, not here.
GroupsReads each joined group's recent posts from the index.Logistic regression on the light features; keeps 165.Busy groups crowd out small ones, so each group is capped at 30.
PagesReads followed pages' recent posts.Same model family; keeps 60.A page that posts 20 times a day.
Friends engaged withWalks from the viewer to her 20 strongest ties, then to posts they liked or commented on in the last 48 h.The sum of tie strengths of the friends who engaged.Amplifies whatever her friends already like.
InterestTwo-tower user vector, ANN search over public posts from the last 72 h (video-recommendation/candidate-retrieval).The dot product.Needs history, so it is weak for new viewers.
Fresh localPosts under 2 h old within 5 km.Early-engagement rate.Brand-new posts with no signal yet.

Real systems split the work the same way. X's 2023 write-up says its For You timeline averages half posts from followed accounts and half from outside, the latter found by walking a user-to-post engagement graph and by searching community embeddings (SimClusters). Pinterest's Pixie runs random walks on a graph of 3 billion nodes and 17 billion edges, and one server answers 1,200 requests a second at 60 ms (Eksombatchai et al. 2018). Stoop's six sources are a smaller version of that mix.

Share of a source's engaged posts captured, by slots given (Stoop, illustrative)

  • Groups
  • Pages
  • Interest
Share of a source's engaged posts captured, by slots given (Stoop, illustrative)Pages flatten out after about 60 slots while the interest source is still climbing at 50, so page slots beyond 60 buy less than the same slots given to interest.020%40%60%80%100%050100150200250300pages quotainterest quotaGroupsPagesInterestEngaged posts captured (%)Slots given to the sourceShare of a source's engaged posts captured, by slots given (Stoop, illustrative)Pages flatten out after about 60 slots while the interest source is still climbing at 50, so page slots beyond 60 buy less than the same slots given to interest.020%40%60%80%100%050100150200250300pages quotainterest quotaGroupsPagesInterestEngaged posts captured (%)Slots given to the source
Measured offline on the random-inventory sample: for each quota, the share of that source's engaged posts the light ranker would have kept.
Data
Slots given to the sourceGroups (%)Pages (%)Interest (%)
0000
20no value45no value
2530no value30
40no value70no value
5050no value45
60no value84no value
90no value90no value
10074no value60
150869568
20092no value73
30097no valueno value
  • At 90: pages quota
  • At 50: interest quota

Moving 30 slots away from pages

Assumptions
Engaged posts per request each source could supply if uncapped
groups 1.2 · pages 0.5 · interest 0.6
Groups slope, 150 → 200 slots
6 points per 50 slots
Pages slope, 60 → 90 slots
6 points per 30 slots
Interest slope, 50 → 100 slots
15 points per 50 slots
Working
  1. One more page slot is worth0.5 × 6% ÷ 300.0010 engaged postsfrom Engaged posts per request each source could supply if uncapped and Pages slope, 60 → 90 slots
  2. One more group slot is worth1.2 × 6% ÷ 500.00144 engaged postsfrom Engaged posts per request each source could supply if uncapped and Groups slope, 150 → 200 slots
  3. One more interest slot is worth0.6 × 15% ÷ 500.0018 engaged postsfrom Engaged posts per request each source could supply if uncapped and Interest slope, 50 → 100 slots
  4. Loss: pages 90 → 606 points × 0.50.030 per requestfrom Engaged posts per request each source could supply if uncapped and Pages slope, 60 → 90 slots
  5. Gain: groups 150 → 165, interest 50 → 6515 × 0.00144 + 15 × 0.0018 = 0.0216 + 0.0270.0486 per requestfrom One more group slot is worth and One more interest slot is worth
  6. Net gain0.0486 − 0.030 = 0.0186; × 160M requests~3.0M engaged posts a dayfrom Loss: pages 90 → 60 and Gain: groups 150 → 165, interest 50 → 65
What it means
  • Give the next slot to the source where it brings in the most future engagement, and re-measure after each move, because the slopes flatten as a source grows.
  • Why not all 30 to interest? Its curve is still straight up to 100 slots, so the limit is the guardrail, not the curve. With 65 interest slots, out-of-network impressions are 11 + 11 + 2 = 24% against the 25% cap; at 80 slots interest would reach about 11 × 80 ÷ 65 = 13.5%, taking the total to about 26.5%.
  • Final quotas: friends 240, groups 165, pages 60, friends engaged 60, interest 65, fresh local 10 = 600.
01
Merging sources
Chosen:Per-source light rankers with quotas tuned by marginal recall
  • Pro:Each source stays easy to debug on its own
  • Pro:A slow source costs only its own slots
  • Pro:Quotas stop one source from flooding the candidates
Downside we accept:
  • Con:Quotas need re-tuning as sources change
  • Con:Scores are not comparable across sources
Ruled out:One shared light ranker over the union, top 600

Needs the same features for every source's posts; The source with the most candidates tends to dominate; One model to retrain for any change to any source

Ruled out:No trimming: send everything to the heavy ranker

~940+ heavy scores per request instead of 600, about 57% more cost (940 ÷ 600 = 1.57); Still finds nothing outside the network

The shared-ranker option is not a straw man. Large feeds often trim the union of all sources with one first-pass model, sometimes distilled from the main ranker; serving-architectures/multi-stage-funnels walks through published examples and transfer-learning-and-compression/distillation covers how such a model is trained. It suits a system whose sources share one feature set. Stoop keeps per-source quotas while its sources are few and differ a lot; merging is the natural move once they converge.

Was this section helpful?

Evaluation.

A source is judged by what it adds that no other source would have found.

Recall@600
Offline, main metric. Share of engaged posts, incl. ones reached outside the feed, that were candidates.
Marginal recall
Offline, per source. Recall lost when this source is removed.
Meaningful interactions
Online, A/B test. Comments + shares per viewer, read with 28-day return.
Out-of-network share
Online, guardrail. Share of impressions, at most 25%.

Each source's share of candidates, impressions and engagements (Stoop, illustrative)

  • Candidates
  • Impressions
  • Meaningful interactions
Each source's share of candidates, impressions and engagements (Stoop, illustrative)For Maya, friends supply a third of candidates but 58% of comments and shares, groups supply 34% of candidates but 22% of interactions, and pages turn 10% of candidates into only 3%.010%20%30%40%50%FriendsGroupsPagesFriends engagedInterestFresh local33.3%34.2%10%10%10.8%1.7%44%26%6%11%11%2%58%22%3%9%6%2%SourceShare (%)Each source's share of candidates, impressions and engagements (Stoop, illustrative)For Maya, friends supply a third of candidates but 58% of comments and shares, groups supply 34% of candidates but 22% of interactions, and pages turn 10% of candidates into only 3%.010%20%30%40%50%FriendsGroupsPagesFriends engagedInterestFresh local33.3%34.2%10%10%10.8%1.7%44%26%6%11%11%2%58%22%3%9%6%2%SourceShare (%)
Candidate shares are Maya's: ~200 friend posts (200 ÷ 600 = 33.3%), and the 40 unused friend slots go to groups (165 + 40 = 205; 205 ÷ 600 = 34.2%). When a source's share of engagement sits well below its share of candidates, test a smaller quota. Leave-one-source-out A/B tests give the causal answer.
Data
SourceCandidates (%)Impressions (%)Meaningful interactions (%)
Friends33.34458
Groups34.22622
Pages1063
Friends engaged10119
Interest10.8116
Fresh local1.722

Correlation is not a source's value

The chart is correlational: the ranker decides which candidates get shown, so a source can look weak because the ranker distrusts it, or strong because it duplicates posts another source would have supplied anyway. The causal test is a leave-one-source-out arm, or a halved quota, run for 2 to 4 weeks (a-b-testing/experiment-design). At Stoop, watch the out-of-network sources closely: if one lifts clicks and likes while comments, shares or return fall, it is not earning its slots, so judge them on meaningful interactions and 28-day return, not on clicks. Out-of-network impressions here add up to 11 + 11 + 2 = 24%, just under the 25% guardrail.

Was this section helpful?

Serving and monitoring.

Sources are called in parallel with a deadline. A late source loses its slots; the page never waits.

One Home request when the graph walker is slow

One Home request when the graph walker is slow, as an ordered list of steps:
One Home request when the graph walker is slow10 steps between Source mixer, Friend inbox, Group and page index, Graph walker, Interest retriever, Ranking stack. The steps are listed as text after the diagram.Ranking stackInterest retrieverGraph walkerGroup and page indexFriend inboxSource mixermisses deadlinefilter · dedupe · trim · quotas; walker's 60 → groupsrefs, last 72 h1groups, pages, fresh local2walk 20 ties, 40 ms3ANN top 3004~285 refs5~910 posts + local6300 posts7600 + source tags8
  1. Source mixer → Friend inbox: refs, last 72 h
  2. Source mixer → Group and page index: groups, pages, fresh local
  3. Source mixer → Graph walker: walk 20 ties, 40 ms
  4. Source mixer → Interest retriever: ANN top 300
  5. Friend inbox → Source mixer (reply): ~285 refs
  6. Group and page index → Source mixer (reply): ~910 posts + local
  7. Interest retriever → Source mixer (reply): 300 posts
  8. Note over Graph walker: misses deadline
  9. Note over Source mixer: filter · dedupe · trim · quotas; walker's 60 → groups
  10. Source mixer → Ranking stack: 600 + source tags
StepWhenWhat happens
Score tie strengthNightlyScore every viewer-friend edge; the friends ranker, the graph walker and the engagement model read it.
Embed new public postsEvery few minutesAppend their vectors to the interest index so a post can be found within minutes of being written.
Retrain light rankersDaily per sourceTrain on yesterday's tagged impressions; compare recall on the random-inventory sample before swapping.
Re-tune quotasMonthly or after a source changesRedraw the marginal-recall curves and move slots toward the steepest one, within the out-of-network guardrail.

What to watch for

FailureImpactDetectionMitigationMeanwhile
Walker keeps timing out3Graph walkerFewer posts that close friends engaged withIts share of candidates falls to zero on the per-source dashboardBackfill its slots from groups; cap the walk's depth; alert on the source's share, not only on latency.Viewers still get 600 candidates, with fewer posts their friends engaged with.
Stale index4Interest retrieverOut-of-network picks are old newsOut-of-network candidates are all more than 24 h oldAppend new posts' vectors every few minutes; alert on the age of the newest vector.In-network sources still fill their quotas.
Popularity loop in the friends-engaged sourceThe same viral posts fill every feedThe share of candidates coming from the top 100 posts keeps risingDown-weight posts already engaged with by many ties across the network; reserve exploration slots (exploration-vs-exploitation/exploration-slots).Other sources keep their own quotas.
Inbox gaps for accounts with huge followings1Friend inboxViewers miss big accounts' new postsFriends' posts reach the feed hours late; time from post to first candidate climbsRead large accounts at request time instead of pushing them (news-feed/fan-out-on-write).Posts arrive late, not lost.
Was this section helpful?
Builds on this
Predicting engagement
Read next