Social feed rankingPredicting engagement

100%

Predicting engagement.

Maya's feed has room for about 25 posts and the ranker is holding 600. Before anything can be ordered, it needs five honest numbers for each post: how likely she is to like it, comment, share it, scroll straight past, or hide it. This topic builds the model that produces them, from which impressions become training rows to how a post's first few minutes of reactions are weighed against its author's track record until there are enough impressions for them to speak for themselves.

Intermediate19 minUpdated 1 Oct 2026

Builds on Candidate sources and Shared layers and task heads.

Framing.

The engagement model doesn't decide the order. It supplies the probabilities that the value model will weigh.

From business goal to ML task

LayerDefinition
Business goalViewers find their feed worth coming back to (28-day return), with more comments and shares between people who know each other.
ML objectiveAccurate, calibrated probabilities of five actions for each (viewer, post) impression, judged by per-head normalized entropy and calibration.
ML taskMulti-label classification: one network with five sigmoid heads, run over about 600 candidates per request.
Out of scopeWhich posts become candidates (candidate sources), how five probabilities become one score (value model), and what policy removes or demotes (integrity).

What the model must do

#1For every candidate in a request, return P(like), P(comment), P(share), P(skip) and P(hide) for this viewer.
#2Keep each head calibrated, so its average prediction matches the rate actually observed.
#3Give a sensible answer for a post that is only minutes old.
#4Score a post on its own merits, not on the slot it happened to be shown in last time.
#5Learn from yesterday by tomorrow.

Scale and budget

Scores per day
96B
160M requests × 600 candidates (illustrative)
At peak
3.3M/s
3× the 1.1M/s average
Scoring budget
110 ms
features + model, of a ~350 ms feed request
Training rows a day
616M
after sampling quiet impressions

One request in, five answers per post out

One request in, five answers per post out. The numbered component cards that follow describe each part.
One request in, five answers per post outComponents: 1. Feature hydrator (Fetches viewer, author, pair and post features for 600 candidates, with post counters from the streaming store.), 2. Post counters (Likes, comments, hides and impressions per post, updated within seconds.), 3. Engagement model (One multi-head model five probabilities per candidate in one batched call.), 4. Training pipeline (Joins impressions with actions after a 24 h settle window retrains daily from the last model.).

Dailyafter the 24 h settle window

Per request

post counters

600 feature rows

600 × 5 probabilities

log what was served

join actions

new model daily

600 candidates
from candidate sources

1Feature hydrator
viewer · author · pair · post · context

2Post counters
updated within seconds

3Engagement model
one batched call
position tower off
P(like) · P(comment) · P(share) · P(skip) · P(hide)

Impression log
features and predictions as served

4Training pipeline
616M weighted rows a day

Value scorer
weighs the five (next topic)

Why five numbers and not one score? Each action means something different to Stoop. A comment starts a conversation the author sees; a share carries the post to people who don't follow its author; a hide is the viewer telling you that you got it wrong. The value model gives each action its own weight, and it can only do that if the engagement model keeps them apart and gets each level right. A model that ranks well but says 8% where the truth is 4% doubles that action's pull in the blend. Facebook's News Feed works the same way: it predicts several kinds of engagement for each story and scores each eligible story independently and in parallel (Meta Engineering 2021).

Was this section helpful?

Data.

A training row is one impression, joined with what the viewer did in the next 24 hours.

Five labels per impression (an impression counts only when the post is at least 50% on screen for at least 0.5 s)

HeadPositive whenBase rateNote
likeliked within 24 h5%The most plentiful active signal; the other heads lean on what it teaches the shared layers.
commentcommented within 24 h0.8%High effort and visible to the author; the action Stoop most wants more of.
shareshared to a friend or group within 24 h0.4%Carries the post beyond its original audience.
skipon screen under 2 s50%Defined on every impression, quiet ones included. LinkedIn defined skips as dwell below a threshold and found one threshold worked across update types (LinkedIn Engineering 2020).
hidehid the post or its author0.15%The strongest negative. Kept separate from skip; scrolling past is not the same as objecting.

The visibility rule matters more than it looks. Without it, a post that loaded below the fold and was never seen becomes a negative for every head, and the model learns that posts in slot 20 are unappealing when they were simply unseen. With it, the log only holds posts the viewer had a real chance to act on.

From 4.0B impressions to 616M training rows a day

Assumptions
Impressions a day
4.0B160M requests × 25 posts seen (illustrative)
Impressions with a like, comment, share or hide
6%
Share of quiet impressions kept
10%
Working
  1. Active rows, kept whole4.0B × 0.06240Mfrom Impressions a day and Impressions with a like, comment, share or hide
  2. Quiet rows kept(4.0B − 240M) × 0.10 = 3.76B × 0.10376Mfrom Impressions a day, Active rows, kept whole and Share of quiet impressions kept
  3. Training rows a day240M + 376M616Mfrom Active rows, kept whole and Quiet rows kept
  4. Loss weight on each kept quiet row1 ÷ 0.10, so 376M × 10 = 3.76B, the true count10from Share of quiet impressions kept and Quiet rows kept
  5. Positives a day (like, comment, share, hide)4.0B × 5%, 0.8%, 0.4%, 0.15%200M, 32M, 16M, 6Mfrom Impressions a day · They add to 254M, more than 240M, because some impressions get two actions (a like and a comment).
  6. Like rate the model would learn with no weights200M ÷ 616M≈ 32%from Positives a day (like, comment, share, hide) and Training rows a day
  7. Like rate it learns with weights200M ÷ (240M + 376M × 10) = 200M ÷ 4.0B5%from Positives a day (like, comment, share, hide), Active rows, kept whole, Quiet rows kept and Loss weight on each kept quiet row
What it means
  • Sampling cuts the day's rows by about 85% (616M instead of 4.0B), which is what lets a daily retrain finish.
  • Weighting kept quiet rows in the loss keeps every head's average at its true rate, including skip, whose positives live mostly in the quiet rows.
  • Facebook's ad-click work downsampled negatives and recalibrated afterwards for a single label (He et al. 2014). With five labels on one row, a weight in the loss is the simpler fix.
01
What to do with 3.76B quiet impressions a day
Chosen:Keep 10% and give each kept row a loss weight of 10
  • Pro:Every head's average prediction stays at its true rate, with no step after training
  • Pro:One rule covers five labels, including skip, which is positive on many quiet rows
  • Pro:About 85% fewer rows to read and train on
Downside we accept:
  • Con:Weighted rows make gradients noisier; rare quiet patterns are thinned along with common ones
  • Con:Anyone reading the table must remember the weight column, or every rate they compute is wrong
Ruled out:Keep 10%, then correct the predictions after training

The correction assumes the dropped rows are negatives of that one label. A quiet row is negative for four heads and often positive for skip, so no single correction fits all five (the correction itself is in class-imbalance-and-sampling/resampling)

Ruled out:Train on all 4.0B rows

About 6.5× the rows (4.0B ÷ 616M) for almost no gain on the rare heads, whose positives are all kept anyway; The daily retrain would no longer fit in a day

Labels arrive late. A comment can come hours after the post was seen, so the join waits 24 hours before it writes a row, and each row is written once, after the window closes. The price is that the model serving today has learned from impressions up to about two days ago. Split offline data by time as well: train on days 1–14 and test on day 15, the way the model will actually be used. A random split lets the model peek at the same posts' later reactions.

Was this section helpful?

Features.

The heavy model can afford hundreds of features per candidate. The ones that move it most describe the pair, and the post's first minutes.

FamilyExamplesFreshness
ViewerActivity level, languages, an embedding of the authors engaged with in 28 days, devicedaily batch
AuthorPosting rate, historical likes per 100 impressions, account agedaily batch
Viewer × authorTie strength (from candidate sources), days since the two last interacted, comments exchanged in 90 days, mutual friendsdaily batch
PostType (text, photo, video), text and image embeddings, age in minutes, which source proposed itat creation
Post countersLikes, comments and hides per 100 impressions so far; impressions in the last 10 minutesstreaming, seconds
ContextTime of day, how deep into the session, network type; feed position (training only)per request

A post an hour old has no history, only counters, so those counters are streamed and read in seconds (see feature-pipelines-and-stores/batch-vs-streaming). Two rules keep them honest. First, log the counter values the model saw at serving time into the impression row, and train on those logged values. Counters recomputed later include reactions that happened after the impression, so training would see information serving never has (feature-pipelines-and-stores/training-serving-skew). Second, smooth early rates toward a prior, so a handful of impressions can't masquerade as a trend. The method, and how strong to make the prior, is in feature-engineering/feature-types; the feed-specific choice is the prior itself. Stoop uses the author's own like rate over the last 28 days rather than a site-wide rate, so a neighbour whose posts usually draw a few likes is not judged against a page that draws hundreds, and the post's own count takes over within a few hundred impressions.

One feature here is easy to overlook: which source proposed the post. It lets the model learn, for example, that a post from the interest source needs a stronger match to earn a like than a friend's post does. It also means a change to the source quotas shifts this feature's mix, so retrain after quotas move rather than before.

Was this section helpful?

Model.

One network shares the work of understanding the viewer and the post; then five small towers answer five questions.

Shared bottom, five heads and a position tower

Shared bottom, five heads and a position tower
Shared bottom, five heads and a position towerParts: Inputs, Shared bottom, P(like), P(comment), P(share), P(skip), P(hide), Position tower.

shared representation

adds to each logit

Headsone sigmoid each

P(like)

P(comment)

P(share)

P(skip)

P(hide)

Inputs
viewer · author · pair · post · context

Shared bottom
embeddings + cross layers + MLP

Position tower
position × device
training only

What each part does

PartWhat it doesWhy here
EmbeddingsTurn viewer, author, post-type and language IDs into vectors; the long tail is hashed into shared bucketsThe viewer–author pair space is huge and sparse
Cross layersBuild explicit feature interactions, such as tie strength × post typeCheaper than hoping the MLP discovers every cross (DCN V2, Wang et al. 2021; see feature-engineering/crosses)
Shared MLPOne common representation for all headsShare (16M positives a day) and hide (6M) borrow from like (200M); see multi-task-and-multi-objective/shared-bottoms
Five towersA small MLP and a sigmoid per actionEach action gets its own nonlinearity and its own loss
Position towerPosition × device, turned into a logit added to every head during trainingSoaks up the "shown higher, engaged more" effect so the heads do not have to

Like rate by feed position, relative to position 1 (Stoop, illustrative)

  • Shuffled
  • Learned
Like rate by feed position, relative to position 1 (Stoop, illustrative)The same post gets about 40% fewer likes at position 10 than at position 1 in shuffled sessions. The tower trained on ordinary traffic learns a steeper drop, about 45%, because higher slots also held better posts and part of that relevance lands in the position term.0.50.60.70.80.91246810ShuffledLearnedLike rate vs position 1Feed positionLike rate by feed position, relative to position 1 (Stoop, illustrative)The same post gets about 40% fewer likes at position 10 than at position 1 in shuffled sessions. The tower trained on ordinary traffic learns a steeper drop, about 45%, because higher slots also held better posts and part of that relevance lands in the position term.0.50.60.70.80.91246810ShuffledLearnedLike rate vs position 1Feed position
On ordinary traffic, position and relevance move together: the old ranker put its best guesses at the top, so the tower alone soaks up some relevance and overstates the drop. The measured curve comes from the 0.5% of sessions (800k a day) whose top 10 is shuffled, so position is the only thing that differs, and it is used to check and correct the learned term. How a shuffled bucket measures position is in data-collection-and-labeling/implicit-labels; what its traffic costs is in exploration-vs-exploitation/exploration-slots.
Data
Feed positionShuffledLearned
111
20.8750.84
30.790.76
40.740.7
50.690.66
60.670.63
70.640.6
80.6250.58
90.610.565
100.60.55

Why not just feed position in as another feature? At serving there is no position yet: the order is what you are computing. The simple baseline is to train with position as a feature and serve every candidate the same fixed default (see data-collection-and-labeling/implicit-labels); it works, but the heads are free to entangle position with everything else. A separate shallow tower refines it by keeping the position effect in its own additive term, crossed with device and trained with dropout; YouTube's version and its live comparison against the plain-feature baseline are covered in video-recommendation/ranking (Zhao et al. 2019). Training uses it, so the heads are not credited with engagement the slot caused; serving leaves it out, so every candidate is judged as if shown in the same place. Dropping position on a slice of training rows keeps the heads from leaning on it. One caution: because the old ranker put better posts higher, the tower also absorbs some relevance, which is why the shuffled sessions are needed to check it.

01
Stoop's engagement model
Chosen:One deep multi-head model with cross layers and a position tower
  • Pro:Sparse heads (share, hide) borrow from the plentiful like signal
  • Pro:One feature fetch and one model call per request
  • Pro:Learns pair × post interactions such as close friend × photo
Downside we accept:
  • Con:All five heads retrain and ship together; a regression in one blocks the rest
  • Con:Needs accelerators to serve 3.3M scores a second at peak
Ruled out:Gradient-boosted trees feeding a logistic regression, one per head

Weak on huge sparse IDs such as viewer and author; Five models to keep in step, each fetching its own features

Ruled out:Five separate deep models

Five times the bottom compute per candidate; Share and hide starve on their own few million positives a day

Was this section helpful?

Evaluation.

Each head is judged on its own, and on whether its probabilities are safe to multiply by a weight.

Normalized entropy
Offline, main. Per head, vs the live model, on the next day's data.
Calibration ratio
Offline, guardrail. Mean predicted ÷ observed, 0.97–1.03 per head and slice.
Meaningful interactions
Online, A/B test. Comments + shares per viewer; 28-day return as the north star.
Hides per 1,000 impressions
Online, guardrail. Must not rise.

Normalized entropy (NE, defined in offline-metrics/calibration and worked for click prediction in ad-click-prediction/ctr-model) divides a head's log loss by the loss of always predicting its base rate (He et al. 2014). That division is why Stoop reads NE rather than raw log loss: the five heads have base rates from 0.15% to 50%, and a hide head at 0.15% has a tiny log loss even when it has learned nothing. Each head is compared with the live model's NE on the same next-day data.

A candidate model against the live one, day 15 (Stoop, illustrative)

HeadNE vs liveAUC vs liveCalibration ratioVerdict
like−0.40%+0.15%1.01Better
comment−0.25%+0.10%0.99Better
share−0.10%+0.03%1.02About the same
skip−0.30%+0.12%1.00Better
hide+0.20%+0.01%1.06Blocks the launch: over-predicts hides by 6%, so the value model would push good posts down

Freshness is part of the readout too. Every day of data a model has not seen costs some NE, and how fast that cost grows sets the retraining schedule (monitoring-and-continual-learning/retraining draws the curve and cites He et al.'s daily-versus-weekly result). For Stoop the comment head ages fastest, because what neighbours talk about changes with local news, so it sets the pace: retrain daily. With the 24 h settle window, a daily model always serves on data one to two days old, and a candidate is compared with the live model on the same held-out day, never with a model of a different age.

Report every head on slices, not just the average: new viewers (under 14 days), viewers with fewer than 20 friends, video against text, and each source tag. A model that is better on average but worse on out-of-network posts quietly tilts the feed toward friends without anyone having decided that. And offline gains don't always survive contact with real traffic (Zhao et al. 2019 describe this gap), so every candidate that passes these checks still goes through an A/B test.

Was this section helpful?

Serving and monitoring.

Features for 600 posts are fetched once and scored in one batched call; the model is replaced daily.

Scoring one request

Scoring one request, as an ordered list of steps:
Scoring one request7 steps between Candidate sources, Feature hydrator, Post counters, Engagement model, Value scorer. The steps are listed as text after the diagram.Value scorerEngagement modelPost countersFeature hydratorCandidate sourcesviewer features once, post features per candidate600 post IDs + source tags1counters for 600 posts2live counts3600 rows, no position4counters used, for the log5600 × 5 probabilities6
  1. Candidate sources → Feature hydrator: 600 post IDs + source tags
  2. Feature hydrator → Post counters: counters for 600 posts
  3. Post counters → Feature hydrator (reply): live counts
  4. Note over Feature hydrator: viewer features once, post features per candidate
  5. Feature hydrator → Engagement model: 600 rows, no position
  6. Feature hydrator → Value scorer: counters used, for the log
  7. Engagement model → Value scorer: 600 × 5 probabilities
StepWhenWhat happens
Hydrate featuresper requestViewer and pair features once; post features and counters for all 600 candidates, counters from the streaming store.
Scoreper requestOne batched call returns 600 × 5 probabilities, with the position tower off.
Logper impressionThe features and predictions the model used, the position shown, and the model version.
Join labelsdaily, after 24 hImpressions plus actions become 616M weighted rows.
RetraindailyWarm-start from yesterday's model, compare with the live model on the held-out day, then canary before full rollout (model-deployment/shadow-and-canary).

What to watch for

FailureImpactDetectionMitigationMeanwhile
Counter skew2Post countersYoung posts are over-scored and crowd out older onesThe like head over-predicts on posts under 1 h oldTrain on counters logged at serving time; alert on the gap between served and recomputed values.Older posts still rank, just lower.
Unsettled labels4Training pipelineThe model learns that comments are rarer than they areComment rate in the training data dips on the most recent dayKeep the 24 h settle window; never train on rows younger than it.Yesterday's model keeps serving.
Calibration drift after a UI change (a new reaction button moves likes)3Engagement modelThe blend over- or under-weights one actionThe like head's calibration ratio leaves 0.97–1.03Recalibrate the head or retrain; freeze value-model weight changes until it is back in range.Ranking still works, with one action mispriced.
Position leaks into servingOffline gains vanish onlineScores change when the same candidates are sent in a different orderAssert position is missing at serving; test with shuffled sessions.Feeds still render; quality quietly drops.
Was this section helpful?
Builds on this
Blending into one score
Read next