Predicting engagement.
Maya's feed has room for about 25 posts and the ranker is holding 600. Before anything can be ordered, it needs five honest numbers for each post: how likely she is to like it, comment, share it, scroll straight past, or hide it. This topic builds the model that produces them, from which impressions become training rows to how a post's first few minutes of reactions are weighed against its author's track record until there are enough impressions for them to speak for themselves.
Builds on Candidate sources and Shared layers and task heads.
Framing.
The engagement model doesn't decide the order. It supplies the probabilities that the value model will weigh.
From business goal to ML task
| Layer | Definition |
|---|---|
| Business goal | Viewers find their feed worth coming back to (28-day return), with more comments and shares between people who know each other. |
| ML objective | Accurate, calibrated probabilities of five actions for each (viewer, post) impression, judged by per-head normalized entropy and calibration. |
| ML task | Multi-label classification: one network with five sigmoid heads, run over about 600 candidates per request. |
| Out of scope | Which posts become candidates (candidate sources), how five probabilities become one score (value model), and what policy removes or demotes (integrity). |
What the model must do
Scale and budget
One request in, five answers per post out
Why five numbers and not one score? Each action means something different to Stoop. A comment starts a conversation the author sees; a share carries the post to people who don't follow its author; a hide is the viewer telling you that you got it wrong. The value model gives each action its own weight, and it can only do that if the engagement model keeps them apart and gets each level right. A model that ranks well but says 8% where the truth is 4% doubles that action's pull in the blend. Facebook's News Feed works the same way: it predicts several kinds of engagement for each story and scores each eligible story independently and in parallel (Meta Engineering 2021).
Data.
A training row is one impression, joined with what the viewer did in the next 24 hours.
Five labels per impression (an impression counts only when the post is at least 50% on screen for at least 0.5 s)
| Head | Positive when | Base rate | Note |
|---|---|---|---|
| like | liked within 24 h | 5% | The most plentiful active signal; the other heads lean on what it teaches the shared layers. |
| comment | commented within 24 h | 0.8% | High effort and visible to the author; the action Stoop most wants more of. |
| share | shared to a friend or group within 24 h | 0.4% | Carries the post beyond its original audience. |
| skip | on screen under 2 s | 50% | Defined on every impression, quiet ones included. LinkedIn defined skips as dwell below a threshold and found one threshold worked across update types (LinkedIn Engineering 2020). |
| hide | hid the post or its author | 0.15% | The strongest negative. Kept separate from skip; scrolling past is not the same as objecting. |
The visibility rule matters more than it looks. Without it, a post that loaded below the fold and was never seen becomes a negative for every head, and the model learns that posts in slot 20 are unappealing when they were simply unseen. With it, the log only holds posts the viewer had a real chance to act on.
From 4.0B impressions to 616M training rows a day
- Impressions a day
- 4.0B160M requests × 25 posts seen (illustrative)
- Impressions with a like, comment, share or hide
- 6%
- Share of quiet impressions kept
- 10%
- Active rows, kept whole4.0B × 0.06240Mfrom Impressions a day and Impressions with a like, comment, share or hide
- Quiet rows kept(4.0B − 240M) × 0.10 = 3.76B × 0.10376Mfrom Impressions a day, Active rows, kept whole and Share of quiet impressions kept
- Training rows a day240M + 376M616Mfrom Active rows, kept whole and Quiet rows kept
- Loss weight on each kept quiet row1 ÷ 0.10, so 376M × 10 = 3.76B, the true count10from Share of quiet impressions kept and Quiet rows kept
- Positives a day (like, comment, share, hide)4.0B × 5%, 0.8%, 0.4%, 0.15%200M, 32M, 16M, 6Mfrom Impressions a day · They add to 254M, more than 240M, because some impressions get two actions (a like and a comment).
- Like rate the model would learn with no weights200M ÷ 616M≈ 32%from Positives a day (like, comment, share, hide) and Training rows a day
- Like rate it learns with weights200M ÷ (240M + 376M × 10) = 200M ÷ 4.0B5%from Positives a day (like, comment, share, hide), Active rows, kept whole, Quiet rows kept and Loss weight on each kept quiet row
- Sampling cuts the day's rows by about 85% (616M instead of 4.0B), which is what lets a daily retrain finish.
- Weighting kept quiet rows in the loss keeps every head's average at its true rate, including skip, whose positives live mostly in the quiet rows.
- Facebook's ad-click work downsampled negatives and recalibrated afterwards for a single label (He et al. 2014). With five labels on one row, a weight in the loss is the simpler fix.
- Pro:Every head's average prediction stays at its true rate, with no step after training
- Pro:One rule covers five labels, including skip, which is positive on many quiet rows
- Pro:About 85% fewer rows to read and train on
- Con:Weighted rows make gradients noisier; rare quiet patterns are thinned along with common ones
- Con:Anyone reading the table must remember the weight column, or every rate they compute is wrong
The correction assumes the dropped rows are negatives of that one label. A quiet row is negative for four heads and often positive for skip, so no single correction fits all five (the correction itself is in class-imbalance-and-sampling/resampling)
About 6.5× the rows (4.0B ÷ 616M) for almost no gain on the rare heads, whose positives are all kept anyway; The daily retrain would no longer fit in a day
Labels arrive late. A comment can come hours after the post was seen, so the join waits 24 hours before it writes a row, and each row is written once, after the window closes. The price is that the model serving today has learned from impressions up to about two days ago. Split offline data by time as well: train on days 1–14 and test on day 15, the way the model will actually be used. A random split lets the model peek at the same posts' later reactions.
Features.
The heavy model can afford hundreds of features per candidate. The ones that move it most describe the pair, and the post's first minutes.
| Family | Examples | Freshness |
|---|---|---|
| Viewer | Activity level, languages, an embedding of the authors engaged with in 28 days, device | daily batch |
| Author | Posting rate, historical likes per 100 impressions, account age | daily batch |
| Viewer × author | Tie strength (from candidate sources), days since the two last interacted, comments exchanged in 90 days, mutual friends | daily batch |
| Post | Type (text, photo, video), text and image embeddings, age in minutes, which source proposed it | at creation |
| Post counters | Likes, comments and hides per 100 impressions so far; impressions in the last 10 minutes | streaming, seconds |
| Context | Time of day, how deep into the session, network type; feed position (training only) | per request |
A post an hour old has no history, only counters, so those counters are streamed and read in seconds (see feature-pipelines-and-stores/batch-vs-streaming). Two rules keep them honest. First, log the counter values the model saw at serving time into the impression row, and train on those logged values. Counters recomputed later include reactions that happened after the impression, so training would see information serving never has (feature-pipelines-and-stores/training-serving-skew). Second, smooth early rates toward a prior, so a handful of impressions can't masquerade as a trend. The method, and how strong to make the prior, is in feature-engineering/feature-types; the feed-specific choice is the prior itself. Stoop uses the author's own like rate over the last 28 days rather than a site-wide rate, so a neighbour whose posts usually draw a few likes is not judged against a page that draws hundreds, and the post's own count takes over within a few hundred impressions.
One feature here is easy to overlook: which source proposed the post. It lets the model learn, for example, that a post from the interest source needs a stronger match to earn a like than a friend's post does. It also means a change to the source quotas shifts this feature's mix, so retrain after quotas move rather than before.
Model.
One network shares the work of understanding the viewer and the post; then five small towers answer five questions.
Shared bottom, five heads and a position tower
What each part does
| Part | What it does | Why here |
|---|---|---|
| Embeddings | Turn viewer, author, post-type and language IDs into vectors; the long tail is hashed into shared buckets | The viewer–author pair space is huge and sparse |
| Cross layers | Build explicit feature interactions, such as tie strength × post type | Cheaper than hoping the MLP discovers every cross (DCN V2, Wang et al. 2021; see feature-engineering/crosses) |
| Shared MLP | One common representation for all heads | Share (16M positives a day) and hide (6M) borrow from like (200M); see multi-task-and-multi-objective/shared-bottoms |
| Five towers | A small MLP and a sigmoid per action | Each action gets its own nonlinearity and its own loss |
| Position tower | Position × device, turned into a logit added to every head during training | Soaks up the "shown higher, engaged more" effect so the heads do not have to |
Like rate by feed position, relative to position 1 (Stoop, illustrative)
- Shuffled
- Learned
Data
| Feed position | Shuffled | Learned |
|---|---|---|
| 1 | 1 | 1 |
| 2 | 0.875 | 0.84 |
| 3 | 0.79 | 0.76 |
| 4 | 0.74 | 0.7 |
| 5 | 0.69 | 0.66 |
| 6 | 0.67 | 0.63 |
| 7 | 0.64 | 0.6 |
| 8 | 0.625 | 0.58 |
| 9 | 0.61 | 0.565 |
| 10 | 0.6 | 0.55 |
Why not just feed position in as another feature? At serving there is no position yet: the order is what you are computing. The simple baseline is to train with position as a feature and serve every candidate the same fixed default (see data-collection-and-labeling/implicit-labels); it works, but the heads are free to entangle position with everything else. A separate shallow tower refines it by keeping the position effect in its own additive term, crossed with device and trained with dropout; YouTube's version and its live comparison against the plain-feature baseline are covered in video-recommendation/ranking (Zhao et al. 2019). Training uses it, so the heads are not credited with engagement the slot caused; serving leaves it out, so every candidate is judged as if shown in the same place. Dropping position on a slice of training rows keeps the heads from leaning on it. One caution: because the old ranker put better posts higher, the tower also absorbs some relevance, which is why the shuffled sessions are needed to check it.
- Pro:Sparse heads (share, hide) borrow from the plentiful like signal
- Pro:One feature fetch and one model call per request
- Pro:Learns pair × post interactions such as close friend × photo
- Con:All five heads retrain and ship together; a regression in one blocks the rest
- Con:Needs accelerators to serve 3.3M scores a second at peak
Weak on huge sparse IDs such as viewer and author; Five models to keep in step, each fetching its own features
Five times the bottom compute per candidate; Share and hide starve on their own few million positives a day
Evaluation.
Each head is judged on its own, and on whether its probabilities are safe to multiply by a weight.
Normalized entropy (NE, defined in offline-metrics/calibration and worked for click prediction in ad-click-prediction/ctr-model) divides a head's log loss by the loss of always predicting its base rate (He et al. 2014). That division is why Stoop reads NE rather than raw log loss: the five heads have base rates from 0.15% to 50%, and a hide head at 0.15% has a tiny log loss even when it has learned nothing. Each head is compared with the live model's NE on the same next-day data.
A candidate model against the live one, day 15 (Stoop, illustrative)
| Head | NE vs live | AUC vs live | Calibration ratio | Verdict |
|---|---|---|---|---|
| like | −0.40% | +0.15% | 1.01 | Better |
| comment | −0.25% | +0.10% | 0.99 | Better |
| share | −0.10% | +0.03% | 1.02 | About the same |
| skip | −0.30% | +0.12% | 1.00 | Better |
| hide | +0.20% | +0.01% | 1.06 | Blocks the launch: over-predicts hides by 6%, so the value model would push good posts down |
Freshness is part of the readout too. Every day of data a model has not seen costs some NE, and how fast that cost grows sets the retraining schedule (monitoring-and-continual-learning/retraining draws the curve and cites He et al.'s daily-versus-weekly result). For Stoop the comment head ages fastest, because what neighbours talk about changes with local news, so it sets the pace: retrain daily. With the 24 h settle window, a daily model always serves on data one to two days old, and a candidate is compared with the live model on the same held-out day, never with a model of a different age.
Report every head on slices, not just the average: new viewers (under 14 days), viewers with fewer than 20 friends, video against text, and each source tag. A model that is better on average but worse on out-of-network posts quietly tilts the feed toward friends without anyone having decided that. And offline gains don't always survive contact with real traffic (Zhao et al. 2019 describe this gap), so every candidate that passes these checks still goes through an A/B test.
Serving and monitoring.
Features for 600 posts are fetched once and scored in one batched call; the model is replaced daily.
Scoring one request
- Candidate sources → Feature hydrator: 600 post IDs + source tags
- Feature hydrator → Post counters: counters for 600 posts
- Post counters → Feature hydrator (reply): live counts
- Note over Feature hydrator: viewer features once, post features per candidate
- Feature hydrator → Engagement model: 600 rows, no position
- Feature hydrator → Value scorer: counters used, for the log
- Engagement model → Value scorer: 600 × 5 probabilities
| Step | When | What happens |
|---|---|---|
| Hydrate features | per request | Viewer and pair features once; post features and counters for all 600 candidates, counters from the streaming store. |
| Score | per request | One batched call returns 600 × 5 probabilities, with the position tower off. |
| Log | per impression | The features and predictions the model used, the position shown, and the model version. |
| Join labels | daily, after 24 h | Impressions plus actions become 616M weighted rows. |
| Retrain | daily | Warm-start from yesterday's model, compare with the live model on the held-out day, then canary before full rollout (model-deployment/shadow-and-canary). |
What to watch for
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| Counter skew2Post counters | Young posts are over-scored and crowd out older ones | The like head over-predicts on posts under 1 h old | Train on counters logged at serving time; alert on the gap between served and recomputed values. | Older posts still rank, just lower. |
| Unsettled labels4Training pipeline | The model learns that comments are rarer than they are | Comment rate in the training data dips on the most recent day | Keep the 24 h settle window; never train on rows younger than it. | Yesterday's model keeps serving. |
| Calibration drift after a UI change (a new reaction button moves likes)3Engagement model | The blend over- or under-weights one action | The like head's calibration ratio leaves 0.97–1.03 | Recalibrate the head or retrain; freeze value-model weight changes until it is back in range. | Ranking still works, with one action mispriced. |
| Position leaks into serving | Offline gains vanish online | Scores change when the same candidates are sent in a different order | Assert position is missing at serving; test with shuffled sessions. | Feeds still render; quality quietly drops. |