Social feed rankingBlending into one score

100%

Blending into one score.

Stoop's engagement model says a neighbour's post about a lost cat will get a comment from Maya with 3% probability, and a viral meme will get a like with 20%. Which is worth more depends on what Stoop believes each action does for Maya, for the neighbour who posted, and for whether both come back next month. This topic turns those beliefs into weights, shows how to estimate them without fooling yourself, and explains why a change that looks bad in week one can be right by week twelve.

Advanced22 minUpdated 1 Oct 2026

Builds on Predicting engagement and Blending objectives.

Framing.

The engagement model predicts; the value model decides what those predictions are worth to Stoop.

From business goal to ML task

By the time Stoop's value model runs, the hard prediction work is done. For each of about 600 candidate posts, the engagement model has said how likely the viewer is to like it, comment, share, scroll past, or hide it. The feed still needs a single number per post to sort by. The value model supplies that number, and in doing so it states what Stoop wants the feed to be for.

That makes it a small piece of code with a large job. If the weights say a like is worth as much as a comment, the feed fills with things that are easy to tap on. If they ignore the person who wrote the post, neighbours who post into silence stop posting, and next year there is less to rank. The weights are the one place where the product's idea of a good feed is written down as numbers.

LayerDefinition
Business goalViewers and the people who post both keep coming back: 28-day return for viewers and for authors.
ML objectiveAn ordering of candidates that maximises long-term return, within the hide and integrity guardrails.
ML taskA scoring function over calibrated predictions: a fixed weighted sum whose weights are estimated from experiments and long-term data, plus two small learned inputs (the survey answer and each viewer's action affinity).
Out of scopeBlend forms in general and weight sweeps (multi-task-and-multi-objective/blending); variety rules after scoring (news-feed/ranking); what integrity demotes and why (integrity).

What the value model must do

VR-01Turn each candidate's predictions into one score in the same unit for every post, so 600 candidates can be sorted.
VR-02Price every action by what it does for long-term return, not by how often it happens.
VR-03Count what the author gains from an interaction, not only what the viewer gains.
VR-04Include one signal that engagement misses, the viewer's own verdict on the post.
VR-05Adjust a weight where an action is worth more or less for a group of viewers, without counting a viewer's habits twice.
VR-06Let the integrity stage shrink a post's whole score, whatever its engagement.
VR-07Change weights per surface and experiment arm without retraining any model.

Scale and budget

Terms in the score
7
5 actions, survey, author; a typical V is 0.1 to 1 like-equivalents (illustrative)
Scored per request
600
under 1 ms: simple arithmetic
Weight changes
config
per surface and arm, no retrain
Longest holdout (months)
6
1% of viewers on the weights from the start of the half

Where the value score comes from

Where the value score comes from. The numbered component cards that follow describe each part.
Where the value score comes fromComponents: 1. Engagement model (Five calibrated probabilities per candidate.), 2. Survey model (Predicts the answer to 'Was this post worth your time?' for every candidate.), 3. Weight config (Global weights per surface and experiment arm changed without retraining.), 4. Value scorer (Personalises the weights, adds up the terms, applies the integrity multiplier.).

5 probabilities

1 probability

weights

scales weights

multiplier

600 scores

1Engagement model
5 head probabilities

2Survey model
p(worth your time)

3Weight config
w per surface · arm

Viewer affinity
segment rate ÷ global

Integrity multiplier m
from the integrity topic

4Value scorer
V = m × (Σ w·a·p + author)

Variety rules
news-feed/ranking

Every input to the scorer except the weights is predicted or measured per candidate. The weights are the only product decision in the picture, which is why they sit in config, carry an owner, and change through experiments rather than code pushes.

Was this section helpful?

Data.

The engagement log says what people did. Setting the weights needs data about what those actions led to weeks later.

Where the evidence for a weight comes from

SourceTells youCatch
Weight experiments: arms that raise or lower one weight, run 4–8 weekscausal effect of a weight on returnOnly a handful of weight settings can ever be tried; Cunningham et al. (2024) note experiments cover a small corner of the possible weights.
Long-term holdout: 1% of viewers kept on old weights for 6 monthswhether many changes added upSmall and slow, but the only guard against a string of short-term wins that sum to a long-term loss.
Observational logs: actions this week against return next month, per viewerfine-grained exchange ratesActive viewers do everything more. Without controlling for prior activity, every action looks valuable.
Item surveys: 'Was this post worth your time?' (yes, somewhat, no)what engagement missesSensitive to wording, and only about 20% of those asked answer, not at random.
Author logs: feedback an author received against whether they post againthe author's side of every actionThe viewer's logs never show it, so it has to be joined in by author ID.

Survey labels a day

Assumptions
Impressions a day
4.0B40M viewers × 4 sessions × 25 posts (illustrative)
Survey shown on
1 in 1,000 impressions
Share of prompts answered
20%
Video share of impressions
15%
New posts a day
12M40M × 0.3 (illustrative)
Working
  1. Prompts shown4.0B ÷ 1,0004.0M a dayfrom Impressions a day and Survey shown on
  2. Answers collected4.0M × 0.20800k labels a dayfrom Prompts shown and Share of prompts answered
  3. Answers about video posts800k × 0.15120k a dayfrom Answers collected and Video share of impressions
  4. Answers per new post, on average800k ÷ 12Mabout 0.07from Answers collected and New posts a day
What it means
  • 800k answers a day is plenty to train a survey model on the same features as the engagement heads.
  • Almost no single post ever gets an answer, so the survey model must predict the answer for every candidate; the raw answers are training labels only.
  • Answers differ by content type, and Cunningham et al. (2024) report survey models that did well on some kinds of content and poorly on others. Check calibration per type, and ask more often on the weak types.
Was this section helpful?

Features.

The value model's inputs are the heads' outputs, plus two things the heads don't know: what the viewer tends to do, and who wrote the post.

The terms in Stoop's score

TermInputSign and weight
Likep_like+1. The unit every other weight is priced in.
Commentp_comment+7. High effort, and it lands on the author.
Sharep_share+3. Spreads the post to a friend or a group.
Linger1 − p_skip+0.1. Catches posts people read but don't react to.
Hidep_hide−10. The strongest negative signal a viewer gives.
Worth your timep_survey_yes+0.1 per unit of probability, from the survey model.
Author bonus0.5 × comment term+ when the author received fewer than 2 comments in the last 7 days, so the comment term counts 1.5×.
Integrity multiplierm in (0, 1]Multiplies the whole sum. Set by the integrity topic.

Scaling the weights to each viewer

Meta's 2021 post on Facebook's feed points out a property of the weighted sum itself: an action a person rarely takes already plays a minimal role in their ranking, because its predicted probability is close to 0. The engagement model does that personalising. Maya has shared twice in three months, so her p_share is already small; no extra weight is needed to quiet it.

Stoop adds one per-viewer factor for a different reason: the same action can be worth more or less return depending on who takes it (Meta's post also notes that some people express themselves more through likes than comments). Stoop runs the exchange-rate estimate separately for viewer activity segments. The affinity aᵥ,ₖ is the segment's rate for action k divided by the global rate, clipped to 0.25 to 2. In Maya's segment, frequent commenters who rarely share, a comment is worth 1.6 times the global rate because it usually opens a thread with neighbours, and a share 0.25 times because it is mostly a forwarded link that brings nobody back.

The trap is setting affinity from how often a viewer takes the action. Maya's p_comment of 0.03, against a 0.8% average, already says she comments a lot; multiplying it by her comment rate as well would count the same habit twice and tilt her feed toward posts that beg for replies. So affinity comes from worth, not frequency, and the segment estimates control for prior activity just as the global ones do. It applies only to like, comment and share: hides are too rare for stable per-segment estimates, so the hide keeps its global −10. New viewers get 1 on every action until they have 30 days of history and a segment.

01
How the author's side enters the score
Chosen:A bonus on the comment term for authors who received little feedback
  • Pro:Targets the authors most likely to stop posting
  • Pro:One number (1.5×) that is easy to test in an arm
  • Pro:Costs nothing for authors who already get plenty of replies
  • Pro:Has a public precedent: X's 2023 heavy-ranker weights put a reply the author engages with far above a plain reply (the weights table is in multi-task-and-multi-objective/blending)
Downside we accept:
  • Con:A threshold (2 comments in 7 days) that authors could in theory game
  • Con:Ignores likes and shares, which also reach the author
Ruled out:Fold the author's gain into every weight

Gives the same boost to a page with 10,000 followers and to a neighbour posting for the first time

Ruled out:Leave authors to a separate producer-side team

The ranking is where attention is handed out; a fix elsewhere arrives after authors have already left

Was this section helpful?

Model.

The value model is a formula. The hard part is the weights, and what they are meant to be: exchange rates to long-term return.

Stoop's value function

# Weights are read from config per surface and experiment arm; this dict is illustrative.
W = {"like": 1.0, "comment": 7.0, "share": 3.0, "linger": 0.1, "hide": -10.0, "survey": 0.1}

def value(p, affinity, author_starved, m):
    v = sum(W[k] * affinity.get(k, 1.0) * p[k] for k in ("like", "comment", "share"))
    v += W["linger"] * (1 - p["skip"])
    v += W["hide"] * p["hide"]           # never scaled by affinity
    v += W["survey"] * p["survey_yes"]
    if author_starved:                    # author got < 2 comments in 7 days
        v += 0.5 * W["comment"] * affinity.get("comment", 1.0) * p["comment"]
    return m * v                          # integrity multiplier, 0 < m <= 1

What a comment is worth at Stoop

Assumptions
Extra 28-day return per like: viewer, author
+0.0015 pp, +0.0005 ppillustrative; activity-controlled estimates checked by weight experiments
Per comment: viewer, author
+0.010 pp, +0.004 pp
Per share: viewer, author
+0.004 pp, +0.002 pp
Per hide: viewer
−0.020 pp
Working
  1. A like, the unit (weight 1)0.0015 + 0.00050.002 ppfrom Extra 28-day return per like: viewer, author
  2. Comment weight(0.010 + 0.004) ÷ 0.0027from Per comment: viewer, author and A like, the unit (weight 1)
  3. Share weight(0.004 + 0.002) ÷ 0.0023from Per share: viewer, author and A like, the unit (weight 1)
  4. Hide weight−0.020 ÷ 0.002−10from Per hide: viewer and A like, the unit (weight 1)
  5. Share of a comment's value that goes to the author0.004 ÷ 0.014≈ 29%from Per comment: viewer, author
  6. Comment weight if only the viewer counted0.010 ÷ 0.0015≈ 6.7from Per comment: viewer, author and Extra 28-day return per like: viewer, author
What it means
  • Counting the author moves a comment from about 6.7 likes to 7. The gap looks small, but it is the part that keeps quiet neighbours posting, and it is larger still for authors who get little feedback (the 1.5× bonus).
  • Linger and the survey term have no clean per-action return estimate, so their weights (0.1 each) were set by a weight experiment with arms at 0.05, 0.1 and 0.2, read on 28-day return. The same estimate run per viewer segment gives the affinities.
  • Each of these numbers has wide error bars. Cunningham et al. (2024) describe retention-maximising weights as hard to pin down because experiments are few and logs are biased, so treat them as starting estimates to check with a weight experiment.

Equal contribution vs exchange rates (weights relative to like)

ActionBase rateEqual contributionExchange rateReading
Like5%11—
Comment0.8%6.257Close: both value comments highly.
Share0.4%12.53Equal contribution overprices shares because they are rare, not because they bring people back.
Hide0.15%−33−10Equal contribution punishes a hide more than three times as hard.

Equal contribution (taught in multi-task-and-multi-objective/blending) prices rarity; an exchange rate prices return. The table shows where that matters at Stoop: shares are rare, so equal contribution gives them 12.5 when the return data say 3. Cunningham et al. (2024) report that active, high-effort engagement such as commenting tends to be worth more for retention than passive signals, which fits comments keeping their high weight. They also cite YouTube's 2012 switch from optimising clicks to clicks plus watch time: clicks dropped in the short term while long-term retention rose.

Stoop is not alone in writing value this way. Instagram's Explore (Meta AI, 2019) described its final score as weighted predictions of likes and saves minus a weighted prediction of negative actions, the same shape with a different list of terms.

Scoring two of Maya's candidates, term by term

Assumptions
Lost-cat post from a neighbour: like, comment, share, skip, hide, survey-yes
0.06, 0.03, 0.02, 0.35, 0.002, 0.55the neighbour got 1 comment this week, so the author bonus applies
Meme from a large page: like, comment, share, skip, hide, survey-yes
0.20, 0.004, 0.03, 0.25, 0.012, 0.30
Maya's affinity: like, comment, share
1.0, 1.6, 0.25her segment's exchange rate ÷ the global one
Weights: like, comment, share, linger, hide, survey
1, 7, 3, 0.1, −10, 0.1
Working
  1. Lost cat, base terms0.06 + 7×1.6×0.03 + 3×0.25×0.02 + 0.1×0.65 − 10×0.002 + 0.1×0.55 = 0.060 + 0.336 + 0.015 + 0.065 − 0.020 + 0.0550.511from Lost-cat post from a neighbour: like, comment, share, skip, hide, survey-yes, Maya's affinity: like, comment, share and Weights: like, comment, share, linger, hide, survey
  2. Lost cat, with the author bonus (m = 1)0.511 + 0.5 × 0.336≈ 0.68from Lost cat, base terms
  3. Meme (no bonus, m = 1)0.20 + 7×1.6×0.004 + 3×0.25×0.03 + 0.1×0.75 − 10×0.012 + 0.1×0.30 = 0.200 + 0.045 + 0.023 + 0.075 − 0.120 + 0.030≈ 0.25from Meme from a large page: like, comment, share, skip, hide, survey-yes, Maya's affinity: like, comment, share and Weights: like, comment, share, linger, hide, survey
  4. Share of the lost cat's positive score from the comment and bonus terms(0.336 + 0.168) ÷ (0.699)≈ 72%from Lost cat, base terms and Lost cat, with the author bonus (m = 1)
What it means
  • The meme is more than three times as likely to be liked, yet the lost-cat post scores about 2.7 times higher, almost entirely because of the comment it is likely to draw and who that comment helps.
  • The meme's hide term (−0.12) cancels more than half its like term. Hides are the reason viral posts do not simply win.
  • Scores like these, about 0.1 to 1 like-equivalents per impression, are the normal range of V at Stoop.
01
How Stoop sets the weights
Chosen:Exchange rates from activity-controlled logs, confirmed by 4–8 week weight experiments, with a 6-month holdout
  • Pro:Every weight has a meaning: a comment is worth 7 likes of return
  • Pro:Experiments check the biased observational estimates
  • Pro:The holdout catches slow drift that no single test sees
Downside we accept:
  • Con:Slow
  • Con:Each experiment tests one or two weights
  • Con:The estimates carry wide error bars
Ruled out:Equal contribution, then tune with 2-week A/B tests

Prices rarity, not return; 2-week tests reward whatever raises activity now, such as shares of outrage

Ruled out:Learn the whole value function on a 28-day-return label

The 28-day label arrives a month late; Hides the trade-offs from the product team

Was this section helpful?

Evaluation.

A weight change is judged on months, not days, because short-term and long-term effects can point in opposite directions.

Adding the survey term: weekly active viewers vs control (Stoop, illustrative)

Adding the survey term: weekly active viewers vs control (Stoop, illustrative)The survey term looks like a 0.6% loss in week one and a 0.4% gain by week twelve; a two-week test would have killed it.2-week test-0.8%-0.6%-0.4%-0.2%00.2%0.4%0.6%024681012ControlSurvey term armWeekly active viewers vs control (%)Weeks since launchAdding the survey term: weekly active viewers vs control (Stoop, illustrative)The survey term looks like a 0.6% loss in week one and a 0.4% gain by week twelve; a two-week test would have killed it.2-wee…-0.8%-0.6%-0.4%-0.2%00.2%0.4%0.6%024681012ControlSurvey term …Weekly active viewers vs control (%)Weeks since launch
Early on, the survey term moves up posts that draw fewer quick reactions, so viewers get fewer reply notifications and open the app slightly less. After YouTube demoted videos it classified as 'trashy' or tabloid-style, watch time was 0.5% lower at three weeks, then recovered and passed the holdout by three months (Goodrow, quoted in Cunningham et al. 2024).
Data
Weeks since launchSurvey term arm (%)
1-0.6
2-0.4
4-0.1
60.1
80.25
120.4
  • Control: Weekly active viewers vs control (%) = 0
  • 2-week test: Weeks since launch from 0 to 2
Replay agreement
Offline, replay. Logged candidates re-scored: how often survey-yes posts reach the top 25.
28-day return
Online, main. Viewers and authors, read after 8+ weeks.
Meaningful interactions
Online, secondary. Comments + shares per viewer.
Hides per 1,000 impressions
Online, guardrail. Read with integrity prevalence.

Replay before you experiment

The engagement model's probabilities are logged with every impression (see engagement-model), so a new weight setting can be tried offline first: re-score the 600 logged candidates of a sample of requests, see how the top 25 changes, and compare the new top 25 against the survey answers and hides logged for those posts. Replay catches cheap mistakes, such as one term swamping the others or a sign flipped in config. It cannot measure return, because the viewer never saw the new order.

The experiment then decides. Zinkevich's Rule #39 is the reminder that a launch metric only stands in for the long-term goal, so a weight arm is read on 28-day return after at least 8 weeks, with the 6-month holdout as the final check. How the arms are compared against guardrails is in success-metrics/guardrails and a-b-testing/experiment-design; friends in different arms affecting each other is in a-b-testing/pitfalls.

Was this section helpful?

Serving and monitoring.

The weights live in config; what to watch is how much of the score each term ends up carrying.

StepWhenWhat happens
Read weightsPer requestGlobal weights for this surface and experiment arm from config.
PersonalisePer requestMultiply like, comment and share weights by the affinities of the viewer's segment, refreshed with the exchange rates.
ScorePer requestV for 600 candidates, times each one's integrity multiplier.
LogPer impressionEach term's contribution, not only V, so term shares can be tracked and replayed.
Re-estimateQuarterlyRefresh the exchange rates; test any change in a weight experiment.

Share of the average score carried by each positive term (Stoop, one week, illustrative)

  • Like
  • Comment
  • Share
  • Linger
  • Worth your time
Share of the average score carried by each positive term (Stoop, one week, illustrative)On Saturday the comment term jumps from about 28% to 44% of the score overnight, which no weight change explains.020%40%60%80%100%MonTueWedThuFriSatSunShare of positive score (%)DayShare of the average score carried by each positive term (Stoop, one week, illustrative)On Saturday the comment term jumps from about 28% to 44% of the score overnight, which no weight change explains.020%40%60%80%100%MonTueWedThuFriSatSunShare of positive score (%)Day
The comment head was retrained on Friday and over-predicts by about 2×; the weights did not change, their inputs did. Only positive terms are stacked; the hide term takes back about 7.5% of the positive total on a normal day. Alert when any term's share moves more than 5 points in a day, and check the engagement model's calibration monitor.
Data
DayLike (%)Comment (%)Share (%)Linger (%)Worth your time (%)
Mon252862516
Tue242962516
Wed252862516
Thu252862615
Fri252862516
Sat204451912
Sun204451912

Where the weekday shares come from

Assumptions
Average impression: p_like, p_comment, p_share, 1 − p_skip, p_survey_yes, p_hide
0.05, 0.008, 0.004, 0.50, 0.32, 0.0015illustrative; base rates from the engagement model (skip 50%)
Working
  1. Weighted terms: like, comment, share, linger, survey1×0.05, 7×0.008, 3×0.004, 0.1×0.50, 0.1×0.320.050, 0.056, 0.012, 0.050, 0.032from Average impression: p_like, p_comment, p_share, 1 − p_skip, p_survey_yes, p_hide
  2. Positive total0.050 + 0.056 + 0.012 + 0.050 + 0.0320.200from Weighted terms: like, comment, share, linger, survey
  3. Shares on a normal dayeach term ÷ 0.20025%, 28%, 6%, 25%, 16%from Weighted terms: like, comment, share, linger, survey and Positive total
  4. Hide term against the positive total10 × 0.0015 ÷ 0.2007.5%from Average impression: p_like, p_comment, p_share, 1 − p_skip, p_survey_yes, p_hide and Positive total
  5. Comment share if the head doubles its predictions0.112 ÷ (0.200 + 0.056)≈ 44%from Weighted terms: like, comment, share, linger, survey and Positive total

What to watch for

FailureImpactDetectionMitigationMeanwhile
One term takes over4Value scorerThe feed tilts toward one action's postsA term's share of the score jumps overnightAlert on term shares; freeze weight changes until the head's calibration is back (engagement-model).Hides and integrity still push bad posts down.
Comment baitAuthors learn that 'Comment YES if…' paysComments per post rise while predicted and answered 'worth your time' fallIntegrity's engagement-bait demotion; watch the survey term against the comment term.The survey term partly offsets the bait.
Survey model weak on one content type2Survey modelThat type is over- or under-rankedIts calibration differs by post typeCheck calibration per type; ask more surveys on the weak type.Other terms still carry most of the score.
Short-term tuning drifts from the north star3Weight configReturn slowly falls while tests look greenEach change wins its 2-week test, yet the 6-month holdout shows no gainJudge weight changes on 8+ weeks; keep the holdout running.The holdout shows the size of the drift.
Feedback loop on the weightsPosts the weights favour get more impressions, so the next exchange-rate estimate sees more of themEstimated rates drift toward the current weights quarter after quarterEstimate from weight experiments and exploration traffic, not only logs (exploration-vs-exploitation/feedback-loops).Today's weights keep serving.
Was this section helpful?
Builds on this
Integrity signals
Read next