Blending into one score.
Stoop's engagement model says a neighbour's post about a lost cat will get a comment from Maya with 3% probability, and a viral meme will get a like with 20%. Which is worth more depends on what Stoop believes each action does for Maya, for the neighbour who posted, and for whether both come back next month. This topic turns those beliefs into weights, shows how to estimate them without fooling yourself, and explains why a change that looks bad in week one can be right by week twelve.
Builds on Predicting engagement and Blending objectives.
Framing.
The engagement model predicts; the value model decides what those predictions are worth to Stoop.
From business goal to ML task
By the time Stoop's value model runs, the hard prediction work is done. For each of about 600 candidate posts, the engagement model has said how likely the viewer is to like it, comment, share, scroll past, or hide it. The feed still needs a single number per post to sort by. The value model supplies that number, and in doing so it states what Stoop wants the feed to be for.
That makes it a small piece of code with a large job. If the weights say a like is worth as much as a comment, the feed fills with things that are easy to tap on. If they ignore the person who wrote the post, neighbours who post into silence stop posting, and next year there is less to rank. The weights are the one place where the product's idea of a good feed is written down as numbers.
| Layer | Definition |
|---|---|
| Business goal | Viewers and the people who post both keep coming back: 28-day return for viewers and for authors. |
| ML objective | An ordering of candidates that maximises long-term return, within the hide and integrity guardrails. |
| ML task | A scoring function over calibrated predictions: a fixed weighted sum whose weights are estimated from experiments and long-term data, plus two small learned inputs (the survey answer and each viewer's action affinity). |
| Out of scope | Blend forms in general and weight sweeps (multi-task-and-multi-objective/blending); variety rules after scoring (news-feed/ranking); what integrity demotes and why (integrity). |
What the value model must do
Scale and budget
Where the value score comes from
Every input to the scorer except the weights is predicted or measured per candidate. The weights are the only product decision in the picture, which is why they sit in config, carry an owner, and change through experiments rather than code pushes.
Data.
The engagement log says what people did. Setting the weights needs data about what those actions led to weeks later.
Where the evidence for a weight comes from
| Source | Tells you | Catch |
|---|---|---|
| Weight experiments: arms that raise or lower one weight, run 4–8 weeks | causal effect of a weight on return | Only a handful of weight settings can ever be tried; Cunningham et al. (2024) note experiments cover a small corner of the possible weights. |
| Long-term holdout: 1% of viewers kept on old weights for 6 months | whether many changes added up | Small and slow, but the only guard against a string of short-term wins that sum to a long-term loss. |
| Observational logs: actions this week against return next month, per viewer | fine-grained exchange rates | Active viewers do everything more. Without controlling for prior activity, every action looks valuable. |
| Item surveys: 'Was this post worth your time?' (yes, somewhat, no) | what engagement misses | Sensitive to wording, and only about 20% of those asked answer, not at random. |
| Author logs: feedback an author received against whether they post again | the author's side of every action | The viewer's logs never show it, so it has to be joined in by author ID. |
Survey labels a day
- Impressions a day
- 4.0B40M viewers × 4 sessions × 25 posts (illustrative)
- Survey shown on
- 1 in 1,000 impressions
- Share of prompts answered
- 20%
- Video share of impressions
- 15%
- New posts a day
- 12M40M × 0.3 (illustrative)
- Prompts shown4.0B ÷ 1,0004.0M a dayfrom Impressions a day and Survey shown on
- Answers collected4.0M × 0.20800k labels a dayfrom Prompts shown and Share of prompts answered
- Answers about video posts800k × 0.15120k a dayfrom Answers collected and Video share of impressions
- Answers per new post, on average800k ÷ 12Mabout 0.07from Answers collected and New posts a day
- 800k answers a day is plenty to train a survey model on the same features as the engagement heads.
- Almost no single post ever gets an answer, so the survey model must predict the answer for every candidate; the raw answers are training labels only.
- Answers differ by content type, and Cunningham et al. (2024) report survey models that did well on some kinds of content and poorly on others. Check calibration per type, and ask more often on the weak types.
Features.
The value model's inputs are the heads' outputs, plus two things the heads don't know: what the viewer tends to do, and who wrote the post.
The terms in Stoop's score
| Term | Input | Sign and weight |
|---|---|---|
| Like | p_like | +1. The unit every other weight is priced in. |
| Comment | p_comment | +7. High effort, and it lands on the author. |
| Share | p_share | +3. Spreads the post to a friend or a group. |
| Linger | 1 − p_skip | +0.1. Catches posts people read but don't react to. |
| Hide | p_hide | −10. The strongest negative signal a viewer gives. |
| Worth your time | p_survey_yes | +0.1 per unit of probability, from the survey model. |
| Author bonus | 0.5 × comment term | + when the author received fewer than 2 comments in the last 7 days, so the comment term counts 1.5×. |
| Integrity multiplier | m in (0, 1] | Multiplies the whole sum. Set by the integrity topic. |
Scaling the weights to each viewer
Meta's 2021 post on Facebook's feed points out a property of the weighted sum itself: an action a person rarely takes already plays a minimal role in their ranking, because its predicted probability is close to 0. The engagement model does that personalising. Maya has shared twice in three months, so her p_share is already small; no extra weight is needed to quiet it.
Stoop adds one per-viewer factor for a different reason: the same action can be worth more or less return depending on who takes it (Meta's post also notes that some people express themselves more through likes than comments). Stoop runs the exchange-rate estimate separately for viewer activity segments. The affinity aᵥ,ₖ is the segment's rate for action k divided by the global rate, clipped to 0.25 to 2. In Maya's segment, frequent commenters who rarely share, a comment is worth 1.6 times the global rate because it usually opens a thread with neighbours, and a share 0.25 times because it is mostly a forwarded link that brings nobody back.
The trap is setting affinity from how often a viewer takes the action. Maya's p_comment of 0.03, against a 0.8% average, already says she comments a lot; multiplying it by her comment rate as well would count the same habit twice and tilt her feed toward posts that beg for replies. So affinity comes from worth, not frequency, and the segment estimates control for prior activity just as the global ones do. It applies only to like, comment and share: hides are too rare for stable per-segment estimates, so the hide keeps its global −10. New viewers get 1 on every action until they have 30 days of history and a segment.
- Pro:Targets the authors most likely to stop posting
- Pro:One number (1.5×) that is easy to test in an arm
- Pro:Costs nothing for authors who already get plenty of replies
- Pro:Has a public precedent: X's 2023 heavy-ranker weights put a reply the author engages with far above a plain reply (the weights table is in multi-task-and-multi-objective/blending)
- Con:A threshold (2 comments in 7 days) that authors could in theory game
- Con:Ignores likes and shares, which also reach the author
Gives the same boost to a page with 10,000 followers and to a neighbour posting for the first time
The ranking is where attention is handed out; a fix elsewhere arrives after authors have already left
Model.
The value model is a formula. The hard part is the weights, and what they are meant to be: exchange rates to long-term return.
Stoop's value function
# Weights are read from config per surface and experiment arm; this dict is illustrative.
W = {"like": 1.0, "comment": 7.0, "share": 3.0, "linger": 0.1, "hide": -10.0, "survey": 0.1}
def value(p, affinity, author_starved, m):
v = sum(W[k] * affinity.get(k, 1.0) * p[k] for k in ("like", "comment", "share"))
v += W["linger"] * (1 - p["skip"])
v += W["hide"] * p["hide"] # never scaled by affinity
v += W["survey"] * p["survey_yes"]
if author_starved: # author got < 2 comments in 7 days
v += 0.5 * W["comment"] * affinity.get("comment", 1.0) * p["comment"]
return m * v # integrity multiplier, 0 < m <= 1What a comment is worth at Stoop
- Extra 28-day return per like: viewer, author
- +0.0015 pp, +0.0005 ppillustrative; activity-controlled estimates checked by weight experiments
- Per comment: viewer, author
- +0.010 pp, +0.004 pp
- Per share: viewer, author
- +0.004 pp, +0.002 pp
- Per hide: viewer
- −0.020 pp
- A like, the unit (weight 1)0.0015 + 0.00050.002 ppfrom Extra 28-day return per like: viewer, author
- Comment weight(0.010 + 0.004) ÷ 0.0027from Per comment: viewer, author and A like, the unit (weight 1)
- Share weight(0.004 + 0.002) ÷ 0.0023from Per share: viewer, author and A like, the unit (weight 1)
- Hide weight−0.020 ÷ 0.002−10from Per hide: viewer and A like, the unit (weight 1)
- Share of a comment's value that goes to the author0.004 ÷ 0.014≈ 29%from Per comment: viewer, author
- Comment weight if only the viewer counted0.010 ÷ 0.0015≈ 6.7from Per comment: viewer, author and Extra 28-day return per like: viewer, author
- Counting the author moves a comment from about 6.7 likes to 7. The gap looks small, but it is the part that keeps quiet neighbours posting, and it is larger still for authors who get little feedback (the 1.5× bonus).
- Linger and the survey term have no clean per-action return estimate, so their weights (0.1 each) were set by a weight experiment with arms at 0.05, 0.1 and 0.2, read on 28-day return. The same estimate run per viewer segment gives the affinities.
- Each of these numbers has wide error bars. Cunningham et al. (2024) describe retention-maximising weights as hard to pin down because experiments are few and logs are biased, so treat them as starting estimates to check with a weight experiment.
Equal contribution vs exchange rates (weights relative to like)
| Action | Base rate | Equal contribution | Exchange rate | Reading |
|---|---|---|---|---|
| Like | 5% | 1 | 1 | — |
| Comment | 0.8% | 6.25 | 7 | Close: both value comments highly. |
| Share | 0.4% | 12.5 | 3 | Equal contribution overprices shares because they are rare, not because they bring people back. |
| Hide | 0.15% | −33 | −10 | Equal contribution punishes a hide more than three times as hard. |
Equal contribution (taught in multi-task-and-multi-objective/blending) prices rarity; an exchange rate prices return. The table shows where that matters at Stoop: shares are rare, so equal contribution gives them 12.5 when the return data say 3. Cunningham et al. (2024) report that active, high-effort engagement such as commenting tends to be worth more for retention than passive signals, which fits comments keeping their high weight. They also cite YouTube's 2012 switch from optimising clicks to clicks plus watch time: clicks dropped in the short term while long-term retention rose.
Stoop is not alone in writing value this way. Instagram's Explore (Meta AI, 2019) described its final score as weighted predictions of likes and saves minus a weighted prediction of negative actions, the same shape with a different list of terms.
Scoring two of Maya's candidates, term by term
- Lost-cat post from a neighbour: like, comment, share, skip, hide, survey-yes
- 0.06, 0.03, 0.02, 0.35, 0.002, 0.55the neighbour got 1 comment this week, so the author bonus applies
- Meme from a large page: like, comment, share, skip, hide, survey-yes
- 0.20, 0.004, 0.03, 0.25, 0.012, 0.30
- Maya's affinity: like, comment, share
- 1.0, 1.6, 0.25her segment's exchange rate ÷ the global one
- Weights: like, comment, share, linger, hide, survey
- 1, 7, 3, 0.1, −10, 0.1
- Lost cat, base terms0.06 + 7×1.6×0.03 + 3×0.25×0.02 + 0.1×0.65 − 10×0.002 + 0.1×0.55 = 0.060 + 0.336 + 0.015 + 0.065 − 0.020 + 0.0550.511from Lost-cat post from a neighbour: like, comment, share, skip, hide, survey-yes, Maya's affinity: like, comment, share and Weights: like, comment, share, linger, hide, survey
- Lost cat, with the author bonus (m = 1)0.511 + 0.5 × 0.336≈ 0.68from Lost cat, base terms
- Meme (no bonus, m = 1)0.20 + 7×1.6×0.004 + 3×0.25×0.03 + 0.1×0.75 − 10×0.012 + 0.1×0.30 = 0.200 + 0.045 + 0.023 + 0.075 − 0.120 + 0.030≈ 0.25from Meme from a large page: like, comment, share, skip, hide, survey-yes, Maya's affinity: like, comment, share and Weights: like, comment, share, linger, hide, survey
- Share of the lost cat's positive score from the comment and bonus terms(0.336 + 0.168) ÷ (0.699)≈ 72%from Lost cat, base terms and Lost cat, with the author bonus (m = 1)
- The meme is more than three times as likely to be liked, yet the lost-cat post scores about 2.7 times higher, almost entirely because of the comment it is likely to draw and who that comment helps.
- The meme's hide term (−0.12) cancels more than half its like term. Hides are the reason viral posts do not simply win.
- Scores like these, about 0.1 to 1 like-equivalents per impression, are the normal range of V at Stoop.
- Pro:Every weight has a meaning: a comment is worth 7 likes of return
- Pro:Experiments check the biased observational estimates
- Pro:The holdout catches slow drift that no single test sees
- Con:Slow
- Con:Each experiment tests one or two weights
- Con:The estimates carry wide error bars
Prices rarity, not return; 2-week tests reward whatever raises activity now, such as shares of outrage
The 28-day label arrives a month late; Hides the trade-offs from the product team
Evaluation.
A weight change is judged on months, not days, because short-term and long-term effects can point in opposite directions.
Adding the survey term: weekly active viewers vs control (Stoop, illustrative)
Data
| Weeks since launch | Survey term arm (%) |
|---|---|
| 1 | -0.6 |
| 2 | -0.4 |
| 4 | -0.1 |
| 6 | 0.1 |
| 8 | 0.25 |
| 12 | 0.4 |
- Control: Weekly active viewers vs control (%) = 0
- 2-week test: Weeks since launch from 0 to 2
Replay before you experiment
The engagement model's probabilities are logged with every impression (see engagement-model), so a new weight setting can be tried offline first: re-score the 600 logged candidates of a sample of requests, see how the top 25 changes, and compare the new top 25 against the survey answers and hides logged for those posts. Replay catches cheap mistakes, such as one term swamping the others or a sign flipped in config. It cannot measure return, because the viewer never saw the new order.
The experiment then decides. Zinkevich's Rule #39 is the reminder that a launch metric only stands in for the long-term goal, so a weight arm is read on 28-day return after at least 8 weeks, with the 6-month holdout as the final check. How the arms are compared against guardrails is in success-metrics/guardrails and a-b-testing/experiment-design; friends in different arms affecting each other is in a-b-testing/pitfalls.
Serving and monitoring.
The weights live in config; what to watch is how much of the score each term ends up carrying.
| Step | When | What happens |
|---|---|---|
| Read weights | Per request | Global weights for this surface and experiment arm from config. |
| Personalise | Per request | Multiply like, comment and share weights by the affinities of the viewer's segment, refreshed with the exchange rates. |
| Score | Per request | V for 600 candidates, times each one's integrity multiplier. |
| Log | Per impression | Each term's contribution, not only V, so term shares can be tracked and replayed. |
| Re-estimate | Quarterly | Refresh the exchange rates; test any change in a weight experiment. |
Share of the average score carried by each positive term (Stoop, one week, illustrative)
- Like
- Comment
- Share
- Linger
- Worth your time
Data
| Day | Like (%) | Comment (%) | Share (%) | Linger (%) | Worth your time (%) |
|---|---|---|---|---|---|
| Mon | 25 | 28 | 6 | 25 | 16 |
| Tue | 24 | 29 | 6 | 25 | 16 |
| Wed | 25 | 28 | 6 | 25 | 16 |
| Thu | 25 | 28 | 6 | 26 | 15 |
| Fri | 25 | 28 | 6 | 25 | 16 |
| Sat | 20 | 44 | 5 | 19 | 12 |
| Sun | 20 | 44 | 5 | 19 | 12 |
Where the weekday shares come from
- Average impression: p_like, p_comment, p_share, 1 − p_skip, p_survey_yes, p_hide
- 0.05, 0.008, 0.004, 0.50, 0.32, 0.0015illustrative; base rates from the engagement model (skip 50%)
- Weighted terms: like, comment, share, linger, survey1×0.05, 7×0.008, 3×0.004, 0.1×0.50, 0.1×0.320.050, 0.056, 0.012, 0.050, 0.032from Average impression: p_like, p_comment, p_share, 1 − p_skip, p_survey_yes, p_hide
- Positive total0.050 + 0.056 + 0.012 + 0.050 + 0.0320.200from Weighted terms: like, comment, share, linger, survey
- Shares on a normal dayeach term ÷ 0.20025%, 28%, 6%, 25%, 16%from Weighted terms: like, comment, share, linger, survey and Positive total
- Hide term against the positive total10 × 0.0015 ÷ 0.2007.5%from Average impression: p_like, p_comment, p_share, 1 − p_skip, p_survey_yes, p_hide and Positive total
- Comment share if the head doubles its predictions0.112 ÷ (0.200 + 0.056)≈ 44%from Weighted terms: like, comment, share, linger, survey and Positive total
What to watch for
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| One term takes over4Value scorer | The feed tilts toward one action's posts | A term's share of the score jumps overnight | Alert on term shares; freeze weight changes until the head's calibration is back (engagement-model). | Hides and integrity still push bad posts down. |
| Comment bait | Authors learn that 'Comment YES if…' pays | Comments per post rise while predicted and answered 'worth your time' fall | Integrity's engagement-bait demotion; watch the survey term against the comment term. | The survey term partly offsets the bait. |
| Survey model weak on one content type2Survey model | That type is over- or under-ranked | Its calibration differs by post type | Check calibration per type; ask more surveys on the weak type. | Other terms still carry most of the score. |
| Short-term tuning drifts from the north star3Weight config | Return slowly falls while tests look green | Each change wins its 2-week test, yet the 6-month holdout shows no gain | Judge weight changes on 8+ weeks; keep the holdout running. | The holdout shows the size of the drift. |
| Feedback loop on the weights | Posts the weights favour get more impressions, so the next exchange-rate estimate sees more of them | Estimated rates drift toward the current weights quarter after quarter | Estimate from weight experiments and exploration traffic, not only logs (exploration-vs-exploitation/feedback-loops). | Today's weights keep serving. |