Integrity signals.
The posts closest to Stoop's rules are often the ones people react to most, so a feed that only follows engagement drifts toward the line. This topic takes the classifiers the moderation team builds and decides what the feed does with them: which scores remove a post, which shrink its reach and by how much, what happens to a post in the seconds before it has been scored, and how to prove the feed shows less harm without guessing.
Builds on Blending into one score and Multimodal classifiers.
Framing.
Integrity is the part of ranking that isn't allowed to trade against engagement.
From business goal to ML task
Stoop, the fictional neighbourhood app of this card (illustrative numbers throughout), ranks each candidate in two steps: the engagement model guesses what the viewer will do with the post, and the value model turns those guesses into one score V. Neither of them asks whether the post breaks Stoop's rules, or whether it earns its reactions by provoking them. A doctored photo of a street flood that never happened, a headline that hides its point, a post that ends "comment YES if you agree": all three tend to score well on engagement. This stage decides what the feed does about that.
| Layer | Definition |
|---|---|
| Business goal | Viewers trust that Stoop won't show them harmful posts, and people posting near the rules gain nothing by doing it. |
| ML objective | Minimise the prevalence of violating content in feed views, and the reach of borderline content, at a bounded cost to legitimate posts (false demotions). |
| ML task | Per post: turn policy and quality classifier scores (from content-moderation) into a feed state (removed, held, demoted, labelled, eligible) and a multiplier m on the value score. |
| Out of scope | Training the classifiers, writing policy and the review tooling (content-moderation); the rest of the value function (value-model). |
What the integrity stage must do
Meta describes its approach in three verbs: remove what breaks the rules, reduce the spread of what is problematic but allowed, and inform people with context. Stoop uses the same split and places each action at a different point: removal and holds act as filters before any score is computed, reduction multiplies the value score, and labels travel with the post to the screen. Neither filter nor multiplier can be bought back by extra engagement, which is the reason integrity is not just one more negative weight inside the value model.
Scale and budget
Stoop's integrity load (illustrative)
Scoring happens when a post is written; ranking only reads the result
Data.
Integrity uses two kinds of labels: ones that train classifiers, and ones that measure what the feed actually showed.
| Source | Used for | Note |
|---|---|---|
| Policy labels from reviewers | classifier training | Owned by content-moderation. Positives are rare and lean toward whatever got reported, so the classifier sees more of the obvious cases than the subtle ones. |
| User reports and hides | signals · review triggers | Tell you what offended the person reporting, which overlaps with rule-breaking but isn't the same thing. Coordinated mass reports are themselves an attack. |
| Fact-checker ratings | misinformation state | Arrive hours or days after posting, so they act on posts that are already spreading. The rating attaches to every copy of the same claim. |
| Sampled feed views labelled by reviewers | prevalence | Sampled by view, so a post counts as often as it was seen; the sampling and the interval are covered in content-moderation/policy-to-labels. Stoop also stores each sampled post's feed state at the moment it was seen, which splits prevalence by cause. |
| Integrity holdout | long-term effect of demotions | 0.5% of viewers (200,000) get the feed with the reduce multipliers switched off. Removals are never switched off for anyone. |
Where Stoop's violating views come from
- Feed impressions a day
- 4.0B160M requests × 25 posts (illustrative)
- Prevalence this week
- 12 per 10,000 views64,000 labelled views; 95% interval 9 to 15
- Violating views per 10,000, by the post's state when seen
- never caught 5 · before a late catch 4.5 · held, seen by friends 1.5 · unscored 0.5 · after removal 0.5illustrative; pooled over four weeks so each part is stable
- Violating views a day4.0B × 12 ÷ 10,0004.8Mfrom Feed impressions a day and Prevalence this week
- From posts no classifier caught5 ÷ 12≈ 42%, about 2.0M a dayfrom Violating views per 10,000, by the post's state when seen and Prevalence this week
- Seen before a late catch4.5 ÷ 12≈ 38%, about 1.8M a dayfrom Violating views per 10,000, by the post's state when seen and Prevalence this week
- From posts the system did catch, too late or too widely(4.5 + 1.5 + 0.5 + 0.5) ÷ 12 = 7 ÷ 12≈ 58%from Violating views per 10,000, by the post's state when seen and Prevalence this week
- The 42% that no classifier caught is a recall problem, owned by whoever builds the classifiers (content-moderation). The other 58% are posts Stoop did catch, and that part is this topic's to fix: how soon the sweeper notices a post spreading, how far a held or unscored post can reach, and how fast a removal reaches cached feeds.
- Late catches are the biggest single lever in Stoop's own hands. A post that changes character as it spreads collects most of its views in the hours before the catch, so catching it sooner shrinks prevalence without touching any classifier.
Features.
Some signals describe the post, some the author or the link, and some the way the post is spreading.
| Signal | Level | Action | Note |
|---|---|---|---|
| P(violation), per policy (hate, nudity, violence, spam) | post | remove · hold | From the content-moderation classifiers; one score per policy, each with its own thresholds. |
| Borderline closeness | post | reduce, m = 1 − 0.7p | Trained on posts reviewers marked "allowed, but close to the line"; it measures distance to a rule, not a rule broken. |
| Clickbait | post · link | reduce ×0.7 | A headline that withholds its point or oversells it, judged from the headline and the landing page. |
| Engagement bait | post | reduce ×0.3 | "Like if…", "Tag a friend who…". Facebook trained its detector on hundreds of thousands of reviewed posts and exempts genuine requests for help (Meta 2017). |
| Domain quality | domain | reduce ×0.8 | Share of a domain's traffic that comes from Stoop compared with its standing on the wider web, the idea behind Facebook's Click-Gap (Meta 2019). |
| Repeat-offender strikes | author · group | reduce ×0.5 for 30 days | Meta reduces distribution for groups that keep sharing content fact-checkers rated false (Meta 2019). Stoop applies it to every post by the author during the strike. |
| Fact-check rating | post | reduce ×0.2 + label | The inform action: the label is shown with the post and on every share of it. |
| Spread anomaly | post | re-score · review | Shares per impression far above the author's usual rate. Catches posts that were scored at low reach and changed character as they spread. |
Two things set these apart from the features in the engagement model. First, most are computed once per post, author or domain, not per viewer, so the store holds one row per post and ranking reads it without any viewer context. Second, people push back against them. Bait moves into images, misspellings dodge the text model, and a domain buys traffic from elsewhere to improve its ratio. A signal that stops moving the numbers after a month may have been learned by the people it targets, so each one gets its own recall check on fresh reviewed samples (see adversaries).
Model.
Three actions, applied at different points: remove before ranking, reduce inside the score, inform on the screen.
A post's eligibility for the feed
States of3Integrity scores
Only Removed and Held change where a post can go; the other states change how far it goes. p is P(violation); the sweeper re-scores a post before it queues it.
| From → To | Event | Guard | Action | Actor |
|---|---|---|---|---|
| Unscored → Eligible | scored | all scores low | Integrity scorer | |
| Unscored → Demoted (m < 1) | scored | borderline, bait or strikes | Integrity scorer | |
| Unscored → Held for review | scored | 0.70 ≤ p < 0.97 | queue | Integrity scorer |
| Unscored → Removed | scored | p ≥ 0.97 | Integrity scorer | |
| Held for review → Removed | reviewer: violates | reviewer | ||
| Held for review → Eligible | reviewer: fine | reviewer | ||
| Eligible → Held for review | spread spike or reports | queue | sweeper | |
| Demoted (m < 1) → Labelled + demoted | fact-check: false | attach label | ||
| Removed → Eligible | appeal upheld | reviewer | ||
| Eligible → Demoted (m < 1) | author gets a strike | ×0.5, 30 days | ||
| Labelled + demoted → Eligible | rating corrected | remove label | fact-checker | |
| Demoted (m < 1) → Held for review | spread spike or reports | queue | sweeper |
- Unscoredstart
- friends only
- Eligible
- m = 1
- Demoted (m < 1)
- shown to anyone, ranked lower
- Held for review
- friends only, no out-of-network
- Labelled + demoted
- m includes ×0.2
- Removed
- only an upheld appeal brings it back
Ranking weight as posts approach the policy line, before and after demotion (Stoop, illustrative)
- Undemoted
- Demoted
Data
| Borderline score (closeness to the policy line) | Undemoted | Demoted |
|---|---|---|
| 0 | 1 | 1 |
| 0.2 | 1.05 | 0.9 |
| 0.4 | 1.15 | 0.83 |
| 0.6 | 1.3 | 0.75 |
| 0.8 | 1.5 | 0.66 |
| 0.95 | 1.7 | 0.57 |
- policy line: Borderline score (closeness to the policy line) = 1
A borderline post against a plain one
- Plain post's value score V
- 0.60like-equivalents (value-model)
- Borderline post's value score V
- 0.90more engaging
- p_borderline of that post
- 0.8
- Demotion rule
- m = 1 − 0.7 × p
- Multiplier1 − 0.7 × 0.80.44from p_borderline of that post and Demotion rule
- Borderline post after demotion0.90 × 0.440.396from Borderline post's value score V and Multiplier
- Plain post0.60 × 10.60, so the plain post now ranks abovefrom Plain post's value score V
- V the borderline post needs to win0.60 ÷ 0.44> 1.36from Plain post and Multiplier
- Same post, also bait, by an author with a strike0.44 × 0.3 × 0.50.066 → floored at 0.1from Multiplier
- The amount a multiplier removes, V × (1 − m), grows with the post's engagement, so it keeps working however engaging the post is. A fixed penalty subtracted from V would be bought back by any post engaging enough.
- Stacked factors shrink fast. Stoop floors the combined m at 0.1 and logs every factor; a post that deserves less than that should be held or removed, not demoted further.
- Pro:Removal never depends on how engaging a post is
- Pro:Demotion strength is one tunable number per signal
- Pro:Each action is logged and auditable, which appeals need
- Con:Multipliers stack (0.44 × 0.5 × 0.3…), so the combined m needs a floor
- Con:Thresholds per policy must be re-tuned whenever a classifier is retrained
A post engaging enough outweighs the penalty; The penalty's real size changes whenever any other weight changes
Removes legitimate posts that land just past a threshold; No gradation for borderline posts, which by definition don't break the rules
- Pro:Review load is about 133 reviewers at 40 s a post, each capped at 5 h of review a day (content-moderation/review-queues)
- Pro:Held posts still reach friends, so a false hold costs the author little
- Con:Violating posts scored between 0.5 and 0.7 spread until the sweeper catches them
180,000 reviews a day, three times the staff; Queue delay grows, so held posts wait longer for a decision
Evaluation.
The feed's integrity metric is about views, not posts: how much harm people actually saw.
Demotions cost engagement in the short run, which makes them look like the first thing to cut. Facebook ran a "minimal integrity holdout" in which some users got a feed with fewer quality terms; after one month they showed slightly more activity (about +0.4% impressions), and after two years less (Cunningham et al. 2024, section 4). The short test and the long one gave opposite answers. Stoop therefore keeps 0.5% of viewers on a feed without the reduce multipliers, never without removals, reads it every quarter, and judges any change to a multiplier against it rather than against a two-week A/B test. How a guardrail should block a launch is covered in guardrails.
False demotions need a slice view. A classifier trained mostly on one variety of a language tends to over-flag others, so a feed-wide false-demotion rate of 3% can hide 12% for one community's posts. Stoop samples demoted posts per language and per page category every week (see fairness).
Serving and monitoring.
Every post is scored when it is written and re-scored as it spreads; ranking only reads.
One post from creation to removal
- Post service → Integrity scorer: post created
- Note over Ranking stack: unscored: friends only
- Integrity scorer → Integrity scores: bait 0.91, weak domain → demoted
- Ranking stack → Integrity scores: read 600 states
- Integrity scores → Ranking stack (reply): state demoted, m = 0.24
- Note over Integrity scorer: 1 h later: shares 8× normal
- Integrity scorer → Review queue: demoted → held, queue
- Review queue → Integrity scores (reply): violates: removed
- Integrity scores → Ranking stack: removed, drop from caches
| Step | When | What happens |
|---|---|---|
| First score | on write, p99 < 3 s | Text and image classifiers run on every new post, about 140 a second. Video gets a first pass on its frames and a fuller one later. |
| Re-score | on spread, reports, new model | The sweeper re-scores posts whose reach or report rate jumps, and every post from the last 7 days when a classifier is replaced. |
| Read at ranking | per request | One batch read of 600 states and multipliers: 96B reads a day, about 3.3M a second at peak. A missing row counts as unscored. |
| Removal push | within a minute | A removal is pushed to the session caches so a post already sitting in someone's next page is dropped. |
What to watch for
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| Scoring lag during a spike (breaking news)2Integrity scorer | New posts wait longer to reach beyond friends | Share of candidates still unscored rises above 1% | Unscored posts stay friends-only; add capacity; alert on the unscored share. | Friends still see new posts at once. |
| Adversaries adapt | Prevalence climbs for one policy | Text moves into images, misspellings spread; report-to-removal rate falls for one policy | Retrain on fresh reviewed samples; extract text from images before scoring. | User reports and the sweeper still catch spreading posts. |
| Stale removals in cached feeds3Integrity scores | Removed posts are still seen | Removed posts keep getting views for minutes | Push removals to the session cache; count post-removal views in prevalence. | New feed pages already leave the post out. |
| Over-demotion of one community | Legitimate posts in one language reach fewer people | False demotions concentrated in one language or dialect | Per-slice thresholds or retraining; a working appeal path (content-moderation/review-queues). | Posts still reach friends; only reach is cut. |
| Multiplier creep5Ranking stack | Ordinary posts get demoted too | Median m across candidates drifts down as new signals are added | Track each factor's share of total demotion; floor the combined m; retire signals the holdout shows do nothing. | The 0.1 floor keeps any post rankable. |