Social feed rankingIntegrity signals

100%

Integrity signals.

The posts closest to Stoop's rules are often the ones people react to most, so a feed that only follows engagement drifts toward the line. This topic takes the classifiers the moderation team builds and decides what the feed does with them: which scores remove a post, which shrink its reach and by how much, what happens to a post in the seconds before it has been scored, and how to prove the feed shows less harm without guessing.

Advanced19 minUpdated 1 Oct 2026

Builds on Blending into one score and Multimodal classifiers.

Framing.

Integrity is the part of ranking that isn't allowed to trade against engagement.

From business goal to ML task

Stoop, the fictional neighbourhood app of this card (illustrative numbers throughout), ranks each candidate in two steps: the engagement model guesses what the viewer will do with the post, and the value model turns those guesses into one score V. Neither of them asks whether the post breaks Stoop's rules, or whether it earns its reactions by provoking them. A doctored photo of a street flood that never happened, a headline that hides its point, a post that ends "comment YES if you agree": all three tend to score well on engagement. This stage decides what the feed does about that.

LayerDefinition
Business goalViewers trust that Stoop won't show them harmful posts, and people posting near the rules gain nothing by doing it.
ML objectiveMinimise the prevalence of violating content in feed views, and the reach of borderline content, at a bounded cost to legitimate posts (false demotions).
ML taskPer post: turn policy and quality classifier scores (from content-moderation) into a feed state (removed, held, demoted, labelled, eligible) and a multiplier m on the value score.
Out of scopeTraining the classifiers, writing policy and the review tooling (content-moderation); the rest of the value function (value-model).

What the integrity stage must do

#1A post scored as clearly violating never appears in any feed, whatever its predicted engagement.
#2A post the classifiers are unsure about is held out of reach beyond the author's friends until a reviewer decides.
#3Borderline, baity or low-quality posts can still be shown, but reach less, by an amount set per signal.
#4A post that has not been scored yet reaches friends only, so the feed never waits for a classifier.
#5A post is re-scored as it spreads, and a removal reaches feeds already cached within a minute.
#6Every demotion is logged with the factor that caused it, so an author's appeal can be traced.

Meta describes its approach in three verbs: remove what breaks the rules, reduce the spread of what is problematic but allowed, and inform people with context. Stoop uses the same split and places each action at a different point: removal and holds act as filters before any score is computed, reduction multiplies the value score, and labels travel with the post to the screen. Neither filter nor multiplier can be bought back by extra engagement, which is the reason integrity is not just one more negative weight inside the value model.

Scale and budget

Stoop's integrity load (illustrative)

New posts a day
12M
≈ 140/s; 40M viewers × 0.3
Time to first score
3 s
p99 at most, text and images
Prevalence target
10
at most, violating views per 10,000
Read at ranking
600
states per request, one batch read; 96B a day

Scoring happens when a post is written; ranking only reads the result

Scoring happens when a post is written; ranking only reads the result. The numbered component cards that follow describe each part.
Scoring happens when a post is written ranking only reads the resultComponents: 1. Post service (Stores a new post and emits a post-created event.), 2. Integrity scorer (Runs policy and quality classifiers on each new post, and re-scores posts as they spread.), 3. Integrity scores (Per post: classifier scores, reviewer decisions, fact-check ratings, the resulting state and multiplier.), 4. Review queue (Human reviewers (content-moderation/review-queues).), 5. Ranking stack (The integrity check, the engagement model and the value scorer reads this store for every candidate before scoring.).

On every feed request

When a post is written or spreads

post created

scores, state

uncertain

decision

state, m for 600 candidates

1Post service
text · photos · video · link

2Integrity scorer
policy classifiers
clickbait · engagement bait · borderline
domain and author quality

4Review queue
0.70 ≤ P(violation) < 0.97

5Ranking stack
drops removed; held and unscored: friends only
V × m

3Integrity scores
state + multiplier per post

Was this section helpful?

Data.

Integrity uses two kinds of labels: ones that train classifiers, and ones that measure what the feed actually showed.

SourceUsed forNote
Policy labels from reviewersclassifier trainingOwned by content-moderation. Positives are rare and lean toward whatever got reported, so the classifier sees more of the obvious cases than the subtle ones.
User reports and hidessignals · review triggersTell you what offended the person reporting, which overlaps with rule-breaking but isn't the same thing. Coordinated mass reports are themselves an attack.
Fact-checker ratingsmisinformation stateArrive hours or days after posting, so they act on posts that are already spreading. The rating attaches to every copy of the same claim.
Sampled feed views labelled by reviewersprevalenceSampled by view, so a post counts as often as it was seen; the sampling and the interval are covered in content-moderation/policy-to-labels. Stoop also stores each sampled post's feed state at the moment it was seen, which splits prevalence by cause.
Integrity holdoutlong-term effect of demotions0.5% of viewers (200,000) get the feed with the reduce multipliers switched off. Removals are never switched off for anyone.

Where Stoop's violating views come from

Assumptions
Feed impressions a day
4.0B160M requests × 25 posts (illustrative)
Prevalence this week
12 per 10,000 views64,000 labelled views; 95% interval 9 to 15
Violating views per 10,000, by the post's state when seen
never caught 5 · before a late catch 4.5 · held, seen by friends 1.5 · unscored 0.5 · after removal 0.5illustrative; pooled over four weeks so each part is stable
Working
  1. Violating views a day4.0B × 12 ÷ 10,0004.8Mfrom Feed impressions a day and Prevalence this week
  2. From posts no classifier caught5 ÷ 12≈ 42%, about 2.0M a dayfrom Violating views per 10,000, by the post's state when seen and Prevalence this week
  3. Seen before a late catch4.5 ÷ 12≈ 38%, about 1.8M a dayfrom Violating views per 10,000, by the post's state when seen and Prevalence this week
  4. From posts the system did catch, too late or too widely(4.5 + 1.5 + 0.5 + 0.5) ÷ 12 = 7 ÷ 12≈ 58%from Violating views per 10,000, by the post's state when seen and Prevalence this week
What it means
  • The 42% that no classifier caught is a recall problem, owned by whoever builds the classifiers (content-moderation). The other 58% are posts Stoop did catch, and that part is this topic's to fix: how soon the sweeper notices a post spreading, how far a held or unscored post can reach, and how fast a removal reaches cached feeds.
  • Late catches are the biggest single lever in Stoop's own hands. A post that changes character as it spreads collects most of its views in the hours before the catch, so catching it sooner shrinks prevalence without touching any classifier.
Was this section helpful?

Features.

Some signals describe the post, some the author or the link, and some the way the post is spreading.

SignalLevelActionNote
P(violation), per policy (hate, nudity, violence, spam)postremove · holdFrom the content-moderation classifiers; one score per policy, each with its own thresholds.
Borderline closenesspostreduce, m = 1 − 0.7pTrained on posts reviewers marked "allowed, but close to the line"; it measures distance to a rule, not a rule broken.
Clickbaitpost · linkreduce ×0.7A headline that withholds its point or oversells it, judged from the headline and the landing page.
Engagement baitpostreduce ×0.3"Like if…", "Tag a friend who…". Facebook trained its detector on hundreds of thousands of reviewed posts and exempts genuine requests for help (Meta 2017).
Domain qualitydomainreduce ×0.8Share of a domain's traffic that comes from Stoop compared with its standing on the wider web, the idea behind Facebook's Click-Gap (Meta 2019).
Repeat-offender strikesauthor · groupreduce ×0.5 for 30 daysMeta reduces distribution for groups that keep sharing content fact-checkers rated false (Meta 2019). Stoop applies it to every post by the author during the strike.
Fact-check ratingpostreduce ×0.2 + labelThe inform action: the label is shown with the post and on every share of it.
Spread anomalypostre-score · reviewShares per impression far above the author's usual rate. Catches posts that were scored at low reach and changed character as they spread.

Two things set these apart from the features in the engagement model. First, most are computed once per post, author or domain, not per viewer, so the store holds one row per post and ranking reads it without any viewer context. Second, people push back against them. Bait moves into images, misspellings dodge the text model, and a domain buys traffic from elsewhere to improve its ratio. A signal that stops moving the numbers after a month may have been learned by the people it targets, so each one gets its own recall check on fresh reviewed samples (see adversaries).

Was this section helpful?

Model.

Three actions, applied at different points: remove before ranking, reduce inside the score, inform on the screen.

A post's eligibility for the feed

States of3Integrity scores

A post's eligibility for the feed. 6 states, 12 transitions. The table below lists them.
A post's eligibility for the feedThe states of Integrity scores. 6 states, 12 transitions. The table below lists them.

scored [all scores low]

scored [borderline, bait or strikes]

scored [0.70 ≤ p ‹ 0.97] / queue

scored [p ≥ 0.97]

reviewer: violates

reviewer: fine

spread spike or reports / queue

fact-check: false / attach label

appeal upheld

author gets a strike / ×0.5, 30 days

rating corrected / remove label

spread spike or reports / queue

Unscored

Eligible

Demoted (m ‹ 1)

Held for review

Labelled + demoted

Removed

1 step.

Only Removed and Held change where a post can go; the other states change how far it goes. p is P(violation); the sweeper re-scores a post before it queues it.

Transitions of A post's eligibility for the feed
From → ToEventGuardActionActor
Unscored → Eligiblescoredall scores lowIntegrity scorer
Unscored → Demoted (m < 1)scoredborderline, bait or strikesIntegrity scorer
Unscored → Held for reviewscored0.70 ≤ p < 0.97queueIntegrity scorer
Unscored → Removedscoredp ≥ 0.97Integrity scorer
Held for review → Removedreviewer: violatesreviewer
Held for review → Eligiblereviewer: finereviewer
Eligible → Held for reviewspread spike or reportsqueuesweeper
Demoted (m < 1) → Labelled + demotedfact-check: falseattach label
Removed → Eligibleappeal upheldreviewer
Eligible → Demoted (m < 1)author gets a strike×0.5, 30 days
Labelled + demoted → Eligiblerating correctedremove labelfact-checker
Demoted (m < 1) → Held for reviewspread spike or reportsqueuesweeper
Unscoredstart
friends only
Eligible
m = 1
Demoted (m < 1)
shown to anyone, ranked lower
Held for review
friends only, no out-of-network
Labelled + demoted
m includes ×0.2
Removed
only an upheld appeal brings it back

Ranking weight as posts approach the policy line, before and after demotion (Stoop, illustrative)

  • Undemoted
  • Demoted
Ranking weight as posts approach the policy line, before and after demotion (Stoop, illustrative)Ranked on engagement alone, a post near the line carries up to 70% more ranking weight than a typical post; after demotion (engagement × m), its ranking weight falls the closer it gets to the line.0.40.60.811.21.41.61.8200.20.40.60.81policy lineUndemotedDemotedWeight vs a typical postBorderline score (closeness to the policy line)Ranking weight as posts approach the policy line, before and after demotion (Stoop, illustrative)Ranked on engagement alone, a post near the line carries up to 70% more ranking weight than a typical post; after demotion (engagement × m), its ranking weight falls the closer it gets to the line.0.40.60.811.21.41.61.8200.20.40.60.81policy lineUndemotedDemotedWeight vs a typical postBorderline score (closeness to the policy line)
Facebook described the same shape: the nearer a post sits to a rule, the more reactions it tends to draw, including from people who later say they disliked it. Its answer was to cut the distribution of borderline posts (Zuckerberg 2018). The curve and its numbers are Stoop's and illustrative. Both lines are ranking weight relative to a typical post. Demotion does not change how people react to a post once they see it; it changes how often the post is shown. Each dashed point is the solid one times m = 1 − 0.7 × score: at 0.8, 1.5 × 0.44 = 0.66.
Data
Borderline score (closeness to the policy line)UndemotedDemoted
011
0.21.050.9
0.41.150.83
0.61.30.75
0.81.50.66
0.951.70.57
  • policy line: Borderline score (closeness to the policy line) = 1

A borderline post against a plain one

Assumptions
Plain post's value score V
0.60like-equivalents (value-model)
Borderline post's value score V
0.90more engaging
p_borderline of that post
0.8
Demotion rule
m = 1 − 0.7 × p
Working
  1. Multiplier1 − 0.7 × 0.80.44from p_borderline of that post and Demotion rule
  2. Borderline post after demotion0.90 × 0.440.396from Borderline post's value score V and Multiplier
  3. Plain post0.60 × 10.60, so the plain post now ranks abovefrom Plain post's value score V
  4. V the borderline post needs to win0.60 ÷ 0.44> 1.36from Plain post and Multiplier
  5. Same post, also bait, by an author with a strike0.44 × 0.3 × 0.50.066 → floored at 0.1from Multiplier
What it means
  • The amount a multiplier removes, V × (1 − m), grows with the post's engagement, so it keeps working however engaging the post is. A fixed penalty subtracted from V would be bought back by any post engaging enough.
  • Stacked factors shrink fast. Stoop floors the combined m at 0.1 and logs every factor; a post that deserves less than that should be held or removed, not demoted further.
01
How integrity enters ranking
Chosen:Remove and hold as filters before ranking; reduce as a multiplier on V; inform as a label
  • Pro:Removal never depends on how engaging a post is
  • Pro:Demotion strength is one tunable number per signal
  • Pro:Each action is logged and auditable, which appeals need
Downside we accept:
  • Con:Multipliers stack (0.44 × 0.5 × 0.3…), so the combined m needs a floor
  • Con:Thresholds per policy must be re-tuned whenever a classifier is retrained
Ruled out:Integrity as one more negative term in the weighted sum

A post engaging enough outweighs the penalty; The penalty's real size changes whenever any other weight changes

Ruled out:Hard filter on every signal

Removes legitimate posts that land just past a threshold; No gradation for borderline posts, which by definition don't break the rules

02
Where the lower review threshold sits
Chosen:Hold at 0.70 (≈ 0.5% of posts, 60,000 reviews a day)
  • Pro:Review load is about 133 reviewers at 40 s a post, each capped at 5 h of review a day (content-moderation/review-queues)
  • Pro:Held posts still reach friends, so a false hold costs the author little
Downside we accept:
  • Con:Violating posts scored between 0.5 and 0.7 spread until the sweeper catches them
Ruled out:Hold at 0.50 (≈ 1.5% of posts)

180,000 reviews a day, three times the staff; Queue delay grows, so held posts wait longer for a decision

Was this section helpful?

Evaluation.

The feed's integrity metric is about views, not posts: how much harm people actually saw.

Precision at the remove threshold
Offline, main. At least 97%, audited weekly by reviewers.
Prevalence
Online, main. Violating views per 10,000, with its 95% interval.
Borderline reach
Online, secondary. Views of posts scoring ≥ 0.6, per 10,000.
False demotions
Online, guardrail. Sampled demoted posts judged fine, by language.

Demotions cost engagement in the short run, which makes them look like the first thing to cut. Facebook ran a "minimal integrity holdout" in which some users got a feed with fewer quality terms; after one month they showed slightly more activity (about +0.4% impressions), and after two years less (Cunningham et al. 2024, section 4). The short test and the long one gave opposite answers. Stoop therefore keeps 0.5% of viewers on a feed without the reduce multipliers, never without removals, reads it every quarter, and judges any change to a multiplier against it rather than against a two-week A/B test. How a guardrail should block a launch is covered in guardrails.

False demotions need a slice view. A classifier trained mostly on one variety of a language tends to over-flag others, so a feed-wide false-demotion rate of 3% can hide 12% for one community's posts. Stoop samples demoted posts per language and per page category every week (see fairness).

Was this section helpful?

Serving and monitoring.

Every post is scored when it is written and re-scored as it spreads; ranking only reads.

One post from creation to removal

One post from creation to removal, as an ordered list of steps:
One post from creation to removal9 steps between Post service, Integrity scorer, Integrity scores, Review queue, Ranking stack. The steps are listed as text after the diagram.Ranking stackReview queueIntegrity scoresIntegrity scorerPost serviceunscored: friends only1 h later: shares 8× normalpost created1bait 0.91, weak domain → demoted2read 600 states3state demoted, m = 0.244demoted → held, queue5violates: removed6removed, drop from caches7
  1. Post service → Integrity scorer: post created
  2. Note over Ranking stack: unscored: friends only
  3. Integrity scorer → Integrity scores: bait 0.91, weak domain → demoted
  4. Ranking stack → Integrity scores: read 600 states
  5. Integrity scores → Ranking stack (reply): state demoted, m = 0.24
  6. Note over Integrity scorer: 1 h later: shares 8× normal
  7. Integrity scorer → Review queue: demoted → held, queue
  8. Review queue → Integrity scores (reply): violates: removed
  9. Integrity scores → Ranking stack: removed, drop from caches
StepWhenWhat happens
First scoreon write, p99 < 3 sText and image classifiers run on every new post, about 140 a second. Video gets a first pass on its frames and a fuller one later.
Re-scoreon spread, reports, new modelThe sweeper re-scores posts whose reach or report rate jumps, and every post from the last 7 days when a classifier is replaced.
Read at rankingper requestOne batch read of 600 states and multipliers: 96B reads a day, about 3.3M a second at peak. A missing row counts as unscored.
Removal pushwithin a minuteA removal is pushed to the session caches so a post already sitting in someone's next page is dropped.

What to watch for

FailureImpactDetectionMitigationMeanwhile
Scoring lag during a spike (breaking news)2Integrity scorerNew posts wait longer to reach beyond friendsShare of candidates still unscored rises above 1%Unscored posts stay friends-only; add capacity; alert on the unscored share.Friends still see new posts at once.
Adversaries adaptPrevalence climbs for one policyText moves into images, misspellings spread; report-to-removal rate falls for one policyRetrain on fresh reviewed samples; extract text from images before scoring.User reports and the sweeper still catch spreading posts.
Stale removals in cached feeds3Integrity scoresRemoved posts are still seenRemoved posts keep getting views for minutesPush removals to the session cache; count post-removal views in prevalence.New feed pages already leave the post out.
Over-demotion of one communityLegitimate posts in one language reach fewer peopleFalse demotions concentrated in one language or dialectPer-slice thresholds or retraining; a working appeal path (content-moderation/review-queues).Posts still reach friends; only reach is cut.
Multiplier creep5Ranking stackOrdinary posts get demoted tooMedian m across candidates drifts down as new signals are addedTrack each factor's share of total demotion; floor the combined m; retire signals the holdout shows do nothing.The 0.1 floor keeps any post rankable.
Was this section helpful?
Related
From policy to labels
Read next