Class imbalance and samplingWhy imbalance hurts

100%

Why imbalance hurts.

Sort Hatchway's candidate scam models by accuracy and a rule that flags nothing lands above a model that stops 2,400 scams a week. Rare positives don't break the loss function as much as people think; they break the metric, the threshold, the batches and the number of distinct examples the model gets to learn from. This topic separates the real damage from the folklore, so the fix in the next topic targets the right thing.

Beginner14 minUpdated 1 Oct 2026

The idea.

One scam in 500 listings, and a leaderboard that ranks the wrong model first.

The leaderboard that ranks the wrong model first

Hatchway is a marketplace for used furniture and home goods. Sellers post about 2,000,000 new listings a week, and around 0.2% of them later turn out to be scams: a sofa that does not exist, a deposit asked for up front. That is 4,000 scams hidden among 1,996,000 honest listings. A trust team checks listings the model flags, and it has time for 6,000 of them a week.

The team's first leaderboard of candidate models is sorted by accuracy. Near the top sits a baseline rule that flags nothing (99.80%); just below it, a model that stops 2,400 scams a week (99.77%). The ordering is not a rounding accident. Accuracy charges the same for every mistake, and that model trades 2,400 missed scams for 2,994 wrongly flagged sellers, so its total error count is higher: 1,600 misses plus 2,994 false flags is 4,594, against the empty rule's 4,000 misses. When 998 of every 1,000 rows are honest, accuracy is mostly a score for how a model treats honest listings.

A real model does better, and its rates still mislead. Suppose it finds 80% of scams while wrongly flagging 1% of honest listings. One percent sounds harmless, but it is 1% of almost two million: 19,960 honest sellers flagged, next to 3,200 real scams. Only 13.8% of flags are scams, and the 23,160 flags are nearly four times what the team can read. The sheer count of negatives turns every small rate into a big number, and neither accuracy nor the false-positive rate shows it.

One Hatchway week, 4,000 scams among 2,000,000 listings

Accuracy of the flag-nothing rule
99.80%
beats 99.77% for the 60%-recall model (4,000 errors vs 4,594)
Scams the higher-ranked rule catches
0
of 4,000 that week
Precision at 80% recall
13.8%
with 1% of honest listings flagged
Flags per week at 80% recall
23,160
against a review capacity of 6,000

Same week, four ways to call it

Model or ruleScams caughtHonest listings flaggedFlags per weekPrecisionAccuracy
Never flag0 of 4,000 caught0 honest flagged0 flagsprecision undefined99.80% accurate
Recall 80%, FPR 1%3,200 caught19,960 honest flagged23,160 flags13.8% precise98.96% accurate
Recall 60%, FPR 0.15%2,400 caught2,994 honest flagged5,394 flags44.5% precise99.77% accurate
Recall 50%, FPR 0.05%2,000 caught998 honest flagged2,998 flags66.7% precise99.85% accurate

Weekly flags against the review queue

Weekly flags against the review queueOnly the two stricter operating points fit under the 6,000-a-week review queue; the 80% recall model produces 23,160 flags, almost four times the capacity.05k10k15k20k25kRecall 80%Recall 60%Recall 50%Never flag23.2k5.39k3k0Review capacity 6,000Operating pointFlags per weekWeekly flags against the review queueOnly the two stricter operating points fit under the 6,000-a-week review queue; the 80% recall model produces 23,160 flags, almost four times the capacity.05k10k15k20k25kRecall 80%Recall 60%Recall 50%Never flag23.2k5.39k3k0Review capacity 6,000Operating pointFlags per week
Bars are labelled by recall; their false-positive rates are 1%, 0.15% and 0.05%. Accuracy barely moves across these rows (98.96% to 99.85%) and even ranks never-flag above the model that catches 2,400 scams. The flag count against the queue is what decides which row Hatchway can run.
Data
Operating pointflags per week
Recall 80%23,160
Recall 60%5,394
Recall 50%2,998
Never flag0
  • Review capacity 6,000: Flags per week = 6,000
Was this section helpful?

How it works.

Four separate things go wrong when positives are rare, and plain log loss is not one of them.

What rarity actually breaks

1. The metric
Accuracy is an average taken mostly over negatives, and the false-positive rate is taken over negatives only, so both stay flattering while the positives are handled badly. Precision, recall at a fixed flag budget and the PR curve look at the positives directly. The definitions live in the classification-metrics topic; here only the base-rate effect matters.
2. The threshold
A model trained on the natural 0.2% rate is right to give most scams scores like 0.01 or 0.05, because even suspicious listings are usually honest. Almost nothing crosses 0.5, so the default rule flags almost nothing. The cut-off has to be chosen from data and cost, not inherited. Elor and Averbuch-Elor found that tuning it recovered most of what balancing the classes gains.
3. Distinct positives
48,000 scams over 12 weeks sounds like plenty, but they split into sub-types, and a new trick may have 300 examples. The model has little to generalise from, and those rare sub-types are exactly where it misses. King and Zeng make the same point for rare-event data: the information lives in the events. New labelled positives help; copies of old ones add no new pattern.
4. Compute and batches
99.8% of each training pass is spent on honest listings, most already scored near zero. With small mini-batches many contain no scam at all, so a large share of updates see only one class. The chart below counts them.

Log loss is not the part that breaks

Assumptions
Scams in 12 weeks of training data
48,000
Honest listings in the same window
23,952,000
Model's starting score for every listing
0.002the base rate
Working
  1. Gradient of log loss on the logit, one scamp − 1 = 0.002 − 1−0.998from Model's starting score for every listing
  2. Gradient on the logit, one honest listingp − 0+0.002from Model's starting score for every listing
  3. Total pull from the scams48,000 × 0.99847,904from Scams in 12 weeks of training data and Gradient of log loss on the logit, one scam
  4. Total pull from the honest listings23,952,000 × 0.00247,904from Honest listings in the same window and Gradient on the logit, one honest listing
What it means
  • Each scam pulls about 500 times harder than each honest listing, which exactly offsets there being about 500 times fewer of them. At the base rate the two classes pull equally on the intercept, so log loss neither ignores the rare class nor over-predicts the common one; the prior it learns is the true one. This check covers the intercept only; it says nothing about whether the features or splits separate scams from honest listings.
  • With 48,000 scams the small-sample bias King and Zeng describe is negligible overall, but it returns for rare sub-types with a few hundred examples. A model trained at the natural rate is usually close to calibrated, which rebalancing in the resampling topic undoes.

Scams per mini-batch of 512

Scams per mini-batch of 512At a 0.2% positive rate, 35.9% of 512-row batches contain no scam at all and another 36.8% contain exactly one.05%10%15%20%25%30%35%40%012345+35.9%36.8%18.8%6.4%1.6%0.4%Share of batches (%)Scams in the batchScams per mini-batch of 512At a 0.2% positive rate, 35.9% of 512-row batches contain no scam at all and another 36.8% contain exactly one.05%10%15%20%25%30%35%40%012345+35.9%36.8%18.8%6.4%1.6%0.4%Share of batches (%)Scams in the batch
Binomial with n = 512 and p = 0.002; 5+ is what remains after 0 to 4 (bars rounded, so they sum to 99.9%). At batch 4,096 the mean is 8.2 scams and an empty batch has a 0.03% chance (0.998^4,096); at batch 256 it is empty 60% of the time (0.998^256).
Data
Scams in the batchshare of batches (%)
035.9
136.8
218.8
36.4
41.6
5+0.4

Rare positives make small evaluation sets noisy

A quick 100,000-listing sample holds about 200 scams, and an 80% recall read from it is only good to about ±5.5 points; a second model reading 76% is not measurably worse. A full week holds 4,000 scams and narrows that to about ±1.2. Why the count of positives sets the noise, the interval arithmetic and a paired test for comparing two models are in classification metrics.

Two habits follow. Split by time first, so the test set resembles next week: train on the 12-week window, validate on week 13, test on week 14. Both held-out weeks keep the natural rate on their own. Inside the training weeks, any random split or cross-validation used for tuning should stratify by label (scikit-learn's StratifiedKFold) so each fold keeps about 0.2% scams instead of leaving one fold with a handful by chance.

Was this section helpful?

In practice.

Measure at the natural rate, see how precision depends on the base rate, and pick the threshold from the review queue.

Precision at 80% recall, by base rate

  • π = 20%
  • π = 2%
  • π = 0.2%
Precision at 80% recall, by base rateThe same 1% false-positive rate gives 95% precision when a fifth of cases are positive, 62% at 2% and 14% when one in 500 is.020%40%60%80%100%0.00010.0010.010.1FPR 1%π = 20%π = 2%π = 0.2%Precision (%)False-positive ratePrecision at 80% recall, by base rateThe same 1% false-positive rate gives 95% precision when a fifth of cases are positive, 62% at 2% and 14% when one in 500 is.020%40%60%80%100%0.00010.0010.010.1FPR 1%π = 20%π = 2%π = 0.2%Precision (%)False-positive rate
Precision = 0.8π ÷ (0.8π + FPR × (1 − π)), with recall fixed at 80%. Recall and the false-positive rate are each measured inside one class, so they cannot see π; precision mixes the two classes and falls with it. Hatchway sits on the π = 0.2% curve. Why that makes the PR curve more telling than ROC on rare classes is covered in the classification-metrics topic.
Data
False-positive rateπ = 20%π = 2%π = 0.2%
010099.494.1
099.998.284.2
0.00199.594.261.6
0.00398.584.534.8
0.0195.26213.8
0.038735.25.1
0.166.7141.6
  • FPR 1%: False-positive rate = 0.01

Picking the threshold from the review queue

Assumptions
Honest listings per week
1,996,000
Scams per week
4,000
Review capacity
6,000 a week
Operating point read off the validation PR curve
recall 60%, FPR 0.15%
Working
  1. Scams flagged4,000 × 0.602,400from Scams per week and Operating point read off the validation PR curve
  2. Honest listings flagged1,996,000 × 0.00152,994from Honest listings per week and Operating point read off the validation PR curve
  3. Flags per week2,400 + 2,9945,394 (about 10% under capacity)from Scams flagged, Honest listings flagged and Review capacity
  4. Precision2,400 ÷ 5,39444.5%from Scams flagged and Flags per week
  5. Scams not flagged4,000 − 2,4001,600 a weekfrom Scams per week and Scams flagged
What it means
  • The threshold is a business number set by capacity or cost, not a constant like 0.5. The 1,600 missed scams fall to buyer reports and the next model.
  • Re-check it whenever the base rate moves. A scam wave that doubles the rate roughly doubles true positives while false positives stay put, so the same threshold now overflows the queue.

Evaluation rules for rare positives

Split by time first; inside the training weeks, stratify any random tuning split or cross-validation by label.
Keep validation and test sets at the natural rate. Never resample them.
Report PR-AUC, or recall at a fixed precision or flag budget, next to ROC-AUC.
Report counts (flags per week, scams missed), not only rates.
Slice recall by scam type; a rare sub-type can fail completely inside a healthy average.
Size the test set by its positives, not its rows.
Was this section helpful?

Trade-offs.

Imbalance is often a symptom. Decide what to fix first; the chosen option comes first.

01
Hatchway's first move on the scam model
Chosen:Fix the metric and the threshold first
  • Pro:No retraining; ships the same day
  • Pro:Keeps the scores readable as probabilities
  • Pro:For strong GBDT models, Elor and Averbuch-Elor found balancing added little once the threshold was tuned
Downside we accept:
  • Con:Adds no information about rare scam sub-types
  • Con:The threshold must be revisited whenever the base rate or the queue changes
  • Con:Hatchway still downsamples negatives later, for training speed rather than accuracy (next topic)
Ruled out:Resample the training set

Shifts the learned prior, so scores need a correction (the resampling topic); Throwing away negatives loses some of their variety

Ruled out:Reweight the loss or change it

Same prior shift as resampling; One more hyperparameter to tune and log

Ruled out:Collect more positives

Slow and costly; needs labelling or weak supervision (the data-collection topics)

When "imbalance" is really something else

SymptomLikely causeWhere to look
Low recall on one scam typeToo few labelled examples of that typeLabel more of it (data collection and labeling)
Precision falls month after monthThe base rate or the scammers' tactics driftedDrift monitoring; adversaries in fraud detection
Scams look like honest listings on every featureThe classes overlap; the features are too weakFeature engineering
Moderators disagree on the same listingLabel noiseLabel noise (data collection and labeling)

Interview probes

Q1
A parcel-damage detector scans 400,000 parcels a day, and a manager calls its 0.5% false-positive rate 'basically no false alarms'. What do you ask next?
How many parcels are actually damaged. 0.5% of roughly 400,000 good parcels is about 2,000 wrongly held parcels a day; if only 600 are damaged, most holds are false even at perfect recall. Ask for precision and daily counts of holds and misses against what the inspection desk can handle.
Q2
Why not keep the threshold at 0.5?
A model trained at the natural rate gives most positives small scores, correctly, so 0.5 flags almost nothing. Choose the threshold on untouched validation data from capacity or cost, and revisit it when the base rate moves.
Q3
So when would you resample at all?
When training cost is the problem, drop most negatives and correct the scores afterwards; when a weak learner ignores the rare class, rebalancing can help it. For a strong boosted model whose threshold is already tuned, it usually adds little. The resampling topic covers both, with the correction.
Was this section helpful?
Builds on this
Resampling and weighting
Read next