Why imbalance hurts.
Sort Hatchway's candidate scam models by accuracy and a rule that flags nothing lands above a model that stops 2,400 scams a week. Rare positives don't break the loss function as much as people think; they break the metric, the threshold, the batches and the number of distinct examples the model gets to learn from. This topic separates the real damage from the folklore, so the fix in the next topic targets the right thing.
The idea.
One scam in 500 listings, and a leaderboard that ranks the wrong model first.
The leaderboard that ranks the wrong model first
Hatchway is a marketplace for used furniture and home goods. Sellers post about 2,000,000 new listings a week, and around 0.2% of them later turn out to be scams: a sofa that does not exist, a deposit asked for up front. That is 4,000 scams hidden among 1,996,000 honest listings. A trust team checks listings the model flags, and it has time for 6,000 of them a week.
The team's first leaderboard of candidate models is sorted by accuracy. Near the top sits a baseline rule that flags nothing (99.80%); just below it, a model that stops 2,400 scams a week (99.77%). The ordering is not a rounding accident. Accuracy charges the same for every mistake, and that model trades 2,400 missed scams for 2,994 wrongly flagged sellers, so its total error count is higher: 1,600 misses plus 2,994 false flags is 4,594, against the empty rule's 4,000 misses. When 998 of every 1,000 rows are honest, accuracy is mostly a score for how a model treats honest listings.
A real model does better, and its rates still mislead. Suppose it finds 80% of scams while wrongly flagging 1% of honest listings. One percent sounds harmless, but it is 1% of almost two million: 19,960 honest sellers flagged, next to 3,200 real scams. Only 13.8% of flags are scams, and the 23,160 flags are nearly four times what the team can read. The sheer count of negatives turns every small rate into a big number, and neither accuracy nor the false-positive rate shows it.
One Hatchway week, 4,000 scams among 2,000,000 listings
Same week, four ways to call it
| Model or rule | Scams caught | Honest listings flagged | Flags per week | Precision | Accuracy |
|---|---|---|---|---|---|
| Never flag | 0 of 4,000 caught | 0 honest flagged | 0 flags | precision undefined | 99.80% accurate |
| Recall 80%, FPR 1% | 3,200 caught | 19,960 honest flagged | 23,160 flags | 13.8% precise | 98.96% accurate |
| Recall 60%, FPR 0.15% | 2,400 caught | 2,994 honest flagged | 5,394 flags | 44.5% precise | 99.77% accurate |
| Recall 50%, FPR 0.05% | 2,000 caught | 998 honest flagged | 2,998 flags | 66.7% precise | 99.85% accurate |
Weekly flags against the review queue
Data
| Operating point | flags per week |
|---|---|
| Recall 80% | 23,160 |
| Recall 60% | 5,394 |
| Recall 50% | 2,998 |
| Never flag | 0 |
- Review capacity 6,000: Flags per week = 6,000
How it works.
Four separate things go wrong when positives are rare, and plain log loss is not one of them.
What rarity actually breaks
Log loss is not the part that breaks
- Scams in 12 weeks of training data
- 48,000
- Honest listings in the same window
- 23,952,000
- Model's starting score for every listing
- 0.002the base rate
- Gradient of log loss on the logit, one scamp − 1 = 0.002 − 1−0.998from Model's starting score for every listing
- Gradient on the logit, one honest listingp − 0+0.002from Model's starting score for every listing
- Total pull from the scams48,000 × 0.99847,904from Scams in 12 weeks of training data and Gradient of log loss on the logit, one scam
- Total pull from the honest listings23,952,000 × 0.00247,904from Honest listings in the same window and Gradient on the logit, one honest listing
- Each scam pulls about 500 times harder than each honest listing, which exactly offsets there being about 500 times fewer of them. At the base rate the two classes pull equally on the intercept, so log loss neither ignores the rare class nor over-predicts the common one; the prior it learns is the true one. This check covers the intercept only; it says nothing about whether the features or splits separate scams from honest listings.
- With 48,000 scams the small-sample bias King and Zeng describe is negligible overall, but it returns for rare sub-types with a few hundred examples. A model trained at the natural rate is usually close to calibrated, which rebalancing in the resampling topic undoes.
Scams per mini-batch of 512
Data
| Scams in the batch | share of batches (%) |
|---|---|
| 0 | 35.9 |
| 1 | 36.8 |
| 2 | 18.8 |
| 3 | 6.4 |
| 4 | 1.6 |
| 5+ | 0.4 |
Rare positives make small evaluation sets noisy
A quick 100,000-listing sample holds about 200 scams, and an 80% recall read from it is only good to about ±5.5 points; a second model reading 76% is not measurably worse. A full week holds 4,000 scams and narrows that to about ±1.2. Why the count of positives sets the noise, the interval arithmetic and a paired test for comparing two models are in classification metrics.
Two habits follow. Split by time first, so the test set resembles next week: train on the 12-week window, validate on week 13, test on week 14. Both held-out weeks keep the natural rate on their own. Inside the training weeks, any random split or cross-validation used for tuning should stratify by label (scikit-learn's StratifiedKFold) so each fold keeps about 0.2% scams instead of leaving one fold with a handful by chance.
In practice.
Measure at the natural rate, see how precision depends on the base rate, and pick the threshold from the review queue.
Precision at 80% recall, by base rate
- π = 20%
- π = 2%
- π = 0.2%
Data
| False-positive rate | π = 20% | π = 2% | π = 0.2% |
|---|---|---|---|
| 0 | 100 | 99.4 | 94.1 |
| 0 | 99.9 | 98.2 | 84.2 |
| 0.001 | 99.5 | 94.2 | 61.6 |
| 0.003 | 98.5 | 84.5 | 34.8 |
| 0.01 | 95.2 | 62 | 13.8 |
| 0.03 | 87 | 35.2 | 5.1 |
| 0.1 | 66.7 | 14 | 1.6 |
- FPR 1%: False-positive rate = 0.01
Picking the threshold from the review queue
- Honest listings per week
- 1,996,000
- Scams per week
- 4,000
- Review capacity
- 6,000 a week
- Operating point read off the validation PR curve
- recall 60%, FPR 0.15%
- Scams flagged4,000 × 0.602,400from Scams per week and Operating point read off the validation PR curve
- Honest listings flagged1,996,000 × 0.00152,994from Honest listings per week and Operating point read off the validation PR curve
- Flags per week2,400 + 2,9945,394 (about 10% under capacity)from Scams flagged, Honest listings flagged and Review capacity
- Precision2,400 ÷ 5,39444.5%from Scams flagged and Flags per week
- Scams not flagged4,000 − 2,4001,600 a weekfrom Scams per week and Scams flagged
- The threshold is a business number set by capacity or cost, not a constant like 0.5. The 1,600 missed scams fall to buyer reports and the next model.
- Re-check it whenever the base rate moves. A scam wave that doubles the rate roughly doubles true positives while false positives stay put, so the same threshold now overflows the queue.
Evaluation rules for rare positives
Trade-offs.
Imbalance is often a symptom. Decide what to fix first; the chosen option comes first.
- Pro:No retraining; ships the same day
- Pro:Keeps the scores readable as probabilities
- Pro:For strong GBDT models, Elor and Averbuch-Elor found balancing added little once the threshold was tuned
- Con:Adds no information about rare scam sub-types
- Con:The threshold must be revisited whenever the base rate or the queue changes
- Con:Hatchway still downsamples negatives later, for training speed rather than accuracy (next topic)
Shifts the learned prior, so scores need a correction (the resampling topic); Throwing away negatives loses some of their variety
Same prior shift as resampling; One more hyperparameter to tune and log
Slow and costly; needs labelling or weak supervision (the data-collection topics)
When "imbalance" is really something else
| Symptom | Likely cause | Where to look |
|---|---|---|
| Low recall on one scam type | Too few labelled examples of that type | Label more of it (data collection and labeling) |
| Precision falls month after month | The base rate or the scammers' tactics drifted | Drift monitoring; adversaries in fraud detection |
| Scams look like honest listings on every feature | The classes overlap; the features are too weak | Feature engineering |
| Moderators disagree on the same listing | Label noise | Label noise (data collection and labeling) |