Resampling and weighting.
Training on 24 million listings to learn from 48,000 scams spends most of the compute on negatives the model already gets right. Keeping one negative in 40 makes training 37× cheaper, but the model now believes scams are 7.4% of listings instead of 0.2%. This topic covers the ways to change the class mix (undersampling, oversampling, SMOTE, class weights, focal loss), the one-line correction that puts the real probabilities back, and when to skip all of it and just move the threshold.
Builds on Why imbalance hurts.
The idea.
Four ways to make Hatchway's 4,000 weekly scams count for more in training, and the price each one charges.
Why anyone changes the class mix
Hatchway is a fictional marketplace for secondhand furniture. Its trust team wants a model that scores each new listing for the chance it is a scam: a fake sofa, a deposit request for a table that does not exist. A 12-week training window holds 24,000,000 labelled listings, of which 48,000 were confirmed scams. That is 2 in every 1,000. The two weeks after the window are held back for validation and test.
Training on all of it as logged works, but it is wasteful. Almost every row is a legitimate listing that the model classifies correctly after the first few passes, so most of the compute goes into confirming what the model already knows. A weak learner can also settle on predicting 'not a scam' for everything, a failure accuracy hides, as why imbalance hurts shows.
There are four ways to push back. Use fewer negatives (undersampling). Use more positives, either as exact copies or as synthetic points built between real ones (oversampling, SMOTE). Leave the rows alone and make each positive count for more in the loss (class weights). Or change the loss so that examples the model already gets right count for less, whichever class they are in (focal loss).
Each one has the same side effect. The model learns the class mix it was shown, not the one it will face. A model trained with one scam in 14 rows will tell you a listing is 30% likely to be a scam when the real chance is about 1%. Moderators sorting a queue by score will not notice, but anything that reads the number as a probability will: a threshold, an expected-loss calculation, a dashboard of predicted scam volume.
Per 1,000 training rows, before and after
- legit, 998 items, Legitimate listing, value 998
- scam, 2 items, Real scam, value 2
- Legitimate listing
- Real scam
- Synthetic scam (SMOTE)
- Scam with a larger weight
As it starts. 5 steps follow.
The four families at a glance
| Technique | What changes | Training rows | Probabilities afterwards | Main risk |
|---|---|---|---|---|
| Undersample negatives, 1 in 40 | Fewer negatives | 646,800 rows | Inflated about 37-fold near the base rate; fix with the prior correction | Throws away the variety of rare legitimate listings |
| Oversample positives (copies) | More positive rows | 24.4M rows (10× copies) | Inflated; needs a correction | Overfits to the duplicated rows |
| SMOTE | Synthetic positive rows | 24.4M rows (10× scams) | Inflated; needs a correction | Invents implausible points; cannot handle categories or text |
| Class weights / focal loss | The loss, not the rows | 24.0M rows, unchanged | Inflated (weights) or distorted (focal); focal needs post-hoc recalibration, not a formula | One more hyperparameter; still needs recalibration |
How it works.
Where resampling sits in the pipeline, the correction formula with numbers, then oversampling, SMOTE, class weights and focal loss.
Where resampling happens
Undersampling and the prior correction
Why one formula undoes the sample
Every scam survives the sample and each legitimate listing survives with probability w. So among listings that look alike, the sample holds the same scams but only a fraction w of the honest ones, and the odds of a scam in the sample are the true odds divided by w: odds′ = odds ÷ w. A model fitted to the sample learns odds′. To get back, multiply by w: odds(q) = w × p′ ÷ (1 − p′). Turning odds back into a probability gives q = p′ ÷ (p′ + (1 − p′) ÷ w), and taking logs gives the same thing as a constant shift, logit(q) = logit(p′) + ln w.
The one assumption is that the sample keeps negatives at random with respect to the features. If the keep rate depended on, say, the listing's category, the shift would differ by category and a single ln w would be wrong.
Undoing a 1-in-40 negative sample
- Negative keep rate w
- 0.0251 negative in 40
- Scams kept
- 48,000all of them
- Legitimate listings logged
- 23,952,000
- Raw score of one listing from the downsampled model p′
- 0.30
- Negatives kept23,952,000 × 0.025598,800from Legitimate listings logged and Negative keep rate w
- Training rows48,000 + 598,800646,800 (37× fewer than 24M)from Scams kept and Negatives kept
- Base rate the model sees48,000 ÷ 646,8007.42%from Scams kept and Training rows
- Corrected probability qp′ ÷ (p′ + (1 − p′) ÷ w) = 0.30 ÷ (0.30 + 0.70 ÷ 0.025) = 0.30 ÷ 28.301.06%from Raw score of one listing from the downsampled model p′ and Negative keep rate w
- The same correction as a shift in log-oddslogit(p′) + ln w = −0.847 + (−3.689) = −4.536q = 1 ÷ (1 + e^4.536) = 1.06%from Raw score of one listing from the downsampled model p′ and Negative keep rate w
- Check: the sampled base rate maps back to the real one0.0742 ÷ (0.0742 + 0.9258 ÷ 0.025)0.200%from Base rate the model sees and Negative keep rate w
- He et al. (Facebook ads, 2014) give exactly this recalibration, q = p / (p + (1 − p) / w), and report 0.025 as the best negative keep rate among those they tried for their click model.
- For logistic regression it is King and Zeng's prior correction: subtract ln[((1 − τ)/τ)(ȳ/(1 − ȳ))] from the intercept, where τ is the true rate and ȳ the sampled one. The slopes are already consistent and need nothing.
- Ranking by p′ and by q gives the same order, so ROC-AUC does not move. Only thresholds and consumers that read q as a probability need it.
What a downsampled score really means
- w = 0.025
- w = 0.05
- w = 0.1
- w = 1 (none)
Data
| Raw score from the downsampled model p′ | w = 0.025 (%) | w = 0.05 (%) | w = 0.1 (%) | w = 1 (none) (%) |
|---|---|---|---|---|
| 0.074 | 0.2 | 0.4 | 0.8 | 7.42 |
| 0.1 | 0.28 | 0.55 | 1.1 | 10 |
| 0.3 | 1.06 | 2.1 | 4.11 | 30 |
| 0.5 | 2.44 | 4.76 | 9.09 | 50 |
| 0.7 | 5.51 | 10.4 | 18.9 | 70 |
| 0.9 | 18.4 | 31 | 47.4 | 90 |
| 0.97 | 44.7 | 61.8 | 76.4 | 97 |
- Hatchway base rate 0.2%: Corrected probability q (%) = 0.2
Oversampling and SMOTE
Copies, and points in between
Random oversampling repeats positive rows until the mix looks the way you want. It adds no information, and a deep tree or a big network can simply memorise the repeated rows. SMOTE (Chawla et al., 2002) tries to do better by building new positives. For a scam x_i it picks one of its k = 5 nearest scam neighbours, x_nn, draws a gap uniformly from [0, 1], and emits x_new = x_i + gap × (x_nn − x_i): a point somewhere on the segment between the two.
A worked case with two numeric features. Scam A asks ₹4,000 and comes from a seller account 2 days old. Its neighbour, scam B, asks ₹6,000 from an account 6 days old. With gap 0.3 the synthetic scam asks 4,000 + 0.3 × 2,000 = ₹4,600 from an account 2 + 0.3 × 4 = 3.2 days old. That is a believable listing.
Now the cases where it breaks. The seller's city is Pune for A and Delhi for B; there is no city 30% of the way between them. The listing's text embedding can be interpolated arithmetically, but the midpoint of two scam descriptions is not a description of anything. Two scams whose neighbourhood is full of honest listings produce a synthetic scam in the middle of them, teaching the model that honest listings look suspicious. And SMOTE run before the train and test split finds neighbours in the test set, so test information leaks into training and the offline numbers look better than production will.
Class weights and focal loss
scikit-learn's 'balanced' weights for Hatchway
- Training rows n
- 24,000,000
- Scams
- 48,000
- Legitimate listings
- 23,952,000
- Weight per class
- n ÷ (n_classes × n_class)scikit-learn compute_class_weight
- Weight on each scam24,000,000 ÷ (2 × 48,000)250from Training rows n, Scams and Weight per class
- Weight on each legitimate listing24,000,000 ÷ (2 × 23,952,000)0.501from Training rows n, Legitimate listings and Weight per class
- One scam counts like this many legitimate listings250 ÷ 0.501499from Weight on each scam and Weight on each legitimate listing
- Effective class mix in the loss(48,000 × 250) vs (23,952,000 × 0.501)12M vs 12M, so 50/50from Scams, Legitimate listings, Weight on each scam and Weight on each legitimate listing
- Log-odds correction to get back to 0.2%ln(1 ÷ 499)−6.21from One scam counts like this many legitimate listings
- In expectation this is the same as oversampling scams 499-fold without copying any rows. The learned prior becomes 50%, so the probabilities need the same kind of correction as undersampling.
- King and Zeng's weighted likelihood picks the weights the other way round, w1 = τ/ȳ and w0 = (1 − τ)/(1 − ȳ), so that the weighted sample reproduces the true rate τ and the probabilities come out right without a correction step.
Focal loss shrinks the easy examples
| p_t (probability given to the true class) | Cross-entropy (CE), −log p_t | Focal (FL), γ = 2 | Shrunk by |
|---|---|---|---|
| p_t = 0.5 | CE 0.693 | FL 0.173 | 4× smaller |
| p_t = 0.9 | CE 0.105 | FL 0.00105 | 100× smaller |
| p_t = 0.968 | CE 0.0325 | FL 0.0000333 | about 1,000× smaller |
| p_t = 0.99 | CE 0.0101 | FL 0.00000101 | 10,000× smaller |
Where focal loss comes from, and where it pays
Focal loss multiplies cross-entropy by (1 − p_t)^γ, so the factor is the model's own error raised to a power: FL = −α(1 − p_t)^γ log p_t. The table leaves α out; with γ = 2, an example the model already gets right at p_t = 0.9 counts 100 times less than it would under plain cross-entropy. Lin et al. (2017) built it for one-stage object detectors, which score on the order of 100,000 candidate boxes per image, nearly all of them obvious background. There γ = 2 with α = 0.25 worked best. It targets easy examples of either class, not the minority class as such. Its outputs lose cross-entropy's guarantee of calibration and there is no closed-form prior correction, so recalibrate on natural-rate held-out data (Platt or isotonic). The distortion is not always harmful: Mukhoti et al. (2020) found focal loss often left deep networks better calibrated than cross-entropy did.
For Hatchway's GBDT it is rarely worth the trouble. It matters where the bulk of the examples are trivially easy and swamp the gradient: dense detection, or scoring long candidate lists.
Many rare classes at once
If Hatchway split scams into 30 subtypes, a few common and most rare, inverse-frequency weights would go to extremes: the rarest subtype would get a huge weight from a handful of rows. Cui et al. (2019) weight each class by the inverse of its effective number of samples, (1 − β^n)/(1 − β) for n examples and β just below 1. The effective number grows almost linearly for small n and levels off for large n, because the 10,000th example of a common class mostly repeats what the model has seen, while the 10th example of a rare class still adds something new.
In practice.
What Hatchway actually ships, what it saves, and the mistakes that reach production.
What 1-in-20 downsampling buys Hatchway
- Training rows as logged
- 24,000,000
- GBDT training throughput
- ~1M rows a minuteon the team's training machine (illustrative)
- Negative keep rate
- 0.05all 48,000 scams kept
- Negatives kept23,952,000 × 0.051,197,600from Training rows as logged and Negative keep rate
- Training rows1,197,600 + 48,0001,245,600from Negatives kept
- Training time24M ÷ 1M/min vs 1.2456M ÷ 1M/min24 min → about 1.25 minfrom Training rows as logged, Training rows and GBDT training throughput
- Speed-up24 ÷ 1.2456about 19×from Training time
- A 200-trial hyperparameter search200 × 24 min vs 200 × 1.25 min80 h → about 4.2 hfrom Training time
- Sampled base rate48,000 ÷ 1,245,6003.85%from Training rows
- Log-odds correction applied at servingln 0.05−3.00from Negative keep rate
- A single full-data fit is already short. The saving shows up when fits multiply: a search that would run for more than three days finishes in an afternoon.
- He et al. found that training on a uniform 10% of their data cost only about 1% in normalized entropy, while the negative keep rate had a clear effect of its own. Treat w as a hyperparameter to tune, not a constant to guess.
Resampling traps that ship to production
Where each technique earns its place
Trade-offs.
Hatchway's pick comes first, followed by the alternatives it turned down and why.
- Pro:About 19× cheaper training, which turns a multi-day hyperparameter search into an afternoon
- Pro:Calibrated again after one known shift in log-odds (−3.00)
- Pro:One number, w, to log with the model
- Con:Loses some of the variety among rare legitimate listings
- Con:The correction has to ship with every model version
Keeps only 48,000 of 23,952,000 negatives (0.2%), so most legitimate variety is gone; A large correction, ln(48,000 ÷ 23,952,000) ≈ −6.2, and the high-score tail where the threshold sits, made of legitimate listings that look like scams, is learned from very few negatives
Synthetic points are meaningless for categorical and text features; Training set grows to about 48M rows
No compute saving at all; The loss is balanced 50/50, so scores need the ln(1 ÷ 499) correction as well; Interacts with the booster's regularisation and minimum-leaf settings
Choosing a technique
| Technique | Saves compute | Correction | Handles sparse or categorical features | Best when |
|---|---|---|---|---|
| Undersample negatives + correction | Saves compute, in proportion to w | Add ln w to the logit | Handles any feature; rows are untouched | Negatives are plentiful and training cost matters |
| Random oversampling | No saving; the set grows | Subtract ln k for k copies of each positive | Handles any feature; exact copies | Small data and a weak learner |
| SMOTE | No saving; the set grows | Subtract ln of the positive multiplier; approximate, since new points are not real | Numeric only; interpolation needs continuous features | Small, dense, numeric data and a weak learner |
| Class weights | No saving | Subtract ln of the weight ratio (499 here) | Handles any feature | You cannot drop rows and want one knob |
| Focal loss | No saving | No formula; recalibrate on natural-rate held-out data | Handles any feature | Most examples are trivially easy (detection, long candidate lists) |
| Threshold only, no resampling | No saving | None needed | Handles any feature | A strong learner such as a GBDT; often the right default |