Class imbalance and samplingResampling and weighting

100%

Resampling and weighting.

Training on 24 million listings to learn from 48,000 scams spends most of the compute on negatives the model already gets right. Keeping one negative in 40 makes training 37× cheaper, but the model now believes scams are 7.4% of listings instead of 0.2%. This topic covers the ways to change the class mix (undersampling, oversampling, SMOTE, class weights, focal loss), the one-line correction that puts the real probabilities back, and when to skip all of it and just move the threshold.

Intermediate19 minUpdated 1 Oct 2026

Builds on Why imbalance hurts.

The idea.

Four ways to make Hatchway's 4,000 weekly scams count for more in training, and the price each one charges.

Why anyone changes the class mix

Hatchway is a fictional marketplace for secondhand furniture. Its trust team wants a model that scores each new listing for the chance it is a scam: a fake sofa, a deposit request for a table that does not exist. A 12-week training window holds 24,000,000 labelled listings, of which 48,000 were confirmed scams. That is 2 in every 1,000. The two weeks after the window are held back for validation and test.

Training on all of it as logged works, but it is wasteful. Almost every row is a legitimate listing that the model classifies correctly after the first few passes, so most of the compute goes into confirming what the model already knows. A weak learner can also settle on predicting 'not a scam' for everything, a failure accuracy hides, as why imbalance hurts shows.

There are four ways to push back. Use fewer negatives (undersampling). Use more positives, either as exact copies or as synthetic points built between real ones (oversampling, SMOTE). Leave the rows alone and make each positive count for more in the loss (class weights). Or change the loss so that examples the model already gets right count for less, whichever class they are in (focal loss).

Each one has the same side effect. The model learns the class mix it was shown, not the one it will face. A model trained with one scam in 14 rows will tell you a listing is 30% likely to be a scam when the real chance is about 1%. Moderators sorting a queue by score will not notice, but anything that reads the number as a probability will: a threshold, an expected-loss calculation, a dashboard of predicted scam volume.

Per 1,000 training rows, before and after

Rows per 1,000 logged
  1. legit, 998 items, Legitimate listing, value 998
  2. scam, 2 items, Real scam, value 2
  • Legitimate listing
  • Real scam
  • Synthetic scam (SMOTE)
  • Scam with a larger weight
Start

As it starts. 5 steps follow.

Start from the log as recorded: 998 legitimate listings and 2 scams per 1,000. Step through the four families. Only the first one removes rows, and only the last one leaves the rows exactly as they were.

The four families at a glance

TechniqueWhat changesTraining rowsProbabilities afterwardsMain risk
Undersample negatives, 1 in 40Fewer negatives646,800 rowsInflated about 37-fold near the base rate; fix with the prior correctionThrows away the variety of rare legitimate listings
Oversample positives (copies)More positive rows24.4M rows (10× copies)Inflated; needs a correctionOverfits to the duplicated rows
SMOTESynthetic positive rows24.4M rows (10× scams)Inflated; needs a correctionInvents implausible points; cannot handle categories or text
Class weights / focal lossThe loss, not the rows24.0M rows, unchangedInflated (weights) or distorted (focal); focal needs post-hoc recalibration, not a formulaOne more hyperparameter; still needs recalibration
Was this section helpful?

How it works.

Where resampling sits in the pipeline, the correction formula with numbers, then oversampling, SMOTE, class weights and focal loss.

Where resampling happens

Labelled listings (14 weeks)
split by time
Train on weeks 1–12 (the only rows resampled)
all scams kept
Keep 1 negative in 40 (w = 0.025)
fit
Train the GBDT
correct scores
Prior correction with the logged w
tune on week 13
Choose the threshold on untouched week 13 (at most 6,000 flags a week)
report once
Test on week 14 at the natural rate
The data is split by time first. Only the 12 training weeks are resampled; the correction and the threshold are fitted on week 13 at the natural 0.2% rate, and the test week, 14, is never touched. This walk-through uses w = 0.025; Hatchway ships 0.05.
Rules
Store w with the model artifact. The serving code needs it to apply the correction.
1

Undersampling and the prior correction

Why one formula undoes the sample

Every scam survives the sample and each legitimate listing survives with probability w. So among listings that look alike, the sample holds the same scams but only a fraction w of the honest ones, and the odds of a scam in the sample are the true odds divided by w: odds′ = odds ÷ w. A model fitted to the sample learns odds′. To get back, multiply by w: odds(q) = w × p′ ÷ (1 − p′). Turning odds back into a probability gives q = p′ ÷ (p′ + (1 − p′) ÷ w), and taking logs gives the same thing as a constant shift, logit(q) = logit(p′) + ln w.

The one assumption is that the sample keeps negatives at random with respect to the features. If the keep rate depended on, say, the listing's category, the shift would differ by category and a single ln w would be wrong.

Undoing a 1-in-40 negative sample

Assumptions
Negative keep rate w
0.0251 negative in 40
Scams kept
48,000all of them
Legitimate listings logged
23,952,000
Raw score of one listing from the downsampled model p′
0.30
Working
  1. Negatives kept23,952,000 × 0.025598,800from Legitimate listings logged and Negative keep rate w
  2. Training rows48,000 + 598,800646,800 (37× fewer than 24M)from Scams kept and Negatives kept
  3. Base rate the model sees48,000 ÷ 646,8007.42%from Scams kept and Training rows
  4. Corrected probability qp′ ÷ (p′ + (1 − p′) ÷ w) = 0.30 ÷ (0.30 + 0.70 ÷ 0.025) = 0.30 ÷ 28.301.06%from Raw score of one listing from the downsampled model p′ and Negative keep rate w
  5. The same correction as a shift in log-oddslogit(p′) + ln w = −0.847 + (−3.689) = −4.536q = 1 ÷ (1 + e^4.536) = 1.06%from Raw score of one listing from the downsampled model p′ and Negative keep rate w
  6. Check: the sampled base rate maps back to the real one0.0742 ÷ (0.0742 + 0.9258 ÷ 0.025)0.200%from Base rate the model sees and Negative keep rate w
What it means
  • He et al. (Facebook ads, 2014) give exactly this recalibration, q = p / (p + (1 − p) / w), and report 0.025 as the best negative keep rate among those they tried for their click model.
  • For logistic regression it is King and Zeng's prior correction: subtract ln[((1 − τ)/τ)(ȳ/(1 − ȳ))] from the intercept, where τ is the true rate and ȳ the sampled one. The slopes are already consistent and need nothing.
  • Ranking by p′ and by q gives the same order, so ROC-AUC does not move. Only thresholds and consumers that read q as a probability need it.

What a downsampled score really means

  • w = 0.025
  • w = 0.05
  • w = 0.1
  • w = 1 (none)
What a downsampled score really meansWith 1 negative in 40 kept, a raw score of 0.5 means only a 2.4% chance of a scam, and the sampled base rate maps back to 0.2%.0.1%1%10%100%00.20.40.60.81Hatchway base rate 0.2%w = 0.025w = 0.05w = 0.1w = 1 (none)Corrected probability q (%)Raw score from the downsampled model p′What a downsampled score really meansWith 1 negative in 40 kept, a raw score of 0.5 means only a 2.4% chance of a scam, and the sampled base rate maps back to 0.2%.0.1%1%10%100%00.20.40.60.81Hatchway base rate 0.2%w = 0.025w = 0.05w = 0.1w = 1 (none)Corrected probability q (%)Raw score from the downsampled model p′
q = p′ / (p′ + (1 − p′) / w) for three keep rates. The dashed line is a model trained without sampling, where the score is already the probability. Every curve starts at p′ = 0.0742, the base rate of the 1-in-40 sample. Log scale on the corrected axis.
Data
Raw score from the downsampled model p′w = 0.025 (%)w = 0.05 (%)w = 0.1 (%)w = 1 (none) (%)
0.0740.20.40.87.42
0.10.280.551.110
0.31.062.14.1130
0.52.444.769.0950
0.75.5110.418.970
0.918.43147.490
0.9744.761.876.497
  • Hatchway base rate 0.2%: Corrected probability q (%) = 0.2
2

Oversampling and SMOTE

Copies, and points in between

Random oversampling repeats positive rows until the mix looks the way you want. It adds no information, and a deep tree or a big network can simply memorise the repeated rows. SMOTE (Chawla et al., 2002) tries to do better by building new positives. For a scam x_i it picks one of its k = 5 nearest scam neighbours, x_nn, draws a gap uniformly from [0, 1], and emits x_new = x_i + gap × (x_nn − x_i): a point somewhere on the segment between the two.

A worked case with two numeric features. Scam A asks ₹4,000 and comes from a seller account 2 days old. Its neighbour, scam B, asks ₹6,000 from an account 6 days old. With gap 0.3 the synthetic scam asks 4,000 + 0.3 × 2,000 = ₹4,600 from an account 2 + 0.3 × 4 = 3.2 days old. That is a believable listing.

Now the cases where it breaks. The seller's city is Pune for A and Delhi for B; there is no city 30% of the way between them. The listing's text embedding can be interpolated arithmetically, but the midpoint of two scam descriptions is not a description of anything. Two scams whose neighbourhood is full of honest listings produce a synthetic scam in the middle of them, teaching the model that honest listings look suspicious. And SMOTE run before the train and test split finds neighbours in the test set, so test information leaks into training and the offline numbers look better than production will.

3

Class weights and focal loss

scikit-learn's 'balanced' weights for Hatchway

Assumptions
Training rows n
24,000,000
Scams
48,000
Legitimate listings
23,952,000
Weight per class
n ÷ (n_classes × n_class)scikit-learn compute_class_weight
Working
  1. Weight on each scam24,000,000 ÷ (2 × 48,000)250from Training rows n, Scams and Weight per class
  2. Weight on each legitimate listing24,000,000 ÷ (2 × 23,952,000)0.501from Training rows n, Legitimate listings and Weight per class
  3. One scam counts like this many legitimate listings250 ÷ 0.501499from Weight on each scam and Weight on each legitimate listing
  4. Effective class mix in the loss(48,000 × 250) vs (23,952,000 × 0.501)12M vs 12M, so 50/50from Scams, Legitimate listings, Weight on each scam and Weight on each legitimate listing
  5. Log-odds correction to get back to 0.2%ln(1 ÷ 499)−6.21from One scam counts like this many legitimate listings
What it means
  • In expectation this is the same as oversampling scams 499-fold without copying any rows. The learned prior becomes 50%, so the probabilities need the same kind of correction as undersampling.
  • King and Zeng's weighted likelihood picks the weights the other way round, w1 = τ/ȳ and w0 = (1 − τ)/(1 − ȳ), so that the weighted sample reproduces the true rate τ and the probabilities come out right without a correction step.

Focal loss shrinks the easy examples

p_t (probability given to the true class)Cross-entropy (CE), −log p_tFocal (FL), γ = 2Shrunk by
p_t = 0.5CE 0.693FL 0.1734× smaller
p_t = 0.9CE 0.105FL 0.00105100× smaller
p_t = 0.968CE 0.0325FL 0.0000333about 1,000× smaller
p_t = 0.99CE 0.0101FL 0.0000010110,000× smaller

Where focal loss comes from, and where it pays

Focal loss multiplies cross-entropy by (1 − p_t)^γ, so the factor is the model's own error raised to a power: FL = −α(1 − p_t)^γ log p_t. The table leaves α out; with γ = 2, an example the model already gets right at p_t = 0.9 counts 100 times less than it would under plain cross-entropy. Lin et al. (2017) built it for one-stage object detectors, which score on the order of 100,000 candidate boxes per image, nearly all of them obvious background. There γ = 2 with α = 0.25 worked best. It targets easy examples of either class, not the minority class as such. Its outputs lose cross-entropy's guarantee of calibration and there is no closed-form prior correction, so recalibrate on natural-rate held-out data (Platt or isotonic). The distortion is not always harmful: Mukhoti et al. (2020) found focal loss often left deep networks better calibrated than cross-entropy did.

For Hatchway's GBDT it is rarely worth the trouble. It matters where the bulk of the examples are trivially easy and swamp the gradient: dense detection, or scoring long candidate lists.

Many rare classes at once

If Hatchway split scams into 30 subtypes, a few common and most rare, inverse-frequency weights would go to extremes: the rarest subtype would get a huge weight from a handful of rows. Cui et al. (2019) weight each class by the inverse of its effective number of samples, (1 − β^n)/(1 − β) for n examples and β just below 1. The effective number grows almost linearly for small n and levels off for large n, because the 10,000th example of a common class mostly repeats what the model has seen, while the 10th example of a rare class still adds something new.

Was this section helpful?

In practice.

What Hatchway actually ships, what it saves, and the mistakes that reach production.

What 1-in-20 downsampling buys Hatchway

Assumptions
Training rows as logged
24,000,000
GBDT training throughput
~1M rows a minuteon the team's training machine (illustrative)
Negative keep rate
0.05all 48,000 scams kept
Working
  1. Negatives kept23,952,000 × 0.051,197,600from Training rows as logged and Negative keep rate
  2. Training rows1,197,600 + 48,0001,245,600from Negatives kept
  3. Training time24M ÷ 1M/min vs 1.2456M ÷ 1M/min24 min → about 1.25 minfrom Training rows as logged, Training rows and GBDT training throughput
  4. Speed-up24 ÷ 1.2456about 19×from Training time
  5. A 200-trial hyperparameter search200 × 24 min vs 200 × 1.25 min80 h → about 4.2 hfrom Training time
  6. Sampled base rate48,000 ÷ 1,245,6003.85%from Training rows
  7. Log-odds correction applied at servingln 0.05−3.00from Negative keep rate
What it means
  • A single full-data fit is already short. The saving shows up when fits multiply: a search that would run for more than three days finishes in an afternoon.
  • He et al. found that training on a uniform 10% of their data cost only about 1% in normalized entropy, while the negative keep rate had a clear effect of its own. Treat w as a hyperparameter to tune, not a constant to guess.

Resampling traps that ship to production

Resample after the train and test split, never before.
Never resample validation or test data.
Store w and any class weights with the model version, and apply the correction in the serving code.
Treat a retrain with a new w as a change to every score.
Check calibration on untouched data after correcting.
Compare the mean predicted probability with the observed scam rate every week.
In a streaming pipeline, choose which negatives to keep with a hash of the listing id, not a random draw.

Where each technique earns its place

Ad click models
Negative downsampling plus the recalibration formula (He et al.). Billions of impressions, few clicks, and a bid that multiplies the probability, so the correction is not optional; the calibrated-bids topic in ad click prediction covers the auction side.
Dense object detection
Focal loss (Lin et al.). Around 100,000 candidate boxes per image, almost all easy background, which would otherwise dominate the gradient.
Small tabular data, weak learners
Random oversampling, undersampling or SMOTE can help a decision tree or a linear model that would otherwise ignore the rare class. Elor and Averbuch-Elor found the gains shrink to nothing for strong boosted models with a tuned threshold.
Long-tailed classes
Class-balanced loss by effective number (Cui et al.) when there are many classes of very different sizes, such as image categories or fine-grained abuse types.
Was this section helpful?

Trade-offs.

Hatchway's pick comes first, followed by the alternatives it turned down and why.

01
How Hatchway trains the scam model
Chosen:Keep 1 negative in 20, prior-correct, tune the threshold on natural-rate validation
  • Pro:About 19× cheaper training, which turns a multi-day hyperparameter search into an afternoon
  • Pro:Calibrated again after one known shift in log-odds (−3.00)
  • Pro:One number, w, to log with the model
Downside we accept:
  • Con:Loses some of the variety among rare legitimate listings
  • Con:The correction has to ship with every model version
Ruled out:Undersample to 50/50

Keeps only 48,000 of 23,952,000 negatives (0.2%), so most legitimate variety is gone; A large correction, ln(48,000 ÷ 23,952,000) ≈ −6.2, and the high-score tail where the threshold sits, made of legitimate listings that look like scams, is learned from very few negatives

Ruled out:SMOTE to 50/50

Synthetic points are meaningless for categorical and text features; Training set grows to about 48M rows

Ruled out:class_weight = 'balanced'

No compute saving at all; The loss is balanced 50/50, so scores need the ln(1 ÷ 499) correction as well; Interacts with the booster's regularisation and minimum-leaf settings

Choosing a technique

TechniqueSaves computeCorrectionHandles sparse or categorical featuresBest when
Undersample negatives + correctionSaves compute, in proportion to wAdd ln w to the logitHandles any feature; rows are untouchedNegatives are plentiful and training cost matters
Random oversamplingNo saving; the set growsSubtract ln k for k copies of each positiveHandles any feature; exact copiesSmall data and a weak learner
SMOTENo saving; the set growsSubtract ln of the positive multiplier; approximate, since new points are not realNumeric only; interpolation needs continuous featuresSmall, dense, numeric data and a weak learner
Class weightsNo savingSubtract ln of the weight ratio (499 here)Handles any featureYou cannot drop rows and want one knob
Focal lossNo savingNo formula; recalibrate on natural-rate held-out dataHandles any featureMost examples are trivially easy (detection, long candidate lists)
Threshold only, no resamplingNo savingNone neededHandles any featureA strong learner such as a GBDT; often the right default
Was this section helpful?
Next in Core
Negative sampling
Read next