Label noise.
On Loftmarket the seller picks the category, so "Sneakers" sometimes holds boots and "Handbags" sometimes holds wallets. Training on those labels mostly works, until the model memorises the mistakes, and scoring against them quietly understates how good the model is.
Builds on Human labeling.
The idea.
Seller-picked categories, and the three kinds of wrong they come in.
Where Loftmarket's labels come from
Every Loftmarket listing gets a category from its seller, chosen from a menu while they upload photos. That choice is free and plentiful, so it became the label for the category classifier, which suggests a category to new sellers and routes listings into browse pages. Nobody checks it. Most sellers pick correctly, but a hurried seller files ankle boots under Sneakers, a seller chasing views files a wallet under Handbags because more buyers browse there, and a few pick the first option on the menu.
Think of whatever produced the labels as a model of its own, with an error pattern you never chose. A useful first question is not how many labels are wrong but what the errors depend on, because the three answers call for different fixes.
Three kinds of wrong (the Frénay and Verleysen taxonomy)
| Kind | Depends on | Loftmarket example | How much it hurts |
|---|---|---|---|
| Completely at random (NCAR) | Nothing | A labeler on the review queue mis-taps about 1 time in 100, whatever the item | Mild: spread evenly, so it mostly averages out given enough data |
| At random given the class (NAR) | The true class | Boots land in Sneakers far more often than in Phones; wallets drift into Handbags | Moves the boundary between similar classes; described by a class-to-class flip table |
| Not at random (NNAR) | The item itself | Dark photos of brown leather sneakers get filed as Boots | Worst: the wrong labels sit exactly where the model is already unsure |
How wrong labels get in
Seller choice is only one entry point. Vendor labelers get tired late in a shift. A guideline changes between versions and old labels are never revisited. Proxy labels built from behaviour are wrong in their own ways (see implicit labels). A label model built from rules has its own error rate (see weak supervision). And some noise is not noise at all but a bug in the pipeline that builds the table, which the in-practice section covers.
What wrong test labels do to a measured score
- Truly 100%
- Truly 95%
- Truly 90%
Data
| Wrong test labels (%) | Truly 100% | Truly 95% | Truly 90% |
|---|---|---|---|
| 0 | 100 | 95 | 90 |
| 5 | 95 | 90.5 | 86 |
| 10 | 90 | 86 | 82 |
| 15 | 85 | 81.5 | 78 |
| 20 | 80 | 77 | 74 |
The same model, scored on clean and on noisy test labels
How it works.
What noise does to training, and how to find the wrong labels with the model's own help.
Clean first, then the mistakes
A large enough network can learn anything, including nonsense. Zhang and colleagues replaced every CIFAR-10 label with a random one and standard image networks still reached zero training error; weight decay and dropout did not stop them. So given enough epochs, the category classifier will memorise that listing #88213, an obvious boot, is Sneakers.
The order matters, though. Arpit and colleagues showed that networks pick up the patterns shared by many consistent examples first and memorise the odd ones out later. A wrongly labeled boot contradicts thousands of correctly labeled boots, so for a while the model keeps predicting Boots for it and its training loss stays high. Two practical tools follow: stop training on a clean validation set before memorisation sets in, and treat a stubbornly high per-example loss as a hint that the label, not the model, is wrong. Co-teaching (Han et al.) builds on the second: two networks each pick the lowest-loss examples in a batch and train the other one on them, so neither feeds on its own mistakes.
Training accuracy on clean and on flipped listings
- Clean
- Flipped
- Validation
Data
| Epoch | Clean (%) | Flipped (%) | Validation (%) |
|---|---|---|---|
| 0 | 12 | 10 | 12 |
| 2 | 55 | 6 | 52 |
| 4 | 80 | 4 | 74 |
| 6 | 92 | 4 | 83 |
| 8 | 97 | 5 | 86 |
| 10 | 98 | 8 | 86 |
| 12 | no value | 15 | 85 |
| 14 | 98.5 | 25 | 83.5 |
| 16 | no value | 38 | 82 |
| 18 | 99 | 50 | 80.5 |
| 20 | no value | 61 | 79.5 |
| 22 | 99.3 | 70 | no value |
| 24 | no value | 77 | 78 |
| 26 | 99.5 | 83 | no value |
| 28 | no value | 87 | no value |
| 30 | 99.6 | 90 | 77 |
- stop: Epoch from 8 to 12
Finding wrong labels: confident learning
Ask a model trained without each listing what it thinks the listing is, and compare that with what the seller said, using a separate bar for each class (Northcutt, Jiang and Chuang 2021).
Step 1: split the listings into five folds, train on four and predict the fifth, five times over. Every listing now has class probabilities from a model that never saw its label. That matters: a model that trained on #88213 may already have memorised Sneakers for it.
Step 2: set one threshold per class, the average probability the model gives that class across the listings labeled with it. Some classes are easy (Phones) and some are muddled (Sneakers and Boots), so a single 0.5 cut-off would flag too much in one and too little in the other.
Step 3: for each listing, find the classes whose probability clears their own threshold and take the most probable of them as its likely true class. Count it in the cell (given label, likely label). That table is the confident joint. Listings off the diagonal are candidate errors; listings that clear no threshold are left out.
Is listing #88213 mislabeled?
- Seller's category
- Sneakers
- Out-of-fold probabilities (Sneakers, Boots, Sandals)
- 0.12, 0.81, 0.07
- Per-class thresholds (Sneakers, Boots, Sandals)
- 0.72, 0.68, 0.75mean probability of each class over listings given that label (illustrative)
- Does Sneakers clear its threshold?0.12 ≥ 0.72Nofrom Out-of-fold probabilities (Sneakers, Boots, Sandals) and Per-class thresholds (Sneakers, Boots, Sandals)
- Does Boots clear its threshold?0.81 ≥ 0.68Yesfrom Out-of-fold probabilities (Sneakers, Boots, Sandals) and Per-class thresholds (Sneakers, Boots, Sandals)
- Does Sandals clear its threshold?0.07 ≥ 0.75Nofrom Out-of-fold probabilities (Sneakers, Boots, Sandals) and Per-class thresholds (Sneakers, Boots, Sandals)
- Cell in the confident jointgiven Sneakers, likely Boots (the only class that cleared)Off-diagonal, flaggedfrom Does Sneakers clear its threshold?, Does Boots clear its threshold?, Does Sandals clear its threshold? and Seller's category
- Flag it and send it to a person; don't trust the seller or the model outright. A model that is 81% sure is still wrong about 1 time in 5 at that confidence, if it is calibrated.
Confident joint for two footwear classes (illustrative)
| Given label (listings) | Likely Sneakers | Likely Boots | Likely Sandals | Below every threshold |
|---|---|---|---|---|
| Sneakers (4,000) | 3,640 | 290 | 20 | 50 |
| Boots (2,500) | 180 | 2,280 | 5 | 35 |
From counts to an estimated noise rate
- Sneakers row
- 3,640 + 290 + 20 counted, 50 below threshold, 4,000 in all
- Boots row
- 180 + 2,280 + 5 counted, 35 below threshold, 2,500 in all
- Off-diagonal listings (flagged for review)290 + 20 + 180 + 5495 of 6,500from Sneakers row and Boots row
- Share of counted Sneakers that look like something else(290 + 20) ÷ 3,950≈ 7.8%from Sneakers row
- Estimated wrong Sneakers labels, scaled to the whole row7.8% × 4,000≈ 314from Share of counted Sneakers that look like something else
- Share of counted Boots that look like something else(180 + 5) ÷ 2,465≈ 7.5%from Boots row
- The two big off-diagonal cells point both ways (290 Sneakers that look like Boots, 180 Boots that look like Sneakers): sellers confuse this pair in both directions, which is class-dependent (NAR) noise.
- Scaling each row to its full count is the paper's calibration step; it keeps the listings that cleared no threshold from shrinking the estimate.
In practice.
Cleaning without deleting the hard cases, protecting the test set, and the bugs that look like noise.
Loftmarket's label-cleaning loop
Flagged is not the same as wrong
When Northcutt, Athalye and Mueller put confident-learning flags from ten benchmark test sets in front of five crowd workers each, 51% of the flagged items were confirmed as errors on average. The rest were correct labels the model found hard, or items where both labels applied. Deleting every flagged Loftmarket listing would therefore throw away roughly half good, hard examples: suede sneakers that look like boots are exactly what the model needs to see.
So flagged listings go to the relabel queue, where a reviewer looks at the photo and picks a category, or "neither", before seeing anything else. Only after that answer are the seller's category and the model's guess shown, so the reviewer can take a second look at a disagreement. The mechanics of that queue (guidelines, repeated labels, agreement) belong to human labeling, which gives the same rule for model pre-labels: a reviewer who sees the model's guess first tends to anchor on it and pass the model's blind spots straight back into the labels.
Review the flags, not the whole table
- Flagged listings
- 495
- All footwear listings
- 6,500
- Cost per reviewed listing
- $0.243 labels at $0.08 each (20 s at $15/h, as in human labeling)
- Share of flags that are real errors
- ~51%the average in Northcutt et al. 2021; Loftmarket's rate must be measured
- Reviewing only the flags495 × $0.24≈ $119from Flagged listings and Cost per reviewed listing
- Reviewing every listing6,500 × $0.24$1,560from All footwear listings and Cost per reviewed listing
- Labels corrected by reviewing the flags495 × 0.51≈ 252from Flagged listings and Share of flags that are real errors
- Cost per corrected label$119 ÷ 252≈ $0.47from Reviewing only the flags and Labels corrected by reviewing the flags
- Reviewing everything costs 13× more and mostly re-confirms labels that were already right; it is worth it only for the test set.
- Listings where the model makes the same mistake as the seller land on the diagonal and are never flagged, so a small random audit (say 300 listings) is still needed to estimate what the loop misses.
Training labels can be noisy; test labels cannot
Noise in the test set does not just lower every score by the same amount. Northcutt, Athalye and Mueller estimate at least 3.3% wrong test labels on average across ten widely used benchmarks, and at least 6% (2,916 images) in the ImageNet validation set. Removing those items left model rankings unchanged, but on the items that could be corrected, judging against the corrected labels reorders models: on ImageNet, ResNet-18 overtakes ResNet-50 once the share of originally mislabeled test items rises about 6 points; on CIFAR-10, VGG-11 overtakes VGG-19 at about 5 points. Larger models had learned to agree with the systematic errors.
Loftmarket's rule follows from that: the 2,000-listing test set is labeled by two category experts with a third settling disagreements, and the cleaning loop never edits it. If the loop were allowed to fix test labels using the model's own opinion, the test set would drift toward agreeing with the model and every score would look better than it is. How the scores themselves are defined belongs to classification metrics.
Bugs that look like label noise
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| A join on the wrong key shifts categories to the neighbouring listing ID6Label pipeline | A random-looking flip rate appears overnight, across all classes | Sanity checks between label and features (a 'Phones' listing with a shoe-sized image embedding); the flag rate jumps between two daily builds | Join on the listing ID and its version together; fail the build if the flag rate moves more than a set amount | Train from the last good snapshot |
| On the day a taxonomy change goes live, the nightly job still maps categories with the previous day's snapshot6Label pipeline | Listings created just after 00:00 UTC that day get a retired or wrong category, which looks like a burst of seller mistakes | Plot the flag rate by listing-creation hour; a cliff at 00:00 UTC is a snapshot bug, not seller noise | Map with the taxonomy version in force when the listing was created, and stamp that version on the label | Hold that day's new listings out of training until they are remapped |
| Labels written under guideline v1 mixed with v2, which moved slippers from Sandals to a new Slippers class1Training labels | The same kind of item carries two labels; it looks like NAR noise between the two classes | Stamp every label with its guideline and taxonomy version; count disagreements per version | Relabel or map the old version before training, or train only on one version | Train on the v2 labels alone, a smaller but consistent set |
| The same item is reposted and lands in both training and test5Test set | The test score rewards memorisation and hides the effect of noise | Near-duplicate hashing of photos and titles across the split | Split by seller and item cluster, not by listing ID | Report scores on the test items with no near-duplicate in training |
| The taxonomy merges "Trainers" into Sneakers but old rows keep the retired name6Label pipeline | A class that should be empty keeps receiving examples and steals probability | Watch the label distribution per class between builds; a retired class with new rows is a bug | Apply the taxonomy mapping in one place and reject unknown categories | Map the retired name to Sneakers when rows are read, until they are rewritten |
Trade-offs.
The chosen option is first; the others stay visible so the reasoning can be checked.
- Pro:Fixes the label itself, so every later model and every metric benefits
- Pro:The confirmed errors show which classes the guideline or the menu confuses
- Pro:Cheap when limited to the flags (about $120 here)
- Con:Needs a review queue and a few days of turnaround
- Con:Misses errors the model agrees with, so a random audit is still needed
About half the flags were correct, hard examples; dropping them teaches the model an easier world; Shrinks small classes most, since they have the fewest examples to spare
Needs calibrated probabilities, or the weights encode the model's own bias; Harder to explain and to audit than a corrected label
Assumes a noise structure; item-dependent noise breaks the flip-matrix assumption; Adds hyperparameters (kept fraction, smoothing ε) that need a clean validation set to tune
When noise is tolerable and when to invest
| Situation | Tolerate it | Invest in cleaning |
|---|---|---|
| Kind of noise | Completely at random, a few percent | Item-dependent, or a class pair that flips both ways |
| Amount of data | Hundreds of thousands of examples per class | Rare classes with a few hundred examples |
| Training setup | Early stopping and regularisation on a clean validation set | Long training of a large model that will memorise |
| Which split | Training labels | The validation and test sets, always |
| Cost of an error | A browse page shows a boot among sneakers | A prohibited item slips through, or the wrong model ships |