Data collection and labelingLabel noise

100%

Label noise.

On Loftmarket the seller picks the category, so "Sneakers" sometimes holds boots and "Handbags" sometimes holds wallets. Training on those labels mostly works, until the model memorises the mistakes, and scoring against them quietly understates how good the model is.

Intermediate17 minUpdated 1 Oct 2026

Builds on Human labeling.

The idea.

Seller-picked categories, and the three kinds of wrong they come in.

Where Loftmarket's labels come from

Every Loftmarket listing gets a category from its seller, chosen from a menu while they upload photos. That choice is free and plentiful, so it became the label for the category classifier, which suggests a category to new sellers and routes listings into browse pages. Nobody checks it. Most sellers pick correctly, but a hurried seller files ankle boots under Sneakers, a seller chasing views files a wallet under Handbags because more buyers browse there, and a few pick the first option on the menu.

Think of whatever produced the labels as a model of its own, with an error pattern you never chose. A useful first question is not how many labels are wrong but what the errors depend on, because the three answers call for different fixes.

Three kinds of wrong (the Frénay and Verleysen taxonomy)

KindDepends onLoftmarket exampleHow much it hurts
Completely at random (NCAR)NothingA labeler on the review queue mis-taps about 1 time in 100, whatever the itemMild: spread evenly, so it mostly averages out given enough data
At random given the class (NAR)The true classBoots land in Sneakers far more often than in Phones; wallets drift into HandbagsMoves the boundary between similar classes; described by a class-to-class flip table
Not at random (NNAR)The item itselfDark photos of brown leather sneakers get filed as BootsWorst: the wrong labels sit exactly where the model is already unsure

How wrong labels get in

Seller choice is only one entry point. Vendor labelers get tired late in a shift. A guideline changes between versions and old labels are never revisited. Proxy labels built from behaviour are wrong in their own ways (see implicit labels). A label model built from rules has its own error rate (see weak supervision). And some noise is not noise at all but a bug in the pipeline that builds the table, which the in-practice section covers.

What wrong test labels do to a measured score

  • Truly 100%
  • Truly 95%
  • Truly 90%
What wrong test labels do to a measured scoreWith 5% of test labels wrong, a 95%-accurate model scores 90.5%, and not even a perfect model can score above 95%.70%75%80%85%90%95%100%05%10%15%20%Truly 100%Truly 95%Truly 90%Measured accuracy (%)Wrong test labels (%)What wrong test labels do to a measured scoreWith 5% of test labels wrong, a 95%-accurate model scores 90.5%, and not even a perfect model can score above 95%.70%75%80%85%90%95%100%05%10%15%20%Truly 100%Truly 95%Truly 90%Measured accuracy (%)Wrong test labels (%)
Each line is a model's true accuracy. Binary label (is this listing Sneakers or not?), test-label errors independent of the model's errors, so measured = a(1 − e) + (1 − a)e, where a is true accuracy and e the test-label error rate. Worked: 0.95 × 0.95 + 0.05 × 0.05 = 0.905. With many classes the second term shrinks toward zero, so measured ≈ a(1 − e). Errors that are correlated with the model's can also reorder models (next section).
Data
Wrong test labels (%)Truly 100%Truly 95%Truly 90%
01009590
59590.586
10908682
158581.578
20807774

The same model, scored on clean and on noisy test labels

True accuracy
95%
scored against expert labels
Measured with 5% wrong labels
90.5%
0.95 × 0.95 + 0.05 × 0.05
Best score a perfectly correct model can get
95%
it disagrees with every wrong label; this holds under independent label errors, and a model that copies systematic errors can score higher
Was this section helpful?

How it works.

What noise does to training, and how to find the wrong labels with the model's own help.

Clean first, then the mistakes

A large enough network can learn anything, including nonsense. Zhang and colleagues replaced every CIFAR-10 label with a random one and standard image networks still reached zero training error; weight decay and dropout did not stop them. So given enough epochs, the category classifier will memorise that listing #88213, an obvious boot, is Sneakers.

The order matters, though. Arpit and colleagues showed that networks pick up the patterns shared by many consistent examples first and memorise the odd ones out later. A wrongly labeled boot contradicts thousands of correctly labeled boots, so for a while the model keeps predicting Boots for it and its training loss stays high. Two practical tools follow: stop training on a clean validation set before memorisation sets in, and treat a stubbornly high per-example loss as a hint that the label, not the model, is wrong. Co-teaching (Han et al.) builds on the second: two networks each pick the lowest-loss examples in a batch and train the other one on them, so neither feeds on its own mistakes.

Training accuracy on clean and on flipped listings

  • Clean
  • Flipped
  • Validation
Training accuracy on clean and on flipped listingsThe network fits the clean listings within about 8 epochs and only then starts memorising the flipped ones, so stopping near epoch 10 keeps most of the benefit and little of the noise.stop020%40%60%80%100%051015202530CleanFlippedValidationAccuracy (%)EpochTraining accuracy on clean and on flipped listingsThe network fits the clean listings within about 8 epochs and only then starts memorising the flipped ones, so stopping near epoch 10 keeps most of the benefit and little of the noise.stop020%40%60%80%100%051015202530CleanFlippedValidationAccuracy (%)Epoch
Illustrative curves with the shape reported by Arpit et al. 2017, for a run with 8% of training labels flipped. Clean and Flipped are training accuracy on the two parts of the training set; on the flipped part it means agreeing with the wrong label. Validation is scored on expert-checked labels.
Data
EpochClean (%)Flipped (%)Validation (%)
0121012
255652
480474
692483
897586
1098886
12no value1585
1498.52583.5
16no value3882
18995080.5
20no value6179.5
2299.370no value
24no value7778
2699.583no value
28no value87no value
3099.69077
  • stop: Epoch from 8 to 12

Finding wrong labels: confident learning

Ask a model trained without each listing what it thinks the listing is, and compare that with what the seller said, using a separate bar for each class (Northcutt, Jiang and Chuang 2021).

Step 1: split the listings into five folds, train on four and predict the fifth, five times over. Every listing now has class probabilities from a model that never saw its label. That matters: a model that trained on #88213 may already have memorised Sneakers for it.

Step 2: set one threshold per class, the average probability the model gives that class across the listings labeled with it. Some classes are easy (Phones) and some are muddled (Sneakers and Boots), so a single 0.5 cut-off would flag too much in one and too little in the other.

Step 3: for each listing, find the classes whose probability clears their own threshold and take the most probable of them as its likely true class. Count it in the cell (given label, likely label). That table is the confident joint. Listings off the diagonal are candidate errors; listings that clear no threshold are left out.

Is listing #88213 mislabeled?

Assumptions
Seller's category
Sneakers
Out-of-fold probabilities (Sneakers, Boots, Sandals)
0.12, 0.81, 0.07
Per-class thresholds (Sneakers, Boots, Sandals)
0.72, 0.68, 0.75mean probability of each class over listings given that label (illustrative)
Working
  1. Does Sneakers clear its threshold?0.12 ≥ 0.72Nofrom Out-of-fold probabilities (Sneakers, Boots, Sandals) and Per-class thresholds (Sneakers, Boots, Sandals)
  2. Does Boots clear its threshold?0.81 ≥ 0.68Yesfrom Out-of-fold probabilities (Sneakers, Boots, Sandals) and Per-class thresholds (Sneakers, Boots, Sandals)
  3. Does Sandals clear its threshold?0.07 ≥ 0.75Nofrom Out-of-fold probabilities (Sneakers, Boots, Sandals) and Per-class thresholds (Sneakers, Boots, Sandals)
  4. Cell in the confident jointgiven Sneakers, likely Boots (the only class that cleared)Off-diagonal, flaggedfrom Does Sneakers clear its threshold?, Does Boots clear its threshold?, Does Sandals clear its threshold? and Seller's category
What it means
  • Flag it and send it to a person; don't trust the seller or the model outright. A model that is 81% sure is still wrong about 1 time in 5 at that confidence, if it is calibrated.

Confident joint for two footwear classes (illustrative)

Given label (listings)Likely SneakersLikely BootsLikely SandalsBelow every threshold
Sneakers (4,000)3,6402902050
Boots (2,500)1802,280535

From counts to an estimated noise rate

Assumptions
Sneakers row
3,640 + 290 + 20 counted, 50 below threshold, 4,000 in all
Boots row
180 + 2,280 + 5 counted, 35 below threshold, 2,500 in all
Working
  1. Off-diagonal listings (flagged for review)290 + 20 + 180 + 5495 of 6,500from Sneakers row and Boots row
  2. Share of counted Sneakers that look like something else(290 + 20) ÷ 3,950≈ 7.8%from Sneakers row
  3. Estimated wrong Sneakers labels, scaled to the whole row7.8% × 4,000≈ 314from Share of counted Sneakers that look like something else
  4. Share of counted Boots that look like something else(180 + 5) ÷ 2,465≈ 7.5%from Boots row
What it means
  • The two big off-diagonal cells point both ways (290 Sneakers that look like Boots, 180 Boots that look like Sneakers): sellers confuse this pair in both directions, which is class-dependent (NAR) noise.
  • Scaling each row to its full count is the paper's calibration step; it keeps the listings that cleared no threshold from shrinking the estimate.
Was this section helpful?

In practice.

Cleaning without deleting the hard cases, protecting the test set, and the bugs that look like noise.

Loftmarket's label-cleaning loop

Loftmarket's label-cleaning loop. The numbered component cards that follow describe each part.
Loftmarket's label-cleaning loopComponents: 1. Training labels (About 6,500 footwear listings with the category their seller picked. The labels the model learns from some are wrong.), 2. Cross-validation scorer (Trains five models, each on four fifths of the data, and scores every listing with the model that never saw it.), 3. Confident joint (Computes a threshold per class, counts each listing in a (given label, likely label) cell and ranks the off-diagonal ones.), 4. Relabel queue (Human reviewers see the photo and pick a category (or "neither") first only then are the seller's category and the model's guess shown for a second look.), 5. Test set (2,000 listings labeled by two category experts, every disagreement settled by a third. Never cleaned by the model it is used to judge.).

all listings

probabilities

flagged only

corrected labels

Never touched by the loop

5Test set
2,000 listings
2 experts + tie-break

1Training labels
6,500 footwear listings
seller-picked

2Cross-validation scorer
5 folds
out-of-fold probabilities

3Confident joint
per-class thresholds
495 off-diagonal

4Relabel queue
photo first
then seller label + model guess
~500 reviews

Flagged is not the same as wrong

When Northcutt, Athalye and Mueller put confident-learning flags from ten benchmark test sets in front of five crowd workers each, 51% of the flagged items were confirmed as errors on average. The rest were correct labels the model found hard, or items where both labels applied. Deleting every flagged Loftmarket listing would therefore throw away roughly half good, hard examples: suede sneakers that look like boots are exactly what the model needs to see.

So flagged listings go to the relabel queue, where a reviewer looks at the photo and picks a category, or "neither", before seeing anything else. Only after that answer are the seller's category and the model's guess shown, so the reviewer can take a second look at a disagreement. The mechanics of that queue (guidelines, repeated labels, agreement) belong to human labeling, which gives the same rule for model pre-labels: a reviewer who sees the model's guess first tends to anchor on it and pass the model's blind spots straight back into the labels.

Review the flags, not the whole table

Assumptions
Flagged listings
495
All footwear listings
6,500
Cost per reviewed listing
$0.243 labels at $0.08 each (20 s at $15/h, as in human labeling)
Share of flags that are real errors
~51%the average in Northcutt et al. 2021; Loftmarket's rate must be measured
Working
  1. Reviewing only the flags495 × $0.24≈ $119from Flagged listings and Cost per reviewed listing
  2. Reviewing every listing6,500 × $0.24$1,560from All footwear listings and Cost per reviewed listing
  3. Labels corrected by reviewing the flags495 × 0.51≈ 252from Flagged listings and Share of flags that are real errors
  4. Cost per corrected label$119 ÷ 252≈ $0.47from Reviewing only the flags and Labels corrected by reviewing the flags
What it means
  • Reviewing everything costs 13× more and mostly re-confirms labels that were already right; it is worth it only for the test set.
  • Listings where the model makes the same mistake as the seller land on the diagonal and are never flagged, so a small random audit (say 300 listings) is still needed to estimate what the loop misses.

Training labels can be noisy; test labels cannot

Noise in the test set does not just lower every score by the same amount. Northcutt, Athalye and Mueller estimate at least 3.3% wrong test labels on average across ten widely used benchmarks, and at least 6% (2,916 images) in the ImageNet validation set. Removing those items left model rankings unchanged, but on the items that could be corrected, judging against the corrected labels reorders models: on ImageNet, ResNet-18 overtakes ResNet-50 once the share of originally mislabeled test items rises about 6 points; on CIFAR-10, VGG-11 overtakes VGG-19 at about 5 points. Larger models had learned to agree with the systematic errors.

Loftmarket's rule follows from that: the 2,000-listing test set is labeled by two category experts with a third settling disagreements, and the cleaning loop never edits it. If the loop were allowed to fix test labels using the model's own opinion, the test set would drift toward agreeing with the model and every score would look better than it is. How the scores themselves are defined belongs to classification metrics.

Bugs that look like label noise

FailureImpactDetectionMitigationMeanwhile
A join on the wrong key shifts categories to the neighbouring listing ID6Label pipelineA random-looking flip rate appears overnight, across all classesSanity checks between label and features (a 'Phones' listing with a shoe-sized image embedding); the flag rate jumps between two daily buildsJoin on the listing ID and its version together; fail the build if the flag rate moves more than a set amountTrain from the last good snapshot
On the day a taxonomy change goes live, the nightly job still maps categories with the previous day's snapshot6Label pipelineListings created just after 00:00 UTC that day get a retired or wrong category, which looks like a burst of seller mistakesPlot the flag rate by listing-creation hour; a cliff at 00:00 UTC is a snapshot bug, not seller noiseMap with the taxonomy version in force when the listing was created, and stamp that version on the labelHold that day's new listings out of training until they are remapped
Labels written under guideline v1 mixed with v2, which moved slippers from Sandals to a new Slippers class1Training labelsThe same kind of item carries two labels; it looks like NAR noise between the two classesStamp every label with its guideline and taxonomy version; count disagreements per versionRelabel or map the old version before training, or train only on one versionTrain on the v2 labels alone, a smaller but consistent set
The same item is reposted and lands in both training and test5Test setThe test score rewards memorisation and hides the effect of noiseNear-duplicate hashing of photos and titles across the splitSplit by seller and item cluster, not by listing IDReport scores on the test items with no near-duplicate in training
The taxonomy merges "Trainers" into Sneakers but old rows keep the retired name6Label pipelineA class that should be empty keeps receiving examples and steals probabilityWatch the label distribution per class between builds; a retired class with new rows is a bugApply the taxonomy mapping in one place and reject unknown categoriesMap the retired name to Sneakers when rows are read, until they are rewritten
Was this section helpful?

Trade-offs.

The chosen option is first; the others stay visible so the reasoning can be checked.

01
What to do with the 495 flagged footwear listings
Chosen:Relabel the flagged listings
  • Pro:Fixes the label itself, so every later model and every metric benefits
  • Pro:The confirmed errors show which classes the guideline or the menu confuses
  • Pro:Cheap when limited to the flags (about $120 here)
Downside we accept:
  • Con:Needs a review queue and a few days of turnaround
  • Con:Misses errors the model agrees with, so a random audit is still needed
Ruled out:Drop the flagged listings

About half the flags were correct, hard examples; dropping them teaches the model an easier world; Shrinks small classes most, since they have the fewest examples to spare

Ruled out:Down-weight or use soft labels

Needs calibrated probabilities, or the weights encode the model's own bias; Harder to explain and to audit than a corrected label

Ruled out:Train through the noise (robust training)

Assumes a noise structure; item-dependent noise breaks the flip-matrix assumption; Adds hyperparameters (kept fraction, smoothing ε) that need a clean validation set to tune

When noise is tolerable and when to invest

SituationTolerate itInvest in cleaning
Kind of noiseCompletely at random, a few percentItem-dependent, or a class pair that flips both ways
Amount of dataHundreds of thousands of examples per classRare classes with a few hundred examples
Training setupEarly stopping and regularisation on a clean validation setLong training of a large model that will memorise
Which splitTraining labelsThe validation and test sets, always
Cost of an errorA browse page shows a boot among sneakersA prohibited item slips through, or the wrong model ships
Was this section helpful?
Builds on this
Weak supervision
Read next