Data collection and labelingHuman labeling

100%

Human labeling.

Loftmarket's prohibited-item classifier can't learn from clicks: buyers happily click a knife listing or a recalled crib. Its labels come from people reading a policy. Every design choice here trades money per label against how much each label can be trusted.

Beginner20 minUpdated 1 Oct 2026

The idea.

A policy, a pile of listings, and people paid to decide which is which.

Why this model needs people

Loftmarket is a fictional second-hand marketplace for phones, sneakers, handbags and furniture. Its search ranker learns from clicks and purchases, because a click really does say something about relevance. Its prohibited-item classifier cannot. A hunting knife, a recalled baby crib or a fake designer bag gets clicked as eagerly as anything else, often more. The question the model must answer is 'does this listing break our policy?', and no amount of user behaviour answers it.

So someone has to read the policy, look at a listing and decide. The policy lists 14 prohibited categories. Loftmarket's trust-and-safety team has a few policy experts who know it cold; they cost about $60 an hour and there are not many of them. A labeling vendor supplies dozens of trained people at about $15 an hour who are quick but make more mistakes. The whole job is spending expert hours only where they change the answer.

Three levers control cost and quality: who labels an item, how many people label it, and which items get labeled at all. The rest of this topic pulls each one.

Labeling 40,000 listings

Assumptions
Listings to label
40,000
Time per vendor judgment
20 sabout 180 judgments an hour (illustrative)
Vendor labeler cost
$15 per hour
Gold questions mixed into every queue
5% extra judgmentslistings with a known expert answer, hidden in the queue to score labelers
Items that get two extra labels
20%first label low-confidence or disagrees with the model
Items that end in expert adjudication
3% (1,200), 90 s eachan expert settles listings the labelers could not agree on
Policy expert cost
$60 per hour
Working
  1. Vendor judgments, selective repeats40,000 × (1 + 0.2 × 2) × 1.0558,800from Listings to label, Items that get two extra labels and Gold questions mixed into every queue
  2. Vendor hours58,800 × 20 s ÷ 3,600≈ 327 hfrom Vendor judgments, selective repeats and Time per vendor judgment
  3. Vendor cost327 h × $15≈ $4,900from Vendor hours and Vendor labeler cost
  4. Expert adjudication1,200 × 90 s = 30 h; 30 h × $60$1,800from Items that end in expert adjudication and Policy expert cost
  5. Total, selective repeats$4,900 + $1,800≈ $6,700from Vendor cost and Expert adjudication
  6. Alternative, three labels on every item40,000 × 3 × 1.05 = 126,000 judgments = 700 h × $15 = $10,500; + $1,800≈ $12,300from Listings to label, Gold questions mixed into every queue, Time per vendor judgment, Vendor labeler cost and Expert adjudication
What it means
  • Asking for extra labels only where the first one looks shaky costs about 55% of labeling everything three times, and most items were never in doubt anyway.
  • Sheng, Provost and Ipeirotis (2008) reached the same conclusion in general: choosing which items to relabel usually beats relabeling all of them.
  • Expert time is under a third of the bill but is the scarce part; 30 hours is roughly one expert's week.

Who can supply the label

SourceCost per labelAccuracyThroughputUse at Loftmarket
Policy expertsHighest (~$1.50 at 90 s and $60/h)Highest; they wrote the rulesTens of items an hour, a few peopleAdjudication, the evaluation set, guideline examples
Vendor team with training~$0.08 at 20 s and $15/hGood on clear cases, weaker on edge casesThousands of items a dayMost training labels
Open crowd (anonymous workers)LowestVaries widely; needs redundancy and gold checksVery highNot used: listings can contain personal data, and the policy needs training
Users (the report button)FreeLow precision: people report rivals and bad sellersOnly where users lookA signal that puts listings into the queue, not a label
Model pre-label, human confirmsCheapest human option per itemInherits the model's blind spots through anchoringHighest per labelerOnly with the pre-label hidden until the labeler answers

How many amateurs make an expert?

Snow, O'Connor, Jurafsky and Ng (EMNLP 2008) paid Mechanical Turk workers to redo five language-annotation tasks that experts had already labeled. On the affect-recognition task (seven emotion and valence ratings), about four non-expert labels per item, averaged together, matched one expert's agreement with the other experts. Some subtasks needed only two, and one (fear) never got there. At 2008 prices that meant at least 875 expert-equivalent labels per US dollar. Treat the prices as history, and the ratio as a result for that kind of task: simple judgments with clear instructions. A policy call about whether a replica prop gun counts as a weapon is not that kind of task, which is why Loftmarket keeps experts for the hard cases.

Was this section helpful?

How it works.

The guideline, measuring agreement, combining several labels, and what happens to one listing in the queue.

The guideline is the spec

Labelers can only be as consistent as the document they follow. Loftmarket's guideline has these parts for each of the 14 categories.

A decision rule per category, with a one-line definition.
At least three clear positive examples and three tricky negatives per category.
An explicit 'unsure / needs policy' answer instead of forcing a guess.
An order of precedence for items that match two categories.
A version number stamped on every label.

Agreement beyond chance (Cohen's κ)

Assumptions
Listings both vendor labelers judged
200from an active-learning batch, which is rich in suspicious listings; that is why both call far more prohibited than the marketplace's roughly 3%
Listings where they gave the same answer
170
Labeler A says prohibited on
30%
Labeler B says prohibited on
25%
Rare-class case: both say prohibited on
5%, raw agreement 96%
Working
  1. Observed agreement, p_o170 ÷ 2000.85from Listings both vendor labelers judged and Listings where they gave the same answer
  2. Agreement expected by chance, p_e0.30 × 0.25 + 0.70 × 0.750.60from Labeler A says prohibited on and Labeler B says prohibited on · Both say yes by luck, plus both say no by luck, if each answered at random at their own rate.
  3. Cohen's κ(0.85 − 0.60) ÷ (1 − 0.60)0.625from Observed agreement, p_o and Agreement expected by chance, p_e
  4. Chance agreement for the rare class0.05² + 0.95²0.905from Rare-class case: both say prohibited on
  5. κ for the rare class(0.96 − 0.905) ÷ (1 − 0.905)≈ 0.58from Rare-class case: both say prohibited on and Chance agreement for the rare class
What it means
  • κ answers the question raw agreement cannot: how much better than two people guessing at their own rates? Cohen defined it in 1960 as (p_o − p_e) ÷ (1 − p_e).
  • If each labeler says allowed on 95% of listings, they would agree about 90% of the time by chance alone (0.95² + 0.05² = 0.905), so 96% raw agreement is only a κ of about 0.58.
  • Report κ per category, not one number for the whole guideline; with more than two labelers per item, Krippendorff's α plays the same role.

Accuracy of a majority vote

  • p = 0.6
  • p = 0.7
  • p = 0.8
  • p = 0.9
Accuracy of a majority voteGoing from one label to three buys the most; for 90%-accurate labelers more than three adds almost nothing, for 60%-accurate ones even nine is not enough.0.50.60.70.80.9124680.784p = 0.6p = 0.7p = 0.8p = 0.9Accuracy of the majority labelLabels per itemAccuracy of a majority voteGoing from one label to three buys the most; for 90%-accurate labelers more than three adds almost nothing, for 60%-accurate ones even nine is not enough.0.50.60.70.80.9124680.784p = 0.6p = 0.7p = 0.8p = 0.9Accuracy of the majority labelLabels per item
Each curve is one labeler accuracy p. The value is the chance that more than half of n independent labels are correct, from the binomial distribution (computed, not measured). It assumes errors are independent. A confusing guideline makes labelers wrong together, and then the curves flatten; Sheng, Provost and Ipeirotis (2008) show the same dependence on labeler quality and the shrinking gain per extra label.
Data
Labels per itemp = 0.6p = 0.7p = 0.8p = 0.9
10.60.70.80.9
30.6480.7840.8960.972
50.6830.8370.9420.991
70.710.8740.9670.997
90.7330.9010.980.999
  • At 3: 0.784

Better than one person, one vote

The arithmetic for three labelers at p = 0.7: the majority is right if all three are right (0.7³ = 0.343) or exactly two are (3 × 0.7² × 0.3 = 0.441), so 0.784 in total. The same sum at p = 0.4 gives 0.352, lower than one label on its own. Voting amplifies whatever side of 50% the labelers sit on, so it rescues decent labelers and sinks bad ones.

Plain majority vote also treats a careful labeler and a careless one the same. Dawid and Skene (1979) fixed that: estimate a confusion matrix for every labeler (how often each true class gets each answer) together with the likely true label of every item, alternating the two steps by expectation maximization until they settle. A labeler who calls every replica 'counterfeit' then counts for less on replicas, and one who is reliable on weapons keeps full weight there. The same idea, applied to heuristic rules instead of people, is the label model in the weak-supervision topic.

Gold questions measure labelers directly instead of inferring it. Five per cent of every queue is listings the experts have already decided, indistinguishable from the rest. Loftmarket's rule: a labeler whose rolling gold accuracy drops under 85% is paused and retrained, and their labels since the last good check are sent for another look.

One listing's path through the label queue

States of6Label aggregator

One listing's path through the label queue. 7 states, 8 transitions. The table below lists them.
One listing's path through the label queueThe states of Label aggregator. 7 states, 8 transitions. The table below lists them.

labeled

confident [model agrees, gold ≥ 90%]

low confidence or model disagrees

3 labels [≥ 2 match]

no majority or any 'unsure'

expert decides

expert: policy is silent

guideline v+1 published / relabel

Queued

One label

Needs 2 more labels

Agreed

Disputed

Adjudicated by an expert

Guideline gap

2 steps. A used iPhone 12 with a clear photo and a normal price.

What the label aggregator does with a single listing. Most listings leave after one label; the rest collect more votes, and the ones that still split go to an expert. A listing the policy does not cover sends the guideline, not the labeler, back for repair.

Transitions of One listing's path through the label queue
From → ToEventGuardActionActor
Queued → One labellabeledVendor labeler
One label → Agreedconfidentmodel agrees, gold ≥ 90%Aggregator
One label → Needs 2 more labelslow confidence or model disagreesAggregator
Needs 2 more labels → Agreed3 labels≥ 2 matchAggregator
Needs 2 more labels → Disputedno majority or any 'unsure'Aggregator
Disputed → Adjudicated by an expertexpert decidesPolicy expert
Disputed → Guideline gapexpert: policy is silentPolicy expert
Guideline gap → Queuedguideline v+1 publishedrelabelPolicy expert
Queuedstart
Agreedend
Adjudicated by an expertend
Guideline gaperror
Sent back to the policy owners
Was this section helpful?

In practice.

Letting the model pick what gets labeled, and closing the loop with production.

The labeling loop at Loftmarket

The labeling loop at Loftmarket. The numbered component cards that follow describe each part.
The labeling loop at LoftmarketComponents: 1. Unlabeled listings (Every live and recently removed Loftmarket listing that has no trusted label yet, with its title, photos, description and seller.), 2. Active-learning picker (Scores the pool with the current classifier each week and chooses the next batch to label, mixing uncertain, diverse and random listings.), 3. Label queue (Holds listings waiting for a judgment, with gold items mixed in, and hands each to a labeler who has not seen it.), 4. Vendor labelers (A contracted team trained on the guideline. Fast and affordable each person makes about 180 judgments an hour.), 5. Policy experts (A handful of trust-and-safety staff who own the guideline, settle disputes and label the evaluation set.), 6. Label aggregator (Combines judgments into one label per listing, scores labelers against gold items and decides when an item needs more labels.), 7. Classifier (The prohibited-item classifier: a prohibited-or-not score plus a category head over the 14 categories. Retrained weekly on the aggregated labels, then used to score the pool for the next batch.).

candidates

weekly batch

tasks

judgments

disputed listings

guideline, adjudicated labels

gold answers

final labels

scores on the pool

flagged listings

1Unlabeled listings

2Active-learning picker
uncertain + diverse + 20% random
2,000 listings a week

3Label queue
5% gold items mixed in

4Vendor labelers

5Policy experts
guideline owners
adjudication, eval set

6Label aggregator
gold checks
votes, Dawid–Skene weights

7Classifier
retrained weekly

User reports
report button

What goes round

Each Monday the classifier scores every unlabeled listing. The picker chooses 2,000 of them, the vendor team labels them during the week, the aggregator turns judgments into labels, and the classifier is retrained on the larger set. Next week's picks come from the new model, so each batch goes after whatever confuses the model now rather than what confused it a month ago. Settles (2010) calls this pool-based active learning: the learner sees a large unlabeled pool and asks for labels on the items it expects to learn the most from.

How the picker chooses (after Settles 2010)

StrategyPicksWatch out
Least confidentListings whose most likely class has the lowest probabilityIgnores how the rest of the probability is spread
MarginListings with the smallest gap between the top two classesGood for telling similar categories apart (weapon vs replica)
EntropyListings with the most spread-out distribution; on the category head that is 15 classes (14 categories plus allowed)On the yes/no prohibited score all three of these pick the same item: the one scored near 0.5
Query by committeeListings on which several models trained on different samples disagree mostCosts several models; the committee must be genuinely varied
Diversity (batch mode)Cluster the uncertain listings and take a few from each clusterNeeded when picking 2,000 at once, or the batch is 2,000 copies of one seller's reposted listing
RandomA uniform sample of the poolThe control arm; also the only unbiased sample

Where active learning bites back

It can work well. Settles' survey shows a text classifier separating baseball from hockey posts that reaches 81% accuracy after 30 uncertainty-chosen labels, against 73% after 30 random ones (logistic regression, 10-fold cross-validation), and the active curve stays above the random one for all of the first 100 labels. The survey is equally clear about the failure cases.

  • Outliers look uncertain. A listing in Welsh, or a blurry photo with no text, sits near 0.5 because the model has never seen anything like it, not because it is on the policy boundary. Labeling a thousand of them teaches little.
  • Picking the top 2,000 by uncertainty from one model snapshot tends to pick near-duplicates, which is why batches need a diversity step; a naive batch can do worse than random.
  • Labels chosen this way are a biased sample of the marketplace. Measuring accuracy on them tells you how the model does on its hardest cases, not on what buyers see. Loftmarket keeps 2,000 uniformly random listings labeled by experts, never trains on them, and measures the model there.
  • The labelers are not a perfect oracle. The items the model finds hardest are often the ones people find hardest too, so active batches have lower agreement and need more repeat labels. Mixing about 20% random items into each weekly batch keeps the training set anchored to the real distribution.

Operating rules

Show the pre-label last
If the labeler sees 'model says: allowed' before deciding, they tend to agree with it, and the model's mistakes come back as labels. Ask first, reveal after; or run a slice with the pre-label hidden and compare the answers to measure the pull.
Protect the labelers
Weapons and violent listings are unpleasant to look at all day. Cap exposure per shift, rotate people across queues and provide support; the review-queues topic in content moderation covers this in depth.
Redact before the vendor sees it
Listings carry phone numbers, addresses and faces. Strip or blur them before tasks leave the company, and put retention and access limits in the vendor contract (see privacy-and-compliance, consent and retention).
Record who, which version, when
Every label stores the labeler, the guideline version and a timestamp. When a labeler is found to be failing gold checks, or version 7 turns out wrong, the affected labels can be pulled in one query instead of discarding the whole batch.
Was this section helpful?

Trade-offs.

The chosen option is first; the others stay visible so the reasoning can be checked.

01
How many labels each listing gets
Chosen:One label, with selective repeats and expert adjudication
  • Pro:About $6,700 for 40,000 listings, roughly 55% of the three-labels plan
  • Pro:Repeats go where the first label or the model is shaky, which is where they change the answer
  • Pro:Gold checks catch a drifting labeler within days
Downside we accept:
  • Con:An easy-looking listing labeled wrong by one person, with the model agreeing, never gets a second look
  • Con:Needs a calibrated model and per-labeler gold accuracy before it works
Ruled out:Three labels on everything

About $12,300 for the same 40,000 listings; Most of the extra judgments confirm answers nobody doubted; Does nothing when labelers share a misunderstanding of the guideline

Ruled out:Experts only

40,000 × 90 s = 1,000 expert hours, about $60,000 and months of a small team's time; Pulls experts away from the guideline and the evaluation set

Ruled out:One label, no checks

No way to know how wrong the labels are; A weak or careless labeler contaminates the whole set

02
Which listings get labeled
Chosen:Uncertainty with diversity, plus a 20% random slice
  • Pro:Spends most of each batch where the model is confused
  • Pro:Diversity stops a batch filling with near-duplicates
  • Pro:The random slice keeps the training data close to the real mix
Downside we accept:
  • Con:Needs the model to score the whole pool every week
  • Con:Hard items have lower agreement, so more of them need repeat labels
Ruled out:Random only

With about 3% of listings prohibited, most of each batch is easy allowed listings the model already gets right

Ruled out:Uncertainty only

Pulls in outliers and near-duplicates; The training set drifts away from what buyers actually see

When people are the wrong tool

Human labels are the right answer when the question is a judgment and you can afford a few thousand of them a week. They are the wrong first move when you need labels on day one or ten times more than the budget allows; then heuristics and a label model get you started, as the weak-supervision topic shows, with a small hand-labeled set kept for checking them. When a label source already exists but is known to be unreliable, such as the category sellers pick for their own listings, the job is finding and fixing the wrong ones, which is the label-noise topic. Human labeling is what both of those still lean on for their test sets. And before paying for any of it, Zinkevich's Rules of Machine Learning (#27) suggests the cheapest step: have human raters measure how often the unwanted behaviour happens, so you know the problem is big enough to label for.

Was this section helpful?
Builds on this
Label noise
Read next