Human labeling.
Loftmarket's prohibited-item classifier can't learn from clicks: buyers happily click a knife listing or a recalled crib. Its labels come from people reading a policy. Every design choice here trades money per label against how much each label can be trusted.
The idea.
A policy, a pile of listings, and people paid to decide which is which.
Why this model needs people
Loftmarket is a fictional second-hand marketplace for phones, sneakers, handbags and furniture. Its search ranker learns from clicks and purchases, because a click really does say something about relevance. Its prohibited-item classifier cannot. A hunting knife, a recalled baby crib or a fake designer bag gets clicked as eagerly as anything else, often more. The question the model must answer is 'does this listing break our policy?', and no amount of user behaviour answers it.
So someone has to read the policy, look at a listing and decide. The policy lists 14 prohibited categories. Loftmarket's trust-and-safety team has a few policy experts who know it cold; they cost about $60 an hour and there are not many of them. A labeling vendor supplies dozens of trained people at about $15 an hour who are quick but make more mistakes. The whole job is spending expert hours only where they change the answer.
Three levers control cost and quality: who labels an item, how many people label it, and which items get labeled at all. The rest of this topic pulls each one.
Labeling 40,000 listings
- Listings to label
- 40,000
- Time per vendor judgment
- 20 sabout 180 judgments an hour (illustrative)
- Vendor labeler cost
- $15 per hour
- Gold questions mixed into every queue
- 5% extra judgmentslistings with a known expert answer, hidden in the queue to score labelers
- Items that get two extra labels
- 20%first label low-confidence or disagrees with the model
- Items that end in expert adjudication
- 3% (1,200), 90 s eachan expert settles listings the labelers could not agree on
- Policy expert cost
- $60 per hour
- Vendor judgments, selective repeats40,000 × (1 + 0.2 × 2) × 1.0558,800from Listings to label, Items that get two extra labels and Gold questions mixed into every queue
- Vendor hours58,800 × 20 s ÷ 3,600≈ 327 hfrom Vendor judgments, selective repeats and Time per vendor judgment
- Vendor cost327 h × $15≈ $4,900from Vendor hours and Vendor labeler cost
- Expert adjudication1,200 × 90 s = 30 h; 30 h × $60$1,800from Items that end in expert adjudication and Policy expert cost
- Total, selective repeats$4,900 + $1,800≈ $6,700from Vendor cost and Expert adjudication
- Alternative, three labels on every item40,000 × 3 × 1.05 = 126,000 judgments = 700 h × $15 = $10,500; + $1,800≈ $12,300from Listings to label, Gold questions mixed into every queue, Time per vendor judgment, Vendor labeler cost and Expert adjudication
- Asking for extra labels only where the first one looks shaky costs about 55% of labeling everything three times, and most items were never in doubt anyway.
- Sheng, Provost and Ipeirotis (2008) reached the same conclusion in general: choosing which items to relabel usually beats relabeling all of them.
- Expert time is under a third of the bill but is the scarce part; 30 hours is roughly one expert's week.
Who can supply the label
| Source | Cost per label | Accuracy | Throughput | Use at Loftmarket |
|---|---|---|---|---|
| Policy experts | Highest (~$1.50 at 90 s and $60/h) | Highest; they wrote the rules | Tens of items an hour, a few people | Adjudication, the evaluation set, guideline examples |
| Vendor team with training | ~$0.08 at 20 s and $15/h | Good on clear cases, weaker on edge cases | Thousands of items a day | Most training labels |
| Open crowd (anonymous workers) | Lowest | Varies widely; needs redundancy and gold checks | Very high | Not used: listings can contain personal data, and the policy needs training |
| Users (the report button) | Free | Low precision: people report rivals and bad sellers | Only where users look | A signal that puts listings into the queue, not a label |
| Model pre-label, human confirms | Cheapest human option per item | Inherits the model's blind spots through anchoring | Highest per labeler | Only with the pre-label hidden until the labeler answers |
How many amateurs make an expert?
Snow, O'Connor, Jurafsky and Ng (EMNLP 2008) paid Mechanical Turk workers to redo five language-annotation tasks that experts had already labeled. On the affect-recognition task (seven emotion and valence ratings), about four non-expert labels per item, averaged together, matched one expert's agreement with the other experts. Some subtasks needed only two, and one (fear) never got there. At 2008 prices that meant at least 875 expert-equivalent labels per US dollar. Treat the prices as history, and the ratio as a result for that kind of task: simple judgments with clear instructions. A policy call about whether a replica prop gun counts as a weapon is not that kind of task, which is why Loftmarket keeps experts for the hard cases.
How it works.
The guideline, measuring agreement, combining several labels, and what happens to one listing in the queue.
The guideline is the spec
Labelers can only be as consistent as the document they follow. Loftmarket's guideline has these parts for each of the 14 categories.
Agreement beyond chance (Cohen's κ)
- Listings both vendor labelers judged
- 200from an active-learning batch, which is rich in suspicious listings; that is why both call far more prohibited than the marketplace's roughly 3%
- Listings where they gave the same answer
- 170
- Labeler A says prohibited on
- 30%
- Labeler B says prohibited on
- 25%
- Rare-class case: both say prohibited on
- 5%, raw agreement 96%
- Observed agreement, p_o170 ÷ 2000.85from Listings both vendor labelers judged and Listings where they gave the same answer
- Agreement expected by chance, p_e0.30 × 0.25 + 0.70 × 0.750.60from Labeler A says prohibited on and Labeler B says prohibited on · Both say yes by luck, plus both say no by luck, if each answered at random at their own rate.
- Cohen's κ(0.85 − 0.60) ÷ (1 − 0.60)0.625from Observed agreement, p_o and Agreement expected by chance, p_e
- Chance agreement for the rare class0.05² + 0.95²0.905from Rare-class case: both say prohibited on
- κ for the rare class(0.96 − 0.905) ÷ (1 − 0.905)≈ 0.58from Rare-class case: both say prohibited on and Chance agreement for the rare class
- κ answers the question raw agreement cannot: how much better than two people guessing at their own rates? Cohen defined it in 1960 as (p_o − p_e) ÷ (1 − p_e).
- If each labeler says allowed on 95% of listings, they would agree about 90% of the time by chance alone (0.95² + 0.05² = 0.905), so 96% raw agreement is only a κ of about 0.58.
- Report κ per category, not one number for the whole guideline; with more than two labelers per item, Krippendorff's α plays the same role.
Accuracy of a majority vote
- p = 0.6
- p = 0.7
- p = 0.8
- p = 0.9
Data
| Labels per item | p = 0.6 | p = 0.7 | p = 0.8 | p = 0.9 |
|---|---|---|---|---|
| 1 | 0.6 | 0.7 | 0.8 | 0.9 |
| 3 | 0.648 | 0.784 | 0.896 | 0.972 |
| 5 | 0.683 | 0.837 | 0.942 | 0.991 |
| 7 | 0.71 | 0.874 | 0.967 | 0.997 |
| 9 | 0.733 | 0.901 | 0.98 | 0.999 |
- At 3: 0.784
Better than one person, one vote
The arithmetic for three labelers at p = 0.7: the majority is right if all three are right (0.7³ = 0.343) or exactly two are (3 × 0.7² × 0.3 = 0.441), so 0.784 in total. The same sum at p = 0.4 gives 0.352, lower than one label on its own. Voting amplifies whatever side of 50% the labelers sit on, so it rescues decent labelers and sinks bad ones.
Plain majority vote also treats a careful labeler and a careless one the same. Dawid and Skene (1979) fixed that: estimate a confusion matrix for every labeler (how often each true class gets each answer) together with the likely true label of every item, alternating the two steps by expectation maximization until they settle. A labeler who calls every replica 'counterfeit' then counts for less on replicas, and one who is reliable on weapons keeps full weight there. The same idea, applied to heuristic rules instead of people, is the label model in the weak-supervision topic.
Gold questions measure labelers directly instead of inferring it. Five per cent of every queue is listings the experts have already decided, indistinguishable from the rest. Loftmarket's rule: a labeler whose rolling gold accuracy drops under 85% is paused and retrained, and their labels since the last good check are sent for another look.
One listing's path through the label queue
States of6Label aggregator
What the label aggregator does with a single listing. Most listings leave after one label; the rest collect more votes, and the ones that still split go to an expert. A listing the policy does not cover sends the guideline, not the labeler, back for repair.
| From → To | Event | Guard | Action | Actor |
|---|---|---|---|---|
| Queued → One label | labeled | Vendor labeler | ||
| One label → Agreed | confident | model agrees, gold ≥ 90% | Aggregator | |
| One label → Needs 2 more labels | low confidence or model disagrees | Aggregator | ||
| Needs 2 more labels → Agreed | 3 labels | ≥ 2 match | Aggregator | |
| Needs 2 more labels → Disputed | no majority or any 'unsure' | Aggregator | ||
| Disputed → Adjudicated by an expert | expert decides | Policy expert | ||
| Disputed → Guideline gap | expert: policy is silent | Policy expert | ||
| Guideline gap → Queued | guideline v+1 published | relabel | Policy expert |
- Queuedstart
- Agreedend
- Adjudicated by an expertend
- Guideline gaperror
- Sent back to the policy owners
In practice.
Letting the model pick what gets labeled, and closing the loop with production.
The labeling loop at Loftmarket
What goes round
Each Monday the classifier scores every unlabeled listing. The picker chooses 2,000 of them, the vendor team labels them during the week, the aggregator turns judgments into labels, and the classifier is retrained on the larger set. Next week's picks come from the new model, so each batch goes after whatever confuses the model now rather than what confused it a month ago. Settles (2010) calls this pool-based active learning: the learner sees a large unlabeled pool and asks for labels on the items it expects to learn the most from.
How the picker chooses (after Settles 2010)
| Strategy | Picks | Watch out |
|---|---|---|
| Least confident | Listings whose most likely class has the lowest probability | Ignores how the rest of the probability is spread |
| Margin | Listings with the smallest gap between the top two classes | Good for telling similar categories apart (weapon vs replica) |
| Entropy | Listings with the most spread-out distribution; on the category head that is 15 classes (14 categories plus allowed) | On the yes/no prohibited score all three of these pick the same item: the one scored near 0.5 |
| Query by committee | Listings on which several models trained on different samples disagree most | Costs several models; the committee must be genuinely varied |
| Diversity (batch mode) | Cluster the uncertain listings and take a few from each cluster | Needed when picking 2,000 at once, or the batch is 2,000 copies of one seller's reposted listing |
| Random | A uniform sample of the pool | The control arm; also the only unbiased sample |
Where active learning bites back
It can work well. Settles' survey shows a text classifier separating baseball from hockey posts that reaches 81% accuracy after 30 uncertainty-chosen labels, against 73% after 30 random ones (logistic regression, 10-fold cross-validation), and the active curve stays above the random one for all of the first 100 labels. The survey is equally clear about the failure cases.
- Outliers look uncertain. A listing in Welsh, or a blurry photo with no text, sits near 0.5 because the model has never seen anything like it, not because it is on the policy boundary. Labeling a thousand of them teaches little.
- Picking the top 2,000 by uncertainty from one model snapshot tends to pick near-duplicates, which is why batches need a diversity step; a naive batch can do worse than random.
- Labels chosen this way are a biased sample of the marketplace. Measuring accuracy on them tells you how the model does on its hardest cases, not on what buyers see. Loftmarket keeps 2,000 uniformly random listings labeled by experts, never trains on them, and measures the model there.
- The labelers are not a perfect oracle. The items the model finds hardest are often the ones people find hardest too, so active batches have lower agreement and need more repeat labels. Mixing about 20% random items into each weekly batch keeps the training set anchored to the real distribution.
Operating rules
Trade-offs.
The chosen option is first; the others stay visible so the reasoning can be checked.
- Pro:About $6,700 for 40,000 listings, roughly 55% of the three-labels plan
- Pro:Repeats go where the first label or the model is shaky, which is where they change the answer
- Pro:Gold checks catch a drifting labeler within days
- Con:An easy-looking listing labeled wrong by one person, with the model agreeing, never gets a second look
- Con:Needs a calibrated model and per-labeler gold accuracy before it works
About $12,300 for the same 40,000 listings; Most of the extra judgments confirm answers nobody doubted; Does nothing when labelers share a misunderstanding of the guideline
40,000 × 90 s = 1,000 expert hours, about $60,000 and months of a small team's time; Pulls experts away from the guideline and the evaluation set
No way to know how wrong the labels are; A weak or careless labeler contaminates the whole set
- Pro:Spends most of each batch where the model is confused
- Pro:Diversity stops a batch filling with near-duplicates
- Pro:The random slice keeps the training data close to the real mix
- Con:Needs the model to score the whole pool every week
- Con:Hard items have lower agreement, so more of them need repeat labels
With about 3% of listings prohibited, most of each batch is easy allowed listings the model already gets right
Pulls in outliers and near-duplicates; The training set drifts away from what buyers actually see
When people are the wrong tool
Human labels are the right answer when the question is a judgment and you can afford a few thousand of them a week. They are the wrong first move when you need labels on day one or ten times more than the budget allows; then heuristics and a label model get you started, as the weak-supervision topic shows, with a small hand-labeled set kept for checking them. When a label source already exists but is known to be unreliable, such as the category sellers pick for their own listings, the job is finding and fixing the wrong ones, which is the label-noise topic. Human labeling is what both of those still lean on for their test sets. And before paying for any of it, Zinkevich's Rules of Machine Learning (#27) suggests the cheapest step: have human raters measure how often the unwanted behaviour happens, so you know the problem is big enough to label for.