Weak supervision.
Loftmarket opens a luxury handbags category on Monday and must catch counterfeits from day one, with 2 million listings and not one label. A policy team can write a handful of rough rules in a few days. Weak supervision turns those rules, which overlap, disagree and abstain, into probabilistic labels good enough to train on.
Builds on Human labeling and Label noise.
The idea.
A new category, no labels, and six rules that are each partly right.
Rules first, because there is nothing else
Loftmarket's luxury handbags category goes live with 2 million listings copied in from sellers' other categories. Somewhere between 3% and 8% of them are fakes, the policy team guesses, and nobody has marked which. Hand-labeling enough of them to train a classifier would take weeks, and fakes are rare, so a random batch of 20,000 would hold only about 1,000 counterfeits. The prohibited-item classifier from human labeling does list counterfeits among its 14 categories, but its training labels hold few luxury bags, and telling a good fake from the real thing takes a closer look at stitching and serials than its 20-second judgments allow.
What the policy team does have is knowledge. They know that sellers of fakes use certain words, price far below resale value, and often open an account the week they list. They know that a verified receipt and a long clean selling record point the other way. And the brand partners send takedown notices with photos of listings already proven fake. Each of these hunches becomes a labeling function (LF): a few lines of code that look at one listing and vote counterfeit, vote authentic, or abstain because the rule has nothing to say.
No single LF is good enough to ship. Some fire on a sliver of listings, some are wrong a third of the time, and two of them can vote opposite ways on the same bag. Weak supervision is the machinery that decides, without any ground truth, which LF to trust how much, and then hands a clean-enough training set to an ordinary classifier.
Six labeling functions for counterfeit handbags
| LF | Votes | Fires on (coverage) | Precision on the dev set |
|---|---|---|---|
| kw_replica | Title or description says "replica", "1:1" or "mirror quality" → counterfeit | 2% | 0.92 |
| price_low | Price below 25% of the brand's median resale price → counterfeit | 4% | 0.71 |
| new_seller_brand | Account younger than 7 days and a luxury brand in the title → counterfeit | 5% | 0.45 |
| takedown_hash | A photo matches a brand partner's takedown archive (distant supervision) → counterfeit | 1.5% | 0.97 |
| receipt | Seller uploaded a receipt that passed verification → authentic | 13% | 0.995 |
| trusted_seller | At least 50 completed sales and under 1% disputed → authentic | 24% | 0.99 |
Coverage is the share of the 2 million listings on which an LF casts a vote. Every LF here is one-sided: it only ever votes one way. So the dev number is a precision: of the dev listings where it fires, the share whose expert label matches its vote (all numbers illustrative). Read precision against the 5% base rate. new_seller_brand is right less than half the time, yet a listing it flags is 9 times as likely to be fake as an average listing. receipt at 0.995 is useful only because 0.5% fakes is ten times cleaner than the category's 5%.
At least one LF fires on 41% of listings, about 820,000. On the other 59%, every rule abstains, so rules alone would leave most of the category unjudged. The counterfeit rules overlap heavily, mostly on the same fakes, which is why their catches (coverage × precision, about 8% of listings summed over the four) add up to more than the 5% of listings that are fake.
Six listings through the label matrix
- replica, 1 item, Counterfeit vote or label
- price, 1 item, Counterfeit vote or label
- new, 1 item, Counterfeit vote or label
- hash, 1 item, Abstain
- receipt, 1 item, Abstain
- trusted, 1 item, Abstain
- outcome, 2 items, Not decided yet
- replica, 1 item, Abstain
- price, 1 item, Abstain
- new, 1 item, Abstain
- hash, 1 item, Abstain
- receipt, 1 item, Authentic vote or label
- trusted, 1 item, Authentic vote or label
- outcome, 2 items, Not decided yet
- replica, 1 item, Abstain
- price, 1 item, Counterfeit vote or label
- new, 1 item, Abstain
- hash, 1 item, Abstain
- receipt, 1 item, Abstain
- trusted, 1 item, Authentic vote or label
- outcome, 2 items, Not decided yet
- replica, 1 item, Abstain
- price, 1 item, Abstain
- new, 1 item, Abstain
- hash, 1 item, Abstain
- receipt, 1 item, Abstain
- trusted, 1 item, Abstain
- outcome, 2 items, Not decided yet
- replica, 1 item, Abstain
- price, 1 item, Abstain
- new, 1 item, Counterfeit vote or label
- hash, 1 item, Abstain
- receipt, 1 item, Abstain
- trusted, 1 item, Abstain
- outcome, 2 items, Not decided yet
- replica, 1 item, Abstain
- price, 1 item, Abstain
- new, 1 item, Abstain
- hash, 1 item, Counterfeit vote or label
- receipt, 1 item, Abstain
- trusted, 1 item, Abstain
- outcome, 2 items, Not decided yet
- Counterfeit vote or label
- Authentic vote or label
- Abstain
- Not decided yet
- No label (tie or no votes)
As it starts. 3 steps follow.
How it works.
From rules to a label matrix, from the matrix to probabilities, and from probabilities to a model that ships.
The weak supervision pipeline for counterfeit handbags
Three stages, three different jobs
Applying the LFs is a batch job: run six functions over 2 million listings and store a sparse matrix of votes. The label model never sees a listing's photos or text, only that matrix. Its job is to answer two questions with no ground truth at all: how accurate is each LF, and given a row of votes, how likely is this listing to be fake? The end model is a normal classifier that learns from those probabilities and the listing's full features.
The first question sounds impossible, and it is the heart of the method. The trick is that LFs which are each right more often than chance agree with each other more than chance would predict, and how much each pair agrees depends on how accurate both members are. Given enough overlap, the agreement rates pin the accuracies down. The idea is older than weak supervision: Dawid and Skene (1979) estimated each human rater's error rates from the raters' votes alone, with no answer key. A label model does the same for programs instead of people (the human labeling topic covers the original).
Estimating LF accuracy with no labels (the triplet trick)
- Votes and the true class
- +1 counterfeit, −1 authenticWrite each LF's vote as λ and the unknown class as y.
- Three LFs that vote both ways
- serial_format, price_band, logo_checkTwo-sided LFs, for this example only. Each is assumed to be equally accurate on both classes, and their errors independent once y is known.
- M12 = average of λ1·λ2 where both fire
- 0.16LF1 and LF2 agree 58% of the time: 2 × 0.58 − 1
- M13 = average of λ1·λ3
- 0.32agree 66% of the time
- M23 = average of λ2·λ3
- 0.08agree 54% of the time
- Why the moments factory² = 1, so λi·λj = (λi·y)(λj·y); independence given y makes the average a_i · a_j, with a_i = average of λi·y = 2·acc_i − 1M12 = a1a2, M13 = a1a3, M23 = a2a3from Votes and the true class and Three LFs that vote both ways
- Solve for LF1's qualitya1 = √(M12 × M13 ÷ M23) = √(0.16 × 0.32 ÷ 0.08) = √0.64a1 = 0.8 → acc1 = (1 + 0.8) ÷ 2 = 0.90from Why the moments factor, M12 = average of λ1·λ2 where both fire, M13 = average of λ1·λ3 and M23 = average of λ2·λ3
- Then LF2a2 = M12 ÷ a1 = 0.16 ÷ 0.8a2 = 0.2 → acc2 = 0.60from Solve for LF1's quality and M12 = average of λ1·λ2 where both fire
- Then LF3a3 = M13 ÷ a1 = 0.32 ÷ 0.8a3 = 0.4 → acc3 = 0.70from Solve for LF1's quality and M13 = average of λ1·λ3
- Check against the third momenta2 × a3 = 0.2 × 0.40.08 = M23 ✓from Then LF2, Then LF3 and M23 = average of λ2·λ3
- Agreement rates alone reveal each LF's accuracy, up to a sign; assuming every LF beats a coin flip picks the positive root.
- This closed-form triplet idea is what FlyingSquid (Fu et al. 2020) uses, about 170× faster on average than earlier label models; Snorkel fits a similar model by gradient-based optimisation (Ratner et al. 2016, 2017).
- It needs overlap. An LF that never fires alongside two others cannot be placed, which is why the dev set still matters.
One listing, three disagreeing votes
- LF1 (acc 0.90)
- votes counterfeit
- LF2 (acc 0.60)
- votes authentic
- LF3 (acc 0.70)
- votes authentic
- Share of listings that are counterfeit
- 5%illustrative; the label model can estimate it too
- Weight of a vote: log-odds of its accuracy, ln(acc ÷ (1 − acc))ln 9, ln 1.5, ln 2.332.20, 0.41, 0.85from LF1 (acc 0.90), LF2 (acc 0.60) and LF3 (acc 0.70)
- Unweighted majority vote1 counterfeit vs 2 authenticauthenticfrom LF1 (acc 0.90), LF2 (acc 0.60) and LF3 (acc 0.70)
- Weighted, starting from a 50/50 prior+2.20 − 0.41 − 0.85 = +0.94; P = 1 ÷ (1 + e^−0.94)P(counterfeit) = 0.72from Weight of a vote: log-odds of its accuracy, ln(acc ÷ (1 − acc))
- Weighted, starting from the 5% priorln(0.05 ÷ 0.95) = −2.94; −2.94 + 0.94 = −2.00; P = 1 ÷ (1 + e^2.00)P(counterfeit) = 0.12from Weighted, starting from a 50/50 prior and Share of listings that are counterfeit
- The label model outputs a probability, not a vote. Class balance moves it as much as the LF weights do.
- This is a naive-Bayes combination under the same independence assumption as the triplet trick. Real label models also learn how often each LF fires and can model LFs that are correlated.
- The end model trains on these soft labels, so a listing at 0.12 teaches it much less than one at 0.99.
Weighting the one-sided rules from their dev precision
- Share of listings that are counterfeit
- 5%Also the rate at which the dev precisions were measured.
- price_low precision (votes counterfeit)
- 0.71
- trusted_seller precision (votes authentic)
- 0.99So 1% of the listings it vouches for are fake.
- Starting log-odds of counterfeitln(0.05 ÷ 0.95)−2.94 (odds 1 in 19)from Share of listings that are counterfeit
- price_low weight: its log-odds minus the prior's, since its precision already contains the base rateln(0.71 ÷ 0.29) − (−2.94) = 0.90 + 2.94+3.84 (odds × 46)from price_low precision (votes counterfeit) and Starting log-odds of counterfeit
- trusted_seller weight, from the 1% of fakes among its votesln(0.01 ÷ 0.99) − (−2.94) = −4.60 + 2.94−1.65 (odds ÷ 5.2)from trusted_seller precision (votes authentic) and Starting log-odds of counterfeit
- Listing L3: price_low says counterfeit, trusted_seller says authentic−2.94 + 3.84 − 1.65 = −0.75; P = 1 ÷ (1 + e^0.75)P(counterfeit) = 0.32from Starting log-odds of counterfeit, price_low weight: its log-odds minus the prior's, since its precision already contains the base rate and trusted_seller weight, from the 1% of fakes among its votes
- Check: a lone price_low vote−2.94 + 3.84 = +0.90; P = 1 ÷ (1 + e^−0.90)0.71, its own precision ✓from Starting log-odds of counterfeit and price_low weight: its log-odds minus the prior's, since its precision already contains the base rate
- A one-sided rule's dev number is a precision, and precision already includes the base rate. Adding the prior again on top of ln(precision ÷ (1 − precision)) counts it twice; subtract it from the weight instead.
- The two-sided example above is different: there accuracy is measured separately on each class, so its log-odds is the weight and the prior is added once.
- This sketch treats an abstain as saying nothing. A full label model also learns how often each rule fires on each class, and a rule that catches most fakes makes its silence mildly reassuring.
The same rule in a category with fewer fakes
- takedown_hash
- price_low
Data
| Share of listings that are counterfeit (%) | takedown_hash | price_low |
|---|---|---|
| 0.1 | 0.381 | 0.044 |
| 0.2 | 0.552 | 0.085 |
| 0.5 | 0.755 | 0.189 |
| 1 | 0.861 | 0.32 |
| 2 | 0.926 | 0.487 |
| 5 | 0.97 | 0.71 |
| 10 | 0.986 | 0.838 |
| 20 | 0.994 | 0.921 |
- P = 0.5: P(counterfeit) = 0.5
- At 5: Handbags (5%)
Why train a model at all
The label model can only label rows where some LF fired: 820,000 of the 2 million listings. The end model is trained on those rows, but it learns from everything in them: the photos, the wording, the seller's history. So it picks up signals no rule mentions, such as a counterfeiter who writes "rep" or "UA quality" instead of "replica", or stitching that looks like the fakes in the takedown matches, and it scores the 59% of listings where every rule abstained.
The end model uses a noise-aware loss: its training target is the probability (0.32, 0.45, 0.97), not a hard 0 or 1, so confident rows count for more. In the Snorkel evaluation (Ratner et al. 2017), end models trained this way averaged 132% better than models trained by distant supervision, an earlier heuristic baseline, and came within 3.60% on average of models trained on large hand-labeled sets.
The label model is not always worth it. The same paper notes that when LFs rarely overlap, there is almost nothing to weigh, and when many LFs vote on every row, an unweighted majority is already accurate; the learned weights help most in between. Measure majority vote on the dev set as a baseline before trusting the fancier answer.
In practice.
Writing LFs like code, borrowing what the company already knows, and the other ways out of zero labels.
LF report after the third iteration
| LF | Coverage | Overlaps | Conflicts | Dev precision | Action |
|---|---|---|---|---|---|
| kw_replica | 2% | 1.4% | 0.1% | 0.92 | Keep; added a negation guard after dev errors on "not a replica" and "no replicas, 100% genuine" |
| price_low | 4% | 2.6% | 1.0% | 0.71 | Narrow: skip listings marked "for parts" or "damaged", which are cheap and genuine |
| new_seller_brand | 5% | 3.0% | 1.2% | 0.45 | Split by brand: 0.72 on the three most-faked brands (2% coverage), 0.27 on the rest (3%); keep only the first part |
| takedown_hash | 1.5% | 1.0% | 0.1% | 0.97 | Keep; refresh the archive weekly |
| receipt | 13% | 3.4% | 0.9% | 0.995 | Keep; audit a sample of verified receipts monthly for forgeries |
| trusted_seller | 24% | 5.6% | 1.5% | 0.99 | Keep; exclude accounts with a password reset in the last 30 days (takeovers) |
LFs are code, so treat them like code
Coverage is where an LF fires, overlaps where it fires alongside another LF, and conflicts where they disagree (all as a share of the 2 million listings, illustrative). An LF with low precision is not automatically bad: new_seller_brand at 0.45 is wrong more often than right, but a listing it flags is still 9 times as likely to be fake as an average one, so it earns a real weight. What an LF must not be is wrong in a systematic way the model cannot see, like price_low firing on every broken bag.
So LFs live in a repository, are reviewed like any change, carry a version, and are tested on the dev set every time they change. The dev set is 600 listings the policy experts labeled: 300 at random and 50 drawn from wherever each LF fires, so every LF has enough examples to measure. Labeling a random 600 alone would give takedown_hash about 9 examples, too few to say anything.
How much 50 dev examples tell you about one LF
- Dev listings where the LF fires
- 50
- price_low measured precision
- 0.71
- takedown_hash measured precision
- 0.97
- 95% margin for price_low1.96 × √(0.71 × 0.29 ÷ 50) = 1.96 × 0.064± 0.13from Dev listings where the LF fires and price_low measured precision
- 95% range for takedown_hash (Wilson interval; the ± formula above breaks down this close to 1)(0.97 + 1.96² ÷ 100 ± 1.96 × √(0.97 × 0.03 ÷ 50 + 1.96² ÷ 10,000)) ÷ (1 + 1.96² ÷ 50)0.88 to 0.99from Dev listings where the LF fires and takedown_hash measured precision
- What that range does to takedown_hash's likelihood ratio at a 5% base rate(0.88 ÷ 0.12) ÷ (0.05 ÷ 0.95) = 139; (0.97 ÷ 0.03) ÷ (0.05 ÷ 0.95) = 614a factor of about 4.4from 95% range for takedown_hash (Wilson interval; the ± formula above breaks down this close to 1)
- Dev precision is a rough guide for choosing which LFs to fix, not a precise weight. The label model's own estimates use hundreds of thousands of unlabeled rows.
- The naive ± margin would claim 0.92 to 1.02 for takedown_hash, which is impossible above 1 and too narrow below. Near 0 or 1, use the Wilson interval or label more examples.
- Fifty examples cannot tell receipt's 0.995 from 0.99: both would usually show 50 out of 50. Precise weights for near-perfect rules need far more labels or the label model's estimate.
LFs can use what the model never will
LFs run offline, over history, so they can read signals that do not exist when a listing goes live: a chargeback filed six weeks after the sale, a moderator's free-text note, a brand's takedown archive the company may not expose to a serving path. The model that ships reads only what is there at listing time: photos, text, price and seller features. At Google, Snorkel DryBell used exactly this split, turning internal resources that could not be served into servable classifiers, with an average 52% gain from those resources and quality comparable to classifiers trained on tens of thousands of hand labels (Bach et al. 2019).
At Loftmarket the next LFs to write are chargeback_fake (buyer refunded with reason "not authentic") and brand_confirmed (brand partner confirms a report). Both arrive weeks late, which is fine for labeling last month's listings and useless for scoring today's.
What the Snorkel papers measured
Other ways out of zero labels
Before trusting the weak labels
Trade-offs.
What day one costs each way, the choice Loftmarket made, and when weak supervision is the wrong tool.
Day one: hand labels or rules
- Listings hand-labeled for training
- 20,000
- Time to judge one handbag from its photos
- 60 sslower than the 20 s used for prohibited items, because authenticity needs close looks at stitching, logos and serials (illustrative)
- Labels per listing
- 2
- Vendor labeler cost
- $15 per hour
- Policy expert cost
- $60 per hour
- Vendor team
- 5 labelers × 30 h a week
- Share of listings that are fake
- 5%
- LF writing and debugging
- 2 policy experts × 3 days of 8 h
- Vendor hours20,000 × 2 × 60 s ÷ 3,600≈ 667 hfrom Listings hand-labeled for training, Time to judge one handbag from its photos and Labels per listing
- Vendor cost667 h × $15≈ $10,000from Vendor hours and Vendor labeler cost
- Expert tie-breaks on 10% of listings at 90 s2,000 × 90 s = 50 h; 50 h × $60$3,000from Listings hand-labeled for training and Policy expert cost
- Hand-labeling plan$10,000 + $3,000; 667 h ÷ 150 h a week≈ $13,000 and 4.4 weeksfrom Vendor cost, Expert tie-breaks on 10% of listings at 90 s and Vendor team
- Counterfeits in that set20,000 × 5%≈ 1,000from Listings hand-labeled for training and Share of listings that are fake
- Weak supervision: 1,000 dev and test listings labeled by experts at 90 s1,000 × 90 s = 25 h; 25 h × $60$1,500from Policy expert cost
- LF writing2 × 3 × 8 h = 48 h; 48 h × $60$2,880from LF writing and debugging and Policy expert cost
- Weak supervision plan$1,500 + $2,880; labeling and LF writing run in parallel≈ $4,400 and about a weekfrom Weak supervision: 1,000 dev and test listings labeled by experts at 90 s and LF writing
- Weak supervision is not free: expert time writing and debugging LFs is two thirds of its cost. It moves the spend from labeling each item to encoding what experts know once, and comes to about a third of the hand-labeling plan.
- The hand-labeled set would be clean but small, with about 1,000 fakes; the weak set is noisy but covers 820,000 listings.
- Pro:Soft labels for 820,000 listings within about a week, for about $4,400
- Pro:Rules are readable, so policy can audit why a listing was labeled fake
- Pro:Can use late or private signals (takedowns, chargebacks) the served model never sees
- Con:Needs people who can write and debug LFs
- Con:Correlated or systematically wrong LFs mislead the label model silently
- Con:The model learns mostly from the 41% of listings the rules cover, and inherits their blind spots
About $13,000 and 4.4 weeks before the first model; Only about 1,000 counterfeits to learn from
A per-call cost and latency on every new listing; Accuracy unknown until a test set exists anyway; Cannot use photos unless the model is multimodal, and cannot use seller history it is not given
Covers 41% of listings; the rest get no score; Conflicts must be settled by hand-written priorities; Counterfeiters route around a fixed rule within days
When weak supervision fits
| Situation | Fits | Poor fit |
|---|---|---|
| Domain knowledge | Experts can state rules of thumb, even rough ones | Nobody can say why an item is positive (subtle visual quality, taste) |
| Unlabeled data | Large pool, hundreds of thousands of items or more | A few thousand items: hand-label them instead |
| Label cost | Labels need scarce experts, or data is private and cannot go to a vendor | Labels are cheap and quick with a crowd |
| Existing records | Takedown lists, knowledge bases, past decisions to borrow from | Nothing to match against |
| Rule overlap | Several LFs fire on the same items, so accuracies can be estimated | Each LF fires on its own slice; nothing to compare |
What goes wrong
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| An LF is systematically wrong on a slice2Labeling functions | price_low labels every cheap, damaged but genuine bag as fake, and the end model learns that damage means counterfeit | Dev-set errors grouped by listing attributes; spot-check the highest-weight wrong labels | Narrow the LF, add a guard, and re-run the label model | The rest of the LFs still vote; accuracy on that slice drops |
| Two LFs are copies of each other4Label model | Their agreement inflates both accuracy estimates, and their evidence is counted twice | Pairwise overlap near their full coverage with almost no conflicts | Merge them, or give the label model the dependency | Probabilities are overconfident but mostly point the right way |
| Sellers adapt to the rules5Counterfeit classifier | kw_replica's coverage falls; new fakes look like the uncovered 59% | Coverage per LF by week; share of expert-confirmed fakes the model missed | Write new LFs from review-queue findings, retrain weekly, move to human labels as they build up | The model keeps catching old patterns and misses new ones |
| The test set leaks into LF development6Hand-labeled dev and test sets | The reported score is optimistic because rules were tuned on it | Access logs on the test set; a score gap between test and a fresh expert sample | Keep test labels in a separate store that LF authors cannot query | The model may be fine; nobody knows how fine |