Data collection and labelingWeak supervision

100%

Weak supervision.

Loftmarket opens a luxury handbags category on Monday and must catch counterfeits from day one, with 2 million listings and not one label. A policy team can write a handful of rough rules in a few days. Weak supervision turns those rules, which overlap, disagree and abstain, into probabilistic labels good enough to train on.

Advanced29 minUpdated 1 Oct 2026

Builds on Human labeling and Label noise.

The idea.

A new category, no labels, and six rules that are each partly right.

Rules first, because there is nothing else

Loftmarket's luxury handbags category goes live with 2 million listings copied in from sellers' other categories. Somewhere between 3% and 8% of them are fakes, the policy team guesses, and nobody has marked which. Hand-labeling enough of them to train a classifier would take weeks, and fakes are rare, so a random batch of 20,000 would hold only about 1,000 counterfeits. The prohibited-item classifier from human labeling does list counterfeits among its 14 categories, but its training labels hold few luxury bags, and telling a good fake from the real thing takes a closer look at stitching and serials than its 20-second judgments allow.

What the policy team does have is knowledge. They know that sellers of fakes use certain words, price far below resale value, and often open an account the week they list. They know that a verified receipt and a long clean selling record point the other way. And the brand partners send takedown notices with photos of listings already proven fake. Each of these hunches becomes a labeling function (LF): a few lines of code that look at one listing and vote counterfeit, vote authentic, or abstain because the rule has nothing to say.

No single LF is good enough to ship. Some fire on a sliver of listings, some are wrong a third of the time, and two of them can vote opposite ways on the same bag. Weak supervision is the machinery that decides, without any ground truth, which LF to trust how much, and then hands a clean-enough training set to an ordinary classifier.

Six labeling functions for counterfeit handbags

LFVotesFires on (coverage)Precision on the dev set
kw_replicaTitle or description says "replica", "1:1" or "mirror quality" → counterfeit2%0.92
price_lowPrice below 25% of the brand's median resale price → counterfeit4%0.71
new_seller_brandAccount younger than 7 days and a luxury brand in the title → counterfeit5%0.45
takedown_hashA photo matches a brand partner's takedown archive (distant supervision) → counterfeit1.5%0.97
receiptSeller uploaded a receipt that passed verification → authentic13%0.995
trusted_sellerAt least 50 completed sales and under 1% disputed → authentic24%0.99

Coverage is the share of the 2 million listings on which an LF casts a vote. Every LF here is one-sided: it only ever votes one way. So the dev number is a precision: of the dev listings where it fires, the share whose expert label matches its vote (all numbers illustrative). Read precision against the 5% base rate. new_seller_brand is right less than half the time, yet a listing it flags is 9 times as likely to be fake as an average listing. receipt at 0.995 is useful only because 0.5% fakes is ten times cleaner than the category's 5%.

At least one LF fires on 41% of listings, about 820,000. On the other 59%, every rule abstains, so rules alone would leave most of the category unjudged. The counterfeit rules overlap heavily, mostly on the same fakes, which is why their catches (coverage × precision, about 8% of listings summed over the four) add up to more than the 5% of listings that are fake.

Six listings through the label matrix

L1
  1. replica, 1 item, Counterfeit vote or label
  2. price, 1 item, Counterfeit vote or label
  3. new, 1 item, Counterfeit vote or label
  4. hash, 1 item, Abstain
  5. receipt, 1 item, Abstain
  6. trusted, 1 item, Abstain
  7. outcome, 2 items, Not decided yet
L2
  1. replica, 1 item, Abstain
  2. price, 1 item, Abstain
  3. new, 1 item, Abstain
  4. hash, 1 item, Abstain
  5. receipt, 1 item, Authentic vote or label
  6. trusted, 1 item, Authentic vote or label
  7. outcome, 2 items, Not decided yet
L3
  1. replica, 1 item, Abstain
  2. price, 1 item, Counterfeit vote or label
  3. new, 1 item, Abstain
  4. hash, 1 item, Abstain
  5. receipt, 1 item, Abstain
  6. trusted, 1 item, Authentic vote or label
  7. outcome, 2 items, Not decided yet
L4
  1. replica, 1 item, Abstain
  2. price, 1 item, Abstain
  3. new, 1 item, Abstain
  4. hash, 1 item, Abstain
  5. receipt, 1 item, Abstain
  6. trusted, 1 item, Abstain
  7. outcome, 2 items, Not decided yet
L5
  1. replica, 1 item, Abstain
  2. price, 1 item, Abstain
  3. new, 1 item, Counterfeit vote or label
  4. hash, 1 item, Abstain
  5. receipt, 1 item, Abstain
  6. trusted, 1 item, Abstain
  7. outcome, 2 items, Not decided yet
L6
  1. replica, 1 item, Abstain
  2. price, 1 item, Abstain
  3. new, 1 item, Abstain
  4. hash, 1 item, Counterfeit vote or label
  5. receipt, 1 item, Abstain
  6. trusted, 1 item, Abstain
  7. outcome, 2 items, Not decided yet
  • Counterfeit vote or label
  • Authentic vote or label
  • Abstain
  • Not decided yet
  • No label (tie or no votes)
Start

As it starts. 3 steps follow.

Rows are listings L1 to L6, columns the six LFs in table order (new is new_seller_brand, hash is takedown_hash), plus an outcome on the right. Orange votes counterfeit, green votes authentic, dotted gray abstains. Step through to compare an unweighted majority vote with the label model's weighted answer, which also knows that about 5% of listings are fake. Probabilities are computed in the next section.
Was this section helpful?

How it works.

From rules to a label matrix, from the matrix to probabilities, and from probabilities to a model that ships.

The weak supervision pipeline for counterfeit handbags

The weak supervision pipeline for counterfeit handbags. The numbered component cards that follow describe each part.
The weak supervision pipeline for counterfeit handbagsComponents: 1. Unlabeled listings (The 2 million luxury handbag listings in the new category, with title, description, photos, price and seller history, and no labels.), 2. Labeling functions (Six small programs written by the policy team. Each looks at one listing and votes counterfeit, votes authentic, or abstains.), 3. Label matrix (One row per listing, one column per labeling function, holding each vote or an abstain. Mostly abstains.), 4. Label model (Estimates how accurate each labeling function is from their agreements and disagreements alone, then turns each row of votes into a probability of counterfeit.), 5. Counterfeit classifier (The model that ships. Trained on the probabilistic labels, it reads photos, text and seller features, so it also scores listings that no rule fires on.), 6. Hand-labeled dev and test sets (600 dev listings for writing and debugging rules and 400 random test listings for the final measurement, all labeled by policy experts.).

apply every LF

votes and abstains

agreements and conflicts

820k soft labels, noise-aware loss

write and debug LFs (dev only)

final score on the 400 test listings

1Unlabeled listings
2M listings
no labels

2Labeling functions
6 rules
vote or abstain

3Label matrix
2M × 6
mostly abstain

4Label model
LF accuracies
P(counterfeit) per row

5Counterfeit classifier
photos
text
seller features

6Hand-labeled dev and test sets
600 dev
400 test

Three stages, three different jobs

Applying the LFs is a batch job: run six functions over 2 million listings and store a sparse matrix of votes. The label model never sees a listing's photos or text, only that matrix. Its job is to answer two questions with no ground truth at all: how accurate is each LF, and given a row of votes, how likely is this listing to be fake? The end model is a normal classifier that learns from those probabilities and the listing's full features.

The first question sounds impossible, and it is the heart of the method. The trick is that LFs which are each right more often than chance agree with each other more than chance would predict, and how much each pair agrees depends on how accurate both members are. Given enough overlap, the agreement rates pin the accuracies down. The idea is older than weak supervision: Dawid and Skene (1979) estimated each human rater's error rates from the raters' votes alone, with no answer key. A label model does the same for programs instead of people (the human labeling topic covers the original).

Estimating LF accuracy with no labels (the triplet trick)

Assumptions
Votes and the true class
+1 counterfeit, −1 authenticWrite each LF's vote as λ and the unknown class as y.
Three LFs that vote both ways
serial_format, price_band, logo_checkTwo-sided LFs, for this example only. Each is assumed to be equally accurate on both classes, and their errors independent once y is known.
M12 = average of λ1·λ2 where both fire
0.16LF1 and LF2 agree 58% of the time: 2 × 0.58 − 1
M13 = average of λ1·λ3
0.32agree 66% of the time
M23 = average of λ2·λ3
0.08agree 54% of the time
Working
  1. Why the moments factory² = 1, so λi·λj = (λi·y)(λj·y); independence given y makes the average a_i · a_j, with a_i = average of λi·y = 2·acc_i − 1M12 = a1a2, M13 = a1a3, M23 = a2a3from Votes and the true class and Three LFs that vote both ways
  2. Solve for LF1's qualitya1 = √(M12 × M13 ÷ M23) = √(0.16 × 0.32 ÷ 0.08) = √0.64a1 = 0.8 → acc1 = (1 + 0.8) ÷ 2 = 0.90from Why the moments factor, M12 = average of λ1·λ2 where both fire, M13 = average of λ1·λ3 and M23 = average of λ2·λ3
  3. Then LF2a2 = M12 ÷ a1 = 0.16 ÷ 0.8a2 = 0.2 → acc2 = 0.60from Solve for LF1's quality and M12 = average of λ1·λ2 where both fire
  4. Then LF3a3 = M13 ÷ a1 = 0.32 ÷ 0.8a3 = 0.4 → acc3 = 0.70from Solve for LF1's quality and M13 = average of λ1·λ3
  5. Check against the third momenta2 × a3 = 0.2 × 0.40.08 = M23 ✓from Then LF2, Then LF3 and M23 = average of λ2·λ3
What it means
  • Agreement rates alone reveal each LF's accuracy, up to a sign; assuming every LF beats a coin flip picks the positive root.
  • This closed-form triplet idea is what FlyingSquid (Fu et al. 2020) uses, about 170× faster on average than earlier label models; Snorkel fits a similar model by gradient-based optimisation (Ratner et al. 2016, 2017).
  • It needs overlap. An LF that never fires alongside two others cannot be placed, which is why the dev set still matters.

One listing, three disagreeing votes

Assumptions
LF1 (acc 0.90)
votes counterfeit
LF2 (acc 0.60)
votes authentic
LF3 (acc 0.70)
votes authentic
Share of listings that are counterfeit
5%illustrative; the label model can estimate it too
Working
  1. Weight of a vote: log-odds of its accuracy, ln(acc ÷ (1 − acc))ln 9, ln 1.5, ln 2.332.20, 0.41, 0.85from LF1 (acc 0.90), LF2 (acc 0.60) and LF3 (acc 0.70)
  2. Unweighted majority vote1 counterfeit vs 2 authenticauthenticfrom LF1 (acc 0.90), LF2 (acc 0.60) and LF3 (acc 0.70)
  3. Weighted, starting from a 50/50 prior+2.20 − 0.41 − 0.85 = +0.94; P = 1 ÷ (1 + e^−0.94)P(counterfeit) = 0.72from Weight of a vote: log-odds of its accuracy, ln(acc ÷ (1 − acc))
  4. Weighted, starting from the 5% priorln(0.05 ÷ 0.95) = −2.94; −2.94 + 0.94 = −2.00; P = 1 ÷ (1 + e^2.00)P(counterfeit) = 0.12from Weighted, starting from a 50/50 prior and Share of listings that are counterfeit
What it means
  • The label model outputs a probability, not a vote. Class balance moves it as much as the LF weights do.
  • This is a naive-Bayes combination under the same independence assumption as the triplet trick. Real label models also learn how often each LF fires and can model LFs that are correlated.
  • The end model trains on these soft labels, so a listing at 0.12 teaches it much less than one at 0.99.

Weighting the one-sided rules from their dev precision

Assumptions
Share of listings that are counterfeit
5%Also the rate at which the dev precisions were measured.
price_low precision (votes counterfeit)
0.71
trusted_seller precision (votes authentic)
0.99So 1% of the listings it vouches for are fake.
Working
  1. Starting log-odds of counterfeitln(0.05 ÷ 0.95)−2.94 (odds 1 in 19)from Share of listings that are counterfeit
  2. price_low weight: its log-odds minus the prior's, since its precision already contains the base rateln(0.71 ÷ 0.29) − (−2.94) = 0.90 + 2.94+3.84 (odds × 46)from price_low precision (votes counterfeit) and Starting log-odds of counterfeit
  3. trusted_seller weight, from the 1% of fakes among its votesln(0.01 ÷ 0.99) − (−2.94) = −4.60 + 2.94−1.65 (odds ÷ 5.2)from trusted_seller precision (votes authentic) and Starting log-odds of counterfeit
  4. Listing L3: price_low says counterfeit, trusted_seller says authentic−2.94 + 3.84 − 1.65 = −0.75; P = 1 ÷ (1 + e^0.75)P(counterfeit) = 0.32from Starting log-odds of counterfeit, price_low weight: its log-odds minus the prior's, since its precision already contains the base rate and trusted_seller weight, from the 1% of fakes among its votes
  5. Check: a lone price_low vote−2.94 + 3.84 = +0.90; P = 1 ÷ (1 + e^−0.90)0.71, its own precision ✓from Starting log-odds of counterfeit and price_low weight: its log-odds minus the prior's, since its precision already contains the base rate
What it means
  • A one-sided rule's dev number is a precision, and precision already includes the base rate. Adding the prior again on top of ln(precision ÷ (1 − precision)) counts it twice; subtract it from the weight instead.
  • The two-sided example above is different: there accuracy is measured separately on each class, so its log-odds is the weight and the prior is added once.
  • This sketch treats an abstain as saying nothing. A full label model also learns how often each rule fires on each class, and a rule that catches most fakes makes its silence mildly reassuring.

The same rule in a category with fewer fakes

  • takedown_hash
  • price_low
The same rule in a category with fewer fakesA lone takedown_hash match stays above 0.5 down to about a 0.2% counterfeit share; a lone price_low vote drops below 0.5 once fakes fall under about 2%.00.20.40.60.810.1%1%10%100%P = 0.5Handbags (5%)takedown_hashprice_lowP(counterfeit)Share of listings that are counterfeit (%)The same rule in a category with fewer fakesA lone takedown_hash match stays above 0.5 down to about a 0.2% counterfeit share; a lone price_low vote drops below 0.5 once fakes fall under about 2%.00.20.40.60.810.1%1%10%100%P = 0.5Handbags (5%)takedown_hashprice_lowP(counterfeit)Share of listings that are counterfeit (%)
P(counterfeit) for a listing whose only vote is takedown_hash or price_low, if Loftmarket reuses the rule in a category with a different share of fakes. Assumes each rule flags the same fraction of fakes and of genuine listings there as in handbags, so its likelihood ratio (614 and 46) carries over; its precision does not. At the handbags rate of 5% each curve passes through the precision measured on the dev set (0.97 and 0.71). Posterior odds = prior odds × likelihood ratio; illustrative.
Data
Share of listings that are counterfeit (%)takedown_hashprice_low
0.10.3810.044
0.20.5520.085
0.50.7550.189
10.8610.32
20.9260.487
50.970.71
100.9860.838
200.9940.921
  • P = 0.5: P(counterfeit) = 0.5
  • At 5: Handbags (5%)

Why train a model at all

The label model can only label rows where some LF fired: 820,000 of the 2 million listings. The end model is trained on those rows, but it learns from everything in them: the photos, the wording, the seller's history. So it picks up signals no rule mentions, such as a counterfeiter who writes "rep" or "UA quality" instead of "replica", or stitching that looks like the fakes in the takedown matches, and it scores the 59% of listings where every rule abstained.

The end model uses a noise-aware loss: its training target is the probability (0.32, 0.45, 0.97), not a hard 0 or 1, so confident rows count for more. In the Snorkel evaluation (Ratner et al. 2017), end models trained this way averaged 132% better than models trained by distant supervision, an earlier heuristic baseline, and came within 3.60% on average of models trained on large hand-labeled sets.

The label model is not always worth it. The same paper notes that when LFs rarely overlap, there is almost nothing to weigh, and when many LFs vote on every row, an unweighted majority is already accurate; the learned weights help most in between. Measure majority vote on the dev set as a baseline before trusting the fancier answer.

Was this section helpful?

In practice.

Writing LFs like code, borrowing what the company already knows, and the other ways out of zero labels.

LF report after the third iteration

LFCoverageOverlapsConflictsDev precisionAction
kw_replica2%1.4%0.1%0.92Keep; added a negation guard after dev errors on "not a replica" and "no replicas, 100% genuine"
price_low4%2.6%1.0%0.71Narrow: skip listings marked "for parts" or "damaged", which are cheap and genuine
new_seller_brand5%3.0%1.2%0.45Split by brand: 0.72 on the three most-faked brands (2% coverage), 0.27 on the rest (3%); keep only the first part
takedown_hash1.5%1.0%0.1%0.97Keep; refresh the archive weekly
receipt13%3.4%0.9%0.995Keep; audit a sample of verified receipts monthly for forgeries
trusted_seller24%5.6%1.5%0.99Keep; exclude accounts with a password reset in the last 30 days (takeovers)

LFs are code, so treat them like code

Coverage is where an LF fires, overlaps where it fires alongside another LF, and conflicts where they disagree (all as a share of the 2 million listings, illustrative). An LF with low precision is not automatically bad: new_seller_brand at 0.45 is wrong more often than right, but a listing it flags is still 9 times as likely to be fake as an average one, so it earns a real weight. What an LF must not be is wrong in a systematic way the model cannot see, like price_low firing on every broken bag.

So LFs live in a repository, are reviewed like any change, carry a version, and are tested on the dev set every time they change. The dev set is 600 listings the policy experts labeled: 300 at random and 50 drawn from wherever each LF fires, so every LF has enough examples to measure. Labeling a random 600 alone would give takedown_hash about 9 examples, too few to say anything.

How much 50 dev examples tell you about one LF

Assumptions
Dev listings where the LF fires
50
price_low measured precision
0.71
takedown_hash measured precision
0.97
Working
  1. 95% margin for price_low1.96 × √(0.71 × 0.29 ÷ 50) = 1.96 × 0.064± 0.13from Dev listings where the LF fires and price_low measured precision
  2. 95% range for takedown_hash (Wilson interval; the ± formula above breaks down this close to 1)(0.97 + 1.96² ÷ 100 ± 1.96 × √(0.97 × 0.03 ÷ 50 + 1.96² ÷ 10,000)) ÷ (1 + 1.96² ÷ 50)0.88 to 0.99from Dev listings where the LF fires and takedown_hash measured precision
  3. What that range does to takedown_hash's likelihood ratio at a 5% base rate(0.88 ÷ 0.12) ÷ (0.05 ÷ 0.95) = 139; (0.97 ÷ 0.03) ÷ (0.05 ÷ 0.95) = 614a factor of about 4.4from 95% range for takedown_hash (Wilson interval; the ± formula above breaks down this close to 1)
What it means
  • Dev precision is a rough guide for choosing which LFs to fix, not a precise weight. The label model's own estimates use hundreds of thousands of unlabeled rows.
  • The naive ± margin would claim 0.92 to 1.02 for takedown_hash, which is impossible above 1 and too narrow below. Near 0 or 1, use the Wilson interval or label more examples.
  • Fifty examples cannot tell receipt's 0.995 from 0.99: both would usually show 50 out of 50. Precise weights for near-perfect rules need far more labels or the label model's estimate.

LFs can use what the model never will

LFs run offline, over history, so they can read signals that do not exist when a listing goes live: a chargeback filed six weeks after the sale, a moderator's free-text note, a brand's takedown archive the company may not expose to a serving path. The model that ships reads only what is there at listing time: photos, text, price and seller features. At Google, Snorkel DryBell used exactly this split, turning internal resources that could not be served into servable classifiers, with an average 52% gain from those resources and quality comparable to classifiers trained on tens of thousands of hand labels (Bach et al. 2019).

At Loftmarket the next LFs to write are chargeback_fake (buyer refunded with reason "not authentic") and brand_confirmed (brand partner confirms a report). Both arrive weeks late, which is fine for labeling last month's listings and useless for scoring today's.

What the Snorkel papers measured

Faster model building
2.8×
vs 7 h of hand labeling, user study (Ratner et al. 2017)
Over distant supervision
+132%
average predictive gain (Ratner et al. 2017)
Gap to large hand-labeled sets
3.60%
average (Ratner et al. 2017)
From non-servable resources
+52%
average, Google (Bach et al. 2019)

Other ways out of zero labels

Distant supervision
Label by matching against a record that already exists for another purpose. Mintz et al. (2009) labeled sentences by looking up entity pairs in Freebase. Here it is the brand takedown archive; its weakness is that the archive covers only what brands chose to report.
Pretrained model plus a few hundred labels
Start from an image and text encoder trained elsewhere and fine-tune it on the 600 dev listings plus whatever the label model produces. It often beats training from scratch on weak labels alone; the fine-tuning topic covers how.
LLM prompts as labeling functions
Ask a large language model "Does this listing read like a counterfeit?" and treat its answer as one more noisy vote whose accuracy the label model estimates, rather than as truth (Smith et al. 2022). Grading outputs with an LLM is a different job, covered in the LLM-as-judge topic.
Self-training
Train on a small labeled set, label the pool with the model's own confident predictions, retrain. Cheap, but its mistakes feed back into its training data, so a bias in the first model gets louder with every round. The teacher-and-student version of the idea is covered in the distillation topic.

Before trusting the weak labels

Keep the 600 dev and 400 test listings expert-labeled, and never look at the test set while writing LFs.
Report coverage and dev accuracy for every LF on every change.
Watch for LFs that copy each other.
Re-run the pipeline when sellers adapt.
Size the test set for the rare class.
Move toward human labels as they accumulate.
Was this section helpful?

Trade-offs.

What day one costs each way, the choice Loftmarket made, and when weak supervision is the wrong tool.

Day one: hand labels or rules

Assumptions
Listings hand-labeled for training
20,000
Time to judge one handbag from its photos
60 sslower than the 20 s used for prohibited items, because authenticity needs close looks at stitching, logos and serials (illustrative)
Labels per listing
2
Vendor labeler cost
$15 per hour
Policy expert cost
$60 per hour
Vendor team
5 labelers × 30 h a week
Share of listings that are fake
5%
LF writing and debugging
2 policy experts × 3 days of 8 h
Working
  1. Vendor hours20,000 × 2 × 60 s ÷ 3,600≈ 667 hfrom Listings hand-labeled for training, Time to judge one handbag from its photos and Labels per listing
  2. Vendor cost667 h × $15≈ $10,000from Vendor hours and Vendor labeler cost
  3. Expert tie-breaks on 10% of listings at 90 s2,000 × 90 s = 50 h; 50 h × $60$3,000from Listings hand-labeled for training and Policy expert cost
  4. Hand-labeling plan$10,000 + $3,000; 667 h ÷ 150 h a week≈ $13,000 and 4.4 weeksfrom Vendor cost, Expert tie-breaks on 10% of listings at 90 s and Vendor team
  5. Counterfeits in that set20,000 × 5%≈ 1,000from Listings hand-labeled for training and Share of listings that are fake
  6. Weak supervision: 1,000 dev and test listings labeled by experts at 90 s1,000 × 90 s = 25 h; 25 h × $60$1,500from Policy expert cost
  7. LF writing2 × 3 × 8 h = 48 h; 48 h × $60$2,880from LF writing and debugging and Policy expert cost
  8. Weak supervision plan$1,500 + $2,880; labeling and LF writing run in parallel≈ $4,400 and about a weekfrom Weak supervision: 1,000 dev and test listings labeled by experts at 90 s and LF writing
What it means
  • Weak supervision is not free: expert time writing and debugging LFs is two thirds of its cost. It moves the spend from labeling each item to encoding what experts know once, and comes to about a third of the hand-labeling plan.
  • The hand-labeled set would be clean but small, with about 1,000 fakes; the weak set is noisy but covers 820,000 listings.
01
How to train the first counterfeit model
Chosen:Weak supervision plus a small expert-labeled dev and test set
  • Pro:Soft labels for 820,000 listings within about a week, for about $4,400
  • Pro:Rules are readable, so policy can audit why a listing was labeled fake
  • Pro:Can use late or private signals (takedowns, chargebacks) the served model never sees
Downside we accept:
  • Con:Needs people who can write and debug LFs
  • Con:Correlated or systematically wrong LFs mislead the label model silently
  • Con:The model learns mostly from the 41% of listings the rules cover, and inherits their blind spots
Ruled out:Hand-label 20,000 listings first

About $13,000 and 4.4 weeks before the first model; Only about 1,000 counterfeits to learn from

Ruled out:Zero-shot LLM classifier in production

A per-call cost and latency on every new listing; Accuracy unknown until a test set exists anyway; Cannot use photos unless the model is multimodal, and cannot use seller history it is not given

Ruled out:Serve the rules directly

Covers 41% of listings; the rest get no score; Conflicts must be settled by hand-written priorities; Counterfeiters route around a fixed rule within days

When weak supervision fits

SituationFitsPoor fit
Domain knowledgeExperts can state rules of thumb, even rough onesNobody can say why an item is positive (subtle visual quality, taste)
Unlabeled dataLarge pool, hundreds of thousands of items or moreA few thousand items: hand-label them instead
Label costLabels need scarce experts, or data is private and cannot go to a vendorLabels are cheap and quick with a crowd
Existing recordsTakedown lists, knowledge bases, past decisions to borrow fromNothing to match against
Rule overlapSeveral LFs fire on the same items, so accuracies can be estimatedEach LF fires on its own slice; nothing to compare

What goes wrong

FailureImpactDetectionMitigationMeanwhile
An LF is systematically wrong on a slice2Labeling functionsprice_low labels every cheap, damaged but genuine bag as fake, and the end model learns that damage means counterfeitDev-set errors grouped by listing attributes; spot-check the highest-weight wrong labelsNarrow the LF, add a guard, and re-run the label modelThe rest of the LFs still vote; accuracy on that slice drops
Two LFs are copies of each other4Label modelTheir agreement inflates both accuracy estimates, and their evidence is counted twicePairwise overlap near their full coverage with almost no conflictsMerge them, or give the label model the dependencyProbabilities are overconfident but mostly point the right way
Sellers adapt to the rules5Counterfeit classifierkw_replica's coverage falls; new fakes look like the uncovered 59%Coverage per LF by week; share of expert-confirmed fakes the model missedWrite new LFs from review-queue findings, retrain weekly, move to human labels as they build upThe model keeps catching old patterns and misses new ones
The test set leaks into LF development6Hand-labeled dev and test setsThe reported score is optimistic because rules were tuned on itAccess logs on the test set; a score gap between test and a fresh expert sampleKeep test labels in a separate store that LF authors cannot queryThe model may be fine; nobody knows how fine
Was this section helpful?
Related
Implicit feedback as labels
Read next