Data collection and labelingImplicit feedback as labels

100%

Implicit feedback as labels.

Nobody at Loftmarket labels search results by hand; buyers do it for free every time they tap, linger, message a seller or pay. The work is choosing which of those actions means "good result", deciding what a missing click means, and undoing the position the ranker itself gave each listing.

Intermediate19 minUpdated 1 Oct 2026

The idea.

One search, twenty listings, and the five things a buyer's actions can tell us that we might call a label.

One buyer, five possible labels

A buyer types "air max 90" into Loftmarket and gets 20 listings. She taps one of them (L4, fourth from the top), reads it for 40 seconds, messages its seller about the sole, and pays for it two days later. Nobody asked her to rate anything, yet the logs now hold five different things a training pipeline could call a positive.

Each one answers its own question. A click says the photo and price looked worth a look. A long stay on the page says the listing held her attention. A message or an offer says she was seriously considering it. A purchase says it was the right item. A purchase she did not return within 14 days of the order adds that the item matched its description. Walking down that list, the signal gets closer to what Loftmarket actually wants, and it also gets rarer and later.

Explicit feedback exists too: star ratings after delivery, a "not relevant" button. It is unambiguous but tiny next to the logs. Covington and colleagues made the same call for YouTube's recommender: thumbs and surveys existed, but they trained on watches because the implicit history was orders of magnitude larger and reached videos deep in the tail where explicit ratings barely exist.

One day of candidate labels at Loftmarket

One day of candidate labels at LoftmarketA purchase label is about 1,600 times rarer than an impression and 48 times rarer than a click, which is why most rankers start from clicks and correct toward purchases.10k100k1M10M100M1GImpressionsSeen ≥ 1 sClicksDwell ≥ 20 sMessages or offersPurchasesKept 14 days240M96M7.2M4.3M600k150k141kEventEvents per dayOne day of candidate labels at LoftmarketA purchase label is about 1,600 times rarer than an impression and 48 times rarer than a click, which is why most rankers start from clicks and correct toward purchases.10k100k1M10M100M1GImpressionsSeen ≥ 1 sClicksDwell ≥ 20 sMessages or offersPurchasesKept 14 days240M96M7.2M4.3M600k150k141kEventEvents per day
Illustrative counts for 12M searches × 20 results. 240M ÷ 150k = 1,600 impressions per purchase; 7.2M ÷ 150k = 48 clicks per purchase. Purchases can arrive up to 7 days after the search; the kept label is known 14 days after the order. Log scale.
Data
Eventevents per day
Impressions240,000,000
Seen ≥ 1 s96,000,000
Clicks7,200,000
Dwell ≥ 20 s4,300,000
Messages or offers600,000
Purchases150,000
Kept 14 days141,000

What each candidate label really measures

LabelQuestion it answersVolume per dayKnown afterMain trap
clickDid the thumbnail and price look worth a tap?7.2MsecondsFlattering photos and low teaser prices earn taps; position inflates it
long dwellDid the page hold attention?4.3Mabout a minuteSlow-loading pages and long descriptions also keep people there
message or offerWas the buyer seriously interested?600kminutes to hoursScammers and hagglers message too
purchaseWas it the right item?150khours to 7 daysRare, late, and driven by price as much as by relevance
kept 14 daysWas it described honestly?141kup to 21 daysToo late to feed a model retrained every day

The top row is the cheapest and the most dangerous. Covington and colleagues report that ranking YouTube videos by click-through rate tends to promote clickbait that people abandon, and that expected watch time tracks real engagement better. Loftmarket's version of clickbait is a stock photo of a new shoe on a worn pair. The choice of label is where the model learns what "good" means, so it deserves the same care as the model.

Was this section helpful?

How it works.

How a tap becomes a training row, what a listing that wasn't tapped means, and how much of a click the position explains.

From a tap to a training row

From a tap to a training row. The numbered component cards that follow describe each part.
From a tap to a training rowComponents: 1. Loftmarket app (The buyer's phone. Shows results and reports what the buyer saw and did, each event tagged with the search's request_id.), 2. Event collector (Receives impression, click and order events from the app and the orders service, drops bot traffic, and appends the rest to the log.), 3. Event log (Append-only record of every impression, click, message and order, partitioned by day.), 4. Label builder (A daily batch job that joins each search's impressions to the clicks and orders that followed, once the search's attribution window has closed.), 5. Training table (One row per (search, listing) that was seen, with its rank, the model version that placed it, and its labels.), 6. Search ranker (Scores listings for a query and returns the top 20, stamping each result with its rank and the model version.).

results + request_id

impression, click, order events

append

events for closed windows

labelled rows

6Search ranker
returns 20 listings
rank + model_version on each

1Loftmarket app
reports viewport time per result

2Event collector
bot filter

3Event log
impressions · clicks · orders

4Label builder
join on request_id
window closes after 7 days

5Training table
one row per seen (search, listing)

The ranker stamps every result with request_id, rank and model_version, and the app sends them back with each event it reports. That small habit is what makes the rest of this topic possible: without the rank on the impression, the position effect can never be corrected later, and Rules of ML #36 needs the rank at training time.

One search, followed until its labels are written

One search, followed until its labels are written, as an ordered list of steps:
One search, followed until its labels are written13 steps between Loftmarket app, Search ranker, Event collector, Event log, Label builder, Training table. The steps are listed as text after the diagram.Training tableLabel builderEvent logEvent collectorSearch rankerLoftmarket app40 s on L4, then a message to the sellerday 7: the window for r81 closesL5–L20 were never scrolled into view: dropped or down-weightedsearch("air max 90")120 listings, request_id r812impression r81 × 20 (rank, viewport ms)3click r81, listing L4 at rank 44order o552 for L4 (day 2)5append6read all r81 events720 impressions, 1 click, 1 order8(r81, L4): purchase = 19(r81): pairs L4 > L1, L4 > L2, L4 > L3 (Click > Skip Above)10
  1. Loftmarket app → Search ranker: search("air max 90")
  2. Search ranker → Loftmarket app (reply): 20 listings, request_id r81
  3. Loftmarket app → Event collector: impression r81 × 20 (rank, viewport ms)
  4. Loftmarket app → Event collector: click r81, listing L4 at rank 4
  5. Note over Loftmarket app: 40 s on L4, then a message to the seller
  6. Loftmarket app → Event collector: order o552 for L4 (day 2)
  7. Event collector → Event log: append
  8. Note over Label builder: day 7: the window for r81 closes
  9. Label builder → Event log: read all r81 events
  10. Event log → Label builder (reply): 20 impressions, 1 click, 1 order
  11. Label builder → Training table: (r81, L4): purchase = 1
  12. Label builder → Training table: (r81): pairs L4 > L1, L4 > L2, L4 > L3 (Click > Skip Above)
  13. Note over Label builder: L5–L20 were never scrolled into view: dropped or down-weighted

What a missing click means

A listing without a click is one of two very different things. Either the buyer looked at it and passed, or the buyer never saw it. Only the first is evidence against it.

Click rules against human judges (Joachims et al. 2005, Table 4, Phase I)

RuleReads asAgrees with judges
Click > Skip AboveA clicked result beats every unclicked result ranked above it80.8% ± 3.6
Last Click > Skip AboveSame, but only for the final click of the search83.1% ± 3.8
Click > Skip PreviousA clicked result beats the unclicked result directly above it82.3% ± 7.3
Click > No Click NextA clicked result beats the unclicked result directly below it84.1% ± 4.9 (70.4% across all Phase II conditions)
Click > Earlier ClickA later click beats an earlier one67.2% ± 12.3 (46.9% in Phase II: no better than a coin)
Two human judgesThe ceiling: how often two people agree with each other89.5%

The table scores each rule by how often its pairs agree with explicit judgments of the result snippets; the ± is the paper's 95% interval, and the ranker produced its normal ordering. The rules that stay reliable when the ranking changes share one assumption: people scan from the top, so a result above a click was almost certainly looked at and passed over. Click > Skip Above holds 79.6% to 88.0% in every Phase II condition, including a reversed ranking. The top Phase I score, Click > No Click Next, is partly an artefact: its pairs agree with the search engine's own order, so a user who always clicked rank 1 would already score 62.4%, and it falls to 70.0% when the ranking is reversed. That turns the r81 search into three trustworthy pairs, L4 over L1, L2 and L3, and says nothing about L5 to L20. A pairwise rule built this way lands within about 6 to 9 points of a second human judge (81% to 83% against 89.5%), far above the 50% of a coin. The absolute reading, "clicked means relevant and unclicked means irrelevant", does not, and the pairs feed straight into the pairwise losses of learning to rank. How many unclicked listings to sample as negatives, and how, belongs to negative sampling.

How much of a click is the position?

Joachims and colleagues tracked searchers' eyes. The first two results were looked at about equally often, yet the first drew far more clicks. To separate relevance from rank they had judges decide which of the top two abstracts was better. When the better one sat at rank 1, 19 of the 20 searches with a single click on the top two went to it. When the better one sat at rank 2, only 2 of 7 did; the other 5 still went to rank 1. People trust the order they are given. The same study found that viewing drops with every rank down, and sharply at the edge of the screen.

A simple model folds both effects into one factor per rank: the click rate at rank k ≈ position factor(k) × P(click | relevance). The factor mixes how often rank k is looked at with how much buyers trust whatever sits there. The classic examination model assumes the click, once a result is seen, depends on relevance alone; trust bias breaks that, since ranks 1 and 2 were seen equally but clicked unequally (a WWW 2019 paper by Agarwal and colleagues models trust as its own term). The randomized bucket below measures the combined factor, which is what the correction needs. That factor, the propensity, belongs to the page layout, not to the listing. Inverse propensity scoring (Joachims, Swaminathan and Schnabel, 2017) weights each click in the loss by 1 ÷ propensity of the rank where it happened, so a click earned deep in the page counts more than one handed to rank 1. In expectation this removes the position effect, provided the propensities are right and every rank has some chance of being seen.

Click-through rate by rank when the top 10 are shuffled

Click-through rate by rank when the top 10 are shuffledWhen order is random, relevance is the same at every rank, so the fall from 6% to 0.7% is pure position: rank 4 draws a third of rank 1's clicks for the same relevance.01%2%3%4%5%6%123456789106%3.7%2.7%2%1.6%1.4%1.2%0.9%0.8%0.7%Click-through rate (%)RankClick-through rate by rank when the top 10 are shuffledWhen order is random, relevance is the same at every rank, so the fall from 6% to 0.7% is pure position: rank 4 draws a third of rank 1's clicks for the same relevance.01%2%3%4%5%6%123456789106%3.7%2.7%2%1.6%1.4%1.2%0.9%0.8%0.7%Click-through rate (%)Rank
Illustrative data from a 0.5% bucket in which the top 10 are shuffled per search. Shuffling lowers click rates overall, since weaker listings land on top; only the ratios matter. Propensity = CTR at k ÷ CTR at rank 1: 1, 0.62, 0.45, 0.33, 0.27, 0.23, 0.20, 0.15, 0.13, 0.12.
Data
RankCTR in the shuffled bucket (%)
16
23.7
32.7
42
51.6
61.4
71.2
80.9
90.8
100.7

Which listing is really better?

Assumptions
Listing A: shown at rank 1, CTR
6.0%
Listing B: shown at rank 4, CTR
2.4%
Propensity at rank 1
1.0from the shuffled bucket
Propensity at rank 4
0.332.0 ÷ 6.0
Propensity at rank 10
0.120.7 ÷ 6.0
Largest weight allowed
5
Working
  1. A's click rate with position removed6.0% ÷ 1.06.0%from Listing A: shown at rank 1, CTR and Propensity at rank 1
  2. B's click rate with position removed2.4% ÷ 0.33≈ 7.3%from Listing B: shown at rank 4, CTR and Propensity at rank 4
  3. IPS weight of one click at rank 41 ÷ 0.33≈ 3.0from Propensity at rank 4
  4. IPS weight of one click at rank 101 ÷ 0.12≈ 8.3, clipped to 5from Propensity at rank 10 and Largest weight allowed · Clipping accepts a little bias to stop a handful of deep clicks from dominating a training batch.
What it means
  • Raw CTR ranks A above B; corrected for position, B is the better listing (7.3% against 6.0% at the same position).
  • Trained on raw clicks, the ranker keeps A on top, A keeps collecting clicks, and the next model 'confirms' the choice. That loop is the subject of feedback loops; here the fix is only to correct the label.
Was this section helpful?

In practice.

Getting propensities without hurting search, labels that arrive late, and the data the model never showed.

Used:Shuffle the top 10 for a small slice of trafficThe cleanest estimate, but shuffled pages are worse pages. Loftmarket limits it to a 0.5% bucket and refreshes it quarterly, when the page layout changes.
Not used:Swap adjacent pairs only (RandPair)Cheaper for users than a full shuffle; Wang et al. (2018) compare it with shuffling the top n. Kept in reserve if the shuffled bucket's revenue cost grows.
Used:Harvest interventions from past A/B testsTwo ranker variants often put the same listing at different ranks for the same query. Agarwal et al. (2019) estimate propensities from exactly that, at no extra cost to buyers.
Used:EM on ordinary click logsWang et al. (2018) fit propensity and relevance jointly from regular logs with regression-based EM. Used as a cross-check on the other two, since it relies on its own modelling assumptions.
Not used:Position as a training feature, fixed at servingRules of ML #36's simpler alternative to weighting: the model learns how much of a click the rank explains, and at serving every candidate gets the same default rank. Keep the position term separate from the listing features. Loftmarket uses the IPS weights instead, not both: stacking the two corrects the same effect twice.

When the labels of search r81 become final

Scenario 1 of 2: As described.

Timeline as a list

When the labels of search r81 become final: 5 lanes, from 0 d to 21 d.

  1. 0 d · Impression · 20 results shown
  2. 0 d · Click · click L4
  3. 0–7 d · all lanes · window: purchase window
  4. 2–16 d · Return window · 14-day return window
  5. 2 d · Order · order o552
  6. 7 d · Label builder · closes: purchase = 1 (deadline, ok)
  7. 16 d · Label builder · final: kept = 1 (deadline, ok)
  8. 21 d · Label builder · latest possible kept label (tick)
Days after the search. Purchases are attributed to a search for 7 days; the kept label needs the 14-day return window on top, so it is final on day 16 here and on day 21 at worst. How to model conversions that arrive late belongs to delayed feedback (Chapelle 2014); this topic only defines the label and its window.

Labels exist only where the model looked

Every row in the training table is a listing the current ranker chose to show. A good listing it always buried on page 3 never gets a click, so it never gets a positive. Position weighting cannot rescue it: its propensity at rank 60 is effectively zero, and 1 ÷ 0 is not a weight. How that gap compounds as each model trains on the last one's choices is the subject of feedback loops, and spending a little traffic on purpose to show unproven listings is exploration slots. Here the question is narrower: where else can labels come from?

Two remedies come from practice. Rules of ML #34 advises that when a model filters things out, you hold back a small share of traffic, around 1%, from the filter so the filtered items still earn honest labels. And Covington and colleagues built YouTube's training examples from all watches, including videos played embedded on other sites, not only from its own recommendations, so new videos had a way in. Loftmarket's equivalent is to also learn from listings reached through category browsing, saved searches and shared links, where the ranker did not decide what the buyer saw.

Rules the label builder enforces

Log the rank and the viewport
Each impression carries rank, request_id, model_version and milliseconds on screen. Without them no propensity can be applied later, and 'never seen' cannot be told apart from 'seen and skipped'.
Equal weight per buyer
Covington et al. generated a fixed number of examples per user so that heavy users did not dominate the loss. Loftmarket caps each buyer at 50 rows a day; a reseller who runs 900 searches a day otherwise teaches the model their taste.
Drop bots before labelling
A price-scraping crawler taps every result in order. Left in, its taps become positives spread evenly over all ranks, which also corrupts the propensity estimate. The collector filters them before the log.
Don't leak the future
Features for a row come only from events before the impression. Covington et al. predict the next watch rather than a randomly held-out one for the same reason: a random hold-out lets the model peek at what the user did afterwards.
Was this section helpful?

Trade-offs.

The chosen option is first; the others stay visible so the reasoning can be checked.

01
The training label for Loftmarket search
Chosen:Position-weighted clicks plus a purchase-weighted term
  • Pro:7.2M clicks a day are enough to retrain daily and cover rare queries
  • Pro:The purchase term pulls the model away from listings that attract taps but never sell
  • Pro:IPS weights keep rank 1's head start out of the label
Downside we accept:
  • Con:Two labels to maintain, with a weight between them that needs tuning against online results
  • Con:Needs a propensity estimate that stays current as the page layout changes
Ruled out:Purchases only

48 times fewer labels than clicks, and up to 7 days late; Small categories and new sellers barely appear in the positives

Ruled out:Clicks only

Rewards clickbait photos and teaser prices; Carries the full position effect unless corrected

Ruled out:Explicit ratings

Orders of magnitude sparser than any logged action; The buyers who rate are not typical buyers, so the sample is skewed

Implicit against explicit feedback

PropertyImplicit (clicks, purchases)Explicit (ratings, relevance button)
VolumeMillions a day at LoftmarketThousands a day
DelaySeconds for clicks, up to 21 days for kept purchasesDays after the order, if ever
BiasPosition, presentation, and whatever the ranker chose to showWho bothers to rate; mood; social pressure
CostFree, apart from logging and joiningFree from users, but prompts annoy them
Tail coverageReaches rare queries and new listings, if they are shown at allAlmost none
MeaningRelative: better than what was skippedCloser to absolute, but noisy between people

What implicit labels cannot tell you

Engagement is not permission. Buyers happily tap a counterfeit bag priced at a tenth of retail, and a scam listing can collect messages all day. No weighting scheme turns those clicks into the answer to "should this be allowed?"; that question needs people reading a policy, covered in human labeling. When you have no labels at all yet, rules and heuristics can bootstrap them, as in weak supervision. And every source, implicit ones included, leaves some rows simply wrong; measuring and cleaning those is label noise. When the same logged signal is used to judge a launch rather than to train, the concerns shift to proxy metrics.

Was this section helpful?
Builds on this
Delayed conversions
Read next