Implicit feedback as labels.
Nobody at Loftmarket labels search results by hand; buyers do it for free every time they tap, linger, message a seller or pay. The work is choosing which of those actions means "good result", deciding what a missing click means, and undoing the position the ranker itself gave each listing.
The idea.
One search, twenty listings, and the five things a buyer's actions can tell us that we might call a label.
One buyer, five possible labels
A buyer types "air max 90" into Loftmarket and gets 20 listings. She taps one of them (L4, fourth from the top), reads it for 40 seconds, messages its seller about the sole, and pays for it two days later. Nobody asked her to rate anything, yet the logs now hold five different things a training pipeline could call a positive.
Each one answers its own question. A click says the photo and price looked worth a look. A long stay on the page says the listing held her attention. A message or an offer says she was seriously considering it. A purchase says it was the right item. A purchase she did not return within 14 days of the order adds that the item matched its description. Walking down that list, the signal gets closer to what Loftmarket actually wants, and it also gets rarer and later.
Explicit feedback exists too: star ratings after delivery, a "not relevant" button. It is unambiguous but tiny next to the logs. Covington and colleagues made the same call for YouTube's recommender: thumbs and surveys existed, but they trained on watches because the implicit history was orders of magnitude larger and reached videos deep in the tail where explicit ratings barely exist.
One day of candidate labels at Loftmarket
Data
| Event | events per day |
|---|---|
| Impressions | 240,000,000 |
| Seen ≥ 1 s | 96,000,000 |
| Clicks | 7,200,000 |
| Dwell ≥ 20 s | 4,300,000 |
| Messages or offers | 600,000 |
| Purchases | 150,000 |
| Kept 14 days | 141,000 |
What each candidate label really measures
| Label | Question it answers | Volume per day | Known after | Main trap |
|---|---|---|---|---|
| click | Did the thumbnail and price look worth a tap? | 7.2M | seconds | Flattering photos and low teaser prices earn taps; position inflates it |
| long dwell | Did the page hold attention? | 4.3M | about a minute | Slow-loading pages and long descriptions also keep people there |
| message or offer | Was the buyer seriously interested? | 600k | minutes to hours | Scammers and hagglers message too |
| purchase | Was it the right item? | 150k | hours to 7 days | Rare, late, and driven by price as much as by relevance |
| kept 14 days | Was it described honestly? | 141k | up to 21 days | Too late to feed a model retrained every day |
The top row is the cheapest and the most dangerous. Covington and colleagues report that ranking YouTube videos by click-through rate tends to promote clickbait that people abandon, and that expected watch time tracks real engagement better. Loftmarket's version of clickbait is a stock photo of a new shoe on a worn pair. The choice of label is where the model learns what "good" means, so it deserves the same care as the model.
How it works.
How a tap becomes a training row, what a listing that wasn't tapped means, and how much of a click the position explains.
From a tap to a training row
The ranker stamps every result with request_id, rank and model_version, and the app sends them back with each event it reports. That small habit is what makes the rest of this topic possible: without the rank on the impression, the position effect can never be corrected later, and Rules of ML #36 needs the rank at training time.
One search, followed until its labels are written
- Loftmarket app → Search ranker: search("air max 90")
- Search ranker → Loftmarket app (reply): 20 listings, request_id r81
- Loftmarket app → Event collector: impression r81 × 20 (rank, viewport ms)
- Loftmarket app → Event collector: click r81, listing L4 at rank 4
- Note over Loftmarket app: 40 s on L4, then a message to the seller
- Loftmarket app → Event collector: order o552 for L4 (day 2)
- Event collector → Event log: append
- Note over Label builder: day 7: the window for r81 closes
- Label builder → Event log: read all r81 events
- Event log → Label builder (reply): 20 impressions, 1 click, 1 order
- Label builder → Training table: (r81, L4): purchase = 1
- Label builder → Training table: (r81): pairs L4 > L1, L4 > L2, L4 > L3 (Click > Skip Above)
- Note over Label builder: L5–L20 were never scrolled into view: dropped or down-weighted
What a missing click means
A listing without a click is one of two very different things. Either the buyer looked at it and passed, or the buyer never saw it. Only the first is evidence against it.
Click rules against human judges (Joachims et al. 2005, Table 4, Phase I)
| Rule | Reads as | Agrees with judges |
|---|---|---|
| Click > Skip Above | A clicked result beats every unclicked result ranked above it | 80.8% ± 3.6 |
| Last Click > Skip Above | Same, but only for the final click of the search | 83.1% ± 3.8 |
| Click > Skip Previous | A clicked result beats the unclicked result directly above it | 82.3% ± 7.3 |
| Click > No Click Next | A clicked result beats the unclicked result directly below it | 84.1% ± 4.9 (70.4% across all Phase II conditions) |
| Click > Earlier Click | A later click beats an earlier one | 67.2% ± 12.3 (46.9% in Phase II: no better than a coin) |
| Two human judges | The ceiling: how often two people agree with each other | 89.5% |
The table scores each rule by how often its pairs agree with explicit judgments of the result snippets; the ± is the paper's 95% interval, and the ranker produced its normal ordering. The rules that stay reliable when the ranking changes share one assumption: people scan from the top, so a result above a click was almost certainly looked at and passed over. Click > Skip Above holds 79.6% to 88.0% in every Phase II condition, including a reversed ranking. The top Phase I score, Click > No Click Next, is partly an artefact: its pairs agree with the search engine's own order, so a user who always clicked rank 1 would already score 62.4%, and it falls to 70.0% when the ranking is reversed. That turns the r81 search into three trustworthy pairs, L4 over L1, L2 and L3, and says nothing about L5 to L20. A pairwise rule built this way lands within about 6 to 9 points of a second human judge (81% to 83% against 89.5%), far above the 50% of a coin. The absolute reading, "clicked means relevant and unclicked means irrelevant", does not, and the pairs feed straight into the pairwise losses of learning to rank. How many unclicked listings to sample as negatives, and how, belongs to negative sampling.
How much of a click is the position?
Joachims and colleagues tracked searchers' eyes. The first two results were looked at about equally often, yet the first drew far more clicks. To separate relevance from rank they had judges decide which of the top two abstracts was better. When the better one sat at rank 1, 19 of the 20 searches with a single click on the top two went to it. When the better one sat at rank 2, only 2 of 7 did; the other 5 still went to rank 1. People trust the order they are given. The same study found that viewing drops with every rank down, and sharply at the edge of the screen.
A simple model folds both effects into one factor per rank: the click rate at rank k ≈ position factor(k) × P(click | relevance). The factor mixes how often rank k is looked at with how much buyers trust whatever sits there. The classic examination model assumes the click, once a result is seen, depends on relevance alone; trust bias breaks that, since ranks 1 and 2 were seen equally but clicked unequally (a WWW 2019 paper by Agarwal and colleagues models trust as its own term). The randomized bucket below measures the combined factor, which is what the correction needs. That factor, the propensity, belongs to the page layout, not to the listing. Inverse propensity scoring (Joachims, Swaminathan and Schnabel, 2017) weights each click in the loss by 1 ÷ propensity of the rank where it happened, so a click earned deep in the page counts more than one handed to rank 1. In expectation this removes the position effect, provided the propensities are right and every rank has some chance of being seen.
Click-through rate by rank when the top 10 are shuffled
Data
| Rank | CTR in the shuffled bucket (%) |
|---|---|
| 1 | 6 |
| 2 | 3.7 |
| 3 | 2.7 |
| 4 | 2 |
| 5 | 1.6 |
| 6 | 1.4 |
| 7 | 1.2 |
| 8 | 0.9 |
| 9 | 0.8 |
| 10 | 0.7 |
Which listing is really better?
- Listing A: shown at rank 1, CTR
- 6.0%
- Listing B: shown at rank 4, CTR
- 2.4%
- Propensity at rank 1
- 1.0from the shuffled bucket
- Propensity at rank 4
- 0.332.0 ÷ 6.0
- Propensity at rank 10
- 0.120.7 ÷ 6.0
- Largest weight allowed
- 5
- A's click rate with position removed6.0% ÷ 1.06.0%from Listing A: shown at rank 1, CTR and Propensity at rank 1
- B's click rate with position removed2.4% ÷ 0.33≈ 7.3%from Listing B: shown at rank 4, CTR and Propensity at rank 4
- IPS weight of one click at rank 41 ÷ 0.33≈ 3.0from Propensity at rank 4
- IPS weight of one click at rank 101 ÷ 0.12≈ 8.3, clipped to 5from Propensity at rank 10 and Largest weight allowed · Clipping accepts a little bias to stop a handful of deep clicks from dominating a training batch.
- Raw CTR ranks A above B; corrected for position, B is the better listing (7.3% against 6.0% at the same position).
- Trained on raw clicks, the ranker keeps A on top, A keeps collecting clicks, and the next model 'confirms' the choice. That loop is the subject of feedback loops; here the fix is only to correct the label.
In practice.
Getting propensities without hurting search, labels that arrive late, and the data the model never showed.
When the labels of search r81 become final
Scenario 1 of 2: As described.
Timeline as a list
When the labels of search r81 become final: 5 lanes, from 0 d to 21 d.
- 0 d · Impression · 20 results shown
- 0 d · Click · click L4
- 0–7 d · all lanes · window: purchase window
- 2–16 d · Return window · 14-day return window
- 2 d · Order · order o552
- 7 d · Label builder · closes: purchase = 1 (deadline, ok)
- 16 d · Label builder · final: kept = 1 (deadline, ok)
- 21 d · Label builder · latest possible kept label (tick)
Labels exist only where the model looked
Every row in the training table is a listing the current ranker chose to show. A good listing it always buried on page 3 never gets a click, so it never gets a positive. Position weighting cannot rescue it: its propensity at rank 60 is effectively zero, and 1 ÷ 0 is not a weight. How that gap compounds as each model trains on the last one's choices is the subject of feedback loops, and spending a little traffic on purpose to show unproven listings is exploration slots. Here the question is narrower: where else can labels come from?
Two remedies come from practice. Rules of ML #34 advises that when a model filters things out, you hold back a small share of traffic, around 1%, from the filter so the filtered items still earn honest labels. And Covington and colleagues built YouTube's training examples from all watches, including videos played embedded on other sites, not only from its own recommendations, so new videos had a way in. Loftmarket's equivalent is to also learn from listings reached through category browsing, saved searches and shared links, where the ranker did not decide what the buyer saw.
Rules the label builder enforces
Trade-offs.
The chosen option is first; the others stay visible so the reasoning can be checked.
- Pro:7.2M clicks a day are enough to retrain daily and cover rare queries
- Pro:The purchase term pulls the model away from listings that attract taps but never sell
- Pro:IPS weights keep rank 1's head start out of the label
- Con:Two labels to maintain, with a weight between them that needs tuning against online results
- Con:Needs a propensity estimate that stays current as the page layout changes
48 times fewer labels than clicks, and up to 7 days late; Small categories and new sellers barely appear in the positives
Rewards clickbait photos and teaser prices; Carries the full position effect unless corrected
Orders of magnitude sparser than any logged action; The buyers who rate are not typical buyers, so the sample is skewed
Implicit against explicit feedback
| Property | Implicit (clicks, purchases) | Explicit (ratings, relevance button) |
|---|---|---|
| Volume | Millions a day at Loftmarket | Thousands a day |
| Delay | Seconds for clicks, up to 21 days for kept purchases | Days after the order, if ever |
| Bias | Position, presentation, and whatever the ranker chose to show | Who bothers to rate; mood; social pressure |
| Cost | Free, apart from logging and joining | Free from users, but prompts annoy them |
| Tail coverage | Reaches rare queries and new listings, if they are shown at all | Almost none |
| Meaning | Relative: better than what was skipped | Closer to absolute, but noisy between people |
What implicit labels cannot tell you
Engagement is not permission. Buyers happily tap a counterfeit bag priced at a tenth of retail, and a scam listing can collect messages all day. No weighting scheme turns those clicks into the answer to "should this be allowed?"; that question needs people reading a policy, covered in human labeling. When you have no labels at all yet, rules and heuristics can bootstrap them, as in weak supervision. And every source, implicit ones included, leaves some rows simply wrong; measuring and cleaning those is label noise. When the same logged signal is used to judge a launch rather than to train, the concerns shift to proxy metrics.