From business goal to ML objective.
"Help buyers find things" is a wish, not something a model can learn. Between the goal and the model sit three choices: which product decision the prediction changes, which logged event stands in for success, and which constraints the answer must meet. Get them right and the rest of the design follows; get them wrong and a model that looks better offline makes the product worse.
The idea.
Oddments' product lead says buyers are not finding things. That sentence has to become one number a model can learn.
Four rungs from goal to objective
Oddments is a second-hand marketplace: people list sofas, bikes, phones and jackets, buyers scroll a home feed, message the seller and usually meet to hand the item over. The ask from the product lead is "grow completed purchases per weekly active buyer". No model can learn that number directly. It is a weekly total over the whole company, and no single prediction can be credited with it. So the first question is not "which model?" but "which decision could a better prediction improve?" On Oddments the answer is the order of listings in the home feed. Search results, price suggestions and scam review are other decisions, owned by other models.
Once the decision is fixed, the objective can be written as one sentence that names what is predicted, for whom, over what window, and what the prediction is used for: for each buyer and each listing shown in the feed, predict the probability that the buyer messages the seller within 24 hours, and order the feed by it.
The label is the logged event that makes that objective trainable. It has to be observable (it is in the logs), attributable (it can be tied to one impression of one listing) and quick enough to arrive before the next retrain; Google's Rules of ML ask for exactly this of a first objective (rule #13). What must not get worse, such as showing sold items or starving new sellers of views, is written down now as a guardrail and measured later.
The Oddments feed, rung by rung
| Rung | Oddments | Question it answers |
|---|---|---|
| Business goal | More completed purchases per weekly active buyer | Why does the company care? |
| Product decision | The order of listings in the home feed | What will the prediction change? |
| ML objective | P(buyer messages seller within 24 h | listing shown in feed) | What number does the model output? |
| Label | 1 if a message event follows the impression within 24 h, else 0 | Which logged event makes it trainable? |
| Guardrails | Share of sold or stale listings shown; share of impressions going to new sellers | What must not get worse? |
| Out of scope | Search ranking, price suggestions, scam detection | Which neighbouring decisions belong to someone else's model? |
What happens per million feed impressions (illustrative)
Each event is a candidate label, and each sits at a different distance from the goal. Of the 6,000 messages, 900 end in a confirmed purchase (900 ÷ 6,000 = 15%); of the 40,000 clicks, only 2.25% do. The purchase count is itself incomplete: about 45% of sales close in cash and are never recorded, so the real number is nearer 900 ÷ 0.55 ≈ 1,640. Picking between these events is the heart of framing, and the next section tests them one by one. Turning raw events into clean labels (attribution windows, position bias, joins) is its own subject, covered in implicit labels.
How it works.
Ask the questions that expose the decision, list candidate labels and test them, then write down what the model must deliver.
The questions that do the work
Every answer takes options off the table before any model is named. "Nobody reviews it first" means each score reaches a buyer as is, so the hard rules have to be enforced in code that sits outside the model. "A median of three days" means a nightly retrain on purchases would always train on half-finished labels. "Every impression picked by the current rule" means the logs only say how buyers reacted to that rule's choices, not to the listings it skipped. Each of those three answers reappears below: the first as the ordering policy, the second in the label table and the label-completeness rule, the third as the feedback loop in the diagram.
Candidate labels, tested
Where the prediction acts, and where its labels come from
Two things in this picture shape the framing. First, the objective lives only in the model box. Rules that must hold exactly, such as never showing a sold sofa, sit in the ordering policy, where they can change on a Tuesday without a retrain; Rules of ML #15 argues for this split between a policy layer and quality ranking. Second, the logs only contain listings the current system chose to show, so the model learns from its own past choices. Sculley et al. call this a feedback loop, and it is covered in feedback loops.
Five labels for the same goal
| Label | Positives per 1M impressions | Arrives after | Observable and attributable? | What goes wrong if you optimise it |
|---|---|---|---|---|
| Click | 40,000 | seconds | Yes: one tap on one impression | Rewards eye-catching photos and too-good prices that do not sell |
| Save | 12,000 | seconds | Yes | Buyers save to compare, not to buy; a weak link to purchases |
| Message within 24 h | 6,000 | minutes to hours (most within 2 h) | Yes: the thread starts from the listing | Can reward listings that attract questions but not deals, such as vague descriptions |
| Confirmed purchase within 7 days | 900 | median 3 days | Partly: only ~55% of sales are confirmed in the app | Missing for cash deals, and biased toward sellers who take in-app payment |
| Purchase value ($) | 900 events, weighted by price | median 3 days | Partly | Pushes expensive items: one $400 sofa counts as much as forty $10 phone cases |
How long until a new city has 100k positives
- New city feed impressions per day
- 500,00020k daily buyers × 25 listings seen
- Click, message, purchase rates
- 4%, 0.6%, 0.09%
- Positives wanted for a first model
- 100,000illustrative size
- All-Oddments impressions per day
- 50,000,0002M daily buyers × 25
- Click positives500,000 × 4% = 20,000/day; 100,000 ÷ 20,0005 daysfrom New city feed impressions per day, Click, message, purchase rates and Positives wanted for a first model
- Message positives500,000 × 0.6% = 3,000/day; 100,000 ÷ 3,000≈ 33 daysfrom New city feed impressions per day, Click, message, purchase rates and Positives wanted for a first model
- Purchase positives500,000 × 0.09% = 450/day; 100,000 ÷ 450, plus a 7-day attribution wait≈ 229 daysfrom New city feed impressions per day, Click, message, purchase rates and Positives wanted for a first model
- Purchase positives across all of Oddments50,000,000 × 0.09% = 45,000/day; 100,000 ÷ 45,000≈ 2.2 daysfrom All-Oddments impressions per day, Click, message, purchase rates and Positives wanted for a first model
- The label closest to the goal can be unusable for most of a year in a small market, while being perfectly fine across the whole marketplace.
- Label rate and label delay are framing inputs, not data-pipeline details. They decide which objective is trainable at all.
- This is why Oddments picks messages, a close proxy that arrives within hours, and keeps purchases for the A/B test.
Folding value into a probability
An objective can carry value without a second model, by weighting the positive examples. YouTube's ranker weights each clicked impression by its watch time and each unclicked one by 1, so the odds the logistic regression learns come out close to expected watch time: odds ≈ E[T](1 + P), and with a small click probability P that is ≈ E[T] (Covington et al. 2016, §4.2). Oddments could weight message positives by listing price to lean toward valuable deals. The cost is visible in the label table: cheap items and the sellers who list them lose reach, so the new-seller guardrail has to be watched more closely. Combining several separate predictions into one score is a different technique, covered in blending objectives.
What the framing commits the model to
| No. | Area | Requirement | Target / measure |
|---|---|---|---|
| 01 | What the score means | Within one feed request, a higher P(message) must mean a listing more worth showing first. Nothing downstream reads the number as a probability, so ordering matters and calibration does not; that changes if anyone starts thresholding it. | ordering quality within a requestno calibration target at launch |
| 02 | Label completeness | A training cut only includes impressions whose 24-hour message window has closed. | cut-off ≥ 24 h before the snapshot Without it, the newest day of impressions looks like it produced no messages. |
| 03 | Who and what gets scored | Buyers with no history and listings posted minutes ago still get a usable score, because both make up a large share of a marketplace feed. | new listingsnew buyers |
| 04 | New-seller reach | The objective must not starve sellers who joined recently. | ≥ 15% of impressions to sellers < 30 days old (illustrative guardrail) |
| 05 | Privacy | No protected attributes as features, and deletion requests are honoured. | features reviewed by legal |
| 06 | Time on the feed path | Scoring a full candidate set fits inside the feed's response budget. | p99 < 50 ms for 1,000 listings (illustrative) The method for sizing a budget like this is in the serving-budget topic of estimation for ML. |
| 07 | Retrain rhythm | Weights and features keep up with what buyers are doing this week. | weekly retrainfeatures minutes old |
| 08 | Spend | Serving cost fits the feed's margin. | CPU only at launch |
In practice.
How a wrong objective shows up in an experiment, what published systems chose instead, and the objective as one sentence.
Same feed, two objectives
- Trained on clicks
- Trained on messages
Data
| Metric | Trained on clicks (%) | Trained on messages (%) |
|---|---|---|
| Clicks | 11 | 2.1 |
| Messages | -3.1 | 6.3 |
| Purchases | -0.7 | 3.8 |
| New sellers | -4.2 | -0.6 |
What a small drop in purchases costs across all of Oddments
- Confirmed purchases per day
- 45,00050M impressions a day across all of Oddments × 0.09%
- Average sale price
- $60
- Fee on in-app payments
- 10%
- Fee revenue per day45,000 × $60 × 10%$270,000from Confirmed purchases per day, Average sale price and Fee on in-app payments
- A 0.7% drop in purchases (the click arm)$270,000 × 0.7%$1,890 a dayfrom Fee revenue per day
- Over a year$1,890 × 365 = $689,850≈ $0.69Mfrom A 0.7% drop in purchases (the click arm)
- The click-trained model looked like an 11% win on its own metric and would have cost about $0.7M a year in fees.
- Only the business metric in an A/B test catches this. Offline click metrics would have rewarded the wrong model.
What real systems optimise
The objective in one sentence
For <population> at <decision point>,
predict <label> within <window>
given <context>,
so that <business metric> goes up
without <guardrail> getting worse.
For buyers opening the home feed,
predict P(message within 24 h)
for each candidate listing,
given buyer history and listing features,
so that purchases per weekly active
buyer go up,
without raising the share of sold
listings shown or cutting new
sellers' reach.
When the objective, not the model, is the problem
Rules of ML #38 describes a late-stage symptom: the team keeps adding features, offline numbers keep improving, and launches stop moving the product metric. At that point the objective has drifted from the goal, and more features will not fix it. Rule #39 adds that launch decisions always weigh several metrics because none of them is the long-term goal itself; the objective is a working stand-in, reviewed at each launch. Choosing the proxy well is the subject of proxy metrics, and whether offline gains carry over to the A/B test is in online vs offline. Choosing and thresholding guardrails such as new-seller reach is in guardrails. Once the objective is written, the next step is naming the shape of the output, in task types.
Trade-offs.
The chosen option is first; the others stay visible so the reasoning can be checked.
- Pro:6,000 positives per 1M impressions, 3,000 a day even in a new city
- Pro:Arrives within hours, so a weekly retrain sees complete labels
- Pro:Tied to one listing, and 15% of messages end in a confirmed purchase
- Con:Rewards listings that draw questions rather than deals
- Con:Needs purchases as the A/B launch criterion to catch that drift
Weakest link to purchases (2.25% of clicks); Invites clickbait photos and fake low prices
900 per 1M impressions; 100,000 ÷ 450 a day ≈ 222 days in a new city, ≈ 229 with the 7-day attribution wait; Median 3-day delay, so recent examples look negative; Missing for cash deals, a biased label
Dominated by a few expensive items; Cheap-item sellers lose reach
- Pro:Easy to debug; one number to explain when the A/B test disagrees
- Pro:Rules of ML #12: early on, most sensible metrics rise together
- Con:Other signals (saves, purchases) only enter as features or guardrails
Three models to train and monitor; The blend weights become a product decision that needs its own experiments (see blending)
How framing goes wrong
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| The objective is a proxy that drifts from the goal3Feed model | Clicks go up and purchases go down | The A/B test shows the trained metric up and the business metric flat or down | Pick a label closer to the goal and make the business metric a launch criterion | The feed still loads and fills, but buyers buy less from it |
| Labels exist only for listings that were shown5Event logs | The model never learns about listings it ranked low, and they stay low | New or low-ranked listings never gain impressions | Reserve a few exploration slots per feed (see exploration slots) | Buyers keep seeing the same established sellers; new listings sit unseen |
| Label delay6Training job | Purchases from the last 7 days are still unknown, so recent examples look negative | The positive rate dips at the end of every training window | Wait out the attribution window, or model the delay (Chapelle 2014; see implicit labels) | Each retrain under-rates the newest listings until their labels arrive |
| Population mismatch | A model trained mostly on heavy users serves everyone | Slice metrics by buyer tenure | Cap examples per user, as YouTube did (Covington §3.4; see implicit labels) | Light and new buyers get a feed tuned to someone else's habits |
| Someone else depends on what the objective used to mean3Feed model | Trust and safety's scam classifier reads P(message) as a feature (lots of interest in a cheap listing is a warning sign); switching to a price-weighted objective quietly shifts that input, and the scam alert volume moves with it | None in the feed metrics; the review team notices more or fewer alerts with no change on their side | List who reads the score before changing the objective, and give them a versioned contract (the wider hidden-cost list is in when not to use ML; Sculley et al. 2015) | The feed itself keeps working; the damage is elsewhere |