Problem framingFrom business goal to ML objective

100%

From business goal to ML objective.

"Help buyers find things" is a wish, not something a model can learn. Between the goal and the model sit three choices: which product decision the prediction changes, which logged event stands in for success, and which constraints the answer must meet. Get them right and the rest of the design follows; get them wrong and a model that looks better offline makes the product worse.

Beginner19 minUpdated 1 Oct 2026

The idea.

Oddments' product lead says buyers are not finding things. That sentence has to become one number a model can learn.

Four rungs from goal to objective

Oddments is a second-hand marketplace: people list sofas, bikes, phones and jackets, buyers scroll a home feed, message the seller and usually meet to hand the item over. The ask from the product lead is "grow completed purchases per weekly active buyer". No model can learn that number directly. It is a weekly total over the whole company, and no single prediction can be credited with it. So the first question is not "which model?" but "which decision could a better prediction improve?" On Oddments the answer is the order of listings in the home feed. Search results, price suggestions and scam review are other decisions, owned by other models.

Once the decision is fixed, the objective can be written as one sentence that names what is predicted, for whom, over what window, and what the prediction is used for: for each buyer and each listing shown in the feed, predict the probability that the buyer messages the seller within 24 hours, and order the feed by it.

The label is the logged event that makes that objective trainable. It has to be observable (it is in the logs), attributable (it can be tied to one impression of one listing) and quick enough to arrive before the next retrain; Google's Rules of ML ask for exactly this of a first objective (rule #13). What must not get worse, such as showing sold items or starving new sellers of views, is written down now as a guardrail and measured later.

The Oddments feed, rung by rung

RungOddmentsQuestion it answers
Business goalMore completed purchases per weekly active buyerWhy does the company care?
Product decisionThe order of listings in the home feedWhat will the prediction change?
ML objectiveP(buyer messages seller within 24 h | listing shown in feed)What number does the model output?
Label1 if a message event follows the impression within 24 h, else 0Which logged event makes it trainable?
GuardrailsShare of sold or stale listings shown; share of impressions going to new sellersWhat must not get worse?
Out of scopeSearch ranking, price suggestions, scam detectionWhich neighbouring decisions belong to someone else's model?

What happens per million feed impressions (illustrative)

Clicks
40,000
4%, within seconds
Saves
12,000
1.2%, within seconds
Messages to the seller
6,000
0.6%, most within 2 h
Confirmed purchases
900
0.09%, median 3 days later

Each event is a candidate label, and each sits at a different distance from the goal. Of the 6,000 messages, 900 end in a confirmed purchase (900 ÷ 6,000 = 15%); of the 40,000 clicks, only 2.25% do. The purchase count is itself incomplete: about 45% of sales close in cash and are never recorded, so the real number is nearer 900 ÷ 0.55 ≈ 1,640. Picking between these events is the heart of framing, and the next section tests them one by one. Turning raw events into clean labels (attribution windows, position bias, joins) is its own subject, covered in implicit labels.

Was this section helpful?

How it works.

Ask the questions that expose the decision, list candidate labels and test them, then write down what the model must deliver.

The questions that do the work

Q1
GoalWhich company number should move if this works, and which experiment will read it?
Purchases per weekly active buyer, read in an A/B test that changes nothing except the feed order.
Q2
DecisionWhat does the score change on screen, and does anyone check it before a buyer sees it?
The order of the home feed, recomputed every time the app opens. Nobody reviews it first.
Q3
LabelWhich logged event, tied to one impression of one listing, would count as success?
Candidates are a click, a save, a message to the seller and a purchase confirmed in the app. Cash deals leave no purchase event.
Q4
Window and delayHow long after the impression does each candidate event land, and how many does a day produce?
Clicks and saves in seconds, messages mostly within two hours, confirmed purchases a median of three days later. Volumes are in the label table below.
Q5
HistoryWho chose what was shown in the logs we would train on?
Six months of feed logs exist, and every impression in them was picked by the current rule, "newest nearby listings first".
Q6
Hard rules and scopeWhat must never happen whatever the score says, and which nearby decisions are not ours?
Sold listings never appear, suspected scams are never boosted, new sellers keep a fair share of views, and no protected attribute is a feature. Search, price suggestions and scam review belong to other models.

Every answer takes options off the table before any model is named. "Nobody reviews it first" means each score reaches a buyer as is, so the hard rules have to be enforced in code that sits outside the model. "A median of three days" means a nightly retrain on purchases would always train on half-finished labels. "Every impression picked by the current rule" means the logs only say how buyers reacted to that rule's choices, not to the listings it skipped. Each of those three answers reappears below: the first as the ordering policy, the second in the label table and the label-completeness rule, the third as the feedback loop in the diagram.

Candidate labels, tested

Where the prediction acts, and where its labels come from

Where the prediction acts, and where its labels come from. The numbered component cards that follow describe each part.
Where the prediction acts, and where its labels come fromComponents: 1. Buyer app (The Oddments app on a buyer's phone. Opens the home feed, shows listings and logs what the buyer does with each one.), 2. Candidate listings (Listings near the buyer that are still for sale about a thousand per feed request.), 3. Feed model (Scores each candidate listing for this buyer. The ML objective lives here and nowhere else.), 4. Ordering policy (Applies business rules to the scores: hide sold listings, cap one seller per screen, boost new listings.), 5. Event logs (Impressions, clicks, saves, messages and purchases, each with a timestamp and the listing it belongs to.), 6. Training job (Joins each impression with the events that followed it into labelled examples, then retrains the feed model.).

Feedback loop: the model picks what it learns from

open feed

~1,000 listings

scores

ordered feed

impressions, messages…

joined to impressions

new weights, weekly

1Buyer app
opens the feed

5Event logs
only shown listings

6Training job
weekly

2Candidate listings
~1
000 per request

3Feed model
the objective lives here

4Ordering policy
hide sold · one seller per screen

Two things in this picture shape the framing. First, the objective lives only in the model box. Rules that must hold exactly, such as never showing a sold sofa, sit in the ordering policy, where they can change on a Tuesday without a retrain; Rules of ML #15 argues for this split between a policy layer and quality ranking. Second, the logs only contain listings the current system chose to show, so the model learns from its own past choices. Sculley et al. call this a feedback loop, and it is covered in feedback loops.

Five labels for the same goal

LabelPositives per 1M impressionsArrives afterObservable and attributable?What goes wrong if you optimise it
Click40,000secondsYes: one tap on one impressionRewards eye-catching photos and too-good prices that do not sell
Save12,000secondsYesBuyers save to compare, not to buy; a weak link to purchases
Message within 24 h6,000minutes to hours (most within 2 h)Yes: the thread starts from the listingCan reward listings that attract questions but not deals, such as vague descriptions
Confirmed purchase within 7 days900median 3 daysPartly: only ~55% of sales are confirmed in the appMissing for cash deals, and biased toward sellers who take in-app payment
Purchase value ($)900 events, weighted by pricemedian 3 daysPartlyPushes expensive items: one $400 sofa counts as much as forty $10 phone cases

How long until a new city has 100k positives

Assumptions
New city feed impressions per day
500,00020k daily buyers × 25 listings seen
Click, message, purchase rates
4%, 0.6%, 0.09%
Positives wanted for a first model
100,000illustrative size
All-Oddments impressions per day
50,000,0002M daily buyers × 25
Working
  1. Click positives500,000 × 4% = 20,000/day; 100,000 ÷ 20,0005 daysfrom New city feed impressions per day, Click, message, purchase rates and Positives wanted for a first model
  2. Message positives500,000 × 0.6% = 3,000/day; 100,000 ÷ 3,000≈ 33 daysfrom New city feed impressions per day, Click, message, purchase rates and Positives wanted for a first model
  3. Purchase positives500,000 × 0.09% = 450/day; 100,000 ÷ 450, plus a 7-day attribution wait≈ 229 daysfrom New city feed impressions per day, Click, message, purchase rates and Positives wanted for a first model
  4. Purchase positives across all of Oddments50,000,000 × 0.09% = 45,000/day; 100,000 ÷ 45,000≈ 2.2 daysfrom All-Oddments impressions per day, Click, message, purchase rates and Positives wanted for a first model
What it means
  • The label closest to the goal can be unusable for most of a year in a small market, while being perfectly fine across the whole marketplace.
  • Label rate and label delay are framing inputs, not data-pipeline details. They decide which objective is trainable at all.
  • This is why Oddments picks messages, a close proxy that arrives within hours, and keeps purchases for the A/B test.

Folding value into a probability

An objective can carry value without a second model, by weighting the positive examples. YouTube's ranker weights each clicked impression by its watch time and each unclicked one by 1, so the odds the logistic regression learns come out close to expected watch time: odds ≈ E[T](1 + P), and with a small click probability P that is ≈ E[T] (Covington et al. 2016, §4.2). Oddments could weight message positives by listing price to lean toward valuable deals. The cost is visible in the label table: cheap items and the sellers who list them lose reach, so the new-seller guardrail has to be watched more closely. Combining several separate predictions into one score is a different technique, covered in blending objectives.

What the framing commits the model to

No.AreaRequirementTarget / measure
01What the score meansWithin one feed request, a higher P(message) must mean a listing more worth showing first. Nothing downstream reads the number as a probability, so ordering matters and calibration does not; that changes if anyone starts thresholding it.
ordering quality within a requestno calibration target at launch
02Label completenessA training cut only includes impressions whose 24-hour message window has closed.
cut-off ≥ 24 h before the snapshot
Without it, the newest day of impressions looks like it produced no messages.
03Who and what gets scoredBuyers with no history and listings posted minutes ago still get a usable score, because both make up a large share of a marketplace feed.
new listingsnew buyers
04New-seller reachThe objective must not starve sellers who joined recently.
≥ 15% of impressions to sellers < 30 days old (illustrative guardrail)
05PrivacyNo protected attributes as features, and deletion requests are honoured.
features reviewed by legal
06Time on the feed pathScoring a full candidate set fits inside the feed's response budget.
p99 < 50 ms for 1,000 listings (illustrative)
The method for sizing a budget like this is in the serving-budget topic of estimation for ML.
07Retrain rhythmWeights and features keep up with what buyers are doing this week.
weekly retrainfeatures minutes old
08SpendServing cost fits the feed's margin.
CPU only at launch
Was this section helpful?

In practice.

How a wrong objective shows up in an experiment, what published systems chose instead, and the objective as one sentence.

Same feed, two objectives

  • Trained on clicks
  • Trained on messages
Same feed, two objectivesThe click-trained feed wins the metric it was trained on, loses purchases and cuts new sellers' impressions by about 4%; the message-trained feed gains clicks, messages and purchases and leaves new sellers' reach almost unchanged.-6%-4%-2%02%4%6%8%10%12%ClicksMessagesPurchasesNew sellers11%-3.1%-0.7%-4.2%2.1%6.3%3.8%-0.6%Change vs current feed (%)MetricSame feed, two objectivesThe click-trained feed wins the metric it was trained on, loses purchases and cuts new sellers' impressions by about 4%; the message-trained feed gains clicks, messages and purchases and leaves new sellers' reach almost unchanged.-6%-4%-2%02%4%6%8%10%12%ClicksMessagesPurchasesNew selle…New sellers11%-3.1%-0.7%-4.2%2.1%6.3%3.8%-0.6%Change vs current feed (%)Metric
Illustrative two-week A/B test on Oddments, each arm measured as a percentage change against the current feed. "New sellers" is the share of impressions going to sellers under 30 days old, the guardrail from the goal-to-objective table: the click arm favours polished photos from established sellers, so newcomers lose reach as well as buyers losing purchases. Covington et al. report the same tap-but-don't-follow-through pull when ranking by click-through rate.
Data
MetricTrained on clicks (%)Trained on messages (%)
Clicks112.1
Messages-3.16.3
Purchases-0.73.8
New sellers-4.2-0.6

What a small drop in purchases costs across all of Oddments

Assumptions
Confirmed purchases per day
45,00050M impressions a day across all of Oddments × 0.09%
Average sale price
$60
Fee on in-app payments
10%
Working
  1. Fee revenue per day45,000 × $60 × 10%$270,000from Confirmed purchases per day, Average sale price and Fee on in-app payments
  2. A 0.7% drop in purchases (the click arm)$270,000 × 0.7%$1,890 a dayfrom Fee revenue per day
  3. Over a year$1,890 × 365 = $689,850≈ $0.69Mfrom A 0.7% drop in purchases (the click arm)
What it means
  • The click-trained model looked like an 11% win on its own metric and would have cost about $0.7M a year in fees.
  • Only the business metric in an A/B test catches this. Offline click metrics would have rewarded the wrong model.

What real systems optimise

YouTube ranking
Expected watch time per impression, not click-through rate. Ranking by clicks promoted deceptive videos that viewers abandoned. The ranker is a weighted logistic regression whose positives are weighted by watch time (Covington et al. 2016, §4.2).
YouTube candidate generation
The label itself was a framing choice. The authors report that the choice of surrogate problem, what exactly the model is asked to predict, moved A/B results far more than offline metrics suggested (§3.4). How they built those labels (predicting the next watch, a fixed number of examples per user) is covered in implicit labels.
Facebook ads
How the output is used decides what it must get right: an auction that consumes the probability needs it calibrated, not only well ordered (He et al. 2014; see calibration in offline metrics). Oddments' feed only sorts, so it sets no calibration target at launch.
Google's rules for a first model
Start with a simple, observable and attributable objective (#13), and keep filtering such as spam removal in a separate policy layer rather than inside the quality score (#15). Early on, most sensible metrics rise together, so the first choice need not be perfect (#12).

The objective in one sentence

For <population> at <decision point>,
predict <label> within <window>
  given <context>,
so that <business metric> goes up
without <guardrail> getting worse.
For buyers opening the home feed,
predict P(message within 24 h)
  for each candidate listing,
  given buyer history and listing features,
so that purchases per weekly active
  buyer go up,
without raising the share of sold
  listings shown or cutting new
  sellers' reach.

When the objective, not the model, is the problem

Rules of ML #38 describes a late-stage symptom: the team keeps adding features, offline numbers keep improving, and launches stop moving the product metric. At that point the objective has drifted from the goal, and more features will not fix it. Rule #39 adds that launch decisions always weigh several metrics because none of them is the long-term goal itself; the objective is a working stand-in, reviewed at each launch. Choosing the proxy well is the subject of proxy metrics, and whether offline gains carry over to the A/B test is in online vs offline. Choosing and thresholding guardrails such as new-seller reach is in guardrails. Once the objective is written, the next step is naming the shape of the output, in task types.

Was this section helpful?

Trade-offs.

The chosen option is first; the others stay visible so the reasoning can be checked.

01
Which label to train the first feed model on
Chosen:Message within 24 h
  • Pro:6,000 positives per 1M impressions, 3,000 a day even in a new city
  • Pro:Arrives within hours, so a weekly retrain sees complete labels
  • Pro:Tied to one listing, and 15% of messages end in a confirmed purchase
Downside we accept:
  • Con:Rewards listings that draw questions rather than deals
  • Con:Needs purchases as the A/B launch criterion to catch that drift
Ruled out:Click

Weakest link to purchases (2.25% of clicks); Invites clickbait photos and fake low prices

Ruled out:Confirmed purchase

900 per 1M impressions; 100,000 ÷ 450 a day ≈ 222 days in a new city, ≈ 229 with the 7-day attribution wait; Median 3-day delay, so recent examples look negative; Missing for cash deals, a biased label

Ruled out:Purchase value

Dominated by a few expensive items; Cheap-item sellers lose reach

02
One objective or several
Chosen:Single objective first
  • Pro:Easy to debug; one number to explain when the A/B test disagrees
  • Pro:Rules of ML #12: early on, most sensible metrics rise together
Downside we accept:
  • Con:Other signals (saves, purchases) only enter as features or guardrails
Ruled out:Blended score of click, message and purchase predictions

Three models to train and monitor; The blend weights become a product decision that needs its own experiments (see blending)

How framing goes wrong

FailureImpactDetectionMitigationMeanwhile
The objective is a proxy that drifts from the goal3Feed modelClicks go up and purchases go downThe A/B test shows the trained metric up and the business metric flat or downPick a label closer to the goal and make the business metric a launch criterionThe feed still loads and fills, but buyers buy less from it
Labels exist only for listings that were shown5Event logsThe model never learns about listings it ranked low, and they stay lowNew or low-ranked listings never gain impressionsReserve a few exploration slots per feed (see exploration slots)Buyers keep seeing the same established sellers; new listings sit unseen
Label delay6Training jobPurchases from the last 7 days are still unknown, so recent examples look negativeThe positive rate dips at the end of every training windowWait out the attribution window, or model the delay (Chapelle 2014; see implicit labels)Each retrain under-rates the newest listings until their labels arrive
Population mismatchA model trained mostly on heavy users serves everyoneSlice metrics by buyer tenureCap examples per user, as YouTube did (Covington §3.4; see implicit labels)Light and new buyers get a feed tuned to someone else's habits
Someone else depends on what the objective used to mean3Feed modelTrust and safety's scam classifier reads P(message) as a feature (lots of interest in a cheap listing is a warning sign); switching to a price-weighted objective quietly shifts that input, and the scam alert volume moves with itNone in the feed metrics; the review team notices more or fewer alerts with no change on their sideList who reads the score before changing the objective, and give them a versioned contract (the wider hidden-cost list is in when not to use ML; Sculley et al. 2015)The feed itself keeps working; the damage is elsewhere
Was this section helpful?
Builds on this
Choosing the ML task
Read next