Problem framingChoosing the ML task

100%

Choosing the ML task.

Once the objective is written down, name the shape of the answer. A probability for one item, a number, an order over a few hundred items, the best thousand out of thirty million, or a paragraph of new text each need different labels, losses, metrics and serving paths. Most real products are several tasks chained together, and the same question can often be posed as more than one of them.

Beginner19 minUpdated 1 Oct 2026

Builds on From business goal to ML objective.

The idea.

Eight asks from one marketplace app, sorted by what the model must hand back.

Oddments is a made-up app for buying and selling used goods, from sofas to phone cases, where most deals end with a cash handover in person. It has 30 million active listings and about 50 million feed impressions a day. Its ML team gets a steady stream of asks. Written in product language they all sound alike ("use ML to make X better"), but the thing each model hands back is different, and so is the program that consumes it.

Read the table row by row and look at the last two columns. A fraud reviewer needs one probability per listing. A pricing screen needs a dollar figure. A feed needs a score it can sort by, and it never shows that score to anyone. The search box needs a short list pulled out of 30 million. The listing form needs fresh text. Naming which of these you are building is the first technical decision in the design.

Eight Oddments problems, by the answer they need

Oddments problemTask typeModel outputWhat the caller does
Is this listing a scam?Binary classificationP(scam) for one listingHolds listings above a threshold for review
Which category is this listing in?Multiclass classificationOne of 40 categories, with a probability for eachPre-fills the category field
What price will it sell for?RegressionA dollar amount, or a rangeSuggests a price to the seller
Which listings go first in the feed?RankingA score per candidate, used only to sortOrders about 1,000 candidates
Which of 30M listings resemble this one?RetrievalThe ~1,000 nearest items in an embedding spaceHands candidates to the ranker
Write a title from the photosGenerationNew text, one token at a timeShows a draft the seller edits
Is this new listing unusual for its seller and category?Anomaly detectionA score of distance from typical listingsFlags for a closer look; needs no labels
Which listings are the same item posted twice?ClusteringGroups of listingsCollapses duplicates in search results

Four decisions made at once

Choosing the type settles four things before any model is picked. The label: a yes or no per listing, a sale price, a (buyer, listing) pair that led to a message, or a reference paragraph. The loss: log loss, squared error, a softmax over many items, or next-token cross-entropy. The offline metric: AUC for a scam flag means nothing for a feed, where the order of the top ten is what counts. And the serving shape: one cheap call per item, one call scoring a thousand items per request, an index lookup, or a loop that emits a token every few tens of milliseconds. Once the type is fixed, most later choices in the design stop being open questions.

Was this section helpful?

How it works.

Six questions pick the type; each type comes with its own label, loss and serving shape; and many problems can be posed more than one way.

Picking the type

Ask the questions in this order. The early ones are about the size and kind of output, the later ones about whether labels exist.

Six questions to a task type

Six questions to a task type
Six questions to a task typeParts: Is the answer new content (text, image, code)?, Must it choose from more items than you can score per request (say > 10k)?, Is the answer an order over several items?, Do you have labels for the answer?, Looking for groups, or for the odd one out?, Is the label a number or a category?, Generation, Retrieval, Ranking, Clustering, Anomaly detection, Regression, Classification.

yes

no

yes

no

yes

no

no

yes

groups

odd one out

number

category

Is the answer new content (text, image, code)?

Must it choose from more items than you can score per request (say > 10k)?

Is the answer an order over several items?

Do you have labels for the answer?

Looking for groups, or for the odd one out?

Is the label a number or a category?

Generation
Draft a listing title from photos

Retrieval
then rank the survivors
1,000 look-alikes out of 30M

Ranking
Order ~1,000 feed candidates

Clustering
Group duplicate listings

Anomaly detection
Unusual new listings

Regression
Expected sale price

Classification
binary, multiclass or multilabel
Scam or not; category

The tree gives a first answer, not a final one. The second question is really about latency and cost: "more than you can score" depends on how expensive the ranker is, which the estimate below works out. The fourth question is about today: a problem with no labels now (unusual listings) can gain them once reviewers start marking alerts, and then a classifier becomes possible. Google's framing guide draws the line for the last split: classification when the product acts differently at a few fixed cut-offs, regression when the cut-offs are set later in code, because a regression model has no idea which small differences in its output the product cares about.

Labels, losses, metrics and serving, per type

TaskA label looks likeTypical lossOffline metric (pointer)Serving shape
ClassificationA class per example (scam / not scam)cross-entropy (log loss)AUC; precision and recall at the chosen threshold; log lossOne call per item, a few milliseconds
RegressionA number per example (the sale price)squared or absolute error; pinball (quantile) loss for rangesMAE, RMSE; how often the true value falls inside the rangeOne call per item
RankingRelevance per (query, item), often a click or a messagepointwise log loss, or a pairwise / listwise lossNDCG@k, recall@kScore ~1,000 items per request
Retrieval(query, item that was engaged with) pairssoftmax over sampled negatives, or in-batch negatives with a popularity correction (Yi et al.)recall@kEmbed the query, then a nearest-neighbour lookup
GenerationReference outputs, or human preference pairsnext-token cross-entropyHuman or model-graded quality, plus task-specific checksMany sequential steps, one per token
Anomaly detectionNone, or a handful of confirmed casesno label loss; distance from typical listings (isolation-forest path length, autoencoder reconstruction error)Share of the top alerts that reviewers confirmA score per event; threshold set by review capacity
ClusteringNonepairwise similarity above a threshold, then connected components (near-duplicate groups)Purity against a small hand-labelled sampleA batch job, not a per-request call

The metric column is a set of pointers; the offline-metrics topics define them. What matters here is that the rows differ in every column. A team that calls the feed "a classification problem" and reports AUC can ship a model that separates messaged from ignored listings well overall and still puts a mediocre sofa above a great bike in the top three slots, which is all a buyer sees.

Price suggestion has more than one answer

The tree stops at the first type that fits, but a seller's price box can be filled four ways, and each puts something different on the seller's screen. Oddments starts with the median sale price of the 20 most similar sold items, and the when-not-ml topic shows why a model only replaces it once listing volume pays for its upkeep. The rows below are the choices for that later model.

Oddments' price suggestion, four ways

FramingSeller seesGood whenWeak when
Regression on sale price$58One number fits the screenIt hides uncertainty; a $5 phone case and a $900 sofa share one error scale unless you predict log price
Quantile regression$45–$70 (10th to 90th percentile)The screen shows a rangeThe tails are only trustworthy with enough sales per category
Classification into price bands$50–$75The app behaves differently per band (a fee tier, a badge)Band edges are arbitrary, and moving one means retraining
Retrieval of similar sold itemsSimilar bikes sold for $40–$65Sellers trust concrete examplesNeeds enough sold neighbours; quality rests on them. The day-one heuristic is this row with a median on top

Reframings in published systems

YouTube's candidate generator is the best-known example. Covington and colleagues posed "which video is watched next?" as classification with one class per video, then served it without ever running that classifier's softmax: the user vector and the video vectors go into a nearest-neighbour search instead. Trained as classification, served as retrieval. How that works, and why it led to today's two-tower models, is in two-tower retrieval.

Ranking is the other common case. Most rankers are trained pointwise, as a classifier of P(click) or P(message) per item, and the product uses only the order of the scores (the learning-to-rank topic covers pairwise and listwise losses). An ads ranker is stricter: the auction multiplies the predicted click probability by the bid, so it must be calibrated, not just well ordered (He et al. 2014). What calibration means and how it is measured is in calibration.

Why 30M listings force a retrieval stage

Assumptions
Active listings
30,000,000
Ranker cost per listing
20 µs on one coreillustrative gradient-boosted model
Retrieval of 1,000 candidates
≈ 5 msillustrative approximate nearest-neighbour lookup
Candidates kept
1,000
Working
  1. Score every listing for one request30,000,000 × 20 µs600 sfrom Active listings and Ranker cost per listing
  2. Same work spread over 100 cores600 s ÷ 1006 sfrom Score every listing for one request
  3. Rank only the retrieved candidates1,000 × 20 µs20 msfrom Candidates kept and Ranker cost per listing
  4. Retrieval plus ranking5 ms + 20 ms≈ 25 msfrom Retrieval of 1,000 candidates and Rank only the retrieved candidates
  5. Work saved per request600 s ÷ 25 ms24,000×from Score every listing for one request and Retrieval plus ranking
What it means
  • "More than you can score" in the tree is a latency question: even 100 cores cannot score the whole catalogue inside a feed's budget.
  • Retrieval trades a little recall (a good listing the index misses is never ranked) for a 24,000× cut in work. The full funnel is the multi-stage-funnels topic.
Was this section helpful?

In practice.

Each Oddments feature is a chain of tasks. Follow the chains and the hidden wait, the review budget and the right first question for a new ticket all fall out.

Oddments features are chains

Home feed
Retrieval (an embedding lookup plus a "nearby and new" rule) → ranking (pointwise P(message within 24 h)) → policy rules that hide sold items and cap one seller per screen. The video-recommendation topics walk through the same funnel at a larger scale.
Search
Query classification (which category does "trek 7.2" mean?) → retrieval (keyword and embedding matches) → ranking. The first step is a small multiclass classifier; the query-understanding topic covers it.
Scam defence
A classifier trained on reported scams, an anomaly score that catches patterns nobody has reported yet, and a human review queue that turns alerts into new labels. Rules of ML #15 keeps this filtering apart from feed quality ranking.
Listing assistant
Generation (a title and description drafted from the photos) → a classification guardrail (prohibited items, phone numbers or emails in the text) → the seller edits and posts. The generator never publishes on its own.

The Listing assistant's draft hides behind the seller

Assumptions
Title plus description
136 output tokensillustrative
Decode time per token
25 msillustrative hosted model
Seller's time on the rest of the form after the first photo
≈ 45 smore photos, condition, pickup area (illustrative)
Working
  1. Generate the full draft136 × 25 ms3.4 sfrom Title plus description and Decode time per token
  2. Draft time as a share of the seller's form time3.4 s ÷ 45 s≈ 8%from Generate the full draft and Seller's time on the rest of the form after the first photo
What it means
  • Against a person filling in a form, 3.4 s is small, so Oddments starts the draft as soon as the first photo uploads and streams it into the form while the seller carries on.
  • The plan breaks on the bulk-upload path: a seller who picks ten photos at once and taps post within a few seconds waits on the draft. Tokens come out one after another, so a faster server shortens each step but not the sequence; the llm-inference topics cover the serving side.

Where Oddments spends request time

Where Oddments spends request timeScoring all 30M listings takes 600,000 ms and a generated draft 3,400 ms, while a retrieval lookup takes 5 ms and ranking 1,000 candidates 20 ms.1 ms10 ms100 ms1 s10 s100 s1,000 sScore all 30MWrite a draftRank 1,000Retrieve 1,000600 s3.4 s20 ms5 msTaskMilliseconds per request (ms)Where Oddments spends request timeScoring all 30M listings takes 600,000 ms and a generated draft 3,400 ms, while a retrieval lookup takes 5 ms and ranking 1,000 candidates 20 ms.1 ms100 ms10 s1,000 sScore all 30MWrite a draftRank 1,000Retrieve 1,000600 s3.4 s20 ms5 msTaskMilliseconds per request (ms)
Oddments' illustrative numbers from the two estimates on this page, log scale: the ranker scoring all 30M listings on one core, a 136-token listing draft on a hosted model, the ranker on 1,000 candidates, and retrieving those 1,000 from 30M. Clustering is left out: it runs as a nightly batch job.
Data
Taskms per request
Score all 30M600,000
Write a draft3,400
Rank 1,00020
Retrieve 1,0005

The Scam defence chain's human link sets the threshold

Assumptions
New listings a month (big cities)
400,000
Reviewers
5
Checks per reviewer per day
80about 6 minutes each over an 8-hour shift (illustrative)
Working
  1. New listings a day400,000 ÷ 30≈ 13,300from New listings a month (big cities)
  2. Alerts the team can check a day5 × 80400from Reviewers and Checks per reviewer per day
  3. Share of new listings that can be flagged400 ÷ 13,300≈ 3%from Alerts the team can check a day and New listings a day
What it means
  • An anomaly score has no natural cut-off. The threshold is whatever sends about 400 new listings a day, the top 3% by score, to people who can actually look at them.
  • Every reviewed alert is a label. After three months at 400 a day (400 × 90), about 36,000 labelled listings exist, and a supervised classifier becomes an option.

New tickets, and the question that settles each

Oddments ticket as writtenThe question that settles itType
Sellers keep underpricing phonesDoes the listing form show one figure, or a low-to-high band?Regression (quantile, for a band)
Buyers say the feed is full of things they would never buyDoes anyone see a number, or only the order of the ~1,000 candidates?Ranking
Support wants fewer fake listings reaching buyersAre there reviewed reports to learn from, and who acts on a listing above the cut-off?Binary classification
The 'more like this' strip shows unrelated junkIs the answer chosen from all 30M listings or from a short list already fetched?Retrieval
New sellers abandon the listing formWill they accept text they edit before posting, and what checks it first?Generation
One account posted 200 listings in an hour last nightHas anyone ever marked cases like this, or is 'odd' all we have?Anomaly detection
Search shows the same sofa five timesIs the output a group of listings, rebuilt nightly, rather than a verdict per request?Clustering

None of these tickets names a task, and none of the settling questions is about wording. Each asks what lands on a screen or in a queue, and whether anyone has ever labelled such cases; those are the two things the tree above turns on.

Was this section helpful?

Trade-offs.

The chosen framing is listed first; the others stay visible so the reasoning can be checked.

01
Price suggestion, once a model pays its way
Chosen:Quantile regression (a range)
  • Pro:Honest about uncertainty; sellers see a band such as $45–$70
  • Pro:Wide ranges flag listings the model knows little about
  • Pro:The cut-off for any badge stays in code and can move without retraining
Downside we accept:
  • Con:Needs enough sales per category for the 10th and 90th percentiles to mean something
  • Con:Two or three outputs to train and monitor instead of one
Ruled out:Single-number regression

One figure looks more certain than it is; On raw price, errors on sofas swamp errors on phone cases; needs log price

Ruled out:Classification into price bands

Band edges are arbitrary; moving one means relabelling and retraining; A $74 and a $76 item land in different bands with no sense of how close they were

Ruled out:Retrieval of similar sold items

Depends on good item embeddings and on enough sold neighbours; Rare items (a vintage synthesiser) have no close neighbours

02
Scam listings
Chosen:Classifier plus an anomaly score (as a feature and as a separate alert stream)
  • Pro:The classifier catches known scam patterns with high precision
  • Pro:The anomaly stream catches new patterns before anyone has labelled them
  • Pro:Reviewed alerts become labels for the next classifier
Downside we accept:
  • Con:Two models and two thresholds to own
  • Con:The anomaly stream's volume must be tuned to review capacity
Ruled out:Classifier only

Blind to new schemes that have no labels yet; Scammers adapt to whatever it has learned (the fraud-detection adversaries topic)

Ruled out:Anomaly detection only

Many alerts are unusual but innocent (a seller clearing a flat before moving); Precision stays low because it never learns what a scam looks like

03
Feed ranking loss
Chosen:Pointwise classification of P(message)
  • Pro:Log loss keeps the score close to a probability, so later blending or a threshold can reuse it; at launch the policy only sorts by it
  • Pro:Simple to debug; Rules of ML #14 favours an interpretable, calibrated start
  • Pro:Standard tooling, and log loss is easy to track day to day
Downside we accept:
  • Con:Optimises each item on its own, not the order of the top slots
Ruled out:Pairwise or listwise ranking loss

The score is not a calibrated probability; blending it with other predictions or thresholding it needs a separate calibration step; Harder to debug when a single listing's score looks wrong

What goes wrong when the type is wrong

FailureImpactDetectionMitigationMeanwhile
Ranking framed as classification with a hard thresholdThe feed shows the "yes" listings in an arbitrary order, so the best one may sit in slot 40NDCG@10 is poor even though AUC looks goodUse the score to order the candidates, and evaluate the top of the listBuyers scroll further and message less
Regression on raw price across all categoriesError is dominated by expensive items; cheap items get suggestions that are off by 100% or moreBreak error down by price decile and by categoryPredict log price, use a per-category or quantile lossSellers of cheap items ignore the suggestion
Anomaly detection kept where labels now existThe review queue fills with activity that is unusual but fineReviewers dismiss most alerts; the dismiss rate climbs week on weekTrain a classifier on the reviewed alerts and keep the anomaly score as one of its featuresReal scams wait longer in a noisy queue
Generation used where retrieval would doA help answer invents a refund policy that does not existGrounded-answer checks find claims that match no help-centre articleRetrieve the relevant article and quote it (the grounded-answers topic)Buyers act on a wrong policy and support tickets rise
Was this section helpful?
Builds on this
Numeric, categorical and text features
Read next