Choosing the ML task.
Once the objective is written down, name the shape of the answer. A probability for one item, a number, an order over a few hundred items, the best thousand out of thirty million, or a paragraph of new text each need different labels, losses, metrics and serving paths. Most real products are several tasks chained together, and the same question can often be posed as more than one of them.
Builds on From business goal to ML objective.
The idea.
Eight asks from one marketplace app, sorted by what the model must hand back.
Oddments is a made-up app for buying and selling used goods, from sofas to phone cases, where most deals end with a cash handover in person. It has 30 million active listings and about 50 million feed impressions a day. Its ML team gets a steady stream of asks. Written in product language they all sound alike ("use ML to make X better"), but the thing each model hands back is different, and so is the program that consumes it.
Read the table row by row and look at the last two columns. A fraud reviewer needs one probability per listing. A pricing screen needs a dollar figure. A feed needs a score it can sort by, and it never shows that score to anyone. The search box needs a short list pulled out of 30 million. The listing form needs fresh text. Naming which of these you are building is the first technical decision in the design.
Eight Oddments problems, by the answer they need
| Oddments problem | Task type | Model output | What the caller does |
|---|---|---|---|
| Is this listing a scam? | Binary classification | P(scam) for one listing | Holds listings above a threshold for review |
| Which category is this listing in? | Multiclass classification | One of 40 categories, with a probability for each | Pre-fills the category field |
| What price will it sell for? | Regression | A dollar amount, or a range | Suggests a price to the seller |
| Which listings go first in the feed? | Ranking | A score per candidate, used only to sort | Orders about 1,000 candidates |
| Which of 30M listings resemble this one? | Retrieval | The ~1,000 nearest items in an embedding space | Hands candidates to the ranker |
| Write a title from the photos | Generation | New text, one token at a time | Shows a draft the seller edits |
| Is this new listing unusual for its seller and category? | Anomaly detection | A score of distance from typical listings | Flags for a closer look; needs no labels |
| Which listings are the same item posted twice? | Clustering | Groups of listings | Collapses duplicates in search results |
Four decisions made at once
Choosing the type settles four things before any model is picked. The label: a yes or no per listing, a sale price, a (buyer, listing) pair that led to a message, or a reference paragraph. The loss: log loss, squared error, a softmax over many items, or next-token cross-entropy. The offline metric: AUC for a scam flag means nothing for a feed, where the order of the top ten is what counts. And the serving shape: one cheap call per item, one call scoring a thousand items per request, an index lookup, or a loop that emits a token every few tens of milliseconds. Once the type is fixed, most later choices in the design stop being open questions.
How it works.
Six questions pick the type; each type comes with its own label, loss and serving shape; and many problems can be posed more than one way.
Picking the type
Ask the questions in this order. The early ones are about the size and kind of output, the later ones about whether labels exist.
Six questions to a task type
The tree gives a first answer, not a final one. The second question is really about latency and cost: "more than you can score" depends on how expensive the ranker is, which the estimate below works out. The fourth question is about today: a problem with no labels now (unusual listings) can gain them once reviewers start marking alerts, and then a classifier becomes possible. Google's framing guide draws the line for the last split: classification when the product acts differently at a few fixed cut-offs, regression when the cut-offs are set later in code, because a regression model has no idea which small differences in its output the product cares about.
Labels, losses, metrics and serving, per type
| Task | A label looks like | Typical loss | Offline metric (pointer) | Serving shape |
|---|---|---|---|---|
| Classification | A class per example (scam / not scam) | cross-entropy (log loss) | AUC; precision and recall at the chosen threshold; log loss | One call per item, a few milliseconds |
| Regression | A number per example (the sale price) | squared or absolute error; pinball (quantile) loss for ranges | MAE, RMSE; how often the true value falls inside the range | One call per item |
| Ranking | Relevance per (query, item), often a click or a message | pointwise log loss, or a pairwise / listwise loss | NDCG@k, recall@k | Score ~1,000 items per request |
| Retrieval | (query, item that was engaged with) pairs | softmax over sampled negatives, or in-batch negatives with a popularity correction (Yi et al.) | recall@k | Embed the query, then a nearest-neighbour lookup |
| Generation | Reference outputs, or human preference pairs | next-token cross-entropy | Human or model-graded quality, plus task-specific checks | Many sequential steps, one per token |
| Anomaly detection | None, or a handful of confirmed cases | no label loss; distance from typical listings (isolation-forest path length, autoencoder reconstruction error) | Share of the top alerts that reviewers confirm | A score per event; threshold set by review capacity |
| Clustering | None | pairwise similarity above a threshold, then connected components (near-duplicate groups) | Purity against a small hand-labelled sample | A batch job, not a per-request call |
The metric column is a set of pointers; the offline-metrics topics define them. What matters here is that the rows differ in every column. A team that calls the feed "a classification problem" and reports AUC can ship a model that separates messaged from ignored listings well overall and still puts a mediocre sofa above a great bike in the top three slots, which is all a buyer sees.
Price suggestion has more than one answer
The tree stops at the first type that fits, but a seller's price box can be filled four ways, and each puts something different on the seller's screen. Oddments starts with the median sale price of the 20 most similar sold items, and the when-not-ml topic shows why a model only replaces it once listing volume pays for its upkeep. The rows below are the choices for that later model.
Oddments' price suggestion, four ways
| Framing | Seller sees | Good when | Weak when |
|---|---|---|---|
| Regression on sale price | $58 | One number fits the screen | It hides uncertainty; a $5 phone case and a $900 sofa share one error scale unless you predict log price |
| Quantile regression | $45–$70 (10th to 90th percentile) | The screen shows a range | The tails are only trustworthy with enough sales per category |
| Classification into price bands | $50–$75 | The app behaves differently per band (a fee tier, a badge) | Band edges are arbitrary, and moving one means retraining |
| Retrieval of similar sold items | Similar bikes sold for $40–$65 | Sellers trust concrete examples | Needs enough sold neighbours; quality rests on them. The day-one heuristic is this row with a median on top |
Reframings in published systems
YouTube's candidate generator is the best-known example. Covington and colleagues posed "which video is watched next?" as classification with one class per video, then served it without ever running that classifier's softmax: the user vector and the video vectors go into a nearest-neighbour search instead. Trained as classification, served as retrieval. How that works, and why it led to today's two-tower models, is in two-tower retrieval.
Ranking is the other common case. Most rankers are trained pointwise, as a classifier of P(click) or P(message) per item, and the product uses only the order of the scores (the learning-to-rank topic covers pairwise and listwise losses). An ads ranker is stricter: the auction multiplies the predicted click probability by the bid, so it must be calibrated, not just well ordered (He et al. 2014). What calibration means and how it is measured is in calibration.
Why 30M listings force a retrieval stage
- Active listings
- 30,000,000
- Ranker cost per listing
- 20 µs on one coreillustrative gradient-boosted model
- Retrieval of 1,000 candidates
- ≈ 5 msillustrative approximate nearest-neighbour lookup
- Candidates kept
- 1,000
- Score every listing for one request30,000,000 × 20 µs600 sfrom Active listings and Ranker cost per listing
- Same work spread over 100 cores600 s ÷ 1006 sfrom Score every listing for one request
- Rank only the retrieved candidates1,000 × 20 µs20 msfrom Candidates kept and Ranker cost per listing
- Retrieval plus ranking5 ms + 20 ms≈ 25 msfrom Retrieval of 1,000 candidates and Rank only the retrieved candidates
- Work saved per request600 s ÷ 25 ms24,000×from Score every listing for one request and Retrieval plus ranking
- "More than you can score" in the tree is a latency question: even 100 cores cannot score the whole catalogue inside a feed's budget.
- Retrieval trades a little recall (a good listing the index misses is never ranked) for a 24,000× cut in work. The full funnel is the multi-stage-funnels topic.
In practice.
Each Oddments feature is a chain of tasks. Follow the chains and the hidden wait, the review budget and the right first question for a new ticket all fall out.
Oddments features are chains
The Listing assistant's draft hides behind the seller
- Title plus description
- 136 output tokensillustrative
- Decode time per token
- 25 msillustrative hosted model
- Seller's time on the rest of the form after the first photo
- ≈ 45 smore photos, condition, pickup area (illustrative)
- Generate the full draft136 × 25 ms3.4 sfrom Title plus description and Decode time per token
- Draft time as a share of the seller's form time3.4 s ÷ 45 s≈ 8%from Generate the full draft and Seller's time on the rest of the form after the first photo
- Against a person filling in a form, 3.4 s is small, so Oddments starts the draft as soon as the first photo uploads and streams it into the form while the seller carries on.
- The plan breaks on the bulk-upload path: a seller who picks ten photos at once and taps post within a few seconds waits on the draft. Tokens come out one after another, so a faster server shortens each step but not the sequence; the llm-inference topics cover the serving side.
Where Oddments spends request time
Data
| Task | ms per request |
|---|---|
| Score all 30M | 600,000 |
| Write a draft | 3,400 |
| Rank 1,000 | 20 |
| Retrieve 1,000 | 5 |
The Scam defence chain's human link sets the threshold
- New listings a month (big cities)
- 400,000
- Reviewers
- 5
- Checks per reviewer per day
- 80about 6 minutes each over an 8-hour shift (illustrative)
- New listings a day400,000 ÷ 30≈ 13,300from New listings a month (big cities)
- Alerts the team can check a day5 × 80400from Reviewers and Checks per reviewer per day
- Share of new listings that can be flagged400 ÷ 13,300≈ 3%from Alerts the team can check a day and New listings a day
- An anomaly score has no natural cut-off. The threshold is whatever sends about 400 new listings a day, the top 3% by score, to people who can actually look at them.
- Every reviewed alert is a label. After three months at 400 a day (400 × 90), about 36,000 labelled listings exist, and a supervised classifier becomes an option.
New tickets, and the question that settles each
| Oddments ticket as written | The question that settles it | Type |
|---|---|---|
| Sellers keep underpricing phones | Does the listing form show one figure, or a low-to-high band? | Regression (quantile, for a band) |
| Buyers say the feed is full of things they would never buy | Does anyone see a number, or only the order of the ~1,000 candidates? | Ranking |
| Support wants fewer fake listings reaching buyers | Are there reviewed reports to learn from, and who acts on a listing above the cut-off? | Binary classification |
| The 'more like this' strip shows unrelated junk | Is the answer chosen from all 30M listings or from a short list already fetched? | Retrieval |
| New sellers abandon the listing form | Will they accept text they edit before posting, and what checks it first? | Generation |
| One account posted 200 listings in an hour last night | Has anyone ever marked cases like this, or is 'odd' all we have? | Anomaly detection |
| Search shows the same sofa five times | Is the output a group of listings, rebuilt nightly, rather than a verdict per request? | Clustering |
None of these tickets names a task, and none of the settling questions is about wording. Each asks what lands on a screen or in a queue, and whether anyone has ever labelled such cases; those are the two things the tree above turns on.
Trade-offs.
The chosen framing is listed first; the others stay visible so the reasoning can be checked.
- Pro:Honest about uncertainty; sellers see a band such as $45–$70
- Pro:Wide ranges flag listings the model knows little about
- Pro:The cut-off for any badge stays in code and can move without retraining
- Con:Needs enough sales per category for the 10th and 90th percentiles to mean something
- Con:Two or three outputs to train and monitor instead of one
One figure looks more certain than it is; On raw price, errors on sofas swamp errors on phone cases; needs log price
Band edges are arbitrary; moving one means relabelling and retraining; A $74 and a $76 item land in different bands with no sense of how close they were
Depends on good item embeddings and on enough sold neighbours; Rare items (a vintage synthesiser) have no close neighbours
- Pro:The classifier catches known scam patterns with high precision
- Pro:The anomaly stream catches new patterns before anyone has labelled them
- Pro:Reviewed alerts become labels for the next classifier
- Con:Two models and two thresholds to own
- Con:The anomaly stream's volume must be tuned to review capacity
Blind to new schemes that have no labels yet; Scammers adapt to whatever it has learned (the fraud-detection adversaries topic)
Many alerts are unusual but innocent (a seller clearing a flat before moving); Precision stays low because it never learns what a scam looks like
- Pro:Log loss keeps the score close to a probability, so later blending or a threshold can reuse it; at launch the policy only sorts by it
- Pro:Simple to debug; Rules of ML #14 favours an interpretable, calibrated start
- Pro:Standard tooling, and log loss is easy to track day to day
- Con:Optimises each item on its own, not the order of the top slots
The score is not a calibrated probability; blending it with other predictions or thresholding it needs a separate calibration step; Harder to debug when a single listing's score looks wrong
What goes wrong when the type is wrong
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| Ranking framed as classification with a hard threshold | The feed shows the "yes" listings in an arbitrary order, so the best one may sit in slot 40 | NDCG@10 is poor even though AUC looks good | Use the score to order the candidates, and evaluate the top of the list | Buyers scroll further and message less |
| Regression on raw price across all categories | Error is dominated by expensive items; cheap items get suggestions that are off by 100% or more | Break error down by price decile and by category | Predict log price, use a per-category or quantile loss | Sellers of cheap items ignore the suggestion |
| Anomaly detection kept where labels now exist | The review queue fills with activity that is unusual but fine | Reviewers dismiss most alerts; the dismiss rate climbs week on week | Train a classifier on the reviewed alerts and keep the anomaly score as one of its features | Real scams wait longer in a noisy queue |
| Generation used where retrieval would do | A help answer invents a refund policy that does not exist | Grounded-answer checks find claims that match no help-centre article | Retrieve the relevant article and quote it (the grounded-answers topic) | Buyers act on a wrong policy and support tickets rise |