Problem framingWhen not to use ML

100%

When not to use ML.

The strongest answer to "design an ML system for this" is sometimes "not yet" or "not at all". Rules win when the answer is known, exact algorithms win when it is computable, humans win when errors are costly and rare, and heuristics win while there is no data. Know the signs, price the model honestly, and ship the simple version first so that the model, if it comes, has data and a baseline to beat.

Beginner20 minUpdated 3 Oct 2026

Builds on From business goal to ML objective.

The idea.

In one quarter, eight requests reach the Oddments ML team. Only two of them get a model straight away, and in one of those a person still makes the final call.

The comparison is never against nothing

Oddments is a fictional second-hand marketplace: people list bikes, sofas and phones, buyers browse a feed, and most deals close in person. Its ML team gets requests from every corner of the company, and each one arrives phrased as a model: "predict the shipping fee", "learn the best time to send the welcome email", "detect stolen bikes".

The useful question is not whether a model could produce an answer. Almost anything can be learned approximately. The question is whether it would beat the best thing you could ship without one: a rule, a lookup table, a solver, a person, or a quick experiment. That simple option is the benchmark, and the model has to clear it by more than the engineers, compute, labels and monitoring it will consume every month for as long as it runs.

Run the eight requests through that test and they sort themselves. Five are answered without ML, one starts as a heuristic and earns a model later, one pairs a classifier with rules and reviewers, and only the home feed is a model pure and simple.

One quarter of requests, triaged

RequestAnswerWhy
Shipping fee by distance and sizeRuleThe fee schedule is a business decision someone already wrote down; a model trained on past fees would only learn it back, with errors
Delivery routes for Oddments couriersSolverThe best route is computable with a vehicle-routing solver; prediction helps only with inputs such as travel times
"Sort search by price"Not a modelThe buyer picked the order; no prediction changes what happens
Prohibited items (weapons, medicines)Rules + model + humanExact terms are caught by rules, a classifier routes suspicious photos to review, and a person gives the verdict
Price suggestion for new listingsHeuristic, then modelThe median sale price of the 20 most similar sold items works today; a model needs months of sales to learn from
Stolen bikesNot ML (yet)No labels exist. Check frame numbers against police registers and add a report button; once reviewer verdicts on reports pile up, this becomes a binary classifier (see task-types)
Welcome-email send timeExperimentThree fixed times in an A/B test settle it; a per-user model would not repay its upkeep at this size
Home feed ordering (established cities)ModelPersonal, changes every day, and millions of labelled events arrive daily

The quarter in four numbers

Requests
8
all phrased as "build a model"
Answered without ML
5
rule, solver, nothing, lookups, experiment
Heuristic first
1
price suggestion
Model from day one
2
home feed in established cities; prohibited items, with reviewers

Ask before any model

Q-01Would a prediction change a decision that matters?
Q-02Is the right answer already a rule someone could write down?
Q-03Can it be computed exactly, by search, optimisation or a lookup?
Q-04Are there labels or feedback, in volume, soon enough to train on?
Q-05Can the product live with the model's mistakes, or will a person or a rule catch them?
Was this section helpful?

How it works.

A path any request can walk down, then an honest bill for the model that comes out at the bottom.

Does this request need a model?

Does this request need a model?
Does this request need a model?Parts: Would the answer change a decision?, Is the answer a known rule?, Is it computable exactly?, Are there labels or feedback in volume?, Are errors cheap, or caught by a person or rule?, Do not build it, Write the rule, Use an algorithm or solver, Ship a heuristic and log everything, Model assists a human, Build the model.

no

yes

yes

no

yes

no

no

yes

no

yes

Would the answer change a decision?

Is the answer a known rule?

Is it computable exactly?
search
optimisation
lookup

Are there labels or feedback in volume?

Are errors cheap, or caught by a person or rule?

Do not build it
sort by price

Write the rule
shipping fee

Use an algorithm or solver
courier routes
stolen-bike frame lookups
ML may still predict its inputs

Ship a heuristic and log everything
launch prices
or A/B test fixed options: welcome email
Rules of ML #1, #2

Model assists a human
prohibited items

Build the model
home feed

Reading the path

The order matters. The first question removes requests where nothing downstream would act on the answer; "sort by price" dies there because the buyer already chose. The next two catch problems that have an exact answer: a fee table is a rule, and a delivery route is an optimisation a solver finds without guessing. A solver can still take a learned input, such as a predicted travel time per road segment, and that is a much smaller model than one that tries to learn routing end to end. Stolen bikes leave at the same point: a frame number either is or is not on a police register.

Only then does data come up. A request with no labels is not a failed ML project; it is a heuristic project that should record what a model will later need. Zinkevich's first two rules say exactly this: launching without ML is fine, and the metrics and logging come before the first model. The welcome email exits here too, as an experiment rather than a rule: with no per-user labels worth modelling, an A/B test of three fixed send times is the heuristic. The stolen-bike report button also belongs here, since it exists to collect labels for later. The last question is about the cost of a wrong answer. Where one mistake bans an honest seller, the model ranks and a person decides.

Where ML tends to earn its place, and where it doesn't

Fits: personal and changing
Each buyer's feed differs and shifts daily as listings sell. Nobody can write a rule per buyer per day; a model learns one from behaviour.
Fits: many fuzzy signals
Telling a phone from a phone case in a photo, or a real listing from a copied one, depends on hundreds of weak cues. No person can write that rule down.
Poor fit: must be predictable or explainable
Refund eligibility and fees must give the same answer every time, with a reason a user or regulator can check. A rule does that by construction.
Poor fit: rare, costly mistakes with no net
Banning a seller costs them their income if wrong, and the cases are few. A model can put the riskiest accounts first in a queue; a person makes the call.

Is the price model worth it?

Assumptions
New listings a month
400,000big cities (illustrative)
Average sale price
$60
Oddments fee
10% of the salecharged only on in-app payments
Sale-through with no suggestion
30%illustrative
Sale-through with the median-of-similar heuristic
33%illustrative
Sale-through with a learned model
35%illustrative
Model upkeep a month
$35,0002 engineers ≈ $30,000 + $5,000 compute (illustrative)
Share of sales paid in-app
55%the rest close in cash and pay no fee
New listings a month a year later
700,000illustrative growth
Working
  1. Heuristic's fee revenue over no suggestion400,000 × $60 × 3 pts × 10% × 55%$39,600 a monthfrom New listings a month, Average sale price, Oddments fee, Sale-through with no suggestion, Sale-through with the median-of-similar heuristic and Share of sales paid in-app
  2. Model's extra fee revenue over the heuristic400,000 × $60 × 2 pts × 10% × 55%$26,400 a monthfrom New listings a month, Average sale price, Oddments fee, Sale-through with the median-of-similar heuristic, Sale-through with a learned model and Share of sales paid in-app
  3. Model's net value today$26,400 − $35,000−$8,600 a monthfrom Model's extra fee revenue over the heuristic and Model upkeep a month
  4. Same, if cash sales also paid the fee (hypothetical)$26,400 ÷ 55% − $35,000 = $48,000 − $35,000+$13,000 a monthfrom Model's extra fee revenue over the heuristic, Share of sales paid in-app and Model upkeep a month
  5. Listings a month needed to cover upkeep$35,000 ÷ ($60 × 2 pts × 10% × 55%) = $35,000 ÷ $0.066≈ 530,000from Model upkeep a month, Average sale price, Oddments fee and Share of sales paid in-app
  6. Model's net value a year later700,000 × $0.066 − $35,000 = $46,200 − $35,000+$11,200 a monthfrom New listings a month a year later, Listings a month needed to cover upkeep and Model upkeep a month
What it means
  • Most of the value comes from the heuristic: $39,600 of the $66,000 that heuristic and model bring together, or 60%. Rules of ML makes the same point in general: a heuristic often gets a large share of what ML would.
  • At today's volume the model loses $8,600 a month. It would look profitable only if cash sales paid a fee, which they do not. Check that kind of assumption before staffing the project.
  • Upkeep is roughly fixed, so the model pays only past a volume (about 530,000 listings a month here). Compare it with the heuristic, never with doing nothing.

The model's extra revenue against its upkeep

The model's extra revenue against its upkeepThe model's extra fee revenue grows with listing volume and only passes its $35,000 monthly upkeep at about 530,000 new listings a month; today's 400,000 falls short.010k20k30k40k50k60k0200k400k600k800kUpkeep $35kBreak-even 530kToday −$8.6k+$11.2kModel gainExtra fee revenue a month ($)New listings a monthThe model's extra revenue against its upkeepThe model's extra fee revenue grows with listing volume and only passes its $35,000 monthly upkeep at about 530,000 new listings a month; today's 400,000 falls short.010k20k30k40k50k60k0200k400k600k800kUpkeep $35kBreak-even 530kToday −$8.6k+$11.2kModel gainExtra fee revenue a month ($)New listings a month
Illustrative: 2 points of extra sale-through over the heuristic, $60 average price, 10% fee on the 55% of sales paid in-app, so $0.066 per listing. Upkeep stays about the same whatever the volume, so below the break-even the heuristic keeps the job. The two marked points are today's 400,000 listings (−$8.6k a month net of upkeep) and 700,000 a year later (+$11.2k).
Data
New listings a monthModel gain
00
100,0006,600
200,00013,200
300,00019,800
400,00026,400
500,00033,000
600,00039,600
700,00046,200
800,00052,800
  • Upkeep $35k: Extra fee revenue a month ($) = 35,000
  • Break-even 530k: New listings a month = 530,303
  • At 400,000: Today −$8.6k
  • At 700,000: +$11.2k

What the upkeep line hides

$35,000 a month sounds precise; it is usually an underestimate. Sculley and colleagues at Google drew a real ML system as a small box of model code surrounded by much larger boxes: data collection, feature extraction, verification, serving, monitoring. They estimate a mature system may be at most 5% ML code and at least 95% glue. Their paper names the costs that land on the team after launch:

  • CACE, changing anything changes everything: adding one feature to the price model moves every suggestion, so no change is local.
  • Undeclared consumers: the payments team starts reading suggested prices for fraud checks, and now a model tweak breaks their system.
  • Feedback loops: sellers accept the suggestion, so tomorrow's training prices are yesterday's suggestions. How such loops form in general, and how holdouts break them, is in feedback loops.
  • Pipeline jungles: scrapers, joins and backfills grow one fix at a time until nobody can rebuild the training set from scratch.

A rule has almost none of these costs. Retraining, drift checks and on-call for models are covered in monitoring-and-continual-learning.

Was this section helpful?

In practice.

The usual route from no model to a model, what it looks like after a city launch, and how to say "not ML" in an interview.

How a feature earns a model

How a feature earns a model. 5 states, 7 transitions. The table below lists them.
How a feature earns a model5 states, 7 transitions. The table below lists them.

metrics and logs shipped

enough labels [heuristic plateaued]

A/B win / heuristic becomes a feature

A/B loss or upkeep › value

weekly retrain

drift or product change

more data or scale

Rule or heuristic

Heuristic + metrics + logs

First simple model

Richer model

Back to heuristic

3 steps.

Most features start at the top. The arrows back to a heuristic are normal outcomes, not failures: the heuristic is kept alive precisely so there is somewhere to go back to. The first three arrows follow Rules of ML #2 (metrics first), #3 (ML over a complex heuristic) and #7 (heuristics become features).

Transitions of How a feature earns a model
From → ToEventGuardAction
Rule or heuristic → Heuristic + metrics + logsmetrics and logs shipped
Heuristic + metrics + logs → First simple modelenough labelsheuristic plateaued
First simple model → Richer modelA/B winheuristic becomes a feature
First simple model → Back to heuristicA/B loss or upkeep > value
Richer model → Richer modelweekly retrain
Richer model → Back to heuristicdrift or product change
Back to heuristic → Heuristic + metrics + logsmore data or scale
Rule or heuristicstart

Feed purchases per 100 sessions after a city launch

  • Heuristic
  • Model
Feed purchases per 100 sessions after a city launchThe heuristic serves the new city for the first two months; the first model starts below it and only overtakes once labels accumulate.logging22.22.42.62.833.23.405101520first model shipsHeuristicModelPurchases per 100 feed sessionsWeeks since launchFeed purchases per 100 sessions after a city launchThe heuristic serves the new city for the first two months; the first model starts below it and only overtakes once labels accumulate.logging22.22.42.62.833.23.405101520first model shipsHeuristicModelPurchases per 100 feed sessionsWeeks since launch
Illustrative numbers. The heuristic is "newest nearby listings first"; the model is the learned ranker, and the shaded weeks only log labels. Even the feed, a model in the established cities, starts on a heuristic in a city with no history. The heuristic slowly degrades as the catalogue grows and "newest" gets noisier. Eight weeks of the new city's ~3,000 messages a day is about 170,000 positives, past the 100,000 the goal-to-objective topic budgets for a first model. A first model losing to the heuristic at launch is common, which is why the heuristic stays as the baseline and the fallback.
Data
Weeks since launchHeuristicModel
02.6no value
42.6no value
82.52.3
10no value2.6
122.52.9
162.43.1
202.43.2
242.33.3
  • first model ships: Weeks since launch = 8
  • logging: Weeks since launch from 0 to 8

What published guidance and practice say

Rules of ML, #1 and #3
Zinkevich tells teams not to be afraid of launching without ML when they have no data, and estimates a heuristic gets roughly half of the gain a model eventually would. But once a heuristic grows into a tangle of special cases, switch to ML: a complicated hand-written rule is harder to maintain than a model.
Google's problem-framing guide
It asks for the non-ML solution first, as the yardstick, and lists what the data must be before ML is worth it: plentiful, consistent, trusted, available when the prediction is made, correct and representative of real traffic.
People + AI Guidebook
PAIR separates automation, where the system acts alone, from augmentation, where it helps a person. Its signs that AI is a poor fit include users needing predictability or full transparency, and mistakes that are very expensive.
The Netflix Prize
Netflix ran two of the Prize's algorithms in production, but decided the extra accuracy of the full winning blend was not worth the engineering needed to serve it. A better offline score is not, by itself, a reason to ship.

Saying "not ML" in an interview

Q1
PolicyThe interviewer asks for an ML system that sets delivery fees.
Say the fee is a pricing policy, set by the business and kept as a rule so it is predictable. Offer ML where there is genuine uncertainty, such as predicted pickup time, and feed that into the rule or show it to the buyer.
Q2
No data"There's no labelled data."
Propose a heuristic launch with logging and a feedback control (a report or thumbs-down button). Name the label it produces, estimate how many arrive per week, and say when a first model becomes viable.
Q3
High stakes"The decision is banning sellers."
The model ranks accounts for reviewers; a person decides and the seller can appeal. Measure how often reviewers agree with the top of the queue and how many bans are reversed.
Q4
Too simple?"Isn't a heuristic too simple for this interview?"
The heuristic is the baseline the model must beat in the A/B test and the fallback when the model breaks. Designing it well is part of the ML design; model-selection/baselines covers how to build and compare one.
Was this section helpful?

Trade-offs.

Two Oddments launches where the plainer design won, then the ways a model wrapped in rules goes stale.

01
Catching prohibited items
Chosen:Rules + classifier + human review
  • Pro:Exact banned terms are caught the same way every time and are easy to explain
  • Pro:A photo classifier finds what text rules miss (no keywords, misspellings, pictures only)
  • Pro:A reviewer makes the final call, so a false positive costs a delay, not a wrongful removal
Downside we accept:
  • Con:Three systems to run, and a review queue that must be staffed
  • Con:Reviewer throughput caps how aggressive the classifier's threshold can be
Ruled out:Rules only

Blind to photos and deliberate misspellings; Sellers learn the list and route around it within days

Ruled out:Classifier only, auto-remove

Unpredictable to honest sellers, with no clear reason or appeal; Adversaries probe it until they find what passes (see fraud-detection/adversaries)

02
Launching the price suggestion
Chosen:Median-of-similar heuristic now, model later
  • Pro:Ships in about two weeks and delivers the larger share of the value ($39,600 a month)
  • Pro:Creates the log of suggestions, chosen prices and sales that a model will train on
  • Pro:Stays as the baseline in every later A/B test and as the fallback
Downside we accept:
  • Con:Leaves the model's extra 2 points unclaimed until volume passes about 530,000 listings a month
  • Con:Needs its own care (how "similar" is defined, stale sales)
Ruled out:Wait and launch a model

Months with no suggestion at all; No baseline to prove the model is worth its $35,000 a month

Ruled out:No suggestion

Gives up the 3-point sale-through lift the heuristic brings

What goes wrong

FailureImpactDetectionMitigationMeanwhile
Rule layers pile up around the model ("if category is bikes, multiply by 0.9")1Rules engineNobody can say which layer set a given price; each fix breaks another caseMore patches in the rules engine than inputs to the modelFold stable rules into the model as features or into a single policy layer; delete dead ones (Rules of ML #3, #7)Suggestions still appear, but drift from what either the model or the rules intended
A model shipped with no heuristic baseline2Price modelNobody can say what the model adds or whether its upkeep is justifiedThe A/B test compares the model against no suggestionKeep the heuristic as the control arm and as the fallback pathIf the model fails there is nothing to fall back to
Direct feedback loop: suggested prices become the prices the model learns from4Event logsThe model stops learning what buyers would pay and learns what it suggestedThe spread of final prices narrows toward the suggestions over timeHold out a share of listings with no suggestion and train on them (Sculley et al. 2015, direct feedback loops)Suggestions look confident while slowly going stale
A high-stakes decision gets automated3Review teamHonest sellers are removed with no human look and no clear reasonAppeal rate and the share of appeals that reverse the decisionThe model ranks the queue, a person decides (PAIR suggests augmenting rather than automating when stakes are high)Review takes longer; nobody is removed by the model alone
CACE: one new input shifts every prediction2Price modelAn unrelated change (a new category taxonomy) moves prices across the whole catalogueMetrics move after a change that was not meant to touch pricingIsolate sub-models where possible and watch prediction distributions for unexpected shifts (Sculley et al. 2015); keep the previous version ready to serveThe previous model version keeps serving while the change is investigated
Was this section helpful?
Builds on this
Baselines
Read next