When not to use ML.
The strongest answer to "design an ML system for this" is sometimes "not yet" or "not at all". Rules win when the answer is known, exact algorithms win when it is computable, humans win when errors are costly and rare, and heuristics win while there is no data. Know the signs, price the model honestly, and ship the simple version first so that the model, if it comes, has data and a baseline to beat.
Builds on From business goal to ML objective.
The idea.
In one quarter, eight requests reach the Oddments ML team. Only two of them get a model straight away, and in one of those a person still makes the final call.
The comparison is never against nothing
Oddments is a fictional second-hand marketplace: people list bikes, sofas and phones, buyers browse a feed, and most deals close in person. Its ML team gets requests from every corner of the company, and each one arrives phrased as a model: "predict the shipping fee", "learn the best time to send the welcome email", "detect stolen bikes".
The useful question is not whether a model could produce an answer. Almost anything can be learned approximately. The question is whether it would beat the best thing you could ship without one: a rule, a lookup table, a solver, a person, or a quick experiment. That simple option is the benchmark, and the model has to clear it by more than the engineers, compute, labels and monitoring it will consume every month for as long as it runs.
Run the eight requests through that test and they sort themselves. Five are answered without ML, one starts as a heuristic and earns a model later, one pairs a classifier with rules and reviewers, and only the home feed is a model pure and simple.
One quarter of requests, triaged
| Request | Answer | Why |
|---|---|---|
| Shipping fee by distance and size | Rule | The fee schedule is a business decision someone already wrote down; a model trained on past fees would only learn it back, with errors |
| Delivery routes for Oddments couriers | Solver | The best route is computable with a vehicle-routing solver; prediction helps only with inputs such as travel times |
| "Sort search by price" | Not a model | The buyer picked the order; no prediction changes what happens |
| Prohibited items (weapons, medicines) | Rules + model + human | Exact terms are caught by rules, a classifier routes suspicious photos to review, and a person gives the verdict |
| Price suggestion for new listings | Heuristic, then model | The median sale price of the 20 most similar sold items works today; a model needs months of sales to learn from |
| Stolen bikes | Not ML (yet) | No labels exist. Check frame numbers against police registers and add a report button; once reviewer verdicts on reports pile up, this becomes a binary classifier (see task-types) |
| Welcome-email send time | Experiment | Three fixed times in an A/B test settle it; a per-user model would not repay its upkeep at this size |
| Home feed ordering (established cities) | Model | Personal, changes every day, and millions of labelled events arrive daily |
The quarter in four numbers
Ask before any model
How it works.
A path any request can walk down, then an honest bill for the model that comes out at the bottom.
Does this request need a model?
Reading the path
The order matters. The first question removes requests where nothing downstream would act on the answer; "sort by price" dies there because the buyer already chose. The next two catch problems that have an exact answer: a fee table is a rule, and a delivery route is an optimisation a solver finds without guessing. A solver can still take a learned input, such as a predicted travel time per road segment, and that is a much smaller model than one that tries to learn routing end to end. Stolen bikes leave at the same point: a frame number either is or is not on a police register.
Only then does data come up. A request with no labels is not a failed ML project; it is a heuristic project that should record what a model will later need. Zinkevich's first two rules say exactly this: launching without ML is fine, and the metrics and logging come before the first model. The welcome email exits here too, as an experiment rather than a rule: with no per-user labels worth modelling, an A/B test of three fixed send times is the heuristic. The stolen-bike report button also belongs here, since it exists to collect labels for later. The last question is about the cost of a wrong answer. Where one mistake bans an honest seller, the model ranks and a person decides.
Where ML tends to earn its place, and where it doesn't
Is the price model worth it?
- New listings a month
- 400,000big cities (illustrative)
- Average sale price
- $60
- Oddments fee
- 10% of the salecharged only on in-app payments
- Sale-through with no suggestion
- 30%illustrative
- Sale-through with the median-of-similar heuristic
- 33%illustrative
- Sale-through with a learned model
- 35%illustrative
- Model upkeep a month
- $35,0002 engineers ≈ $30,000 + $5,000 compute (illustrative)
- Share of sales paid in-app
- 55%the rest close in cash and pay no fee
- New listings a month a year later
- 700,000illustrative growth
- Heuristic's fee revenue over no suggestion400,000 × $60 × 3 pts × 10% × 55%$39,600 a monthfrom New listings a month, Average sale price, Oddments fee, Sale-through with no suggestion, Sale-through with the median-of-similar heuristic and Share of sales paid in-app
- Model's extra fee revenue over the heuristic400,000 × $60 × 2 pts × 10% × 55%$26,400 a monthfrom New listings a month, Average sale price, Oddments fee, Sale-through with the median-of-similar heuristic, Sale-through with a learned model and Share of sales paid in-app
- Model's net value today$26,400 − $35,000−$8,600 a monthfrom Model's extra fee revenue over the heuristic and Model upkeep a month
- Same, if cash sales also paid the fee (hypothetical)$26,400 ÷ 55% − $35,000 = $48,000 − $35,000+$13,000 a monthfrom Model's extra fee revenue over the heuristic, Share of sales paid in-app and Model upkeep a month
- Listings a month needed to cover upkeep$35,000 ÷ ($60 × 2 pts × 10% × 55%) = $35,000 ÷ $0.066≈ 530,000from Model upkeep a month, Average sale price, Oddments fee and Share of sales paid in-app
- Model's net value a year later700,000 × $0.066 − $35,000 = $46,200 − $35,000+$11,200 a monthfrom New listings a month a year later, Listings a month needed to cover upkeep and Model upkeep a month
- Most of the value comes from the heuristic: $39,600 of the $66,000 that heuristic and model bring together, or 60%. Rules of ML makes the same point in general: a heuristic often gets a large share of what ML would.
- At today's volume the model loses $8,600 a month. It would look profitable only if cash sales paid a fee, which they do not. Check that kind of assumption before staffing the project.
- Upkeep is roughly fixed, so the model pays only past a volume (about 530,000 listings a month here). Compare it with the heuristic, never with doing nothing.
The model's extra revenue against its upkeep
Data
| New listings a month | Model gain |
|---|---|
| 0 | 0 |
| 100,000 | 6,600 |
| 200,000 | 13,200 |
| 300,000 | 19,800 |
| 400,000 | 26,400 |
| 500,000 | 33,000 |
| 600,000 | 39,600 |
| 700,000 | 46,200 |
| 800,000 | 52,800 |
- Upkeep $35k: Extra fee revenue a month ($) = 35,000
- Break-even 530k: New listings a month = 530,303
- At 400,000: Today −$8.6k
- At 700,000: +$11.2k
What the upkeep line hides
$35,000 a month sounds precise; it is usually an underestimate. Sculley and colleagues at Google drew a real ML system as a small box of model code surrounded by much larger boxes: data collection, feature extraction, verification, serving, monitoring. They estimate a mature system may be at most 5% ML code and at least 95% glue. Their paper names the costs that land on the team after launch:
- CACE, changing anything changes everything: adding one feature to the price model moves every suggestion, so no change is local.
- Undeclared consumers: the payments team starts reading suggested prices for fraud checks, and now a model tweak breaks their system.
- Feedback loops: sellers accept the suggestion, so tomorrow's training prices are yesterday's suggestions. How such loops form in general, and how holdouts break them, is in feedback loops.
- Pipeline jungles: scrapers, joins and backfills grow one fix at a time until nobody can rebuild the training set from scratch.
A rule has almost none of these costs. Retraining, drift checks and on-call for models are covered in monitoring-and-continual-learning.
In practice.
The usual route from no model to a model, what it looks like after a city launch, and how to say "not ML" in an interview.
How a feature earns a model
Most features start at the top. The arrows back to a heuristic are normal outcomes, not failures: the heuristic is kept alive precisely so there is somewhere to go back to. The first three arrows follow Rules of ML #2 (metrics first), #3 (ML over a complex heuristic) and #7 (heuristics become features).
| From → To | Event | Guard | Action |
|---|---|---|---|
| Rule or heuristic → Heuristic + metrics + logs | metrics and logs shipped | ||
| Heuristic + metrics + logs → First simple model | enough labels | heuristic plateaued | |
| First simple model → Richer model | A/B win | heuristic becomes a feature | |
| First simple model → Back to heuristic | A/B loss or upkeep > value | ||
| Richer model → Richer model | weekly retrain | ||
| Richer model → Back to heuristic | drift or product change | ||
| Back to heuristic → Heuristic + metrics + logs | more data or scale |
- Rule or heuristicstart
Feed purchases per 100 sessions after a city launch
- Heuristic
- Model
Data
| Weeks since launch | Heuristic | Model |
|---|---|---|
| 0 | 2.6 | no value |
| 4 | 2.6 | no value |
| 8 | 2.5 | 2.3 |
| 10 | no value | 2.6 |
| 12 | 2.5 | 2.9 |
| 16 | 2.4 | 3.1 |
| 20 | 2.4 | 3.2 |
| 24 | 2.3 | 3.3 |
- first model ships: Weeks since launch = 8
- logging: Weeks since launch from 0 to 8
What published guidance and practice say
Saying "not ML" in an interview
Trade-offs.
Two Oddments launches where the plainer design won, then the ways a model wrapped in rules goes stale.
- Pro:Exact banned terms are caught the same way every time and are easy to explain
- Pro:A photo classifier finds what text rules miss (no keywords, misspellings, pictures only)
- Pro:A reviewer makes the final call, so a false positive costs a delay, not a wrongful removal
- Con:Three systems to run, and a review queue that must be staffed
- Con:Reviewer throughput caps how aggressive the classifier's threshold can be
Blind to photos and deliberate misspellings; Sellers learn the list and route around it within days
Unpredictable to honest sellers, with no clear reason or appeal; Adversaries probe it until they find what passes (see fraud-detection/adversaries)
- Pro:Ships in about two weeks and delivers the larger share of the value ($39,600 a month)
- Pro:Creates the log of suggestions, chosen prices and sales that a model will train on
- Pro:Stays as the baseline in every later A/B test and as the fallback
- Con:Leaves the model's extra 2 points unclaimed until volume passes about 530,000 listings a month
- Con:Needs its own care (how "similar" is defined, stale sales)
Months with no suggestion at all; No baseline to prove the model is worth its $35,000 a month
Gives up the 3-point sale-through lift the heuristic brings
What goes wrong
| Failure | Impact | Detection | Mitigation | Meanwhile |
|---|---|---|---|---|
| Rule layers pile up around the model ("if category is bikes, multiply by 0.9")1Rules engine | Nobody can say which layer set a given price; each fix breaks another case | More patches in the rules engine than inputs to the model | Fold stable rules into the model as features or into a single policy layer; delete dead ones (Rules of ML #3, #7) | Suggestions still appear, but drift from what either the model or the rules intended |
| A model shipped with no heuristic baseline2Price model | Nobody can say what the model adds or whether its upkeep is justified | The A/B test compares the model against no suggestion | Keep the heuristic as the control arm and as the fallback path | If the model fails there is nothing to fall back to |
| Direct feedback loop: suggested prices become the prices the model learns from4Event logs | The model stops learning what buyers would pay and learns what it suggested | The spread of final prices narrows toward the suggestions over time | Hold out a share of listings with no suggestion and train on them (Sculley et al. 2015, direct feedback loops) | Suggestions look confident while slowly going stale |
| A high-stakes decision gets automated3Review team | Honest sellers are removed with no human look and no clear reason | Appeal rate and the share of appeals that reverse the decision | The model ranks the queue, a person decides (PAIR suggests augmenting rather than automating when stakes are high) | Review takes longer; nobody is removed by the model alone |
| CACE: one new input shifts every prediction2Price model | An unrelated change (a new category taxonomy) moves prices across the whole catalogue | Metrics move after a change that was not meant to touch pricing | Isolate sub-models where possible and watch prediction distributions for unexpected shifts (Sculley et al. 2015); keep the previous version ready to serve | The previous model version keeps serving while the change is investigated |