Machine learning system design interview

Seven weeks from framing a business goal as a prediction to serving and watching a model in production.

The interview

You turn a product goal, such as better recommendations, catching fraud or answering questions over documents, into an ML system: the prediction to make, the data and features, the model, how to evaluate it, and how to serve and monitor it.

Your route

7 weeks, 6 phases

The 45 minutes

A way to spend the time that leaves room for the part that’s hard. Practise against it until it feels natural.

  1. 0–6 minFramingThe business goal, and the prediction that serves it.
  2. 6–11 minMetricsOffline metrics, online metrics and how they relate.
  3. 11–20 minData and featuresWhere labels come from, what to feature, what could leak.
  4. 20–29 minModelA baseline first, then what earns more complexity.
  5. 29–36 minEvaluationOffline tests, then an A/B test, then slices.
  6. 36–43 minServingLatency, scale, and catching drift.
  7. 43–45 minWrap-upRisks and next steps.

What interviewers listen for

  1. 01FramingThe ML task follows from the product goal, not the other way round.
  2. 02Data realismYou know where the labels come from and what they get wrong.
  3. 03BaselinesYou start simple and say what would justify something bigger.
  4. 04ProductionLatency, cost and drift are part of the design, not an afterthought.

Week by week

  1. Week 1

    Framing and metrics

    Turn a vague goal into a prediction you can measure.

    You’ll be able to explain

    • Choosing what to predict
    • Offline versus online metrics
    • Sizing data, training and serving

    PractisePick three features of apps you use and write the prediction behind each, with one offline and one online metric.

    Check yourselfYour offline metric went up and engagement went down. What happened?

  2. Week 2

    Data and features

    Build training data you’d trust, and features that exist at serving time.

    You’ll be able to explain

    • Collecting and labelling data
    • Class imbalance
    • Feature engineering
    • Feature stores and training-serving skew
    • Privacy constraints

    PractiseList the features for a fraud model, and mark the ones that wouldn’t be available when the prediction is made.

    Check yourselfWhere could label leakage creep into this training set?

  3. Weeks 3–4

    Models

    Match the model to the problem, starting from a baseline.

    You’ll be able to explain

    • Choosing a model family
    • Two-stage retrieval and ranking
    • Text and image models
    • Many objectives at once
    • Exploration versus exploitation

    PractiseFor a recommendation feed, describe a baseline you could ship in a week, then the model you’d move to and what would prove it was worth it.

    Check yourselfWhy split recommendation into retrieval and ranking?

  4. Week 5

    Evaluation

    Prove a model is better, and find where it isn’t.

    You’ll be able to explain

    • Offline metrics for ranking and classification
    • A/B tests and their traps
    • Slices and fairness
    • Evaluating generative output

    PractisePlan an A/B test for a new ranking model: the metric, the guardrails, the sample size and how long it runs.

    Check yourselfThe new model wins overall but loses for new users. Do you ship it?

  5. Week 6

    Serving and MLOps

    Serve predictions fast and cheaply, and know when the model goes stale.

    You’ll be able to explain

    • Batch, online and streaming serving
    • Serving large language models
    • Safe rollouts
    • Monitoring and retraining

    PractiseGive a ranking service a latency budget of 100 ms and split it across its stages.

    Check yourselfWhat tells you the model has drifted before your users do?

  6. Week 7

    Product designs and mocks

    Run the full interview on the problems that come up most.

    You’ll be able to explain

    • A repeatable order: framing, metrics, data, model, evaluation, serving
    • Spending time where this product’s risk is

    PractiseTwo timed mock interviews, one ranking problem and one with generative AI. Score yourself against the four signals.

    Check yourselfDid your design say how the model would be monitored after launch?

Common pitfalls

  • Naming a model before framing the problem.
  • Metrics that don’t connect to the product goal.
  • Features that won’t exist at prediction time.
  • No baseline to compare against.
  • Forgetting latency, cost and monitoring.