AI agent system design interview

Six weeks from the agent loop to designing a coding or research agent that is safe, measurable and affordable.

The interview

You design an agent that does real work, such as resolving support tickets, writing code or researching a question: how much it may do on its own, its tools, its context and memory, its guardrails, and how you’ll know it works.

Your route

6 weeks, 6 phases

The 45 minutes

A way to spend the time that leaves room for the part that’s hard. Practise against it until it feels natural.

  1. 0–6 minTask and autonomyWhat the agent does, and how much it may do unsupervised.
  2. 6–12 minAgent loopThe pattern, the models and how a run starts and ends.
  3. 12–20 minToolsWhat it can call, with what inputs, and what can go wrong.
  4. 20–28 minContext and memoryWhat goes in the window, what is remembered, what is retrieved.
  5. 28–36 minSafetyGuardrails, permissions and where a human approves.
  6. 36–43 minEvals and costHow you measure it, and what a run costs.
  7. 43–45 minWrap-upRisks and next steps.

What interviewers listen for

  1. 01AutonomyThe agent gets only as much freedom as the task and its risks justify.
  2. 02Tool contractsTools have clear inputs, outputs and failure modes, and are safe to retry.
  3. 03Context budgetYou decide what fills the context window, and what is left out.
  4. 04MeasurementYou can say how you’d know it works, and what it costs per task.

Week by week

  1. Week 1

    Agent foundations

    Know what makes something an agent, and when not to build one.

    You’ll be able to explain

    • Workflows versus agents
    • The loop: plan, act, observe
    • Reasoning and planning strategies

    PractiseTake three tasks you do at work and decide for each: a script, a workflow or an agent.

    Check yourselfWhen is a fixed workflow better than an agent?

  2. Week 2

    Tools

    Give an agent tools it can use correctly and safely.

    You’ll be able to explain

    • Designing tool interfaces
    • Tool protocols
    • Operating browsers and computers

    PractiseWrite the interface of a “refund order” tool: its inputs, its errors, and what makes it safe to call twice.

    Check yourselfWhat should happen when a tool call times out halfway?

  3. Week 3

    Context and memory

    Fill the context window with what the next step needs, and nothing else.

    You’ll be able to explain

    • Context engineering
    • Short- and long-term memory
    • Retrieval over documents

    PractiseBudget the context window of a support agent: system prompt, history, retrieved documents and tool results.

    Check yourselfWhat does your agent forget between sessions, and should it?

  4. Week 4

    Orchestration

    Coordinate steps, and agents, without losing the thread.

    You’ll be able to explain

    • Workflow patterns
    • Multi-agent designs
    • Shared state and hand-offs

    PractiseRedesign a single agent as an orchestrator with workers, and list what you gained and what you lost.

    Check yourselfWhen does adding a second agent make a system worse?

  5. Week 5

    Safety and operations

    Keep the agent safe, observable and affordable in production.

    You’ll be able to explain

    • Guardrails
    • Human approval points
    • Prompt injection and permissions
    • Evaluating agents
    • Tracing, cost and latency

    PractiseWrite five test cases for an agent, including one that tries to inject instructions through a document.

    Check yourselfWhich actions should never run without a person approving them?

  6. Week 6

    Agent designs and mocks

    Run the full interview on the agents that come up most.

    You’ll be able to explain

    • A repeatable order: task, loop, tools, context, safety, evals
    • Matching autonomy to risk

    PractiseTwo timed mock interviews, one support agent and one coding agent. Score yourself against the four signals.

    Check yourselfCould you explain what your agent costs per task, and how you’d bring it down?

Common pitfalls

  • Reaching for an agent when a workflow would do.
  • Tools with vague inputs and no error handling.
  • Stuffing everything into the context window.
  • No human checkpoint on irreversible actions.
  • No evals, so no way to know if a change helped.