AI agents4 min read
AI Agent Design Patterns: From Workflows to Multi-Agent
The building blocks of agentic systems explained: prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer, agents and multi-agent.
If you want to build with AI agents — or just understand the products you use — it helps to know the handful of patterns almost every system is made of. This guide follows the framework in Anthropic’s influential article Building effective agents, which drew on work with dozens of teams, and adds lessons from its later multi-agent research.
Workflows versus agents
The article makes a distinction that clears up a lot of confusion:
- Workflows are systems where models and tools are orchestrated through predefined code paths. You decide the steps; the model does the work inside each step.
- Agents are systems where the model dynamically directs its own process and tool use, deciding the steps as it goes.
Both are “agentic systems”. Workflows are predictable and cheaper; agents are flexible but cost more and can compound errors. The central advice: start simple and add complexity only when it demonstrably improves results. Often a single model call with good context and examples is enough.
The building block: the augmented LLM
Every pattern starts from a model enhanced with retrieval (it can look things up), tools (it can take actions) and memory (it can keep information between steps). The Model Context Protocol is one standard way to plug tools and data into a model. Whatever you use, give the model a clear, well-documented interface.
Five workflow patterns
1. Prompt chaining
Split a task into a fixed sequence of steps, each using the previous output. You can add checks between steps — “gates” — to stop early if something is wrong.
Use it when the task decomposes cleanly into subtasks. Example: write an outline, check it meets the criteria, then write the document from it.
2. Routing
Classify the input first, then send it to a specialised prompt, tool or model.
Use it when inputs fall into distinct categories that are best handled differently. Example: route customer messages to refund, technical or general flows; send easy questions to a small, cheap model and hard ones to a more capable one.
3. Parallelization
Run several model calls at once and combine the results. Two variants:
- Sectioning: split the task into independent parts, such as one call answering the user while another screens for inappropriate content.
- Voting: run the same task several times to get diverse attempts, such as several reviews of code for vulnerabilities, flagging anything most of them find.
Use it when parts are independent (for speed) or when several perspectives raise confidence.
4. Orchestrator-workers
A central model breaks the task down dynamically and delegates subtasks to worker calls, then combines their results.
Use it when you cannot predict the subtasks in advance. Example: a coding change whose affected files depend on the request.
5. Evaluator-optimizer
One model produces a result; another evaluates it against criteria and gives feedback; the loop repeats until the result passes.
Use it when you have clear evaluation criteria and iteration measurably helps. Example: literary translation, where a reviewer catches nuances the first pass missed.
Agents
When a task is open-ended and the number of steps is unpredictable, a full agent — a model using tools in a loop based on feedback from its environment — is the right tool. The implementation is often simple; the hard parts are tool design, clear goals and stopping conditions, and testing in sandboxed environments with guardrails, because errors and costs can accumulate.
Good agents pause for human input at checkpoints or when blocked, and their plans should be visible so people can follow and correct them.
Multi-agent systems
A multi-agent system is the orchestrator-workers pattern with agents in each role. In Anthropic’s research feature, a lead agent plans the research and starts several sub-agents that search in parallel; it then combines their findings. On Anthropic’s internal research evaluation, this setup outperformed a single agent by 90.2%.
The cost is significant: Anthropic reported that agents use about 4× more tokens than chat, and multi-agent systems about 15× more. Multi-agent designs work best for valuable tasks that split into independent parts, and poorly when all agents need the same context or depend heavily on each other — which is the case for most coding tasks.
Three principles for building agents
- Simplicity. Keep the design as simple as the task allows.
- Transparency. Show the agent’s planning steps so users can understand and correct it.
- A well-crafted agent-computer interface. Put as much care into tool descriptions and parameters as into a user interface.
Tool design tips
Tools are how an agent perceives and acts, and the article’s appendix argues they deserve as much attention as prompts:
- Describe each tool as you would to a new colleague: purpose, parameters, examples, edge cases.
- Choose formats that are natural for the model, without fiddly overhead such as counting lines or escaping.
- Make mistakes hard (“poka-yoke”): for example, require absolute file paths if relative ones confuse the model.
- Test tools with many example inputs and watch where the model goes wrong.
When building its coding agent for the SWE-bench benchmark, Anthropic reports spending more time optimising the tools than the overall prompt.
How to choose
Ask, in order: Can one good prompt do it? Can a fixed chain or router do it? Do the steps really need to be decided on the fly? Is the task valuable enough for the extra cost? Can errors be caught? Only climb to agents — and then to multiple agents — when the answers justify it, and measure results with a set of test cases (evals) at every step.
Frequently asked questions
Do I need an agent framework?
Not to start. Many of these patterns take only a few lines of code with a model’s API directly. If you use a framework, make sure you understand what it does under the hood.
Which pattern should I try first?
The simplest one that could work: usually a single well-prompted call, then prompt chaining or routing. Move to agents only when steps cannot be planned in advance.
Are multi-agent systems always better?
No. They help when work can be split into independent parts, like research across many sources, but they use far more tokens and do poorly when every part needs the same context, as in most coding tasks.


