This page collects concrete examples of multi-agent systems: documented LLM-based systems from research and practice, typical patterns for companies, and examples from classic agent research. For each example you will find its structure, its roles and the lesson you can take from it for your own projects. The concepts and architectures behind them are explained on the page Multi-agent systems.
Documented LLM-based systems
1. Research system with orchestrator and subagents (Anthropic)
Anthropic has described how Claude's research feature is built: a lead agent plans the research, starts several subagents that search different aspects in parallel, and combines their results. A further step checks the citations.
- Pattern: Supervisor with parallel specialists
- Result according to Anthropic: Clearly better results than a single agent on broad research tasks, but a multiple of the token consumption
- Lesson: Multi-agent pays off for tasks that parallelize well and whose value justifies the higher cost
2. Software development as a simulated company (MetaGPT, ChatDev)
MetaGPT models a software company: product manager, architect, project manager, developer and QA work according to fixed standard operating procedures (SOPs) and hand over structured documents such as requirements, system design and code. ChatDev takes a similar approach with roles in a virtual company.
- Pattern: Fixed pipeline with roles
- Lesson: Structured handoffs (documents with a fixed format) work better than free-form chats between agents
3. Coding agents with subagents
Tools such as Claude Code or OpenHands use subagents for well-defined subtasks, such as searching code, writing tests or reviewing changes. The main agent keeps the overview, while the subagents work with their own, smaller context.
- Pattern: Main agent with specialized subagents
- Lesson: Separate contexts prevent a single agent from being overloaded with information
4. Customer service with handoffs
The OpenAI Agents SDK shows a customer service setup in its examples in which a triage agent hands requests over to specialized agents, for example for seat selection, flight status or FAQ. Each specialist has its own tools and rules; guardrails check inputs and outputs.
- Pattern: Network with handoffs
- Lesson: Clear responsibilities and handoff rules make the behavior testable
Documented cases from companies
The examples above come from the makers of the models and frameworks. Published cases show how agents run inside companies. The results are statements by the companies or their technology vendors and have not been independently verified.
- Allianz, claims handling: In Australia seven specialised agents (planner, cyber, coverage, weather, fraud, payout and audit) prepare claims under AUD 500, and a human approves the payout. According to Allianz, processing and settlement take 80 % less time, one day at most instead of several days (Allianz).
- Commerzbank, customer service: According to Microsoft, the assistant Ava resolves 75 % of more than 30,000 conversations a month on its own (Microsoft, vendor statement).
- Klarna, customer service: The AI assistant handled 2.3 million conversations in its first month and resolved requests in under 2 instead of 11 minutes (Klarna). From May 2025 Klarna put more people back into customer service because quality had suffered (Fortune). The case shows why quality measurement and a handover to humans belong in the design from the start.
More cases with industry, function, level of autonomy and result are collected in the Atlas of Agentic Organization (in German), our sister project.
Patterns for companies
The following examples show typical setups. They are deliberately small, because two to four agents are enough for most tasks.
Quote preparation in sales
| Agent | Task | Tools |
|---|---|---|
| Research | Collect customer data, previous quotes and call notes | CRM, document storage |
| Pricing | Compile line items and prices according to the price list | ERP, price list |
| Writing | Draft the cover letter and scope of services | Templates |
| Review | Check completeness, discount limits and wording | Rule set |
A person approves the quote. You can implement this with CrewAI (a crew with four roles) or LangGraph (a graph with approval via interrupt), for example.
Incoming invoices in accounting
- Extraction agent reads the invoice (PDF, e-invoice) and converts it into a fixed data format.
- Matching agent compares it with the purchase order and goods receipt.
- Coding agent suggests the account and cost center.
- Approval: Discrepancies and amounts above a threshold go to a person.
Here a fixed workflow with individual agent steps is often more robust than free collaboration, for example in n8n or with CrewAI Flows.
Support triage
- Triage agent determines category, urgency and language.
- Knowledge agent searches for matching articles and drafts a suggestion.
- Technical agent checks logs or system status when there are error messages.
- Human sends or corrects the reply.
Market monitoring
Several research agents monitor competitors, trade press and tender portals in parallel. A summary agent produces a weekly report with sources. The pattern is a smaller-scale version of the research system from example 1.
Examples from classic agent research
Multi-agent systems existed long before language models. These examples mostly use rule-based or learning agents without an LLM, but they show the same basic problems of coordination and negotiation.
| Domain | Structure | What you learn from it |
|---|---|---|
| Logistics | Vehicle, warehouse and order agents negotiate routes, for example via auctions | Decentralized negotiation responds flexibly to disruptions |
| Energy grids | Producers, storage and consumers trade on local electricity markets | Many small decisions can stabilize an overall system |
| Robotics | Swarms of robots share exploration or transport tasks | Simple local rules produce complex overall behavior |
| Traffic simulation | Each vehicle is an agent with its own goal | Simulation helps test measures before implementing them |
| Games | Non-player characters with their own goals and knowledge | Believable behavior emerges from interaction |
There are dedicated tools for this kind of system, such as JADE (Java) or Mesa (Python, for agent-based simulation).
What the examples have in common
- Clear roles: Each agent has a narrowly defined task.
- Structured handoffs: Results are passed on in fixed formats, not as free text.
- Stop rules: There is an upper limit on steps or rounds.
- Human control: A person decides on actions with external impact.
- Measurement: Quality and cost are compared against a single agent or the existing process.
Frequently asked questions
How many agents make sense?
Two to four for most business tasks. Each additional agent increases cost, run time and sources of error at handoff.
Do I need a framework for these examples?
A framework helps for the LLM-based examples, such as LangGraph, CrewAI, the OpenAI Agents SDK or the Microsoft Agent Framework. Simpler variants can also be built with no-code platforms. See AI agent comparison and No-Code AI Agents.
Are multi-agent systems always better than a single agent?
No. They pay off for subtasks that can be clearly separated or run in parallel. In many cases, a single agent with good tools is cheaper and more reliable.
Sources
- Anthropic: How we built our multi-agent research system
- MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework (arXiv)
- ChatDev: Communicative Agents for Software Development (arXiv)
- OpenAI Agents SDK: Handoffs
- Claude Agent SDK: Subagents
- Mesa: agent-based modeling in Python
- JADE: Java Agent Development Framework
- Agentic Organization: Atlas (in German)
