Skip to content

Magazine

How to Build an AI Agent: Step-by-Step Guide

How to build an AI agent: set a goal, pick no-code or a framework, build tools and prompts, test. With code example, costs and common pitfalls.

Editorial team · Reviewed

This guide takes you from the first idea to a running AI agent. You define a goal, choose between a no-code platform and a framework, build the agent in six steps and estimate the costs. At the end you will find a complete example and the most common pitfalls. For the basics of what an agent is, see What is an AI agent?.

Step 1: Define the goal and success criteria

Start with a single task that costs time today and whose result you can check. Write it down in one sentence:

"The agent reads new support requests, assigns each one to one of five categories and suggests a reply based on our knowledge base."

Then decide how you will measure success:

  • Quality: e.g. 90% correct categories across 50 test cases
  • Time: e.g. suggested reply in under 30 seconds
  • Cost: e.g. under 5 cents per request
  • Limits: What must the agent not do? (e.g. never send a reply without approval)

The narrower the goal, the more reliable the agent. "Automate our customer service" is not a goal; it is a program made up of many agents and workflows.

Step 2: Choose your tooling: no-code or framework

Criterion No-code platform Framework (code)
Examples n8n, Dify, Make, Zapier Agents, Microsoft Copilot Studio LangGraph, OpenAI Agents SDK, CrewAI, Pydantic AI, Microsoft Agent Framework, Google ADK
Getting started Hours Days
Control over the workflow Medium High
Testing and versioning Limited With standard developer tools (Git, CI, unit tests)
Integrations Many ready-made connectors Build them yourself or connect via MCP
Operations Vendor cloud or self-hosting Your own infrastructure or managed services

Rule of thumb: If the task mainly connects systems (email, CRM, spreadsheets) and the logic is manageable, start with no-code. If you need branching workflows, your own data models, tests or several agents working together, use a framework. For a comparison of frameworks, see the AI agent comparison; for an overview of platforms, see No-code AI agents.

Step 3: Choose a model

For most agents, a current mid-range model from one of the major providers (Anthropic, OpenAI, Google) or a strong open model is enough. Pay attention to:

  • Reliable tool calls: Test with your real tools, not with benchmarks.
  • Context length: Is it enough for knowledge base excerpts and conversation history?
  • Data location: For personal data you need a data processing agreement and ideally EU hosting, for example through Azure, AWS Bedrock or Google Vertex AI.
  • Cost per million tokens for input and output.

Many teams use two models: a cheap one for simple steps such as classification and a stronger one for planning and writing replies.

Step 4: Define the tools

Tools are functions the agent is allowed to call. Follow these rules:

  • Few tools: Start with three to five. Each additional tool increases the error rate when the model picks one.
  • Clear descriptions: Write the name, purpose and parameters so that a new colleague would understand them.
  • Narrow permissions: Read access first. Write actions (sending an email, changing a record) only with approval.
  • Structured return values: JSON instead of free text, so the model can evaluate the results reliably.

If an MCP server already exists for your system, you can connect it directly instead of writing your own tools.

Step 5: Write the instructions

The system prompt describes the role, workflow, limits and output format. A template:

You are a support assistant for the accounting software X.
Workflow:
1. Determine the category (invoice, login, export, bug, other).
2. Use search_kb to find matching articles.
3. Write a suggested reply in an informal tone, no more than 120 words.
Rules:
- Do not invent features that are not in the knowledge base.
- If you find nothing, set needs_human to true.
Output: JSON with category, answer, sources, needs_human.

Step 6: Test, observe, improve

  1. Collect 30 to 50 real examples with their expected solution.
  2. Have the agent process all cases and compare the results automatically or by spot check.
  3. Look at the traces: Which tools were called, and in what order? Where did the model take a wrong turn?
  4. Change only one thing at a time (prompt, tool description, model) and test again.

For tracing, you can use LangSmith, Langfuse, Arize Phoenix or the built-in tracing of the OpenAI Agents SDK. In no-code tools, the execution history helps.

Example: support agent with the OpenAI Agents SDK

The following example uses the OpenAI Agents SDK in Python. The knowledge base here is just a function that you replace with your real search.

pip install openai-agents
export OPENAI_API_KEY=...
from pydantic import BaseModel
from agents import Agent, Runner, function_tool

class Antwort(BaseModel):
    category: str
    answer: str
    sources: list[str]
    needs_human: bool

@function_tool
def search_kb(query: str) -> list[dict]:
    """Searches the knowledge base and returns matching articles with title and URL."""
    # Plug in your real search here (e.g. a vector database or full-text search)
    return [{"title": "Export an invoice as PDF", "url": "https://example.com/kb/export"}]

support_agent = Agent(
    name="Support Assistant",
    instructions=open("system_prompt.txt", encoding="utf-8").read(),
    tools=[search_kb],
    output_type=Antwort,
)

result = Runner.run_sync(support_agent, "How do I get my invoice as a PDF?")
print(result.final_output)

If you do not specify a model, the SDK uses a default model; you set it yourself with the model parameter. With output_type you enforce a fixed output format that you can pass straight into your ticketing system.

The same without code: In n8n, you build a workflow with a trigger (new email or ticket), an AI Agent node with a chat model and a tool node for the knowledge base (e.g. HTTP Request or Vector Store). You write the suggested reply back into the ticket and have a person approve it.

What does an AI agent cost?

Costs come from three components:

Item What it depends on Order of magnitude (example)
Model usage Tokens per task times price per token A few cents per support request with a mid-range model
Platform or hosting License, executions, servers n8n Cloud from €20 per month with annual billing, self-hosting on your own server
Development and maintenance Complexity, tests, integrations Usually a few person-days for a first agent

Calculate model costs before you start: an agent that calls the model five times per request and sends 4,000 tokens of context each time quickly uses 20,000 input tokens. Multiply that by your expected volume and the current price on your provider's pricing page. Prompt caching and smaller models for individual steps reduce costs significantly.

Common pitfalls

  • Goal too broad: The agent is supposed to do everything and ends up doing nothing reliably. Split the task.
  • Too many tools: From around ten tools onward, the accuracy of tool selection drops. Group them or distribute them across subagents.
  • No test cases: Without a fixed set of examples, you only notice regressions in production.
  • No stop condition: Set a maximum number of steps, otherwise the agent runs in loops.
  • Write access without approval: Have people confirm critical actions.
  • Prompt injection: Content from emails or websites can contain instructions. Treat it as data, not as commands, and limit permissions.
  • No monitoring: Without traces and a cost overview, an expensive mistake only shows up on the invoice.

Next steps

Sources

Author: Multi-Agent Navigator editorial team

Last reviewed:

Market notes

Get the next analysis by email

New articles and changes in the directory, at most once a week.