This guide takes you from the first idea to a running AI agent. You define a goal, choose between a no-code platform and a framework, build the agent in six steps and estimate the costs. At the end you will find a complete example and the most common pitfalls. For the basics of what an agent is, see What is an AI agent?.
Step 1: Define the goal and success criteria
Start with a single task that costs time today and whose result you can check. Write it down in one sentence:
"The agent reads new support requests, assigns each one to one of five categories and suggests a reply based on our knowledge base."
Then decide how you will measure success:
- Quality: e.g. 90% correct categories across 50 test cases
- Time: e.g. suggested reply in under 30 seconds
- Cost: e.g. under 5 cents per request
- Limits: What must the agent not do? (e.g. never send a reply without approval)
The narrower the goal, the more reliable the agent. "Automate our customer service" is not a goal; it is a program made up of many agents and workflows.
Step 2: Choose your tooling: no-code or framework
| Criterion | No-code platform | Framework (code) |
|---|---|---|
| Examples | n8n, Dify, Make, Zapier Agents, Microsoft Copilot Studio | LangGraph, OpenAI Agents SDK, CrewAI, Pydantic AI, Microsoft Agent Framework, Google ADK |
| Getting started | Hours | Days |
| Control over the workflow | Medium | High |
| Testing and versioning | Limited | With standard developer tools (Git, CI, unit tests) |
| Integrations | Many ready-made connectors | Build them yourself or connect via MCP |
| Operations | Vendor cloud or self-hosting | Your own infrastructure or managed services |
Rule of thumb: If the task mainly connects systems (email, CRM, spreadsheets) and the logic is manageable, start with no-code. If you need branching workflows, your own data models, tests or several agents working together, use a framework. For a comparison of frameworks, see the AI agent comparison; for an overview of platforms, see No-code AI agents.
Step 3: Choose a model
For most agents, a current mid-range model from one of the major providers (Anthropic, OpenAI, Google) or a strong open model is enough. Pay attention to:
- Reliable tool calls: Test with your real tools, not with benchmarks.
- Context length: Is it enough for knowledge base excerpts and conversation history?
- Data location: For personal data you need a data processing agreement and ideally EU hosting, for example through Azure, AWS Bedrock or Google Vertex AI.
- Cost per million tokens for input and output.
Many teams use two models: a cheap one for simple steps such as classification and a stronger one for planning and writing replies.
Step 4: Define the tools
Tools are functions the agent is allowed to call. Follow these rules:
- Few tools: Start with three to five. Each additional tool increases the error rate when the model picks one.
- Clear descriptions: Write the name, purpose and parameters so that a new colleague would understand them.
- Narrow permissions: Read access first. Write actions (sending an email, changing a record) only with approval.
- Structured return values: JSON instead of free text, so the model can evaluate the results reliably.
If an MCP server already exists for your system, you can connect it directly instead of writing your own tools.
Step 5: Write the instructions
The system prompt describes the role, workflow, limits and output format. A template:
You are a support assistant for the accounting software X.
Workflow:
1. Determine the category (invoice, login, export, bug, other).
2. Use search_kb to find matching articles.
3. Write a suggested reply in an informal tone, no more than 120 words.
Rules:
- Do not invent features that are not in the knowledge base.
- If you find nothing, set needs_human to true.
Output: JSON with category, answer, sources, needs_human.
Step 6: Test, observe, improve
- Collect 30 to 50 real examples with their expected solution.
- Have the agent process all cases and compare the results automatically or by spot check.
- Look at the traces: Which tools were called, and in what order? Where did the model take a wrong turn?
- Change only one thing at a time (prompt, tool description, model) and test again.
For tracing, you can use LangSmith, Langfuse, Arize Phoenix or the built-in tracing of the OpenAI Agents SDK. In no-code tools, the execution history helps.
Example: support agent with the OpenAI Agents SDK
The following example uses the OpenAI Agents SDK in Python. The knowledge base here is just a function that you replace with your real search.
pip install openai-agents
export OPENAI_API_KEY=...
from pydantic import BaseModel
from agents import Agent, Runner, function_tool
class Antwort(BaseModel):
category: str
answer: str
sources: list[str]
needs_human: bool
@function_tool
def search_kb(query: str) -> list[dict]:
"""Searches the knowledge base and returns matching articles with title and URL."""
# Plug in your real search here (e.g. a vector database or full-text search)
return [{"title": "Export an invoice as PDF", "url": "https://example.com/kb/export"}]
support_agent = Agent(
name="Support Assistant",
instructions=open("system_prompt.txt", encoding="utf-8").read(),
tools=[search_kb],
output_type=Antwort,
)
result = Runner.run_sync(support_agent, "How do I get my invoice as a PDF?")
print(result.final_output)
If you do not specify a model, the SDK uses a default model; you set it yourself with the model parameter. With output_type you enforce a fixed output format that you can pass straight into your ticketing system.
The same without code: In n8n, you build a workflow with a trigger (new email or ticket), an AI Agent node with a chat model and a tool node for the knowledge base (e.g. HTTP Request or Vector Store). You write the suggested reply back into the ticket and have a person approve it.
What does an AI agent cost?
Costs come from three components:
| Item | What it depends on | Order of magnitude (example) |
|---|---|---|
| Model usage | Tokens per task times price per token | A few cents per support request with a mid-range model |
| Platform or hosting | License, executions, servers | n8n Cloud from €20 per month with annual billing, self-hosting on your own server |
| Development and maintenance | Complexity, tests, integrations | Usually a few person-days for a first agent |
Calculate model costs before you start: an agent that calls the model five times per request and sends 4,000 tokens of context each time quickly uses 20,000 input tokens. Multiply that by your expected volume and the current price on your provider's pricing page. Prompt caching and smaller models for individual steps reduce costs significantly.
Common pitfalls
- Goal too broad: The agent is supposed to do everything and ends up doing nothing reliably. Split the task.
- Too many tools: From around ten tools onward, the accuracy of tool selection drops. Group them or distribute them across subagents.
- No test cases: Without a fixed set of examples, you only notice regressions in production.
- No stop condition: Set a maximum number of steps, otherwise the agent runs in loops.
- Write access without approval: Have people confirm critical actions.
- Prompt injection: Content from emails or websites can contain instructions. Treat it as data, not as commands, and limit permissions.
- No monitoring: Without traces and a cost overview, an expensive mistake only shows up on the invoice.
Next steps
- For several agents working together: Multi-agent systems and examples
- For use in a company: AI agents for businesses
- For choosing a framework: AI agent comparison or the System Finder
