Skip to content
estudIA

AI agents4 min read

Your First Agent With the Claude API, Step by Step

Build a simple tool-using agent in Python: how tools are defined, how the agent loop works and how to add limits and human approval.

In short

  1. Set up the environment. Install the Python SDK and keep your API key in an environment variable.
  2. Define the tools. Describe each tool with a name, a description and its parameters in JSON Schema.
  3. Write the loop. Call the model, run the tools it asks for and send back the results until it finishes.
  4. Add limits. A maximum number of turns, human approval for sensitive actions and a log of what it does.
  5. Test it. Run simple tasks, watch each call and improve the tool descriptions.

An agent is a model that uses tools in a loop until it reaches a goal. It sounds complex, but the idea fits in a few lines of code. In this tutorial we build one in Python with the Claude API: we give it two tools, ask it for a task and watch it decide what to use. By the end you’ll understand what happens inside Claude Code, Codex and the like.

You need some Python and a Claude Console account with an API key.

Step 1: set up the environment

pip install anthropic
export ANTHROPIC_API_KEY="your-key"   # on Windows: setx ANTHROPIC_API_KEY "your-key"

Never paste the key into your code or push it to a repository. Set a spending limit in the console.

Step 2: define the tools

A tool is one of your functions that the model can ask to run. It’s described with a name, a description and its parameters in JSON Schema. The description matters most: it’s all the model sees when deciding when to use it.

import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from the environment

TOOLS = [
    {
        "name": "calculator",
        "description": "Evaluates an arithmetic expression with + - * / and parentheses. Use it for any calculation instead of computing in your head.",
        "input_schema": {
            "type": "object",
            "properties": {"expression": {"type": "string", "description": "For example: (1200 * 0.21) + 15"}},
            "required": ["expression"],
        },
    },
    {
        "name": "save_note",
        "description": "Saves a text note to notes.txt. Use it only when the user asks to save something.",
        "input_schema": {
            "type": "object",
            "properties": {"text": {"type": "string"}},
            "required": ["text"],
        },
    },
]

Step 3: write the real functions

import ast, operator as op

OPS = {ast.Add: op.add, ast.Sub: op.sub, ast.Mult: op.mul, ast.Div: op.truediv, ast.USub: op.neg}

def calculate(expr: str) -> str:
    # Safe evaluation: only numbers and basic operations, never eval().
    def ev(n):
        if isinstance(n, ast.Constant) and isinstance(n.value, (int, float)):
            return n.value
        if isinstance(n, ast.BinOp) and type(n.op) in OPS:
            return OPS[type(n.op)](ev(n.left), ev(n.right))
        if isinstance(n, ast.UnaryOp) and type(n.op) in OPS:
            return OPS[type(n.op)](ev(n.operand))
        raise ValueError("expression not allowed")
    return str(ev(ast.parse(expr, mode="eval").body))

def save_note(text: str) -> str:
    if input(f"Save the note '{text}'? (y/n) ").lower() != "y":
        return "The user declined to save the note."
    with open("notes.txt", "a", encoding="utf-8") as f:
        f.write(text + "\n")
    return "Note saved."

def run_tool(name: str, args: dict) -> str:
    try:
        if name == "calculator":
            return calculate(args["expression"])
        if name == "save_note":
            return save_note(args["text"])
        return f"Unknown tool: {name}"
    except Exception as e:
        return f"Error: {e}"

Note two safety decisions: the calculator doesn’t use eval() (it would run any code) and saving a note asks for human confirmation. These are the minimum guardrails of any agent.

Step 4: the agent loop

This is the heart: call the model, run the tools it asks for, send back the results and repeat until it answers without requesting more tools.

def agent(task: str, max_turns: int = 10) -> str:
    messages = [{"role": "user", "content": task}]
    for _ in range(max_turns):
        response = client.messages.create(
            model="claude-sonnet-5-5",
            max_tokens=4000,
            tools=TOOLS,
            messages=messages,
        )
        if response.stop_reason != "tool_use":
            return "".join(b.text for b in response.content if b.type == "text")

        messages.append({"role": "assistant", "content": response.content})
        results = []
        for block in response.content:
            if block.type == "tool_use":
                print(f"→ {block.name}({block.input})")
                results.append({
                    "type": "tool_result",
                    "tool_use_id": block.id,
                    "content": run_tool(block.name, block.input),
                })
        messages.append({"role": "user", "content": results})
    return "Stopped: the turn limit was reached."

print(agent("A laptop costs €899 before VAT. How much is it with 21% VAT? Save the result in a note."))

When you run it you’ll see something like this: the model asks for the calculator, gets the result, asks to save the note (and you confirm) and finishes with a text answer.

What’s happening

  1. The model receives the task and the list of tools.
  2. If it needs a tool, it responds with stop_reason == "tool_use" and a block with the name and arguments.
  3. Your code runs it and sends back a tool_result with the same tool_use_id.
  4. The model decides whether it needs another tool or can answer.

That’s all there is underneath an agent. “Real” agents add more tools, memory, subagents and context compaction, but the loop is the same.

Step 5: add limits

  • Maximum turns (you already have it) so a bug doesn’t leave it looping and spending tokens.
  • Human approval for any action with side effects: sending, deleting, paying, publishing.
  • Logging of every call: which tool, which arguments and what it returned.
  • Minimal tools: the fewer and more specific, the less can go wrong.

We go deeper in agent safety, and in agent design patterns you’ll see when you need an agent and when a fixed workflow is better.

Next steps

  • Swap claude-sonnet-5-5 for claude-opus-5-5 on harder tasks and compare.
  • Connect real tools through MCP instead of writing them by hand.
  • Try the SDK’s tool runner, which manages this loop for you from decorated functions.

Frequently asked questions

How much does it cost to try?

You pay per token. A simple agent like this costs very little per run with Claude Sonnet 5.5 ($2 per million input tokens and $10 output). Set a spending limit in the console before you start.

Is there a faster way?

Yes: the Python SDK has a tool runner that manages the loop for you from decorated functions. Here we write it by hand so you understand what happens inside.

Glossary terms

Sources

Related articles