Running Automated Tasks with Claude AI Agent: A Hands-On Tutorial from Scratch

After completing this tutorial, you'll have a Claude Agent that can genuinely execute tasks autonomously — one that calls external tools, decides on next steps based on results, and doesn't simply answer a single question and stop.
This isn't one of those "use Claude to write code" articles. What I'm covering here is agent architecture: letting the model decide on its own whether to call a tool, which tool to call, and then feeding the results back in to continue reasoning. In plain terms, you give it a goal and it figures out how to accomplish it — you don't have to stand over its shoulder at every step.
What You'll Need
- Claude API key: Apply at console.anthropic.com — new accounts get free credits to experiment with
- Python 3.10+: This tutorial uses Python, since Anthropic's official SDK is the most complete
- Basic Python knowledge: You don't need to be a senior developer, but you should be comfortable reading functions, dicts, and loops
- A task you want to automate: This tutorial demonstrates "check the weather + decide whether to bring an umbrella" — simple enough to be approachable while still showcasing agent logic
If you're still evaluating whether to run agents on Claude versus another model, take a look at Claude Model Comparison: The Real Differences Between Haiku, Sonnet, and Opus first. For agent tasks, Sonnet is the most cost-effective choice in the majority of situations.
Step 1: Install the SDK and Configure Your Environment
pip install anthropic
export ANTHROPIC_API_KEY="sk-ant-xxxx"
It's recommended to manage your key with a .env file rather than hardcoding it directly in your code.
import anthropic
import os
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
Step 2: Define Your Tools
The core of an agent is tool use. You tell Claude which tools are available and what each one looks like, and Claude decides whether to call one, which one to call, and what parameters to pass.
tools = [
{
"name": "get_weather",
"description": "Query the current weather and precipitation probability for a specified city",
"input_schema": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "City name, e.g. Taipei"
}
},
"required": ["city"]
}
}
]
This schema follows the JSON Schema specification. Once Claude reads it, it knows that calling this tool requires passing in a city string.
Step 3: Implement the Actual Tool Logic
The tool definition simply lets Claude "know" the tool exists — the actual execution is still handled by your Python code. Here we use mock data to simulate a weather API:
def get_weather(city: str) -> dict:
# In a real scenario, replace this with a call to OpenWeatherMap or any weather API
fake_data = {
"Taipei": {"temp": 31, "rain_probability": 75, "condition": "Partly cloudy with showers"},
"Tokyo": {"temp": 28, "rain_probability": 20, "condition": "Clear"},
}
return fake_data.get(city, {"error": "City not found"})
def execute_tool(tool_name: str, tool_input: dict):
if tool_name == "get_weather":
return get_weather(tool_input["city"])
return {"error": "Unknown tool"}
Step 4: Build the Agent Loop
This is the most important part of the entire tutorial. The agent loop works like this: send a message → Claude decides the next step → if a tool call is needed, execute it → feed the result back → continue, until Claude signals it's done.
def run_agent(user_message: str):
messages = [{"role": "user", "content": user_message}]
while True:
response = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=1024,
tools=tools,
messages=messages
)
# Claude signals it's done — return the final text
if response.stop_reason == "end_turn":
for block in response.content:
if hasattr(block, "text"):
return block.text
# Claude wants to call a tool
if response.stop_reason == "tool_use":
# Add Claude's response to messages
messages.append({"role": "assistant", "content": response.content})
# Execute each tool call
tool_results = []
for block in response.content:
if block.type == "tool_use":
result = execute_tool(block.name, block.input)
tool_results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": str(result)
})
# Send the tool results back to Claude
messages.append({"role": "user", "content": tool_results})
Think of it as a back-and-forth ping-pong rally: you serve (user message) → Claude returns it (possibly requesting a tool) → you hit the tool result back → Claude continues, until it decides to stop.
Step 5: Run It and See the Results
result = run_agent("I'm going to Taipei tomorrow — please check the weather and tell me whether I need to bring an umbrella")
print(result)
The output will be something like: "Taipei is currently 31°C with partly cloudy skies and showers, and a 75% chance of rain. It's strongly recommended that you bring an umbrella and consider wearing a waterproof jacket."
Claude decides on its own to call get_weather, translates the result into plain language on its own, and gives its own recommendation — that's the difference between an agent and ordinary Q&A.
Common Pitfalls and How to Avoid Them
Tool results must be converted to strings: The content field of a tool_result must be a string — passing a dict directly will throw an error. The str(result) above is the simplest approach; for production use, json.dumps(result) is the better choice.
Forgetting to add the assistant message back to messages: This is the most common bug. Every time Claude responds — whether or not it makes a tool call — you must first add that round's response.content to messages before continuing the conversation.
A loop with no upper bound: Claude can occasionally get stuck in a "keep calling tools" loop, especially when the task is vaguely defined. It's recommended to add a counter and force a stop after 10 iterations.
Not accounting for token usage: Tool results consume tokens too. If a tool returns a large payload — say, a search result with several thousand words — the context window will fill up quickly. In those cases, you'll need to consider truncating or summarizing the output. On a related note, if API costs are a concern, the analysis of Claude API pricing in How Much Are You Actually Spending on AI Tools Each Month? A Full Breakdown of Major Subscription Costs is worth reading.
Going Further: Chaining Multiple Tools for a More Capable Agent
One tool is enough to illustrate the concept, but real-world scenarios typically require combining multiple tools. You can add more tools to the same tools array, and Claude will decide on its own which tool to call at each step:
search_web: Look up real-time informationread_file/write_file: Read and write local filessend_email: Send notification emailsquery_database: Query a database
Put these together and you have a complete agent pipeline that can "retrieve data → analyze → write a report → send an email." This is also the direction OpenAI is pursuing with its "general-purpose AI Agent" vision — though OpenAI Wants to Build a Universal AI Agent, But There's a Gap Between Can Use and Will Use covers the real-world adoption gap in detail, and it's worth a read.
Final Checklist
- The agent loop terminates correctly (
stop_reason == "end_turn") - The
tool_use_idin each tool call correctly maps to its correspondingtool_result - The messages array maintains the correct role order: user → assistant → user → assistant (no two consecutive messages with the same role)
- A loop limit is in place to prevent infinite loops
- Tool results are in string format
As a next step, consider wrapping this agent in a FastAPI endpoint, or using a framework like LangGraph or LlamaIndex to manage more complex agent state. But before reaching for a framework, get this bare-SDK version running first — that's how you'll truly understand what an agent is actually doing.
Frequently Asked Questions
What's the difference between a Claude Agent and a direct Claude API call?
A direct API call is a single question-and-answer exchange — you ask, it answers, and that's it. Agent mode keeps Claude in a loop where it autonomously decides which tools to call and how many times, stopping only when the task is complete. It's far better suited to tasks that require multi-step reasoning or integration with external data sources.
Do I need a Claude Pro subscription to run a Claude Agent?
No. Agents run through the Anthropic API, which is entirely separate from a Claude.ai subscription. All you need is an API key and you pay per usage — new accounts come with free credits so you can start experimenting right away.
Is there a risk of the agent loop running indefinitely?
There is. This is especially likely when the task description is vague or when a tool keeps returning errors, causing Claude to get stuck in a cycle of repeated calls. The recommended fix is to add a counter inside the while loop and force a break after a fixed number of iterations — say, 10 — then return the current state.
Which Claude model should I use for agents?
The current recommendation is claude-sonnet-4-5, which offers the best balance of speed and capability and has the most stable support for tool use. Haiku is cheaper but tends to make mistakes on complex multi-step tasks; Opus is the most powerful but comes at a high cost, making it appropriate only for scenarios that genuinely require deep reasoning.
Are there ready-made frameworks so I don't have to write the agent loop myself?
Yes — LangGraph and LlamaIndex both support Claude and come with built-in agent architectures. That said, it's strongly recommended that you get a working implementation using the bare SDK first. Understanding the structure of tool_use and messages directly will save you from hitting mysterious bugs later on that you have no idea how to diagnose.
Share
Related articles

Is the ChatGPT Model You're Using Right Now Actually the Best One for You?

Claude vs ChatGPT: Choose Based on Your Use Case, Not Feature Tables

Claude Skills Complete Guide: How to Build, How to Use, and the Mistakes Most People Make

Claude 3.5 Sonnet vs. Opus vs. Haiku: The Most Complete Model Comparison for 2026