How to Use Claude AI Agent? A Hands-On Tutorial from Chat Mode to Automated Task Execution

Have you ever felt that chatting with Claude flows nicely, but always leaves you thinking "is that it"? You ask, it answers, and then you still have to go do the thing yourself.
That's exactly what this article addresses. Agent mode lets Claude do more than just answer you — it actively executes steps, calls tools, and processes data. In plain terms: it helps you get things done, not just talk things through.
By the end of this tutorial, you'll have a Claude agent setup capable of running real tasks, including tool calling, task decomposition, and a prompt structure you can copy and use immediately.
Before You Begin: What You'll Need
- Claude API access (an Anthropic Console account, or via AWS Bedrock / GCP Vertex AI)
- Python 3.10+ or a TypeScript environment (the official SDK supports both)
- A basic understanding of API calls — nothing deep; knowing what a request and response are is enough
- Optional: if you're integrating external services (e.g., search, databases, file systems), you'll need the corresponding API keys
If you're still evaluating the differences between Claude and other models, take a look at this Claude vs. Gemini comparison first. The two differ significantly in agent design philosophy, and choosing the wrong tool means painful adjustments later.
Step 1: Understand How Agent Mode Differs from Ordinary Conversation
The typical way of using Claude is: you ask → it answers → done.
Agent mode adds a loop: Claude decides which tool to call → the tool returns a result → Claude continues reasoning → until the task is complete.
Think of it this way: ordinary conversation is like Googling something yourself; agent mode is like asking someone to do the Googling for you, compile it into a report, send it to you, and then file all the relevant documents away.
In Anthropic's API, this loop is implemented through tool_use content blocks. Claude outputs a tool_use block, your code executes that tool, feeds the result back in, and Claude proceeds to the next step.
Step 2: Define Your Tools (Tool Definition)
Tools are the "capabilities" an agent can call upon. In an API call, you describe each tool's name, purpose, and parameters using JSON Schema.
{
"name": "search_web",
"description": "Search the web for the latest information and return the title and summary of the top three results",
"input_schema": {
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "Search keywords"
}
},
"required": ["query"]
}
}
Writing a good description is critical. Claude relies on this text to decide whether and when to call a given tool. Too vague and it will misuse it; too narrow and it will give up on using it entirely. Explicitly stating "this tool should be called when X" works best.
Step 3: Build the Agent Execution Loop
This is the core structure of the entire agent. You need a while loop that continuously processes Claude's responses until it stops calling tools.
import anthropic
client = anthropic.Anthropic()
tools = [search_web_tool, read_file_tool] # Your defined list of tools
messages = [{"role": "user", "content": user_task}]
while True:
response = client.messages.create(
model="claude-opus-4-5",
max_tokens=4096,
tools=tools,
messages=messages
)
if response.stop_reason == "end_turn":
# Task complete — extract the final reply
print(response.content[-1].text)
break
if response.stop_reason == "tool_use":
# Find all tool_use blocks
tool_results = []
for block in response.content:
if block.type == "tool_use":
result = execute_tool(block.name, block.input)
tool_results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": result
})
# Feed the tool results back into the conversation
messages.append({"role": "assistant", "content": response.content})
messages.append({"role": "user", "content": tool_results})
execute_tool is a dispatch function you implement yourself — it routes by tool name to the actual code or API call.
Step 4: Design a System Prompt That Defines the Agent's Boundaries
Without a solid system prompt, an agent can easily lose its way mid-task or overstep. A few key things to include:
- Role definition: what it is and what its primary task is
- Tool usage rules: when it must use a tool and when it should hold off
- Output format: what the final response should look like
- Stopping condition: what "task complete" means
Example excerpt:
You are a data research assistant. When a user poses a question, you should:
1. First determine whether the latest information needs to be searched (if so, use search_web)
2. After synthesizing the information, output the key points in bullet form, no more than 5 points
3. Do not guess at uncertain information — flag it directly as "requires further verification"
For a deeper look at complete configurations across different task scenarios, this Claude AI Agent Implementation Tutorial has more examples worth referencing.
Step 5: Run Your First Real Task and Observe How It Thinks
Once everything is set up, run a small task to test the full flow. I recommend starting with a single tool, single step, for example:
"Search for the most talked-about AI research papers in September 2026 and list three titles with their abstracts."
While it runs, print out the inputs and outputs of every tool call. You'll start to understand Claude's reasoning rhythm — when it decides to search, when it determines it has enough information, and when it asks for more detail.
This observation process is far more valuable than reading documentation.
Common Mistakes and How to Avoid Them
Tool descriptions are too short → Claude won't know when to use them, resulting in either overuse or complete avoidance. Write at least one sentence per tool: "Call this tool when X."
No iteration limit on the loop → If a tool keeps returning errors, Claude may retry indefinitely. Add a max_iterations counter to protect yourself.
Forgetting to append the assistant's content back into messages → This is the most common bug. Without this step, Claude loses context and behaves erratically.
Defining too many tools at once → Claude's tool selection degrades with too many options. Start with 3–5 tools, confirm the behavior is stable, then expand.
Advanced: Getting the Agent to Handle Multi-Step Tasks
Once the basic flow is working, the next step is task decomposition. You can explicitly instruct Claude in the system prompt to output an execution plan before starting, then carry it out step by step.
Before you begin, output an execution plan in the following format:
Plan:
1. [Step description]
2. [Step description]
...
Then execute each step in order.
This technique has two benefits: it lets you intervene before things go off track, and Claude's consistency improves once it has "spoken" a plan aloud — this relates to chain-of-thought mechanics, but you don't need to dig into the theory. It works, and that's what matters.
It's also worth noting that this article on the fundamental difference between GenAI and traditional AI explains why LLM-based agents reason so differently from rule engines. If you need to explain to your team why this technology is sometimes "unpredictable," that piece will help you make the case clearly.
Post-Completion Checklist
After running your first agent task, verify the following:
- Tools were called correctly (it didn't just hit
end_turnevery time without doing anything) - Tool results were properly fed back into the conversation
- The final output matches the format you defined in the system prompt
- The loop terminated normally (no infinite execution)
From here, you can start thinking about: adding memory (letting the agent retain information across tasks), integrating real external APIs (Notion, GitHub, Slack), or wrapping this agent into a service others can use.
Once agent mode is up and running, your understanding of what AI can actually do jumps up an entire level.
Frequently Asked Questions
How is a Claude agent different from using ChatGPT plugins directly?
The biggest difference is control. ChatGPT plugins are packaged black boxes with limited room for customization. A Claude agent is built by you via the API — tool definitions, execution logic, and loop control are all yours to write. It's the right choice when you need a tailored workflow or integration with your own systems.
Will token consumption spiral out of control when a Claude agent runs multi-step tasks?
Multi-step tasks do accumulate a significant amount of context. The entire conversation history is carried along after every tool call. It's advisable to set a reasonable max_tokens ceiling and periodically summarize intermediate results during long tasks to avoid hitting the context window limit or losing control of costs.
Can I use a Claude agent without knowing how to write Python?
To unlock the full capabilities of an agent, some basic programming knowledge is still necessary. That said, Anthropic and third-party tools (such as LangChain and n8n) offer higher-level abstractions that lower the barrier for non-engineers — though with correspondingly less flexibility.
What if Claude makes a bad decision during agent execution?
The most practical approach is to define a "confirmation mechanism" in the system prompt, requiring Claude to output a plan and wait for human approval before executing irreversible actions (such as deleting files or sending messages). This is called human-in-the-loop, and it's currently one of the best practices for agent deployment.
Which Claude model is best for running agents?
As of September 2026, claude-opus-4-5 is the top choice for complex multi-step tasks — its reasoning capability and tool-use accuracy are the most reliable. For relatively simple tasks where cost control matters, claude-sonnet-4-5 offers solid value. The recommended approach is to validate your workflow with Opus first, then evaluate whether switching models to reduce costs makes sense.
Share
Related articles

How to Use OpenAI Whisper: From Installation to Subtitle Output, All in One Guide

Codex CLI in Practice: Let OpenAI Write Code for You Right in Your Terminal

Fine-tuning vs RAG: Two LLM Customization Approaches — How to Choose?

How to Choose a Vector Database? Comparing Pinecone, Weaviate, and Chroma