AI Tech News HubDaily Updates
AI TechnologySeptember 30, 2026

AI Agent in Practice: Five Steps to Get It Running Your Workflows

A
AI 觀察家
Columnist · 2309 words
AI Agent in Practice: Five Steps to Get It Running Your Workflows

The term "AI Agent" has been beaten to death by 2026, yet I still see plenty of people stuck at the "heard of it, know it's impressive, but have no idea how to actually use it" stage. This article is here to fill that gap — no conceptual fluff, just a complete walkthrough from setup to execution.

What will you have when you're done? An AI Agent capable of automatically completing multi-step tasks like "research a topic → compile a summary → send a report," and you'll have built the entire pipeline yourself — ready to copy and adapt next time.

What You'll Need

Before getting started, make sure you have the following:

  • OpenAI API Key (GPT-4o or above — Agent functionality requires tool-calling capability)
  • An Agent framework: The most widely used options in 2026 are LangChain, AutoGen, or OpenAI's official Assistants API with tools. This article focuses on the OpenAI Assistants API, since it requires no additional framework installation.
  • A basic Python environment (Python 3.10+, with the openai package installed)
  • About 30 minutes of focused time

In plain terms: you're giving the AI a set of "callable tools," then telling it a goal and letting it decide which tools to call, how many times, and how to assemble the results.

Step 1: Create an Assistant and Define Its Role

Open the OpenAI Platform, navigate to the Assistants page, and click "Create."

  • Name: Give it a clear name, such as Research Summarizer
  • Instructions: This is the critical part. Don't just write "you are an assistant" — explicitly tell it the task logic, for example:
You are a research assistant. Upon receiving a topic, you will first use the web_search tool to find the latest information, organize it into 3 key summary points, and then use the send_email tool to send the results to the specified inbox. Report progress after each step is completed.
  • Model: Select gpt-4o — currently the most reliable version for tool calling.

The more specific your instructions, the more predictable the Agent's behavior. A lot of people rush through this step and then spend time wondering "why isn't it doing what I want."

Step 2: Define and Wire Up Tools

An Agent's core capability lies in its ability to call external tools. You'll need to define your tools in JSON Schema format, telling the Agent "what this tool is called, what it does, and what parameters it requires."

tools = [
    {
        "type": "function",
        "function": {
            "name": "web_search",
            "description": "Search the web for the latest information on a given topic",
            "parameters": {
                "type": "object",
                "properties": {
                    "query": {"type": "string", "description": "Search keywords"}
                },
                "required": ["query"]
            }
        }
    }
]

Think of it as writing API documentation for the AI to read — once it reads this spec, it knows when to call the tool and what parameters to pass.

The actual tool function (the logic that does the real scraping or API calls) is something you implement yourself in Python. The Agent's only job is to decide whether to call it and what parameters to pass.

Step 3: Create a Thread and Submit the Task

The Assistants API uses "Threads" to manage conversation state — think of a Thread as a task container.

from openai import OpenAI
client = OpenAI()

thread = client.beta.threads.create()

client.beta.threads.messages.create(
    thread_id=thread.id,
    role="user",
    content="Please research the latest application trends for AI Agents in 2026, compile a summary, and send it to [email protected]"
)

The advantage of Threads is that they retain the entire conversation history. You can pause mid-task and resume later without losing what the Agent has already done.

Step 4: Start a Run and Handle Tool Call Callbacks

This step is the most critical part of the entire Agent workflow — and the one most commonly missed.

run = client.beta.threads.runs.create(
    thread_id=thread.id,
    assistant_id=assistant.id
)

Once a Run is started, you need to poll its status. When the status changes to requires_action, it means the Agent has decided to call a tool but needs you to execute it and return the result:

while run.status != "completed":
    run = client.beta.threads.runs.retrieve(thread_id=thread.id, run_id=run.id)
    if run.status == "requires_action":
        # Retrieve the tools and parameters the Agent wants to call
        tool_calls = run.required_action.submit_tool_outputs.tool_calls
        outputs = []
        for tc in tool_calls:
            result = execute_tool(tc.function.name, tc.function.arguments)
            outputs.append({"tool_call_id": tc.id, "output": result})
        # Return the results to the Agent
        client.beta.threads.runs.submit_tool_outputs(
            thread_id=thread.id, run_id=run.id, tool_outputs=outputs
        )

This cycle of "you run the tool → return the result → Agent continues reasoning" is the underlying mechanism that enables Agents to complete multi-step tasks.

Step 5: Retrieve the Result and Verify the Output

Once the Run is complete, pull the last assistant message from the Thread's messages:

messages = client.beta.threads.messages.list(thread_id=thread.id)
print(messages.data[0].content[0].text.value)

This is the Agent's final output. You should check:

  • Whether it fully executed all the steps you defined
  • Whether the number of tool calls is reasonable (too many suggests your instructions aren't clear enough)
  • Whether the output format matches your expectations

If the results consistently skip a particular step, the instructions are likely too vague — just go back to Step 1 and refine them.

Common Mistakes and How to Avoid Them

Tools never get called: Usually the description isn't precise enough and the Agent doesn't know when to use them. Be more explicit about when the tool should be used — for example, "use this when querying real-time information from after 2024."

Run stays stuck on in_progress for a long time: This often happens when the Agent is in requires_action state but you haven't submitted tool outputs, causing it to stall while waiting for your response. Make sure your polling logic is complete.

Output format varies every time: Explicitly specify the output format in your instructions, or use Structured Outputs (a feature OpenAI introduced in 2025) to enforce JSON responses.

If you later want the Agent to handle document or voice input, check out How to Use OpenAI Whisper? From Installation to Subtitle Output, All in One Guide to integrate speech-to-text into your workflow as well.

Advanced Usage: Connecting More Tools

Once the basic pipeline is running, this architecture can scale indefinitely:

  • Add Code Interpreter: Let the Agent run Python to analyze data on its own, without you writing any additional logic
  • Multi-Agent collaboration: Use AutoGen or LangGraph to have multiple Agents divide the work — one for research, one for writing the report, one for review
  • Memory layer: Integrate a vector database (Pinecone, Qdrant) so the Agent retains knowledge across Threads

If you want to go further and use AI to write code directly from the terminal, Codex CLI Hands-On Tutorial is a natural next read — the two tools are actually quite complementary in their approach.

What You Have After Running Through This

After completing these five steps, you now have an AI Agent that can: read a task, decide on its own which tools to call, and assemble the results of multiple steps.

Next steps you can take:

  1. Replace web_search with the APIs you actually use in your work (internal databases, CRM, Slack notifications)
  2. Rewrite the Instructions to reflect your real business process
  3. Set up a schedule so the Agent runs automatically every day without manual triggering

The real power of an AI Agent isn't how smart it is — it's how deeply you integrate it. Now you know how.

Share

Related articles