DEV Community

Jude Lee
Jude Lee

Posted on AI-assisted

A Survival Guide to Building AI Agents: (1) How AI agents Work

Who this is for: Developers and aspiring developers with no background in AI.

To be honest, I’m not a master of this field, and part of me wondered whether I should be touching a subject with this much technical depth. I’ve tried to keep the explanations as simple as possible. If you have any questions, corrections, or feedback, please feel free to share them.


Introduction

This part provides some useful background for AI agent developers by explaining how AI agents work under the hood.

An AI agent is something that runs on top of an LLM (Large Language Model). So the first thing we need to do is understand LLMs. We don’t need to understand the deep inner workings of an LLM here. Let’s just take a surface-level look at how an LLM operates.

What is an LLM?

An LLM is a model that produces the most likely next token for a given input. It’s called a generative model because it generates text by running this step over and over. We’ll skip the theoretical details here. All you need to know for now is that an LLM repeatedly produces the next token.

How LLM works

OpenAI introduced this architecture under the name GPT, and toward the end of the paper there was a claim, though a tentative one, that the model seemed to reason somewhat like a human (what they called in-context learning). To be honest, I was skeptical myself, and most people didn’t take it seriously either. But in 2022, OpenAI dropped a bombshell by releasing a service called ChatGPT, and it turned out this capability was far more powerful than anyone expected. That is what led to the AI boom we’re living in today.

So what exactly is ChatGPT? It’s the GPT model, tuned so that it can be used in a chat format. Once again, we’ll leave aside the complex theory behind it, such as RLHF. From a developer’s point of view, what matters is this: you give the AI input in the form of chat-style “messages,” and GPT produces its output as chat as well.

As a side note, this is why OpenAI named their API /chat/completions.

So let’s take a closer look at how this “chat,” or “message,” is actually structured.

Agent1

  • System message: Sets the rules and role the LLM should follow throughout the conversation.
  • User message: The question or request the user sends to the LLM.
  • Assistant message: The response the LLM generates in return.

Expressed in JSON, the input looks like this:

[
  {
    "role": "system",
    "content": [
      { "type": "text", "text": "Answer every user question like a pirate." }
    ]
  },
  {
    "role": "user",
    "content": [
      { "type": "text", "text": "What is photosynthesis?" }
    ]
  }
]
Enter fullscreen mode Exit fullscreen mode

And the AI API returns an output like this:

{
  "role": "assistant",
  "content": [
    {
      "type": "text",
      "text": "Arr, matey! Photosynthesis be the way plants turn sunlight, water, and air into food, makin' the very oxygen ye breathe!"
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

(Apologies for the silly example.)

Why is it designed this way? As far as I know, this format was used by early LLMs released on platforms like Hugging Face, and model providers have since adopted similar formats for their APIs. Over time, it has become the de facto standard.

From LLM to AI agents

Once LLMs appeared, many people tried to extend them by connecting them to real-world data, APIs, and other external systems. The standard way to do this is through tools.

The concept of tools emerged shortly after LLMs themselves. The idea is simple: you describe the available tools in the instructions, and the LLM decides to use them when it judges that they are needed.

Agent2

This is called a tool description. A tool description is how you tell the LLM how to call a tool, and it is delivered by inserting it into the system message.

So if you send the following input to the AI API, it will generate a tool call like the one below.

Input:

{
  "messages": [
    {
      "role": "system",
      "content": [
        { "type": "text", "text": "You are a helpful assistant." }
      ]
    },
    {
      "role": "user",
      "content": [
        { "type": "text", "text": "What's the weather like in Seoul right now?" }
      ]
    }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get current weather",
        "parameters": {
          "type": "object",
          "properties": {
            "latitude": { "type": "number" },
            "longitude": { "type": "number" },
            "unit": { "type": "string", "enum": ["celsius", "fahrenheit"] }
          },
          "required": ["latitude", "longitude"]
        }
      }
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

Output:

{
  "role": "assistant",
  "content": null,
  "tool_calls": [
    {
      "id": "call_01",
      "type": "function",
      "function": {
        "name": "get_weather",
        "arguments": "{\"latitude\": 37.57, \"longitude\": 126.98, \"unit\": \"celsius\"}"
      }
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

So how do we handle this output? This is where actual programming begins. For the conversation to continue, you, the developer, have to run the tool yourself and feed the result back in as a tool message.

Agent3

We can write a simple function to handle get_weather.

def get_weather(latitude: float, longitude: float, unit: str = "celsius") -> dict:
    import requests

    resp = requests.get(
        "https://api.open-meteo.com/v1/forecast",
        params={
            "latitude": latitude,
            "longitude": longitude,
            "current": "temperature_2m,weather_code",
            "temperature_unit": unit,
        },
    )
    data = resp.json()["current"]
    return {
        "temperature": data["temperature_2m"],
        "weather_code": data["weather_code"],
        "unit": unit,
    }
Enter fullscreen mode Exit fullscreen mode

Now we call the API one more time to get the assistant’s reply. Keep in mind that the LLM only ever predicts the next token, so the entire conversation keeps getting concatenated. That means we have to send the request with the new messages appended to the existing conversation history.

{
  "messages": [
    {
      "role": "system",
      "content": [
        { "type": "text", "text": "You are a helpful assistant." }
      ]
    },
    {
      "role": "user",
      "content": [
        { "type": "text", "text": "What's the weather like in Seoul right now?" }
      ]
    },
    {
      "role": "assistant",
      "content": null,
      "tool_calls": [
        {
          "id": "call_01",
          "type": "function",
          "function": {
            "name": "get_weather",
            "arguments": "{\"latitude\": 37.57, \"longitude\": 126.98, \"unit\": \"celsius\"}"
          }
        }
      ]
    },
    {
      "role": "tool",
      "tool_call_id": "call_01",
      "content": [
        { "type": "text", "text": "{\"temperature\": 18.3, \"weather_code\": 3, \"unit\": \"celsius\"}" }
      ]
    }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get current weather",
        "parameters": {
          "type": "object",
          "properties": {
            "latitude": { "type": "number" },
            "longitude": { "type": "number" },
            "unit": { "type": "string", "enum": ["celsius", "fahrenheit"] }
          },
          "required": ["latitude", "longitude"]
        }
      }
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

Then you’ll get an output like this:

{
  "role": "assistant",
  "content": [
    { "type": "text", "text": "It's currently 18°C and cloudy in Seoul." }
  ]
}
Enter fullscreen mode Exit fullscreen mode

With this, we have given the LLM a new ability: access to real-time weather information.

To summarize, the LLM interacts with a tool through the following steps:

  • Tool description: "you pass the schema to the API, and the LLM learns what tools are available"
  • Tool calling: the LLM responds with a tool_calls entry instead of a normal answer
  • Tool response: you run the tool and send the result back as a tool message, so the LLM can produce its final answer
Steps When it happens Who produces it Who receives it
Tool description Before the conversation starts (sent with every request) Developer LLM
Tool calling When the LLM decides a tool is needed LLM Developer
Tool response After the developer runs the tool Developer LLM
sequenceDiagram
    participant U as User
    participant D as Your code
    participant L as LLM
    participant T as Tool (get_weather)

    U->>D: "What's the weather like in Seoul?"
    D->>L: messages + tools (tool description)
    L-->>D: assistant message with tool_calls
    D->>T: get_weather(37.57, 126.98, "celsius")
    T-->>D: {"temperature": 18.3, "weather_code": 3, "unit": "celsius"}
    D->>L: messages + tool message
    L-->>D: assistant message "It's 18°C and cloudy in Seoul."
    D-->>U: final answer

In conclusion, to implement a tool, you need to implement two things:

  • Tool description: the JSON schema that tells the LLM what the tool does and how to call it
  • Tool behavior: the actual code that runs when the tool is called

These days, tool frameworks are mature enough that I’d recommend using AI agent development frameworks rather than writing all of this by hand.

With Ailoy, the framework I built, the example above can be written like this:

python
pip install ailoy-py 
Enter fullscreen mode Exit fullscreen mode
import asyncio
import json

from ailoy import AgentBuilder, register_tool


# 1. Implement the tool. In practice, call a real weather API here.
def get_weather(city: str) -> dict:
    return {"city": city, "weather": "sunny", "temperature_c": 23}


# 2. Register it with its description (JSON Schema). The LLM reads this to decide when to use the tool.
weather_tool = register_tool(
    {
        "name": "get_weather",
        "description": "Get the current weather for a given city",
        "parameters": {
            "type": "object",
            "properties": {
                "city": {"type": "string", "description": "The city name, e.g. Seoul"},
            },
            "required": ["city"],
        },
    },
    get_weather,
)


async def main() -> None:
    # 3. Build an agent with the tool attached. (For OpenAI, use "openai/gpt-5.6-luna")
    agent = await (
        AgentBuilder("anthropic/claude-haiku-4-5")
        .instruction("You are a helpful weather assistant.")
        .tool(weather_tool)
        .build()
    )

    # 4. Run it. ailoy loops through tool call -> tool execution -> final answer on its own.
    async for output in agent.run("What's the weather like in Seoul?"):
        message = output["message"]
        if message["role"] == "assistant":
            for call in message.get("tool_calls", []):
                fn = call["function"]
                print(f"[tool call] {fn['name']}({json.dumps(fn['arguments'], ensure_ascii=False)})")
            for part in message["contents"]:
                if part["type"] == "text":
                    print(part["text"])
        elif message["role"] == "tool":
            for part in message["contents"]:
                if part["type"] == "value":
                    print(f"[tool result] {json.dumps(part['value'], ensure_ascii=False)}")


asyncio.run(main())
Enter fullscreen mode Exit fullscreen mode

javascript
npm install @brekkylab/ailoy
Enter fullscreen mode Exit fullscreen mode
// weather.mjs
import ailoy from '@brekkylab/ailoy'

const { AgentBuilder, registerTool } = ailoy

// 1. Register the tool's description (JSON Schema) together with its implementation.
//    In practice, call a real weather API here.
const weatherTool = registerTool(
  {
    name: 'get_weather',
    description: 'Get the current weather for a given city',
    parameters: {
      type: 'object',
      properties: {
        city: { type: 'string', description: 'The city name, e.g. Seoul' },
      },
      required: ['city'],
    },
  },
  async ({ city }) => ({ city, weather: 'sunny', temperature_c: 23 }),
)

// 2. Build an agent with the tool attached. (For OpenAI, use "openai/gpt-5.6-luna")
const agent = await new AgentBuilder('anthropic/claude-haiku-4-5')
  .instruction('You are a helpful weather assistant.')
  .tool(weatherTool)
  .build()

// 3. Run it. ailoy loops through tool call -> tool execution -> final answer on its own.
try {
  for await (const { message } of agent.run("What's the weather like in Seoul?")) {
    if (message.role === 'assistant') {
      for (const call of message.tool_calls ?? []) {
        console.log(`[tool call] ${call.function.name}(${JSON.stringify(call.function.arguments)})`)
      }
      for (const part of message.contents) {
        if (part.type === 'text') console.log(part.text)
      }
    } else if (message.role === 'tool') {
      for (const part of message.contents) {
        if (part.type === 'value') console.log(`[tool result] ${JSON.stringify(part.value)}`)
      }
    }
  }
} finally {
  await agent.close()
}
Enter fullscreen mode Exit fullscreen mode

The basic idea is simple: the LLM chooses a tool, your code runs it, and the result goes back to the LLM. This can happen several times before the agent has enough information to answer the user. That loop is at the heart of the agents we’ve discussed here.

What developers actually need to do

Fortunately (or unfortunately?), the AI agent field is changing at an incredible pace, and even within a year or two, the way we build agents keeps evolving.

In practice, though, you’ll rarely need to implement all of this yourself. It’s useful to understand what happens under the hood, but frameworks already handle most of the work we’ve covered here.

In the next tutorial, I’ll show you how to build a “modern” AI agent without having to put all these pieces together by hand.

Top comments (0)