Who this is for: Developers and aspiring developers with no background in AI.
To be honest, I’m not a master of this field, and part of me wondered whether I should be touching a subject with this much technical depth. I’ve tried to keep the explanations as simple as possible. If you have any questions, corrections, or feedback, please feel free to share them.
Introduction
This part provides some useful background for AI agent developers by explaining how AI agents work under the hood.
An AI agent is something that runs on top of an LLM (Large Language Model). So the first thing we need to do is understand LLMs. We don’t need to understand the deep inner workings of an LLM here. Let’s just take a surface-level look at how an LLM operates.
What is an LLM?
An LLM is a model that produces the most likely next token for a given input. It’s called a generative model because it generates text by running this step over and over. We’ll skip the theoretical details here. All you need to know for now is that an LLM repeatedly produces the next token.
OpenAI introduced this architecture under the name GPT, and toward the end of the paper there was a claim, though a tentative one, that the model seemed to reason somewhat like a human (what they called in-context learning). To be honest, I was skeptical myself, and most people didn’t take it seriously either. But in 2022, OpenAI dropped a bombshell by releasing a service called ChatGPT, and it turned out this capability was far more powerful than anyone expected. That is what led to the AI boom we’re living in today.
So what exactly is ChatGPT? It’s the GPT model, tuned so that it can be used in a chat format. Once again, we’ll leave aside the complex theory behind it, such as RLHF. From a developer’s point of view, what matters is this: you give the AI input in the form of chat-style “messages,” and GPT produces its output as chat as well.
As a side note, this is why OpenAI named their API /chat/completions.
So let’s take a closer look at how this “chat,” or “message,” is actually structured.
- System message: Sets the rules and role the LLM should follow throughout the conversation.
- User message: The question or request the user sends to the LLM.
- Assistant message: The response the LLM generates in return.
Expressed in JSON, the input looks like this:
[
{
"role": "system",
"content": [
{ "type": "text", "text": "Answer every user question like a pirate." }
]
},
{
"role": "user",
"content": [
{ "type": "text", "text": "What is photosynthesis?" }
]
}
]
And the AI API returns an output like this:
{
"role": "assistant",
"content": [
{
"type": "text",
"text": "Arr, matey! Photosynthesis be the way plants turn sunlight, water, and air into food, makin' the very oxygen ye breathe!"
}
]
}
(Apologies for the silly example.)
Why is it designed this way? As far as I know, this format was used by early LLMs released on platforms like Hugging Face, and model providers have since adopted similar formats for their APIs. Over time, it has become the de facto standard.
From LLM to AI agents
Once LLMs appeared, many people tried to extend them by connecting them to real-world data, APIs, and other external systems. The standard way to do this is through tools.
The concept of tools emerged shortly after LLMs themselves. The idea is simple: you describe the available tools in the instructions, and the LLM decides to use them when it judges that they are needed.
This is called a tool description. A tool description is how you tell the LLM how to call a tool, and it is delivered by inserting it into the system message.
So if you send the following input to the AI API, it will generate a tool call like the one below.
Input:
{
"messages": [
{
"role": "system",
"content": [
{ "type": "text", "text": "You are a helpful assistant." }
]
},
{
"role": "user",
"content": [
{ "type": "text", "text": "What's the weather like in Seoul right now?" }
]
}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather",
"parameters": {
"type": "object",
"properties": {
"latitude": { "type": "number" },
"longitude": { "type": "number" },
"unit": { "type": "string", "enum": ["celsius", "fahrenheit"] }
},
"required": ["latitude", "longitude"]
}
}
}
]
}
Output:
{
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_01",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"latitude\": 37.57, \"longitude\": 126.98, \"unit\": \"celsius\"}"
}
}
]
}
So how do we handle this output? This is where actual programming begins. For the conversation to continue, you, the developer, have to run the tool yourself and feed the result back in as a tool message.
We can write a simple function to handle get_weather.
def get_weather(latitude: float, longitude: float, unit: str = "celsius") -> dict:
import requests
resp = requests.get(
"https://api.open-meteo.com/v1/forecast",
params={
"latitude": latitude,
"longitude": longitude,
"current": "temperature_2m,weather_code",
"temperature_unit": unit,
},
)
data = resp.json()["current"]
return {
"temperature": data["temperature_2m"],
"weather_code": data["weather_code"],
"unit": unit,
}
Now we call the API one more time to get the assistant’s reply. Keep in mind that the LLM only ever predicts the next token, so the entire conversation keeps getting concatenated. That means we have to send the request with the new messages appended to the existing conversation history.
{
"messages": [
{
"role": "system",
"content": [
{ "type": "text", "text": "You are a helpful assistant." }
]
},
{
"role": "user",
"content": [
{ "type": "text", "text": "What's the weather like in Seoul right now?" }
]
},
{
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_01",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"latitude\": 37.57, \"longitude\": 126.98, \"unit\": \"celsius\"}"
}
}
]
},
{
"role": "tool",
"tool_call_id": "call_01",
"content": [
{ "type": "text", "text": "{\"temperature\": 18.3, \"weather_code\": 3, \"unit\": \"celsius\"}" }
]
}
],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather",
"parameters": {
"type": "object",
"properties": {
"latitude": { "type": "number" },
"longitude": { "type": "number" },
"unit": { "type": "string", "enum": ["celsius", "fahrenheit"] }
},
"required": ["latitude", "longitude"]
}
}
}
]
}
Then you’ll get an output like this:
{
"role": "assistant",
"content": [
{ "type": "text", "text": "It's currently 18°C and cloudy in Seoul." }
]
}
With this, we have given the LLM a new ability: access to real-time weather information.
To summarize, the LLM interacts with a tool through the following steps:
- Tool description: "you pass the schema to the API, and the LLM learns what tools are available"
- Tool calling: the LLM responds with a
tool_callsentry instead of a normal answer - Tool response: you run the tool and send the result back as a
toolmessage, so the LLM can produce its final answer
| Steps | When it happens | Who produces it | Who receives it |
|---|---|---|---|
| Tool description | Before the conversation starts (sent with every request) | Developer | LLM |
| Tool calling | When the LLM decides a tool is needed | LLM | Developer |
| Tool response | After the developer runs the tool | Developer | LLM |
sequenceDiagram
participant U as User
participant D as Your code
participant L as LLM
participant T as Tool (get_weather)
U->>D: "What's the weather like in Seoul?"
D->>L: messages + tools (tool description)
L-->>D: assistant message with tool_calls
D->>T: get_weather(37.57, 126.98, "celsius")
T-->>D: {"temperature": 18.3, "weather_code": 3, "unit": "celsius"}
D->>L: messages + tool message
L-->>D: assistant message "It's 18°C and cloudy in Seoul."
D-->>U: final answer
In conclusion, to implement a tool, you need to implement two things:
- Tool description: the JSON schema that tells the LLM what the tool does and how to call it
- Tool behavior: the actual code that runs when the tool is called
These days, tool frameworks are mature enough that I’d recommend using AI agent development frameworks rather than writing all of this by hand.
With Ailoy, the framework I built, the example above can be written like this:
python
pip install ailoy-py
import asyncio
import json
from ailoy import AgentBuilder, register_tool
# 1. Implement the tool. In practice, call a real weather API here.
def get_weather(city: str) -> dict:
return {"city": city, "weather": "sunny", "temperature_c": 23}
# 2. Register it with its description (JSON Schema). The LLM reads this to decide when to use the tool.
weather_tool = register_tool(
{
"name": "get_weather",
"description": "Get the current weather for a given city",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "The city name, e.g. Seoul"},
},
"required": ["city"],
},
},
get_weather,
)
async def main() -> None:
# 3. Build an agent with the tool attached. (For OpenAI, use "openai/gpt-5.6-luna")
agent = await (
AgentBuilder("anthropic/claude-haiku-4-5")
.instruction("You are a helpful weather assistant.")
.tool(weather_tool)
.build()
)
# 4. Run it. ailoy loops through tool call -> tool execution -> final answer on its own.
async for output in agent.run("What's the weather like in Seoul?"):
message = output["message"]
if message["role"] == "assistant":
for call in message.get("tool_calls", []):
fn = call["function"]
print(f"[tool call] {fn['name']}({json.dumps(fn['arguments'], ensure_ascii=False)})")
for part in message["contents"]:
if part["type"] == "text":
print(part["text"])
elif message["role"] == "tool":
for part in message["contents"]:
if part["type"] == "value":
print(f"[tool result] {json.dumps(part['value'], ensure_ascii=False)}")
asyncio.run(main())
javascript
npm install @brekkylab/ailoy
// weather.mjs
import ailoy from '@brekkylab/ailoy'
const { AgentBuilder, registerTool } = ailoy
// 1. Register the tool's description (JSON Schema) together with its implementation.
// In practice, call a real weather API here.
const weatherTool = registerTool(
{
name: 'get_weather',
description: 'Get the current weather for a given city',
parameters: {
type: 'object',
properties: {
city: { type: 'string', description: 'The city name, e.g. Seoul' },
},
required: ['city'],
},
},
async ({ city }) => ({ city, weather: 'sunny', temperature_c: 23 }),
)
// 2. Build an agent with the tool attached. (For OpenAI, use "openai/gpt-5.6-luna")
const agent = await new AgentBuilder('anthropic/claude-haiku-4-5')
.instruction('You are a helpful weather assistant.')
.tool(weatherTool)
.build()
// 3. Run it. ailoy loops through tool call -> tool execution -> final answer on its own.
try {
for await (const { message } of agent.run("What's the weather like in Seoul?")) {
if (message.role === 'assistant') {
for (const call of message.tool_calls ?? []) {
console.log(`[tool call] ${call.function.name}(${JSON.stringify(call.function.arguments)})`)
}
for (const part of message.contents) {
if (part.type === 'text') console.log(part.text)
}
} else if (message.role === 'tool') {
for (const part of message.contents) {
if (part.type === 'value') console.log(`[tool result] ${JSON.stringify(part.value)}`)
}
}
}
} finally {
await agent.close()
}
The basic idea is simple: the LLM chooses a tool, your code runs it, and the result goes back to the LLM. This can happen several times before the agent has enough information to answer the user. That loop is at the heart of the agents we’ve discussed here.
What developers actually need to do
Fortunately (or unfortunately?), the AI agent field is changing at an incredible pace, and even within a year or two, the way we build agents keeps evolving.
In practice, though, you’ll rarely need to implement all of this yourself. It’s useful to understand what happens under the hood, but frameworks already handle most of the work we’ve covered here.
In the next tutorial, I’ll show you how to build a “modern” AI agent without having to put all these pieces together by hand.




Top comments (0)