Most "AI chat" tutorials stop at single-shot request/response. You send a message, the model replies, done. That's a chatbot — useful, but limited.
What makes something an agent instead is a loop: the model can call a tool, look at what it got back, and decide whether it needs another tool before it's ready to actually answer you.
In this post I'll walk through how to build that loop using Google Gemini's function-calling API, which has a genuinely free tier — no credit card needed to follow along.
The core loop
Here's the shape of it, stripped down:
async function runAgentLoop(messages, tools) {
for (let step = 0; step < MAX_STEPS; step++) {
const result = await model.generateContent({ messages, tools });
if (result.functionCalls.length === 0) {
return result.text; // model gave a final answer, we're done
}
for (const call of result.functionCalls) {
const toolResult = await runTool(call.name, call.args);
messages.push({ role: "function", content: toolResult });
}
// loop continues - model sees the tool results and decides what's next
}
}
That's genuinely the whole idea. Everything else — auth, streaming, persistence — is infrastructure around this loop, not the loop itself.
Why a step limit matters
Without a cap, a confused model could theoretically call tools in a loop forever. MAX_STEPS is a safety valve, not an optimization — it guarantees the request always terminates.
Streaming tool calls in real time
Rather than making the user wait for the whole loop to finish, you can stream each step as it happens using Server-Sent Events — the user sees "calling calculator... got 564... calling get_current_time..." live, instead of a blank spinner.
I recorded a short walkthrough showing this running live against a real Gemini key, if it's easier to see than read:
Defining a tool
Gemini's function-calling schema is close to standard JSON Schema:
{
name: "calculator",
description: "Evaluate a basic arithmetic expression.",
parameters: {
type: "object",
properties: {
expression: { type: "string" }
},
required: ["expression"]
}
}
The model sees this schema and decides on its own when calling this tool would help answer the user's question — you don't have to detect intent yourself.
Wrapping up
I packaged a complete working version of this — including auth, guest mode, and three working example tools — as a starter kit, if you want a working reference rather than building from scratch:
https://omuphade.gumroad.com/l/pfkmad
Happy to answer questions about any of the architecture decisions in the comments.
Top comments (0)