DEV Community

Dilip V P
Dilip V P

Posted on

AI Tool Calling: The Model Never Runs Your Code

A customer types one sentence into a food app's support chat:

Where is order 4472? Cancel it if no rider is assigned yet.

The app checks the order, cancels it, and replies. Here is the strange part. No code was ever written for that sentence.

An AI model read it, looked at the functions the app offered it, and decided which ones to call. That is called tool calling. Some vendors call it function calling. Either way, it is how a model goes from writing text to doing real work.

This article builds it from zero, starting with the first JSON. Then it covers the loop that nobody tells you how to stop, and the day the model calls five tools and only three answer.

Why code alone fails here

Apps handle fixed requests with fixed code. You tap Place Order, and the same steps run every time. The steps are known, so the code is known.

A support chat is different. Every customer types a different sentence. One types "cancel my order". The next types "do not cancel my order, I changed my mind". A keyword search for "cancel" gets the second sentence exactly wrong, and you cannot write an if for every sentence a human can type.

A language model can read any sentence, but a model alone only produces text. It cannot see order 4472, because the order lives in your database, not inside the model. Ask it where your order is, and you get a polite guess.

So we split the job. Your code does the doing. The model does the deciding.

What a tool is

A tool is a function in your own code that you describe to the model. The description has three parts: a name, an explanation of what the function does, and the inputs it takes. Here is a simplified version:

{
  "name": "get_order",
  "description": "Fetch one order by its ID. Use it before any change to an order.",
  "parameters": { "order_id": "number" }
}
Enter fullscreen mode Exit fullscreen mode

The real APIs describe the inputs with JSON Schema. In OpenAI's Responses API, the schema goes in parameters. In Anthropic's Messages API, it goes in input_schema. The idea is the same in both: a name, a description the model reads to decide, and a schema for the inputs.

The description matters more than it looks. The model chooses a tool by reading it, so "Use it before any change to an order" is an instruction, not a comment.

The request, and what comes back

With every customer sentence, your program sends the model the list of tools it may ask for. In our app, that is get_order, cancel_order and refund_payment.

What comes back is not a sentence. It is structured data, and it is called a tool call:

{ "tool": "get_order", "arguments": { "order_id": 4472 } }
Enter fullscreen mode Exit fullscreen mode

That shape is simplified too. In Anthropic's API, the model returns a tool_use block with an id, a name and an input object. In OpenAI's Responses API, it returns a function_call item with a call_id, a name and an arguments string that holds JSON.

Notice the number. The model pulled 4472 out of the customer's sentence on its own. Nobody wrote a parser for that.

Your program runs it

The model cannot run anything, so your program takes it from there, in whatever language your backend is written in. Ours is Java. Check that the name is on the list, then call the real method:

ToolResult run(ToolCall call) throws TimeoutException {
    long id = call.longArg("order_id");
    return switch (call.name()) {
        case "get_order"    -> ToolResult.ok(call, orders.find(id).toJson());
        case "cancel_order" -> ToolResult.ok(call, orders.cancel(id).toJson());
        default             -> ToolResult.error(call, "Unknown tool: " + call.name());
    };
}
Enter fullscreen mode Exit fullscreen mode

The default branch is your security boundary. A tool call is only text in a fixed shape, and your program decides whether anything happens.

The database answers: order 4472, status preparing, no rider yet. That result goes back to the model with one standing question: what next? There is no rider, and the customer said to cancel in exactly that case. So the reply is one more tool call, cancel_order with order ID 4472.

Inside cancel_order is the code from the message queues episode: one conditional update in the database, and one message on a queue for the restaurant.

UPDATE orders SET status = 'CANCELLED'
WHERE id = 4472 AND rider IS NULL;
Enter fullscreen mode Exit fullscreen mode

Then the model finally replies in plain text: "Done. Your order is cancelled."

The model never ran anything. It asked. Your program did it.

The agent loop, and how it ends

Step back, because this shape has a name. The model reasons. Your program acts. The result comes back as an observation. Then it repeats. Reason, act, observe, repeat. That cycle is the agent loop, and a program wrapped around a model is an agent.

The loop ends when the model replies with plain text instead of a tool call. The text is the signal that it is done. Nothing in the loop itself stops it.

That should make you slightly uncomfortable. Sometimes a model checks the order, checks it again to be sure, cancels, checks that the cancel worked, and checks once more. Every step looks reasonable, and together they wander. Every turn of the loop is one more model request, which costs money and seconds.

So your program adds a fence, a step limit. Count the tool calls, and at a limit, say ten, stop the loop and return a safe reply:

static final int STEP_LIMIT = 10;   // an example; pick yours from the longest honest path

String answer(Conversation chat) {
    int toolCalls = 0;
    while (true) {
        ModelReply reply = model.send(chat);   // the sentence, the tool list, every result so far
        if (reply.toolCalls().isEmpty()) {
            return reply.text();               // plain text: the model is done
        }
        for (ToolCall call : reply.toolCalls()) {
            if (++toolCalls > STEP_LIMIT) {
                return "I could not finish this. A support person will follow up.";
            }
            chat.add(runSafely(call));         // every call gets a result, even a failure
        }
    }
}
Enter fullscreen mode Exit fullscreen mode

That is the pattern under everything that follows. The model decides. Your program sets the boundaries.

Five tools called, three answered

Now a harder customer types: "Cancel order 4472, refund me, and text me when it is done."

Across the loop's turns, the model asks for five tools: get the order, cancel it, start the refund, look up the phone number, and send the SMS.

Remember what a tool is underneath. It is your code calling a real service, such as the database, the payment service or the SMS provider. Real services fail. Today the phone lookup returns an error, and the SMS provider times out. Five tools were called, and three answered.

Your program is holding two dead calls, and it has three honest choices.

Choice one is to retry. A timeout is usually temporary, so retry with backoff: wait 1 second, then 2, then 4. But first, ask one question. Is this tool a read or a write?

get_order is a read, so you can retry it freely, because it changes nothing. refund_payment is a write. If the refund happened and only the reply got lost, a retry pays the customer twice.

Idempotent means doing it twice has the same effect as doing it once. So the refund call carries an idempotency key. Your program creates the key once per refund attempt and saves it with the order before the first call. Every retry sends the same key, and the payment service refuses the duplicate.

Anything that may be retried must be idempotent. And an agent retries.

Choice two is a fallback, which is a second, simpler way to get a result. The SMS tool is down, so queue the message and send it later.

Choice three is to tell the model the truth. Send back a result that says the SMS failed. A failure is one more observation, and the model can reason with it. It replies: "Order cancelled. Refund done. Your text may arrive late."

Both APIs have a place for that truth. In Anthropic's API, the tool_result block has an is_error field, and the docs ask for error messages that say what went wrong. In OpenAI's API, you send a function_call_output whose output describes the failure. In our Java, every exception becomes a result:

ToolResult runSafely(ToolCall call) {
    try {
        return run(call);
    } catch (TimeoutException e) {
        return ToolResult.error(call, call.name() + " timed out. It may or may not have run.");
    } catch (RuntimeException e) {
        return ToolResult.error(call, call.name() + " failed: " + e.getMessage());
    }
}
Enter fullscreen mode Exit fullscreen mode

The timeout message is careful on purpose. After a timeout, your program does not know whether a write happened, so it does not pretend to.

The forbidden move: silence

There is one thing you must never do. Do not send nothing back.

If a call gets no result, the model does what text machines do with a gap: it fills it. It will tell the customer the text was sent, when no text was sent. Given a gap, a model does not show you an error. It shows you confidence.

So every tool call ends in a result, even when the result is "this failed".

RAG is one more tool

Now a different question arrives: "My order came 40 minutes late. Do I get a refund?"

That answer is not in the database. It is in the refund policy, a page of text the model has never seen. So you add one more tool, search_help_pages. Your program searches your own documents by meaning and returns the most relevant passages, and the model answers from them.

That pattern is called retrieval augmented generation, or RAG. From inside the agent, it is just one more tool. It is a read, with a timeout, like the rest. (The refund policy in this example is invented.)

When not to use an agent

Now the opposite case. When you tap Place Order, the steps never change, so no model should choose them. Known steps are a program. A program takes the same path every time, it is testable, it runs in milliseconds, and it costs almost nothing per step.

An agent is not an upgrade to a program. It is a price you pay for one ability: handling requests whose steps cannot be known in advance. So here is the rule:

If you can draw the flowchart before the request arrives, write the program. If the next step depends on what the last step returned, that is agent territory.

Three lines worth keeping

  1. The model never runs anything. It only asks.
  2. Anything that may be retried must be idempotent, and an agent retries.
  3. If you can draw the flowchart before the request arrives, write the program.

One problem from this article is still open. cancel_order wrote to the database, and it put a message on a queue. That is two writes in two systems. One day a service crashes exactly between them. The order is cancelled, and the restaurant never finds out. That crash is the next episode.


The video builds all of this on screen, from the first JSON to the five tools and the forbidden silence: https://youtu.be/3h0EmNKJSEA

The app, the customer, order 4472 and the refund policy are an invented example. The JSON shapes were checked against the current OpenAI and Anthropic documentation on 4 Oct 2026.

Top comments (0)