Debugging ZORGAX: How We Traced an AI Agent Tool-Calling Loop from Open WebUI to Ollama
An evidence-first debugging story, with a reproducible test and an open invitation to contributors.
Building an AI agent is not just about making a language model call tools. It is also about ensuring the model knows when to stop.
While developing ZORGAX, a local DevOps AI agent within the MyZubster ecosystem, I encountered a tool-calling problem: when asked which automations were active, ZORGAX repeatedly invoked list_automations even after the tool returned a successful empty result.
The expected behavior was simple:
- Call
list_automations. - Receive
{"automations":[],"total":0}. - Tell the user that no automations are active.
- Stop.
Instead, the Open WebUI integration sometimes repeated the tool call until interrupted.
We investigated the problem using a custom ZORGAX model derived from Qwen2.5 3B, Ollama, Open WebUI, Docker, and a temporary HTTP diagnostic proxy.
Here is what we discovered.
1. The first clue: direct Ollama worked
Our first experiment bypassed Open WebUI.
We sent a synthetic conversation to Ollama's /api/chat endpoint, including an assistant tool call and a matching tool result.
The result was exactly what we wanted:
Content: Non ci sono automazioni attive al momento.
Tool calls: []
Done: True
ZORGAX understood the empty result and produced a final answer without calling the tool again.
This was useful, but it did not prove Open WebUI was at fault. The direct test used a simplified system prompt and tool schema. We needed to compare the actual integration path.
2. Following the tool result through Open WebUI
A commenter on our earlier Coder Legion article suggested a more precise investigation: capture the request immediately after the first list_automations result and compare it with a working Ollama control.
That changed the direction of our debugging.
We inspected the installed Open WebUI source and identified its tool-call continuation flow.
The middleware transforms tool execution output back into conversation messages, including the assistant's tool_calls and a matching role: tool message.
Rather than relying on source inspection alone, we extracted and exercised the actual convert_output_to_messages() function using a harmless fixture.
The test produced:
[
{
"role": "assistant",
"content": "",
"tool_calls": [
{
"id": "zorgax-test-001",
"type": "function",
"function": {
"name": "list_automations",
"arguments": "{}"
}
}
]
},
{
"role": "tool",
"tool_call_id": "zorgax-test-001",
"content": "{\"automations\":[],\"total\":0}"
}
]
The conversion test passed.
That ruled out a basic failure of this function for our isolated empty-result case. It did not establish that every subsequent network request would preserve the same structure.
3. Adding a safety limit
We also found that the investigated Open WebUI build had a default limit of 256 tool-call iterations.
For debugging, that was far more than we needed.
In our test setup, we configured:
environment:
CHAT_RESPONSE_MAX_TOOL_CALL_ITERATIONS: "3"
After applying the setting, an isolated Open WebUI experiment reproduced the repeated calls and stopped with:
Tool-call iteration limit reached (3)
This was a safety improvement, not a root-cause fix. A single iteration can contain multiple tool calls, so the limit should not be interpreted as an exact cap on individual invocations.
4. Building a disposable diagnostic environment
We created a separate Open WebUI instance rather than modifying the main MyZubster interface.
We then added a harmless custom tool:
"""
title: ZORGAX Automation Test
description: Safe automation-listing test fixture.
version: 1.0.0
"""
class Tools:
def list_automations(self) -> dict:
"""Return a synthetic list of active automations."""
return {
"automations": [],
"total": 0
}
This function never reads real automations, touches a database, or performs an administrative action.
That isolation matters. Debugging an agent with operational tools should not require putting production infrastructure at risk.
5. Capturing the actual Ollama requests
Next, we introduced a temporary HTTP proxy between the diagnostic Open WebUI instance and Ollama.
The proxy forwarded requests while recording only structural metadata: model name, message roles, tool-call IDs, tool names, and whether the empty-result fixture was present.
We deliberately avoided storing complete conversation prompts and authentication data.
The first interesting observation was that one captured request exposed approximately 35 built-in tools to the model.
That was substantially different from the successful direct Ollama test, which exposed only list_automations.
However, this alone did not prove the number of tools caused the loop.
We needed a narrower experiment.
6. The successful end-to-end experiment
We configured ZORGAX Diagnostic to use a single custom tool and disabled the unrelated built-in capabilities.
Then we asked:
Controlla quali automazioni sono attive usando
list_automations. Se non ce ne sono, dimmelo.
ZORGAX invoked the simulated tool and replied that no automations were active.
More importantly, the proxy captured the continuation:
Request 7
Model: zorgax:latest
Tools: [list_automations]
Messages:
user
Request 8
Model: zorgax:latest
Tools: [list_automations]
Messages:
user
assistant -> list_automations
tool -> empty result present
The assistant's call identifier and the tool result were preserved, and the empty result reached the next request destined for Ollama.
For this controlled configuration, the full interaction worked.
That is our most important verified finding: Open WebUI can carry a successful empty tool result through to Ollama and obtain a final response from ZORGAX.
7. What we have not proven
It would be tempting to conclude that exposing too many tools caused the original loop.
But that would be premature.
The successful and unsuccessful experiments did not hold every variable constant. Tool availability was different, and other factors — including prompt instructions, function schemas, model parameters, and sampling behavior — might also affect the result.
We therefore have a strong debugging lead, not a definitive explanation.
Our next objective is a controlled A/B comparison using identical model versions, prompts, tool schemas, and generation options while changing one variable at a time.
We also want regression tests that distinguish three important cases:
- A tool successfully returns an empty result.
- A tool fails with a transport or execution error.
- A model generates invalid tool arguments.
An empty result is valid information. An error is not the same thing.
8. Opening the investigation to the community
We have now published the investigation in the public MyZubster GitHub repository.
Technical reproduction guide:
https://github.com/MyZubster-Ecosystem/myzubster/blob/main/docs/zorgax/TOOL-CALLING-INVESTIGATION-2026-10-09.md
Contributor investigation issue:
https://github.com/MyZubster-Ecosystem/myzubster/issues/1577
Contribution guidelines:
https://github.com/MyZubster-Ecosystem/myzubster/blob/main/CONTRIBUTING.md
Contributors can reproduce the problem on their own machines, improve the fixtures, compare sanitized HTTP payloads, add regression tests, or propose changes to observability and error handling.
No one needs access to the MyZubster VPS or any production credentials.
We welcome small, focused contributions backed by clear evidence.
9. The bigger lesson
A tool-calling loop is not always something you can fix by adding another instruction such as “do not repeat the call.”
The more reliable approach is to inspect the complete sequence:
User request
↓
Assistant tool call
↓
Tool execution
↓
Tool result
↓
Follow-up request to the model
↓
Final response or another tool call
Each boundary needs to preserve the right context.
This investigation taught us to separate three questions:
Did the tool execute?
Did its result reach the next model request?
Did the model make the right decision after receiving it?
Only by testing those questions independently can we narrow down the cause of an agent loop without blaming the wrong component.
What's next for ZORGAX?
ZORGAX is intended to become a dependable local DevOps assistant for monitoring, diagnostics, and carefully authorized infrastructure operations.
The successful single-tool experiment is an important milestone, but the original loop investigation remains open.
We are continuing that work in public so others can reproduce the evidence, challenge our assumptions, and help build a more reliable agent.
If you work with Ollama, Open WebUI, Qwen, or local AI agents, you're welcome to join the investigation.
Build in public. Test with evidence. Keep humans in control.
Top comments (0)