To make an LLM tool loop safe for production in Node.js, you need five things. Cap the number of steps. Give every tool its own timeout. Retry only the failures that are worth retrying, with backoff. Make every tool that changes something idempotent. Log each step as one structured line. Without these, one bad answer from the model can turn into an endless loop, a stuck request or a refund that gets issued twice.
This post walks through each guard in turn, with short snippets you can copy into an existing agent. The examples don't depend on any provider. callModel stands in for whatever SDK you use.
The naive loop and why it breaks
Most tool loops start out like this:
while (true) {
const reply = await callModel(messages);
if (!reply.toolCalls.length) return reply.text;
for (const call of reply.toolCalls) {
const result = await tools[call.name](call.args);
messages.push({ role: 'tool', id: call.id, content: JSON.stringify(result) });
}
}
It works in a demo. In production it breaks in predictable ways:
- The model calls the same search tool over and over because the results never quite answer the question.
- A slow upstream API hangs, so the user's request hangs with it.
- A tool throws, the error escapes the loop and the whole run dies with nothing useful in the logs.
- A network blip triggers a retry, and the "create invoice" tool runs twice.
Each guard below fixes one of these.
1. Cap the steps
A step cap is the cheapest safety feature you can add. Count model turns, not tool calls, and stop cleanly when you reach the limit.
const MAX_STEPS = 8;
async function runAgent(messages, ctx) {
for (let step = 1; step <= MAX_STEPS; step++) {
const reply = await callModel(messages, { signal: AbortSignal.timeout(30000) });
if (!reply.toolCalls.length) return { status: 'done', text: reply.text };
for (const call of reply.toolCalls) {
const result = await runTool(call, { ...ctx, step });
messages.push({ role: 'tool', id: call.id, content: JSON.stringify(result) });
}
}
return { status: 'step_limit', text: 'I could not finish this in the allowed number of steps.' };
}
Here's how to pick the number. Look at the longest legitimate task the agent handles, count its steps and add two or three. If most runs finish in three steps, a cap of 8 catches the runaway loops without cutting off the real work.
Two more limits are worth adding next to the step cap:
-
A wall-clock deadline for the whole run. A run can stay under the step cap and still take far too long. Check
Date.now() - ctx.startedAtat the top of each step. -
Repeat detection. If the model calls the same tool with the same arguments three times, stop and return what you have. Hash
call.name + JSON.stringify(call.args)and count the hashes.
When a run hits a limit, return a status rather than throwing. The caller can decide whether to show a fallback message, hand off to a human or queue the task to try later.
2. Give each tool its own timeout
Tools don't all behave the same. A lookup in your own database should finish in well under a second. A call to a third-party reporting API might reasonably take ten. One global timeout will be too tight for some tools and too loose for others, so put the limit on each tool:
const tools = {
search_orders: { timeoutMs: 3000, retries: 2, handler: searchOrders },
get_shipping_status: { timeoutMs: 8000, retries: 2, handler: getShippingStatus },
create_refund: { timeoutMs: 10000, retries: 1, handler: createRefund },
};
In Node.js 18 and later, AbortSignal.timeout(ms) gives you a signal that aborts itself. Pass it all the way down:
async function getShippingStatus(args, { signal }) {
const res = await fetch('https://carrier.example.com/track/' + args.trackingId, { signal });
if (!res.ok) throw Object.assign(new Error('carrier error'), { status: res.status });
return res.json();
}
The signal only works if the handler actually uses it. Wrapping a call in Promise.race with a timer makes your loop stop waiting, but the request keeps running in the background and holds onto sockets. Pass the signal to fetch, to your database driver and to anything else that accepts one.
The model call needs a timeout too. Slow model responses are one of the most common reasons an agent endpoint hangs.
3. Retry with backoff, but only what is retryable
Retries help with temporary failures and do harm everywhere else. A 400 error will fail again no matter how many times you send it. A 429 or 503 often succeeds a moment later.
const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
function isRetryable(err) {
return err.name === 'TimeoutError' || err.status === 429 || err.status >= 500;
}
async function withRetry(fn, { retries, baseMs = 300 }) {
for (let attempt = 0; ; attempt++) {
try {
return await fn(attempt);
} catch (err) {
if (attempt >= retries || !isRetryable(err)) throw err;
const delay = baseMs * 2 ** attempt + Math.random() * baseMs;
await sleep(delay);
}
}
}
There are three details to get right:
- Jitter. The random part spreads retries out, so a hundred stuck runs don't all hit the recovering service in the same millisecond.
-
A fresh timeout per attempt. Create the
AbortSignalinsidefn, not outside it. A signal that has already fired will make every retry fail immediately. - Keep the retry layers separate. Your SDK might already retry model calls. If you add your own retries on top, three retries become nine without anyone noticing. Know which layer owns retries.
4. Make tools that change something idempotent
This is the guard teams most often skip, and skipping it is the expensive mistake. Suppose the refund API times out after it has already processed the refund. The retry processes it again. Or the model calls the same tool twice in one turn because it lost track.
The fix is an idempotency key that stays the same across retries of the same intended action:
async function createRefund(args, { signal, key }) {
const existing = await db.refunds.findOne({ idempotencyKey: key });
if (existing) return existing;
return payments.refund({ orderId: args.orderId, amount: args.amount, idempotencyKey: key }, { signal });
}
Build the key from the run ID and the tool call ID, for example ctx.runId + ':' + call.id. Every retry of that call reuses the key. A new run gets a new one. If the downstream API supports idempotency keys natively (many payment providers do), pass the key through so the protection holds even if your own database write fails.
The rule is simple: any tool that writes, sends, charges or deletes either takes an idempotency key or doesn't get retried.
5. Return errors to the model, log every step
Here is the wrapper that ties it all together. Errors come back to the model as data, so it can try a different approach, and each attempt produces exactly one log line.
async function runTool(call, ctx) {
const tool = tools[call.name];
if (!tool) return { error: 'unknown_tool', name: call.name };
const key = ctx.runId + ':' + call.id;
const started = Date.now();
try {
const result = await withRetry(
(attempt) => {
logger.info({ runId: ctx.runId, step: ctx.step, tool: call.name, attempt }, 'tool_start');
return tool.handler(call.args, { signal: AbortSignal.timeout(tool.timeoutMs), key });
},
{ retries: tool.retries }
);
logger.info({ runId: ctx.runId, step: ctx.step, tool: call.name, ms: Date.now() - started, ok: true }, 'tool_end');
return result;
} catch (err) {
logger.warn({ runId: ctx.runId, step: ctx.step, tool: call.name, ms: Date.now() - started, ok: false, error: err.name, status: err.status }, 'tool_end');
return { error: err.name, message: 'The tool failed. Try another approach or tell the user.' };
}
}
With a structured logger like pino, these lines answer the questions you will be asked in production. Why did this run take 40 seconds? Which tool is timing out? How often do runs hit the step limit?
Some things to keep out of the logs: raw user messages, full tool results and anything that looks like a credential. Log argument hashes or a few safe fields instead.
A pre-launch checklist
Before an agent goes live, check each of these:
- [ ] A step cap and a wall-clock deadline for the whole run
- [ ] Repeat detection for identical tool calls
- [ ] A per-tool timeout, with the signal passed down to the actual I/O
- [ ] A timeout on the model call itself
- [ ] Retries only for timeouts, 429 and 5xx errors, with jittered backoff
- [ ] One clear owner for retries (SDK or your code, not both)
- [ ] An idempotency key on every tool that writes, sends, charges or deletes
- [ ] Tool errors returned to the model as data, not thrown
- [ ] One structured log line per step, with run ID, step, tool, attempt, duration and outcome
- [ ] A clear final status: done, step_limit, deadline or error
None of this is clever code. On agent projects at Geminate Solutions, these guards go in before any prompt tuning, because they are what keep a bad model turn from turning into a bad day. Once they are in place you can experiment with prompts and tools much more freely.
For the wider picture of structuring agents in Node.js, including tool design and where these guards fit, see the full guide: Building AI agents in Node.js.
Top comments (0)