Function calling (a.k.a. tool use) is what turns an LLM from a text generator into something that can do things: look up an order, create a ticket, charge a card. It's also where most production LLM features quietly fall apart.
The demo works. Then real traffic arrives and you meet:
- The model returns arguments that don't match your schema (
"quantity": "two") - The model calls a tool that doesn't exist
- The provider returns
429or529during a traffic spike - A request hangs for 90 seconds and ties up your workers
- The model gets stuck in a loop calling the same tool forever
- A retry causes the same refund to be issued twice
In this post we'll build a layered resilience strategy in Node.js and TypeScript, where each layer handles a different failure class:
- Schema validation for what the model sends you
- Timeouts and retries for transient failures
- Circuit breakers for sustained provider failures
- Fallbacks so users get something useful when all else fails
- Loop guards and idempotency for the agent loop itself
The mental model: two kinds of failure
Before writing code, separate failures into two buckets, because they need opposite treatment.
| Model-side failures | Infrastructure failures | |
|---|---|---|
| Examples | Bad JSON, wrong types, unknown tool, missing required field | 429, 5xx, timeouts, network resets |
| Cause | The model made a mistake | The provider or network is unhealthy |
| Right response | Tell the model what went wrong and let it retry | Back off, retry, or fail over |
| Counts toward circuit breaker? | No | Yes |
The most common mistake is treating them the same: retrying bad arguments with the identical prompt (it'll fail again), or tripping your circuit breaker because the model sent one malformed payload.
Layer 1: Validate every tool call with Zod
Never trust tool_use input. Even with strict schemas in the prompt, models occasionally produce invalid values. Define each tool once, with a Zod schema as the source of truth, and derive the JSON Schema you send to the model from it.
// tools.ts
import { z } from 'zod';
export interface ToolDef<TInput extends z.ZodTypeAny = z.ZodTypeAny> {
name: string;
description: string;
schema: TInput;
/** Does this tool change state? Controls retry and idempotency behavior. */
sideEffects: boolean;
timeoutMs: number;
run: (input: z.infer<TInput>, ctx: { idempotencyKey: string }) => Promise<unknown>;
}
export const getOrderStatus: ToolDef = {
name: 'get_order_status',
description: 'Look up the current status of a customer order by its ID.',
schema: z.object({
orderId: z.string().regex(/^ORD-\d{6}$/, 'Order IDs look like ORD-123456'),
}),
sideEffects: false,
timeoutMs: 3_000,
run: async ({ orderId }) => db.orders.findStatus(orderId),
};
export const issueRefund: ToolDef = {
name: 'issue_refund',
description: 'Refund an order. Amount is in the smallest currency unit (cents).',
schema: z.object({
orderId: z.string().regex(/^ORD-\d{6}$/),
amountCents: z.number().int().positive().max(50_000),
reason: z.enum(['damaged', 'late', 'wrong_item', 'other']),
}),
sideEffects: true,
timeoutMs: 8_000,
run: async (input, { idempotencyKey }) =>
payments.refund({ ...input, idempotencyKey }),
};
export const registry = new Map(
[getOrderStatus, issueRefund].map((t) => [t.name, t]),
);
Note the details: tight constraints (regex, max, enum) catch nonsense early, and the descriptions in the schema double as documentation for the model.
For the JSON Schema you send to the provider, use z.toJSONSchema() (Zod 4) or the zod-to-json-schema package (Zod 3):
import { z } from 'zod';
export const toolSpecs = [...registry.values()].map((t) => ({
name: t.name,
description: t.description,
input_schema: z.toJSONSchema(t.schema), // Zod 4
}));
Turn validation failures into feedback, not exceptions
When validation fails, don't throw. Return a structured error to the model as the tool result, so it can self-correct on the next turn:
// execute-tool.ts
import { z } from 'zod';
export type ToolOutcome =
| { ok: true; content: unknown }
| { ok: false; kind: 'unknown_tool' | 'invalid_args' | 'tool_error'; message: string };
export async function executeToolCall(
name: string,
rawInput: unknown,
idempotencyKey: string,
): Promise<ToolOutcome> {
const tool = registry.get(name);
if (!tool) {
return {
ok: false,
kind: 'unknown_tool',
message: `Unknown tool "${name}". Available tools: ${[...registry.keys()].join(', ')}`,
};
}
const parsed = tool.schema.safeParse(rawInput);
if (!parsed.success) {
return {
ok: false,
kind: 'invalid_args',
message: `Invalid arguments: ${z.prettifyError(parsed.error)}`, // Zod 4; use error.issues in v3
};
}
try {
const content = await withTimeout(
tool.run(parsed.data, { idempotencyKey }),
tool.timeoutMs,
`Tool ${name} timed out`,
);
return { ok: true, content };
} catch (err) {
return { ok: false, kind: 'tool_error', message: (err as Error).message };
}
}
function withTimeout<T>(p: Promise<T>, ms: number, msg: string): Promise<T> {
let t: NodeJS.Timeout;
const timeout = new Promise<never>((_, rej) => {
t = setTimeout(() => rej(new Error(msg)), ms);
});
return Promise.race([p, timeout]).finally(() => clearTimeout(t));
}
The error message is deliberately specific and actionable. "Order IDs look like ORD-123456" lets the model fix its own mistake in one round trip. This is called a repair loop, and it resolves the large majority of malformed-call failures without any human or fallback involvement.
🔒 Security note: Schema validation is not authorization. If the model asks to refund order
ORD-000042, verify that the current user owns that order inside the tool. Treat model output as untrusted user input.
Layer 2: Timeouts and retries (for infrastructure failures only)
Now the provider call itself. Two rules:
- Always set a timeout. Without one, a hung connection ties up a request slot indefinitely.
- Retry only retryable errors, with exponential backoff and jitter.
// retry.ts
export interface RetryOptions {
retries: number;
baseMs: number;
maxMs: number;
isRetryable: (err: unknown) => boolean;
}
export async function retry<T>(fn: () => Promise<T>, o: RetryOptions): Promise<T> {
let attempt = 0;
for (;;) {
try {
return await fn();
} catch (err) {
if (attempt >= o.retries || !o.isRetryable(err)) throw err;
const exp = Math.min(o.maxMs, o.baseMs * 2 ** attempt);
const delay = Math.random() * exp; // "full jitter"
await new Promise((r) => setTimeout(r, delay));
attempt++;
}
}
}
export function isRetryableProviderError(err: any): boolean {
const status = err?.status ?? err?.statusCode;
if (status === 429 || status === 408 || status === 409) return true;
if (typeof status === 'number' && status >= 500) return true;
// Network-level failures
return ['ECONNRESET', 'ETIMEDOUT', 'ENOTFOUND', 'EAI_AGAIN'].includes(err?.code);
}
Things worth knowing:
-
Honor
Retry-Afteron 429s when the provider sends it, instead of guessing. - Jitter matters. Without it, every client retries at the same instant and re-creates the spike.
- Most official SDKs already retry a couple of times by default. Know what yours does, and don't stack your retries on top blindly. Retries multiply: 3 SDK retries × 3 of yours = 9 attempts.
- Never auto-retry non-idempotent tool executions (see Layer 5).
Layer 3: Circuit breakers
Retries help with blips. But when a provider is down or badly degraded, retrying every request makes things worse: you pile up slow requests, burn through rate limits, and make users wait for a failure that was predictable.
A circuit breaker watches recent failures and, past a threshold, fails fast for a cooldown period instead of calling the provider at all. After the cooldown, it lets a trial request through to test recovery.
failures exceed threshold
CLOSED ───────────────────────────▶ OPEN
▲ │
│ trial succeeds │ after resetTimeout
│ ▼
└─────────────────────────── HALF-OPEN
trial fails → back to OPEN
opossum is a solid, widely used breaker for Node:
// breaker.ts
import CircuitBreaker from 'opossum';
import { retry, isRetryableProviderError } from './retry';
async function callPrimary(params: ModelRequest): Promise<ModelResponse> {
return retry(() => anthropic.messages.create(params, { timeout: 30_000 }), {
retries: 2,
baseMs: 300,
maxMs: 3_000,
isRetryable: isRetryableProviderError,
});
}
export const primaryBreaker = new CircuitBreaker(callPrimary, {
timeout: 35_000, // hard ceiling for one guarded call
errorThresholdPercentage: 50, // open if >=50% of calls fail...
volumeThreshold: 10, // ...once we've seen at least 10 calls in the window
rollingCountTimeout: 30_000, // 30s statistics window
resetTimeout: 20_000, // try again after 20s
// 4xx client errors (except 408/429) are OUR bug, not provider health.
// Returning true means "don't count this as a failure".
errorFilter: (err: any) => {
const s = err?.status;
return typeof s === 'number' && s >= 400 && s < 500 && s !== 408 && s !== 429;
},
});
primaryBreaker.on('open', () => metrics.increment('llm.breaker.open', { provider: 'primary' }));
primaryBreaker.on('halfOpen', () => log.warn('LLM breaker half-open, probing'));
primaryBreaker.on('close', () => log.info('LLM breaker closed, provider recovered'));
Key decisions:
-
What counts as a failure? Timeouts, 5xx, 429s, network errors: yes. Schema validation failures and 400-class errors from your own bad requests: no. Otherwise one buggy prompt can take your whole feature offline. That's what
errorFilteris for. - Volume threshold. Without a minimum sample size, 1 failure out of 2 requests (50%) opens the circuit at 3 AM when traffic is low.
- Per-provider (or even per-model) breakers. Don't share one breaker across providers; you want to fail over, not fail everything.
-
Per-process state.
opossumkeeps state in memory per Node process. That's usually fine (each instance learns quickly), but it means a fleet of 50 pods each independently discovers the outage.
Layer 4: Fallbacks
When the breaker is open or retries are exhausted, what does the user see? Decide this deliberately. A good fallback chain degrades gracefully, from best to worst:
- Secondary model or provider, with the same tools
- Cheaper or smaller model, possibly with fewer tools
- Cached or deterministic answer for common queries
- Graceful "can't do that right now" message, with a way to retry or reach a human
// llm.ts
import CircuitBreaker from 'opossum';
const secondaryBreaker = new CircuitBreaker(callSecondary, {
timeout: 35_000,
errorThresholdPercentage: 50,
volumeThreshold: 10,
resetTimeout: 20_000,
});
export async function callModelResilient(req: ModelRequest): Promise<ModelResponse> {
const chain = [
{ name: 'primary', fire: () => primaryBreaker.fire(req) },
{ name: 'secondary', fire: () => secondaryBreaker.fire(adaptForSecondary(req)) },
];
let lastErr: unknown;
for (const step of chain) {
try {
const res = await step.fire();
metrics.increment('llm.call.ok', { via: step.name });
return res;
} catch (err) {
lastErr = err;
metrics.increment('llm.call.failed', { via: step.name });
// Breaker open => fail fast and move to the next link in the chain
}
}
throw new LlmUnavailableError('All model providers failed', { cause: lastErr });
}
A few cautions about fallbacks:
- Tool-call formats differ between providers. Message shapes, tool schema keys, and how results are returned are not identical. Write an adapter layer so the rest of your code speaks one internal format. Don't let provider-specific types leak through.
- Behavior differs between models. A fallback model might be worse at your tool schemas. Test your tool suite against every model in the chain, not just the primary.
- Consider restricting capabilities in fallback mode. For example, a weaker model might be allowed to read data but not to issue refunds.
- Don't fall back on model-side failures by default. If the model sent bad arguments, repair-loop first. Only escalate to a different model after the repair budget is spent.
And for the final fallback, make it honest and useful:
try {
return await runAgent(userMessage);
} catch (err) {
if (err instanceof LlmUnavailableError) {
return {
type: 'degraded',
message:
"I can't process that automatically right now. I've saved your request, " +
"and you can also check your order status at /orders.",
};
}
throw err;
}
Layer 5: The agent loop (loop guards and idempotency)
Tool use is a loop: the model asks for a tool, you run it, you send the result back, and repeat until the model produces a final answer. Without guardrails, that loop is a liability.
// agent.ts
import { randomUUID } from 'node:crypto';
const MAX_STEPS = 8;
const MAX_REPAIR_ATTEMPTS = 2;
export async function runAgent(userMessage: string, ctx: { requestId: string }) {
const messages: Message[] = [{ role: 'user', content: userMessage }];
const seenCalls = new Map<string, number>(); // detects repeated identical calls
let repairAttempts = 0;
for (let step = 0; step < MAX_STEPS; step++) {
const res = await callModelResilient({
model: process.env.LLM_MODEL!,
max_tokens: 1024,
tools: toolSpecs,
messages,
});
messages.push({ role: 'assistant', content: res.content });
if (res.stop_reason !== 'tool_use') {
return extractText(res); // final answer
}
const toolResults = [];
for (const block of res.content.filter((b) => b.type === 'tool_use')) {
// Loop guard: same tool + same args repeated => the model is stuck
const fingerprint = `${block.name}:${JSON.stringify(block.input)}`;
const count = (seenCalls.get(fingerprint) ?? 0) + 1;
seenCalls.set(fingerprint, count);
if (count > 2) {
throw new AgentLoopError(`Repeated identical call to ${block.name}`);
}
// Deterministic idempotency key: same request + same tool_use id => same key
const idempotencyKey = `${ctx.requestId}:${block.id}`;
const outcome = await executeToolCall(block.name, block.input, idempotencyKey);
if (!outcome.ok && outcome.kind !== 'tool_error') {
if (++repairAttempts > MAX_REPAIR_ATTEMPTS) {
throw new AgentRepairExhaustedError(outcome.message);
}
}
toolResults.push({
type: 'tool_result',
tool_use_id: block.id,
content: JSON.stringify(outcome.ok ? outcome.content : { error: outcome.message }),
is_error: !outcome.ok,
});
}
messages.push({ role: 'user', content: toolResults });
}
throw new AgentLoopError(`Exceeded ${MAX_STEPS} steps without a final answer`);
}
What this loop gives you:
- A hard step cap. An agent that can't finish within N steps shouldn't run forever or burn your token budget.
- A repair budget. The model gets two chances to fix its own mistakes, then you stop and fall back.
- Repeat detection. Identical calls three times in a row means the model is stuck.
-
is_error: trueon failed tool results tells the model the call failed, so it can adjust rather than treat the error text as real data.
Idempotency for side-effecting tools
Retries, timeouts, and double-submits mean a tool will occasionally run twice. For read-only tools, who cares. For issue_refund, it's a real-money bug.
The fix is an idempotency key passed to the downstream system (most payment APIs support this natively). Because we derive the key from requestId + tool_use id, a retry of the same logical call carries the same key, so the downstream service can safely deduplicate.
Also: don't auto-retry a side-effecting tool that timed out unless you know the operation is idempotent. A timeout doesn't mean it didn't run. Instead, check the state ("was the refund created?") or return the ambiguity to the user.
For high-risk tools, consider a confirmation step: the model proposes the action, your app shows it to the user, and only a human click executes it.
Testing resilience (the part everyone skips)
Resilience code that has never been exercised is just hopeful code. Build fault injection into your tests:
// agent.test.ts
import { describe, it, expect, vi } from 'vitest';
describe('agent resilience', () => {
it('repairs invalid tool arguments in one round trip', async () => {
mockModel
.respondWithToolUse('get_order_status', { orderId: '123' }) // invalid
.respondWithToolUse('get_order_status', { orderId: 'ORD-123456' }) // corrected
.respondWithText('Your order has shipped.');
const answer = await runAgent('Where is order 123456?', { requestId: 'r1' });
expect(answer).toContain('shipped');
});
it('fails over to the secondary provider when primary returns 529s', async () => {
mockPrimary.alwaysFail({ status: 529 });
mockSecondary.respondWithText('Hello from the backup.');
const res = await callModelResilient(basicRequest);
expect(extractText(res)).toContain('backup');
});
it('does not issue the same refund twice on retry', async () => {
const calls = await runWithSimulatedTimeoutAfterSuccess('issue_refund');
expect(payments.refund).toHaveBeenCalledTimes(1);
});
it('stops a model stuck calling the same tool', async () => {
mockModel.alwaysRespondWithToolUse('get_order_status', { orderId: 'ORD-123456' });
await expect(runAgent('hi', { requestId: 'r2' })).rejects.toThrow(AgentLoopError);
});
it('opens the breaker after sustained failures and fails fast', async () => {
mockPrimary.alwaysFail({ status: 503 });
for (let i = 0; i < 12; i++) await callModelResilient(basicRequest).catch(() => {});
expect(primaryBreaker.opened).toBe(true);
});
});
In staging, go further: use a proxy (Toxiproxy, or a simple fault-injecting middleware) to add latency, drop connections, and return 429s, and watch what your dashboards and fallbacks actually do.
What to monitor
Instrument each layer so you can tell which failure class is happening:
-
tool.validation_failedby tool name. A spike means a prompt or schema regression. -
tool.repair_succeededvsrepair_exhausted -
llm.call.latency(p50/p95/p99) andllm.call.failedby provider and status -
llm.breaker.statetransitions -
llm.fallback.usedrate. If you're on the fallback 20% of the time, your primary has a problem. -
agent.stepshistogram andagent.loop_guard_tripped - Token usage per request, to catch runaway loops through cost
And alert on rates and trends, not just absolute counts. A creeping validation failure rate after a prompt change is easy to miss otherwise.
Putting it all together
User request
│
▼
Agent loop (step cap, repeat detection)
│
├──▶ callModelResilient
│ ├─ Primary breaker ─▶ retry (backoff + jitter) ─▶ Provider A
│ └─ (open/failed) ───▶ Secondary breaker ───────▶ Provider B
│
├──▶ tool_use returned
│ ├─ Unknown tool / bad args ─▶ error back to model (repair loop)
│ └─ Valid ─▶ timeout + idempotency key ─▶ run tool
│
└──▶ All providers down? ─▶ graceful degraded response
Each layer is small and testable on its own, and together they cover the full set of failure modes.
Key takeaways
- Validate everything the model sends you. Schema-first tools with Zod, and return validation errors to the model so it can repair.
- Separate model failures from infrastructure failures. Only the latter should trigger retries, breakers, and failover.
- Retry with backoff and jitter, and only for retryable errors.
- Use circuit breakers so a sick provider fails fast instead of dragging your app down with it.
- Design the fallback chain on purpose, including what the user sees at the very end.
- Cap the loop and make side effects idempotent. Agents will eventually loop or double-fire; make sure it's harmless.
- Test the failure paths. If you haven't injected faults, you don't know it works.
Resilience isn't about preventing the model from ever being wrong or the provider from ever being down. Both will happen. It's about making sure that when they do, your users see a slightly slower answer or a polite degraded one instead of a stack trace, a double refund, or a hung page.
What failure modes have you hit with LLM tool calling in production? Share them in the comments.
Further reading: opossum (circuit breaker), Zod (schema validation), and your provider's tool-use documentation for the exact message formats.
Top comments (0)