We read 1,662 public posts from builders whose apps broke at or after launch and verified 215 recent cases. AI model and API failures make up 4% of them. Six of the verified cases share one pattern: an outbound or AI call with no timeout, no budget and no fallback. In most of the six, nobody noticed for a while, because the app didn't log model calls or their outcomes.
From the builder's side it shows up as one of these:
- A spinner that runs for minutes, then a 504 from Vercel or a 524 from Cloudflare. On localhost it was fine.
- Users read polite fallback text while the spend dashboard looks normal, because failed calls cost almost nothing.
- An agent or background job burns a week's budget in a day. Then the provider's spend limit trips and every user is cut off at once.
All three have the same cause. The call has no limits of its own, so the only limits that apply are set by other layers.
Reproduce it
Here is a Next.js App Router route using the OpenAI Node SDK with its defaults. Most generated code ships in this shape.
// app/api/summary/route.ts
import OpenAI from "openai";
const openai = new OpenAI(); // defaults: 10 min per attempt, 2 retries
export async function POST(req: Request) {
const { text } = await req.json();
try {
const r = await openai.chat.completions.create({
model: "gpt-5-mini",
messages: [{ role: "user", content: `Summarise this:\n${text}` }],
});
return Response.json({ summary: r.choices[0].message.content });
} catch {
// the user sees this with a 200, and nothing is logged
return Response.json({ summary: "Summary unavailable right now." });
}
}
You don't need a provider outage to see it fail. The SDK reads OPENAI_BASE_URL, so you can point it at a server that accepts connections and never replies:
# terminal 1: a provider that never answers
node -e "require('http').createServer(() => {}).listen(4010)"
# terminal 2
OPENAI_BASE_URL=http://127.0.0.1:4010/v1 OPENAI_API_KEY=test npm run dev
# terminal 3
curl -s -X POST localhost:3000/api/summary \
-H 'content-type: application/json' -d '{"text":"hello"}' -w '\n%{time_total}s\n'
curl waits for at least 15 minutes. On Vercel, the same request is killed at 300 seconds with a 504 FUNCTION_INVOCATION_TIMEOUT before the catch runs, so nothing is logged. Behind Cloudflare the user gets a 524 even sooner.
Now make the stub answer 500 instead: node -e "require('http').createServer((q, s) => { s.statusCode = 500; s.end('{}') }).listen(4010)". After a few seconds of retries the route returns 200 with the fallback sentence. Every monitor you have counts that as a success.
Why it happens
Every layer between the browser and the model has its own clock. In the default setup the clock in your own code is the longest, so an outer layer always runs out first.
| Layer | Default limit | What happens |
|---|---|---|
| Cloudflare proxy | 125 s with no response from the origin | 524 to the user |
| Vercel function (fluid compute) |
maxDuration 300 s |
function terminated, 504 |
Node's fetch (undici) |
300 s waiting for response headers | connection error |
| OpenAI and Anthropic Node SDKs | 10 minutes per attempt | APIConnectionTimeoutError |
Here is what the official docs say about each part:
-
The SDK timeout is per attempt, and retries multiply it. By default, both official Node SDKs retry connection errors, 408, 409, 429 and 5xx twice with a short backoff, and they retry timed-out requests too. That is where the 15 minutes comes from: three attempts, each waiting at least 300 seconds. For non-streaming requests with a large
max_tokens, Anthropic's SDK raises its default timeout above 10 minutes, up to 60. - A non-streamed response sends nothing until the model finishes. Every proxy in front of it sees an idle connection. Anthropic's SDK docs recommend streaming for long requests for this reason. The SDK also throws an error for a non-streaming request expected to run past roughly 10 minutes, unless you set a timeout.
-
The platform ends the process, not the promise. Vercel terminates a function that runs past its
maxDuration. Yourcatchnever runs, and neither does the logging inside it. -
The fallback hides the errors you did catch. A
catchthat returns a 200 turns every provider error into a success in your metrics. Failed calls use few tokens, so the cost chart stays flat. -
Nothing caps spend per user or per run. Provider spend limits are monthly and cover a whole project or workspace. When OpenAI's hard limit is reached, requests return 429 with the code
project_spend_limit_exceeded. Because the SDKs retry every 429, each user's request retries an error that won't clear until the limit resets. Anthropic returns 400invalid_request_errorwhen a workspace spend limit is reached, and retry logic that only looks for 429 won't recognise it.
Catch it before your users do
This suite starts a fake provider on localhost with three behaviours: it never answers, it plays an agent that never finishes, or it returns a spend-limit 429. It imports the wrapper from the next section, and its assertions define what that wrapper has to do. It needs no API key and costs nothing to run.
// lib/ai-guards.test.ts
import http from "node:http";
import type { AddressInfo } from "node:net";
import { afterAll, beforeAll, expect, test } from "vitest";
import { callModel, makeClient, newRun, runAgent } from "./ai";
process.env.OPENAI_API_KEY ??= "test";
let hits = 0;
// One fake provider; the URL prefix picks its behaviour.
const server = http.createServer((req, res) => {
hits++;
if (req.url?.startsWith("/stall/")) return; // accept, never answer
res.setHeader("content-type", "application/json");
if (req.url?.startsWith("/spent/")) {
res.statusCode = 429;
return res.end(JSON.stringify({ error: { code: "project_spend_limit_exceeded", message: "limit reached" } }));
}
res.end(JSON.stringify({ // "/loop/": a model that never says it is done
id: "c1", object: "chat.completion", created: 0, model: "stub",
choices: [{ index: 0, finish_reason: "stop", message: { role: "assistant", content: "keep going" } }],
usage: { prompt_tokens: 500, completion_tokens: 500, total_tokens: 1000 },
}));
});
const client = (mode: string) =>
makeClient(`http://127.0.0.1:${(server.address() as AddressInfo).port}/${mode}/v1`);
const hi = [{ role: "user" as const, content: "hi" }];
beforeAll(() => new Promise<void>((done) => { server.listen(0, "127.0.0.1", done); }));
afterAll(() => { server.closeAllConnections(); server.close(); });
test("a provider that never answers fails inside the deadline", async () => {
const t0 = Date.now();
await expect(callModel(client("stall"), newRun(), hi, 2_000)).rejects.toThrow();
expect(Date.now() - t0).toBeLessThan(3_000);
});
test("an agent that never finishes stops at its step cap", async () => {
const run = newRun(5, 1);
await expect(runAgent(client("loop"), run, "plan my week")).rejects.toThrow("run_budget_exceeded");
expect(run.steps).toBe(5);
});
test("a spend-limit 429 surfaces its code after at most one retry", async () => {
hits = 0;
await expect(callModel(client("spent"), newRun(), hi)).rejects.toMatchObject({
status: 429, code: "project_spend_limit_exceeded",
});
expect(hits).toBeLessThanOrEqual(2);
});
npm i -D vitest
npx vitest run lib/ai-guards
Each test takes a few seconds. Run the suite in CI next to your other tests.
The fix
Give every call its own limits, shorter than every layer outside it, and make failures visible.
// lib/ai.ts
import OpenAI from "openai";
export type Run = { steps: number; maxSteps: number; spentUsd: number; capUsd: number };
export const newRun = (maxSteps = 1, capUsd = 0.05): Run => ({ steps: 0, maxSteps, spentUsd: 0, capUsd });
// 20 s per attempt, one retry. The deadline in callModel caps the total.
export const makeClient = (baseURL?: string) => new OpenAI({ baseURL, timeout: 20_000, maxRetries: 1 });
// Set these from your model's price page; at 0 the spend cap never trips.
const USD_PER_TOKEN = { input: 0, output: 0 };
export async function callModel(
client: OpenAI,
run: Run,
messages: OpenAI.Chat.ChatCompletionMessageParam[],
deadlineMs = 45_000,
) {
if (run.steps >= run.maxSteps || run.spentUsd >= run.capUsd) throw new Error("run_budget_exceeded");
run.steps++;
const res = await client.chat.completions.create(
{ model: "gpt-5-mini", messages, max_completion_tokens: 800 },
{ signal: AbortSignal.timeout(deadlineMs) }, // one clock for the whole call, retries included
);
const u = res.usage;
run.spentUsd += (u?.prompt_tokens ?? 0) * USD_PER_TOKEN.input + (u?.completion_tokens ?? 0) * USD_PER_TOKEN.output;
return res;
}
export async function runAgent(client: OpenAI, run: Run, task: string) {
const messages: OpenAI.Chat.ChatCompletionMessageParam[] = [{ role: "user", content: task }];
for (;;) {
const res = await callModel(client, run, messages);
const text = res.choices[0].message.content ?? "";
if (text.includes("[done]")) return text;
messages.push({ role: "assistant", content: text }, { role: "user", content: "Continue." });
}
}
The route now fails loudly and leaves a record:
// app/api/summary/route.ts
import { callModel, makeClient, newRun } from "@/lib/ai";
const client = makeClient();
export async function POST(req: Request) {
const { text } = await req.json();
const t0 = Date.now();
try {
const r = await callModel(client, newRun(), [{ role: "user", content: `Summarise this:\n${text}` }], 25_000);
console.info(JSON.stringify({ tool: "summary", status: "ok", ms: Date.now() - t0, usage: r.usage }));
return Response.json({ summary: r.choices[0].message.content });
} catch (err: any) {
console.error(JSON.stringify({ tool: "summary", status: "error", code: err?.code ?? err?.name, http: err?.status, ms: Date.now() - t0 }));
return Response.json({ error: "summary_failed" }, { status: 503 });
}
}
Pick deadlineMs below both your proxy's timeout and your function's maxDuration, with time left over to log and respond. If an answer really needs longer, stream it so bytes reach the proxy early, or return a job ID and have the client poll. With streaming, a 200 only means the stream started. Log the call as failed if the stream ends without a finish_reason.
For a per-user daily cap, add up the user's spend for today from the same log rows before creating the run, and pass what's left as capUsd. Once those rows are in a table, this query shows the failures you already have. It sorts by failure rate rather than spend, because failed calls barely move a cost chart:
select tool, count(*) as calls,
avg((status <> 'ok')::int) as failure_rate
from ai_calls
where created_at > now() - interval '7 days'
group by tool
order by failure_rate desc;
Checklist
- [ ] Every provider client sets
timeoutandmaxRetriesexplicitly. - [ ] Every call has one overall deadline (
AbortSignal.timeout), shorter thanmaxDurationand your proxy's timeout. - [ ] Long answers stream, or run as a background job that the client polls.
- [ ] Every agent run has a step cap and a spend cap. Every user has a daily cap, checked before the call.
- [ ] A failed call returns a non-2xx status and writes one log line with tool, status, error code and duration.
- [ ] OpenAI's 429
project_spend_limit_exceededand Anthropic's spend-limit 400 show the user a clear paused state instead of starting a retry loop. - [ ] A provider hard spend limit is set as a backstop, not used as the budget.
- [ ] The three tests above run in CI.
Gemmein (in beta) handles the per-user half of this for you: its AI tools spend or reserve a person's credits before the provider is called, and every call is recorded with who, tool, tokens, credits and outcome.
Which layer in your stack has the shortest clock today, and did you set it yourself?
Top comments (2)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.