DEV Community

Cover image for Your LLM Call Has No Ceiling. TypeScript Is Fine With It. Three CWEs Aren't.
Ofri Peretz
Ofri Peretz

Posted on Originally published at ofriperetz.dev

Your LLM Call Has No Ceiling. TypeScript Is Fine With It. Three CWEs Aren't.

generateText with a model and a prompt and nothing else is a valid call. TypeScript is satisfied. The type system has no opinion about how many tokens come back, how long it takes, or who pays for the answer nobody reads.

Three rules in eslint-plugin-vercel-ai-security fire on a call written that way, and they map to three different CWEs — 770, 400 and 404. (A CWE is a label, not a verdict — what the taxonomy actually claims — and counting CWEs is a proxy metric.) Same missing config object. Three separate ways to lose.

A note on the bound people expect: step count. The SDK defaults stopWhen to stepCountIs(1), so a tool-calling loop does not run away on its own — that hole gets opened deliberately, by raising the ceiling, not by forgetting it. The three below are genuinely unbounded when you say nothing.


1. The output with no cap — CWE-770

await generateText({ model, prompt }); // no maxOutputTokens
Enter fullscreen mode Exit fullscreen mode

Why it gets written: the parameter is optional, and the happy path never needs it.

Why it survives review: output length reads as a quality knob, not a resource bound. But this is billed per token. An unbounded output is an unbounded invoice, and nothing in the diff looks like money.

Fix: set maxOutputTokens to the longest answer you would pay for — there is no correct number, only the difference between a ceiling and none. Note the rename: this was maxTokens in v4. Guidance written against v4 still reads correct in review and bounds nothing.


2. The request that never returns — CWE-400

await generateText({ model, prompt }); // no timeout
Enter fullscreen mode Exit fullscreen mode

Why it gets written: fetch-shaped APIs feel like they time out. This one doesn't — timeout is optional and the SDK sets no default.

Why it survives review: staging latency is fine, so nobody asks what happens when the provider hangs instead of failing.

await generateText({ model, prompt, timeout: { totalMs: 30_000 } });
Enter fullscreen mode Exit fullscreen mode

timeout has been first-class since 6.0.14 and also takes stepMs, chunkMs, toolMs. Below that, hand-roll an AbortController; the rule accepts either, but reach for the parameter first.


3. The stream nobody can cancel — CWE-404

const stream = streamText({ model, prompt }); // no abortSignal
Enter fullscreen mode Exit fullscreen mode

You will file this under polish. Here is why that is wrong: the client is gone, and the server keeps generating — and billing — for a reader who will never see a token of it.

Why shutdown (CWE-404) rather than consumption (400)? The resource was acquired correctly and never released — the handle outlives the request that justified it. A release bug, not an acquisition bug.

const ac = new AbortController();
req.signal.addEventListener("abort", () => ac.abort()); // client disconnected
const stream = streamText({ model, prompt, abortSignal: ac.signal });
Enter fullscreen mode Exit fullscreen mode

The pattern

All three are one defect: an allocation with no ceiling. Tokens, wall-clock, lifetime.

I don't trust an allocation whose ceiling I can't see — the same instinct that keeps me out of a position whose downside I can't draw on one page.

Classic resource-exhaustion review asks whether an attacker can make something loop forever. For an LLM call, the answer is worse — you don't need an attacker. A verbose model and one retry will do it, and the meter runs the entire time.

This is OWASP LLM10, Unbounded Consumption — the full top-10 mapping for this SDK covers the other nine. That list is a separate taxonomy from the web Top 10 — LLM10 has no A-number. For governance rather than linting, see NIST AI RMF. That is why these are security rules, not style rules. It is the move injection across nine interpreters makes too: CWE-770, 400 and 404 are the same sentence in three grammars.


The config

npm install --save-dev eslint-plugin-vercel-ai-security
Enter fullscreen mode Exit fullscreen mode

Node 18+. The peer range is ESLint 8 ∥ 9 ∥ 10, so both config formats work.

// eslint.config.mjs — ESLint 9 · 10
import vercelAi from "eslint-plugin-vercel-ai-security";

export default [
  {
    files: ["**/*.ts"],
    plugins: { "vercel-ai-security": vercelAi },
    rules: {
      "vercel-ai-security/require-max-tokens": "error",
      "vercel-ai-security/require-request-timeout": "warn",
      "vercel-ai-security/require-abort-signal": "warn",
    },
  },
];
Enter fullscreen mode Exit fullscreen mode
// .eslintrc.json — ESLint 8
{
  "plugins": ["vercel-ai-security"],
  "rules": {
    "vercel-ai-security/require-max-tokens": "error",
    "vercel-ai-security/require-request-timeout": "warn",
    "vercel-ai-security/require-abort-signal": "warn"
  }
}
Enter fullscreen mode Exit fullscreen mode

On oxlint, add "jsPlugins": ["eslint-plugin-vercel-ai-security/oxlint"] — alpha, not under semver.

Rule docs: require-max-tokens · require-request-timeout · require-abort-signal


More of these — follow on dev.to.

Has an LLM call ever run longer in production than you expected — and what told you first, the logs or the bill?


Related:

npm · GitHub

Top comments (0)