DEV Community

Cover image for Your Model Upgrade Is a Breaking Change: Build Contract Tests for LLM Providers in TypeScript
Bobby Hall Jr
Bobby Hall Jr

Posted on

Your Model Upgrade Is a Breaking Change: Build Contract Tests for LLM Providers in TypeScript

Most code that calls a model has one line that looks harmless.

model: "claude-sonnet-5"
Enter fullscreen mode Exit fullscreen mode

Changing it feels like a config change.

But that string is part of an API contract.

And this month, the contracts changed.

  • On September 17, Google released Antigravity Agent 09-2026. If you run tools locally or parse function_call steps, "the built-in tools changed": PascalCase parameters, and write_file(path, content) became write_to_file or replace_file_content. The old antigravity-preview-05-2026 "shuts down on October 5, 2026."
  • On September 22, Anthropic launched Claude Opus 5.5. Per the release notes, thinking: {"type": "disabled"} and {"type": "enabled", ...} "return a 400 error." So do tool_choice types any and tool.
  • On September 28, Anthropic launched Claude Sonnet 5.5. The release notes say "Code written for Claude Sonnet 5 can break on Claude Sonnet 5.5 in five ways."
  • On September 29, at DevDay, OpenAI released GPT-6.1 Sol. Its model page says "The none and minimal reasoning efforts are not supported."

Here are Sonnet 5.5's five, in the release notes' words:

  1. "To turn off up-front thinking, send thinking: {"type": "between_tools"} instead of "disabled", at high effort or below."
  2. "Forced tool use (tool_choice types any and tool) returns a 400 error."
  3. "Thinking blocks are tied to the model and the conversation."
  4. "On the Claude API and Google Cloud, the earlier computer_20251124 computer use tool isn't accepted."
  5. "The advisor tool rejects Claude Opus 4.8, Claude Opus 4.7, and Claude Sonnet 5 as advisors."

The What's new page adds one that "alters the response shape without failing any request": text between tool calls comes back in thinking blocks.

A 400 is loud.

An empty progress message is quiet.

Different companies.

Same pattern.

Changing a model string is a dependency upgrade. It deserves a test suite.

So let's build one.

No API key. Both providers are mocks: their request rules follow the docs above, and their replies are made up.

Code: github.com/bobbyhalljr/model-upgrade-gate

Table of Contents

  1. What We Are Building
  2. Project Setup
  3. Step 1: A Neutral Request and Response
  4. Step 2: Mock Two Model Versions
  5. Step 3: Write the Contracts
  6. Step 4: Diff Two Runs
  7. Step 5: Check Known Breaks From the Docs
  8. Step 6: Build the Gate and Run It
  9. Where It Breaks Down
  10. The Bigger Idea

What We Are Building

The model upgrade gate

One check reads the release notes. The other diffs two model versions.

Project Setup

You will need Node.js 18 or newer.

mkdir model-upgrade-gate
cd model-upgrade-gate

npm init -y
npm install --save-dev typescript tsx @types/node
Enter fullscreen mode Exit fullscreen mode

Save the following blocks, in order, as upgrade-gate.ts.

Step 1: A Neutral Request and Response

type Req = {
  prompt: string;
  maxTokens: number;
  thinking?: { type: "adaptive" | "disabled" | "between_tools" };
  toolChoice?: { type: "auto" | "none" | "any" | "tool" };
  tools?: string[];
};

type Block =
  | { type: "text"; text: string }
  | { type: "thinking"; thinking: string }
  | { type: "tool_use"; name: string; input: Record<string, unknown> };

type Stop = "end_turn" | "tool_use" | "max_tokens" | "refusal";
type Ok = { status: 200; stopReason: Stop; content: Block[] };
type Res = Ok | { status: 400; error: string };

type Provider = { model: string; send: (req: Req) => Res };
Enter fullscreen mode Exit fullscreen mode

Your app's shape, not a vendor SDK.

Your contracts should describe your app, not the provider.

Step 2: Mock Two Model Versions

// MOCK PROVIDER. No network, no API key. The two Sonnet 5.5 rejections follow
// Anthropic's docs (Sep 28), and the tool_choice error is quoted from them.
// The other error text and every reply are made up.
const reply = (stopReason: Stop, ...content: Block[]): Ok => ({ status: 200, stopReason, content });

function mockClaude(model: "claude-sonnet-5" | "claude-sonnet-5-5"): Provider {
  const v55 = model === "claude-sonnet-5-5";

  const send = (req: Req): Res => {
    if (v55 && req.thinking?.type === "disabled") {
      return { status: 400, error: 'invalid_request_error: use "between_tools"' };
    }
    if (v55 && ["any", "tool"].includes(req.toolChoice?.type ?? "auto")) {
      return { status: 400, error: 'tool_choice: type "tool" and "any" are not supported for this model.' };
    }
    if (req.prompt.startsWith("[refuse]")) return reply("refusal");
    if (req.maxTokens < 50) return reply("max_tokens", { type: "text", text: "Q3 revenue grew" });

    if (req.tools?.includes("get_weather")) {
      const note = "Checking the forecast first. Then I'll compare it with yesterday.";
      const shown = req.thinking?.type === "between_tools" ? note : "";
      const progress: Block = v55 ? { type: "thinking", thinking: shown } : { type: "text", text: note };
      return reply("tool_use", progress, { type: "tool_use", name: "get_weather", input: { city: "Paris" } });
    }
    if (req.tools?.includes("classify_ticket")) {
      return reply("tool_use", { type: "tool_use", name: "classify_ticket", input: { label: "billing" } });
    }
    return reply("end_turn", { type: "text", text: '{"total": 42.5, "currency": "USD"}' });
  };

  return { model, send };
}
Enter fullscreen mode Exit fullscreen mode

Sonnet 5.5 rejects disabled thinking and forced tool use.

Its progress note also moves into a thinking block. At the default display: "omitted", the docs say its text is empty. With between_tools, it comes back.

A mock that agrees with everything is just a very polite liar.

Step 3: Write the Contracts

type Contract = { name: string; req: Req; check: (res: Ok) => string | null };

const weather: Req = { prompt: "Weather in Paris?", maxTokens: 500, tools: ["get_weather"] };

const contracts: Contract[] = [
  {
    name: "output schema",
    req: { prompt: "Extract the invoice total as JSON", maxTokens: 500 },
    check: ({ content: [first] }) => {
      const data = JSON.parse(first?.type === "text" ? first.text : "null");
      return typeof data?.total === "number" && typeof data?.currency === "string" ? null : "bad JSON shape";
    },
  },
  {
    name: "tool call format",
    req: weather,
    check: (res) => {
      const call = res.content.find((b) => b.type === "tool_use");
      return res.stopReason === "tool_use" && typeof call?.input.city === "string" ? null : "bad tool call";
    },
  },
  {
    name: "progress text between tools",
    req: weather,
    check: ({ content: [first] }) => {
      const shown = first?.type === "text" ? first.text : first?.type === "thinking" ? first.thinking : "";
      return shown ? null : `user sees nothing before the tool call (empty ${first?.type} block)`;
    },
  },
  {
    name: "forced tool use",
    req: { prompt: "Classify this ticket", maxTokens: 200, tools: ["classify_ticket"], toolChoice: { type: "tool" } },
    check: (res) => (res.stopReason === "tool_use" ? null : "no tool call"),
  },
  {
    name: "thinking off (fast path)",
    req: { prompt: "Summarize in one line", maxTokens: 200, thinking: { type: "disabled" } },
    check: () => null, // a 200 is the whole contract
  },
  {
    name: "token limit stop reason",
    req: { prompt: "Write the full quarterly report", maxTokens: 20 },
    check: (res) => (res.stopReason === "max_tokens" ? null : `got ${res.stopReason}`),
  },
  {
    name: "refusal behavior",
    req: { prompt: "[refuse] a request the model declines", maxTokens: 200 },
    check: (res) => (res.stopReason === "refusal" && res.content.length === 0 ? null : "refusal not clean"),
  },
];
Enter fullscreen mode Exit fullscreen mode

Seven promises. Each check returns null or a reason.

The refusal contract follows the docs: a declined request returns HTTP 200 with stop_reason: "refusal".

Step 4: Diff Two Runs

type Result = { name: string; pass: boolean; detail: string };

function runSuite(provider: Provider): Result[] {
  return contracts.map(({ name, req, check }) => {
    const res = provider.send(req);
    if (res.status !== 200) return { name, pass: false, detail: `${res.status} ${res.error}` };
    try {
      const failure = check(res);
      return { name, pass: failure === null, detail: failure ?? "" };
    } catch (err) {
      return { name, pass: false, detail: `threw: ${(err as Error).message}` };
    }
  });
}

function contractDiff(current: Provider, candidate: Provider) {
  const before = runSuite(current);
  const after = runSuite(candidate);
  const broke: string[] = [];

  console.log(`\nContract diff (MOCK ${current.model} -> MOCK ${candidate.model})`);
  after.forEach((a, i) => {
    const status = before[i].pass && !a.pass ? "BROKE" : a.pass ? "same" : "FAIL";
    if (status === "BROKE") broke.push(a.name);
    console.log(`  ${status.padEnd(6)} ${a.name.padEnd(28)} ${a.detail}`.trimEnd());
  });
  return broke;
}
Enter fullscreen mode Exit fullscreen mode

Only one transition matters: passed before, fails now.

Step 5: Check Known Breaks From the Docs

type KnownBreak = { model: string; param: string; bad: string[]; docs: string; source: string };

// From the vendors' docs, checked Oct 3, 2026.
const knownBreaks: KnownBreak[] = [
  { model: "claude-sonnet-5-5", param: "thinking.type", bad: ["disabled"], docs: 'send "between_tools" instead', source: "Claude notes, Sep 28" },
  { model: "claude-sonnet-5-5", param: "tool_choice.type", bad: ["any", "tool"], docs: "returns a 400 error", source: "Claude notes, Sep 28" },
  { model: "claude-opus-5-5", param: "thinking.type", bad: ["disabled", "enabled"], docs: "returns a 400 error", source: "Claude notes, Sep 22" },
  { model: "gpt-6.1-sol", param: "reasoning.effort", bad: ["none", "minimal"], docs: "not supported", source: "OpenAI model page" },
  { model: "antigravity-preview-09-2026", param: "tools", bad: ["write_file", "read_file", "list_files"], docs: "built-in tools changed", source: "Gemini changelog, Sep 17" },
];

const shutdowns: Record<string, string> = { "antigravity-preview-05-2026": "2026-10-05" };

type CallSite = { site: string; from: string; to: string; params: Record<string, string[]> };

function checkKnownBreaks(sites: CallSite[], today: string) {
  let count = 0;
  console.log("\nKnown breaks (from release notes)");
  for (const s of sites) {
    for (const rule of knownBreaks.filter((r) => r.model === s.to)) {
      for (const value of (s.params[rule.param] ?? []).filter((v) => rule.bad.includes(v))) {
        count++;
        console.log(`  BREAK    ${s.site}: ${rule.param}=${value}: ${rule.docs} [${rule.source}]`);
      }
    }
    const end = shutdowns[s.from];
    const days = (Date.parse(end) - Date.parse(today)) / 86_400_000;
    if (end) console.log(`  DEADLINE ${s.site}: ${s.from} shuts down ${end} (${days} days)`);
  }
  if (count === 0) console.log("  no known breaks");
  return count;
}
Enter fullscreen mode Exit fullscreen mode

This is the deprecated-params check. Every row comes from a vendor's docs, with its date.

The DEADLINE line isn't a failure. It's a reason to hurry.

Known breaks, straight from the docs

Step 6: Build the Gate and Run It

// Illustrative call sites in a made-up app.
const S5 = "claude-sonnet-5", S55 = "claude-sonnet-5-5";
const callSites: CallSite[] = [
  { site: "invoice-extractor", from: S5, to: S55, params: {} },
  { site: "ticket-classifier", from: S5, to: S55, params: { "tool_choice.type": ["tool"] } },
  { site: "fast-summary", from: S5, to: S55, params: { "thinking.type": ["disabled"] } },
  { site: "code-agent", from: "gpt-6-sol", to: "gpt-6.1-sol", params: { "reasoning.effort": ["none"] } },
  { site: "file-agent", from: "antigravity-preview-05-2026", to: "antigravity-preview-09-2026", params: { tools: ["write_file"] } },
];

const today = "2026-10-03";
console.log(`Upgrade gate, ${today}`);

const breaks = checkKnownBreaks(callSites, today);
const broke = contractDiff(mockClaude(S5), mockClaude(S55));

const blocked = breaks > 0 || broke.length > 0;
console.log(blocked ? `\nGATE: BLOCKED (${breaks} known breaks, ${broke.length} contract regressions)` : "\nGATE: OPEN");
process.exitCode = blocked ? 1 : 0;
Enter fullscreen mode Exit fullscreen mode

Run it:

npx tsx upgrade-gate.ts
Enter fullscreen mode Exit fullscreen mode

Real output:

Upgrade gate, 2026-10-03

Known breaks (from release notes)
  BREAK    ticket-classifier: tool_choice.type=tool: returns a 400 error [Claude notes, Sep 28]
  BREAK    fast-summary: thinking.type=disabled: send "between_tools" instead [Claude notes, Sep 28]
  BREAK    code-agent: reasoning.effort=none: not supported [OpenAI model page]
  BREAK    file-agent: tools=write_file: built-in tools changed [Gemini changelog, Sep 17]
  DEADLINE file-agent: antigravity-preview-05-2026 shuts down 2026-10-05 (2 days)

Contract diff (MOCK claude-sonnet-5 -> MOCK claude-sonnet-5-5)
  same   output schema
  same   tool call format
  BROKE  progress text between tools  user sees nothing before the tool call (empty thinking block)
  BROKE  forced tool use              400 tool_choice: type "tool" and "any" are not supported for this model.
  BROKE  thinking off (fast path)     400 invalid_request_error: use "between_tools"
  same   token limit stop reason
  same   refusal behavior

GATE: BLOCKED (4 known breaks, 3 contract regressions)
Enter fullscreen mode Exit fullscreen mode

Exit code 1. CI stops.

Look at progress text between tools. No 400. The user just stops seeing progress.

The quiet break is the one a status code will never catch.

The docs name the fixes: between_tools, auto plus strict tool use, low instead of none, and the new Antigravity tool names.

Where It Breaks Down

What a mock can and can't catch

Mocks Drift

I copied the rules by hand. They also differ by platform: computer_20251124 is rejected on the Claude API and Google Cloud, but Sonnet 5.5 still accepts it on Amazon Bedrock.

Run the contracts against the real API before trusting a green gate.

Behavior Isn't a Contract

Anthropic says Sonnet 5.5's "effort levels are recalibrated." A schema check can't see that. Evals can.

State Needs Real Conversations

Thinking blocks are tied to the model, the conversation and the account. On newer accounts, replaying one after editing history can return a 400. Single requests miss that.

The Bigger Idea

We already treat libraries this way.

Pin the version. Read the changelog. Run the tests. Then upgrade.

Models get a string change and a hopeful deploy.

┌──────────────────────────────────────────────┐
│                 Upgrade gate                 │
│                                              │
│  Release notes ──→ Known breaks ──┐          │
│                                   ↓          │
│  Current   ──→ Contracts ──→ Diff ──→ Gate   │
│  Candidate ──→ Contracts ──┘                 │
└──────────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

The model provides capability.

The release notes provide warnings.

The contracts provide expectations.

The diff provides evidence.

The gate provides a decision.

Three vendors, four releases, twelve days. I think upgrade gates become as normal as lockfiles.

That part is prediction, not history.

A new model is a new dependency. Ship it like one.


Your code has a history. Helix makes it understandable.

I'm building Helix so every change, including a model upgrade, comes with evidence: what changed, why, and what it touched.

Connect your GitHub and see what your code knows.

Explore Helix →

Top comments (0)