Your model returned JSON.
It looked fine.
Then production asked for a number and got "high".
Structured outputs are having a moment.
- Anthropic's structured outputs docs (checked October 11, 2026) describe JSON outputs via
output_config.formatand strict tool use. The page lists the failure modes people already know: parsing errors, missing fields, inconsistent types, schema violations that force retries. - OpenAI's structured outputs guide (checked October 11, 2026) frames the feature as schema adherence, not just "valid JSON." It still calls out refusals and max-token cutoffs as cases where the response may not match your schema.
- On October 7, 2026, Anthropic introduced Claude Haiku 5.5 for high-volume, cost-sensitive work: summaries, classification, routing, and similar short tasks. When a small model answers that often, a bad shape is not a one-off. It is a rate.
Different products. Same direction.
The provider can constrain the sample.
Your system still has to decide what to do when the shape is wrong, truncated, or refused.
A schema that passed once is a demo. A gate that rejects bad shapes is a product.
So let's build a tiny version.
By the end, you'll run one command:
npx tsx gate.ts
And watch eight scripted replies hit a validator, take one repair, then accept or reject.
No API key.
No real model.
Just TypeScript.
One honesty note: the "model" here is a fixture table. The confidence numbers and pass rates are example inputs. Swap in your provider's complete call and keep the gate.
Table of Contents
- What We Are Building
- Project Setup
- Step 1: Define the Schema
- Step 2: Validate Without a Framework
- Step 3: Script the Failure Modes
- Step 4: One Repair, Then Decide
- Step 5: Metrics and a Ship Gate
- Where It Breaks Down
- The Bigger Idea
Code: github.com/bobbyhalljr/tiny-schema-gate
What We Are Building
A routing decision with three fields.
action is an enum: route, escalate, or ignore.
target is a string.
confidence is a number between 0 and 1.
The mock model returns valid JSON, truncated JSON, wrong types, a bad enum, extra fields, a missing field, an out-of-range confidence, and one stubborn pair that never repairs.
The gate does three things:
- Parse.
- Validate against the schema.
- If invalid, feed the errors back once. Then accept or reject.
It's also a small version of the question behind Helix: when the system changes, can you show what got better, what got worse, and why?
Project Setup
You will need Node.js 18 or newer.
mkdir tiny-schema-gate
cd tiny-schema-gate
npm init -y
npm install --save-dev typescript tsx @types/node
Save the following blocks, in order, as gate.ts.
Step 1: Define the Schema
// gate.ts: a tiny schema gate for structured LLM output.
// Everything is mocked: the "model" returns scripted strings, not provider calls.
// The fixtures, confidence numbers, and pass rates below are EXAMPLE INPUTS.
// No API key. No network. Run: npx tsx gate.ts
type JsonSchema = {
type: "object";
properties: Record<string, Prop>;
required: string[];
additionalProperties: false;
};
type Prop =
| { type: "string"; enum?: string[] }
| { type: "number"; minimum?: number; maximum?: number };
const ACTIONS = ["route", "escalate", "ignore"] as const;
type Action = (typeof ACTIONS)[number];
type Decision = {
action: Action;
target: string;
confidence: number;
};
const SCHEMA: JsonSchema = {
type: "object",
properties: {
action: { type: "string", enum: [...ACTIONS] },
target: { type: "string" },
confidence: { type: "number", minimum: 0, maximum: 1 },
},
required: ["action", "target", "confidence"],
additionalProperties: false,
};
Three fields. No extras. Enums closed. Numbers bounded.
That is the whole contract the rest of the code will enforce.
If the schema is fuzzy, the gate will be too.
Step 2: Validate Without a Framework
type Issue = { path: string; message: string };
function validate(raw: unknown, schema: JsonSchema): Issue[] {
const issues: Issue[] = [];
if (raw === null || typeof raw !== "object" || Array.isArray(raw)) {
return [{ path: "$", message: "expected an object" }];
}
const obj = raw as Record<string, unknown>;
for (const key of schema.required) {
if (!(key in obj)) issues.push({ path: key, message: "required field missing" });
}
for (const key of Object.keys(obj)) {
if (!(key in schema.properties)) {
issues.push({ path: key, message: "additional property not allowed" });
}
}
for (const [key, prop] of Object.entries(schema.properties)) {
if (!(key in obj)) continue;
const v = obj[key];
if (prop.type === "string") {
if (typeof v !== "string") {
issues.push({ path: key, message: `expected string, got ${typeOf(v)}` });
} else if (prop.enum && !prop.enum.includes(v)) {
issues.push({
path: key,
message: `expected one of [${prop.enum.join(", ")}], got ${JSON.stringify(v)}`,
});
}
} else if (prop.type === "number") {
if (typeof v !== "number" || Number.isNaN(v)) {
issues.push({ path: key, message: `expected number, got ${typeOf(v)}` });
} else {
if (prop.minimum !== undefined && v < prop.minimum) {
issues.push({ path: key, message: `below minimum ${prop.minimum}` });
}
if (prop.maximum !== undefined && v > prop.maximum) {
issues.push({ path: key, message: `above maximum ${prop.maximum}` });
}
}
}
}
return issues;
}
function typeOf(v: unknown): string {
if (v === null) return "null";
if (Array.isArray(v)) return "array";
return typeof v;
}
function parseJson(text: string): { ok: true; value: unknown } | { ok: false; error: string } {
try {
return { ok: true, value: JSON.parse(text) };
} catch (e) {
return { ok: false, error: e instanceof Error ? e.message : "parse failed" };
}
}
No Zod. No Ajv. Just the checks this schema needs.
You can swap in a real JSON Schema library later. The interesting part is not the library. It is what you do with the issues list.
Validation without a decision is just logging.
Step 3: Script the Failure Modes
type Fixture = {
id: string;
label: string;
first: string;
repair?: string;
};
const FIXTURES: Fixture[] = [
{
id: "valid",
label: "valid JSON",
first: '{"action":"route","target":"billing","confidence":0.91}',
},
{
id: "truncated",
label: "truncated JSON",
first: '{"action":"route","target":"billing","confidence":0.8',
repair: '{"action":"route","target":"billing","confidence":0.82}',
},
{
id: "wrong-type",
label: "wrong types",
first: '{"action":"route","target":"support","confidence":"high"}',
repair: '{"action":"route","target":"support","confidence":0.74}',
},
{
id: "bad-enum",
label: "invalid enum",
first: '{"action":"forward","target":"sales","confidence":0.6}',
repair: '{"action":"escalate","target":"sales","confidence":0.6}',
},
{
id: "extra-field",
label: "extra fields",
first: '{"action":"ignore","target":"spam","confidence":0.95,"reason":"promo"}',
repair: '{"action":"ignore","target":"spam","confidence":0.95}',
},
{
id: "missing",
label: "missing required",
first: '{"action":"route","confidence":0.7}',
repair: '{"action":"route","target":"onboarding","confidence":0.7}',
},
{
id: "out-of-range",
label: "confidence out of range",
first: '{"action":"escalate","target":"legal","confidence":1.4}',
repair: '{"action":"escalate","target":"legal","confidence":0.88}',
},
{
id: "stubborn",
label: "repair still invalid",
first: '{"action":"route","target":42,"confidence":0.5}',
repair: '{"action":"route","target":42,"confidence":0.5}',
},
];
function mockComplete(fixture: Fixture, pass: "first" | "repair", _errors?: Issue[]): string {
if (pass === "first") return fixture.first;
return fixture.repair ?? fixture.first;
}
These are the shapes that show up in logs even when the provider promises schema adherence: a truncated stream, a string where a number should be, an enum the model invented, an extra field someone forgot to ban, a missing required key, a confidence above 1, and a repair that repeats the same mistake.
The last one matters. A repair loop without a reject path is an infinite apology.
Step 4: One Repair, Then Decide
type Outcome =
| { status: "accept"; pass: "first" | "repair"; value: Decision }
| { status: "reject"; pass: "first" | "repair"; issues: Issue[] };
function gate(fixture: Fixture): Outcome {
const firstText = mockComplete(fixture, "first");
const firstParsed = parseJson(firstText);
if (!firstParsed.ok) {
const issues: Issue[] = [{ path: "$", message: `JSON parse error: ${firstParsed.error}` }];
return repairOnce(fixture, issues);
}
const firstIssues = validate(firstParsed.value, SCHEMA);
if (firstIssues.length === 0) {
return { status: "accept", pass: "first", value: firstParsed.value as Decision };
}
return repairOnce(fixture, firstIssues);
}
function repairOnce(fixture: Fixture, firstIssues: Issue[]): Outcome {
const repairText = mockComplete(fixture, "repair", firstIssues);
const repaired = parseJson(repairText);
if (!repaired.ok) {
return {
status: "reject",
pass: "repair",
issues: [{ path: "$", message: `JSON parse error: ${repaired.error}` }],
};
}
const repairIssues = validate(repaired.value, SCHEMA);
if (repairIssues.length === 0) {
return { status: "accept", pass: "repair", value: repaired.value as Decision };
}
return { status: "reject", pass: "repair", issues: repairIssues };
}
The sequence is short on purpose.
Parse. Validate. If anything is wrong, ask once with the issues. Validate again. Stop.
In a real stack, mockComplete becomes your provider call, and firstIssues becomes the error payload you send back. Anthropic's docs already list schema violations as a reason people retry. This is that retry, owned by your harness instead of hoped for in the prompt.
One repair is a strategy. Infinite repair is a hang.
Step 5: Metrics and a Ship Gate
type Row = {
id: string;
label: string;
status: "accept" | "reject";
pass: "first" | "repair";
detail: string;
};
const rows: Row[] = FIXTURES.map((f) => {
const out = gate(f);
if (out.status === "accept") {
return {
id: f.id,
label: f.label,
status: "accept",
pass: out.pass,
detail: JSON.stringify(out.value),
};
}
return {
id: f.id,
label: f.label,
status: "reject",
pass: out.pass,
detail: out.issues.map((i) => `${i.path}: ${i.message}`).join("; "),
};
});
const n = rows.length;
const passAt1 = rows.filter((r) => r.status === "accept" && r.pass === "first").length;
const passAtRepair = rows.filter((r) => r.status === "accept").length;
const rejected = rows.filter((r) => r.status === "reject").length;
console.log(`tiny-schema-gate fixtures=${n} (example data)\n`);
console.log("id label result pass detail");
console.log("-".repeat(96));
for (const r of rows) {
const id = r.id.padEnd(13);
const label = r.label.padEnd(25);
const result = r.status.padEnd(8);
const pass = r.pass.padEnd(8);
const detail = r.detail.length > 42 ? r.detail.slice(0, 39) + "..." : r.detail;
console.log(`${id} ${label} ${result} ${pass} ${detail}`);
}
console.log("\n== metrics ==");
console.log(`pass@1 ${(passAt1 / n).toFixed(2)} (${passAt1}/${n})`);
console.log(`pass@repair ${(passAtRepair / n).toFixed(2)} (${passAtRepair}/${n})`);
console.log(`reject rate ${(rejected / n).toFixed(2)} (${rejected}/${n})`);
const shipFloor = 0.75; // example threshold
const ship = passAtRepair / n >= shipFloor;
console.log(`\ngate: ${ship ? "SHIP" : "HOLD"} (pass@repair >= ${shipFloor} required)`);
Three numbers.
pass@1 is how often the first reply was already schema-valid.
pass@repair is how often you had a usable object after one constrained retry.
reject rate is how often you should refuse to act.
The ship floor here is an example: 0.75 after one repair. Pick yours from traffic, not from a blog.
Run it:
npx tsx gate.ts
You should see:
tiny-schema-gate fixtures=8 (example data)
id label result pass detail
------------------------------------------------------------------------------------------------
valid valid JSON accept first {"action":"route","target":"billing","c...
truncated truncated JSON accept repair {"action":"route","target":"billing","c...
wrong-type wrong types accept repair {"action":"route","target":"support","c...
bad-enum invalid enum accept repair {"action":"escalate","target":"sales","...
extra-field extra fields accept repair {"action":"ignore","target":"spam","con...
missing missing required accept repair {"action":"route","target":"onboarding"...
out-of-range confidence out of range accept repair {"action":"escalate","target":"legal","...
stubborn repair still invalid reject repair target: expected string, got number
== metrics ==
pass@1 0.13 (1/8)
pass@repair 0.88 (7/8)
reject rate 0.13 (1/8)
gate: SHIP (pass@repair >= 0.75 required)
Only one fixture passed on the first try.
Six more became valid after the repair.
One stayed wrong. The gate rejected it instead of inventing a target.
That is the whole point.
Accept schema-valid output. Reject the rest. Count both.
Where It Breaks Down
Provider Structured Outputs Are Not Your Gate
Constrained decoding lowers the odds of a bad shape. It does not own your retry budget, your logging, or your decision to stop. OpenAI's guide still lists refusals and incomplete responses. Anthropic still documents schema violations as something applications handle. Keep the provider feature. Keep the gate too.
A Tiny Validator Is Not a Full JSON Schema Engine
This checker covers objects, required keys, enums, number bounds, and additionalProperties: false. It does not do $ref, anyOf, arrays of objects, or regex patterns. Production schemas usually need a real library. The pattern stays the same: issues in, one repair, accept or reject.
Repair Can Hallucinate a "Valid" Lie
Feeding errors back can make the model invent a target that never existed in the ticket. Schema validity is not semantic truth. If the field is high risk, pair the gate with a allowlist or a human step.
Eight Fixtures Are a Smoke Test
pass@1 of 0.13 here is an artifact of how the fixtures were written. On a constrained provider path it should be much higher. Build your matrix from production failures: truncated streams, enum drift, type coercion, extra keys your parser used to ignore.
The Bigger Idea
Structured output is not a prompt trick.
It is a contract between a probabilistic sampler and a deterministic system.
Model reply
↓
Parse JSON
↓
Validate schema ──→ issues
↓
One repair (errors fed back)
↓
Validate again
↓
Accept, or reject and stop
The schema provides the shape.
The validator provides the evidence.
The repair provides one cheap second chance.
The reject path provides the stop.
The metrics provide the ship decision.
Your structured output passed once. Make it earn the next call.
Software should explain itself.
I'm building Helix around this idea: every change should come with the why. What changed, what got better, what got worse, and the evidence behind it.
What critical engineering knowledge is your team losing right now?


Top comments (0)