DEV Community

Cover image for Your Structured Output Passed Once. Build a Tiny Schema Gate in TypeScript.
Bobby Hall Jr
Bobby Hall Jr

Posted on

Your Structured Output Passed Once. Build a Tiny Schema Gate in TypeScript.

Your model returned JSON.

It looked fine.

Then production asked for a number and got "high".

Structured outputs are having a moment.

  • Anthropic's structured outputs docs (checked October 11, 2026) describe JSON outputs via output_config.format and strict tool use. The page lists the failure modes people already know: parsing errors, missing fields, inconsistent types, schema violations that force retries.
  • OpenAI's structured outputs guide (checked October 11, 2026) frames the feature as schema adherence, not just "valid JSON." It still calls out refusals and max-token cutoffs as cases where the response may not match your schema.
  • On October 7, 2026, Anthropic introduced Claude Haiku 5.5 for high-volume, cost-sensitive work: summaries, classification, routing, and similar short tasks. When a small model answers that often, a bad shape is not a one-off. It is a rate.

Different products. Same direction.

The provider can constrain the sample.

Your system still has to decide what to do when the shape is wrong, truncated, or refused.

A schema that passed once is a demo. A gate that rejects bad shapes is a product.

So let's build a tiny version.

By the end, you'll run one command:

npx tsx gate.ts
Enter fullscreen mode Exit fullscreen mode

And watch eight scripted replies hit a validator, take one repair, then accept or reject.

No API key.

No real model.

Just TypeScript.

One honesty note: the "model" here is a fixture table. The confidence numbers and pass rates are example inputs. Swap in your provider's complete call and keep the gate.

Table of Contents

  1. What We Are Building
  2. Project Setup
  3. Step 1: Define the Schema
  4. Step 2: Validate Without a Framework
  5. Step 3: Script the Failure Modes
  6. Step 4: One Repair, Then Decide
  7. Step 5: Metrics and a Ship Gate
  8. Where It Breaks Down
  9. The Bigger Idea

Code: github.com/bobbyhalljr/tiny-schema-gate

What We Are Building

Validate, repair once, accept or reject

A routing decision with three fields.

action is an enum: route, escalate, or ignore.

target is a string.

confidence is a number between 0 and 1.

The mock model returns valid JSON, truncated JSON, wrong types, a bad enum, extra fields, a missing field, an out-of-range confidence, and one stubborn pair that never repairs.

The gate does three things:

  1. Parse.
  2. Validate against the schema.
  3. If invalid, feed the errors back once. Then accept or reject.

It's also a small version of the question behind Helix: when the system changes, can you show what got better, what got worse, and why?

Project Setup

You will need Node.js 18 or newer.

mkdir tiny-schema-gate
cd tiny-schema-gate

npm init -y
npm install --save-dev typescript tsx @types/node
Enter fullscreen mode Exit fullscreen mode

Save the following blocks, in order, as gate.ts.

Step 1: Define the Schema

// gate.ts: a tiny schema gate for structured LLM output.
// Everything is mocked: the "model" returns scripted strings, not provider calls.
// The fixtures, confidence numbers, and pass rates below are EXAMPLE INPUTS.
// No API key. No network. Run: npx tsx gate.ts

type JsonSchema = {
  type: "object";
  properties: Record<string, Prop>;
  required: string[];
  additionalProperties: false;
};

type Prop =
  | { type: "string"; enum?: string[] }
  | { type: "number"; minimum?: number; maximum?: number };

const ACTIONS = ["route", "escalate", "ignore"] as const;
type Action = (typeof ACTIONS)[number];

type Decision = {
  action: Action;
  target: string;
  confidence: number;
};

const SCHEMA: JsonSchema = {
  type: "object",
  properties: {
    action: { type: "string", enum: [...ACTIONS] },
    target: { type: "string" },
    confidence: { type: "number", minimum: 0, maximum: 1 },
  },
  required: ["action", "target", "confidence"],
  additionalProperties: false,
};
Enter fullscreen mode Exit fullscreen mode

Three fields. No extras. Enums closed. Numbers bounded.

That is the whole contract the rest of the code will enforce.

If the schema is fuzzy, the gate will be too.

Step 2: Validate Without a Framework

type Issue = { path: string; message: string };

function validate(raw: unknown, schema: JsonSchema): Issue[] {
  const issues: Issue[] = [];
  if (raw === null || typeof raw !== "object" || Array.isArray(raw)) {
    return [{ path: "$", message: "expected an object" }];
  }
  const obj = raw as Record<string, unknown>;

  for (const key of schema.required) {
    if (!(key in obj)) issues.push({ path: key, message: "required field missing" });
  }

  for (const key of Object.keys(obj)) {
    if (!(key in schema.properties)) {
      issues.push({ path: key, message: "additional property not allowed" });
    }
  }

  for (const [key, prop] of Object.entries(schema.properties)) {
    if (!(key in obj)) continue;
    const v = obj[key];
    if (prop.type === "string") {
      if (typeof v !== "string") {
        issues.push({ path: key, message: `expected string, got ${typeOf(v)}` });
      } else if (prop.enum && !prop.enum.includes(v)) {
        issues.push({
          path: key,
          message: `expected one of [${prop.enum.join(", ")}], got ${JSON.stringify(v)}`,
        });
      }
    } else if (prop.type === "number") {
      if (typeof v !== "number" || Number.isNaN(v)) {
        issues.push({ path: key, message: `expected number, got ${typeOf(v)}` });
      } else {
        if (prop.minimum !== undefined && v < prop.minimum) {
          issues.push({ path: key, message: `below minimum ${prop.minimum}` });
        }
        if (prop.maximum !== undefined && v > prop.maximum) {
          issues.push({ path: key, message: `above maximum ${prop.maximum}` });
        }
      }
    }
  }
  return issues;
}

function typeOf(v: unknown): string {
  if (v === null) return "null";
  if (Array.isArray(v)) return "array";
  return typeof v;
}

function parseJson(text: string): { ok: true; value: unknown } | { ok: false; error: string } {
  try {
    return { ok: true, value: JSON.parse(text) };
  } catch (e) {
    return { ok: false, error: e instanceof Error ? e.message : "parse failed" };
  }
}
Enter fullscreen mode Exit fullscreen mode

No Zod. No Ajv. Just the checks this schema needs.

You can swap in a real JSON Schema library later. The interesting part is not the library. It is what you do with the issues list.

Validation without a decision is just logging.

Step 3: Script the Failure Modes

type Fixture = {
  id: string;
  label: string;
  first: string;
  repair?: string;
};

const FIXTURES: Fixture[] = [
  {
    id: "valid",
    label: "valid JSON",
    first: '{"action":"route","target":"billing","confidence":0.91}',
  },
  {
    id: "truncated",
    label: "truncated JSON",
    first: '{"action":"route","target":"billing","confidence":0.8',
    repair: '{"action":"route","target":"billing","confidence":0.82}',
  },
  {
    id: "wrong-type",
    label: "wrong types",
    first: '{"action":"route","target":"support","confidence":"high"}',
    repair: '{"action":"route","target":"support","confidence":0.74}',
  },
  {
    id: "bad-enum",
    label: "invalid enum",
    first: '{"action":"forward","target":"sales","confidence":0.6}',
    repair: '{"action":"escalate","target":"sales","confidence":0.6}',
  },
  {
    id: "extra-field",
    label: "extra fields",
    first: '{"action":"ignore","target":"spam","confidence":0.95,"reason":"promo"}',
    repair: '{"action":"ignore","target":"spam","confidence":0.95}',
  },
  {
    id: "missing",
    label: "missing required",
    first: '{"action":"route","confidence":0.7}',
    repair: '{"action":"route","target":"onboarding","confidence":0.7}',
  },
  {
    id: "out-of-range",
    label: "confidence out of range",
    first: '{"action":"escalate","target":"legal","confidence":1.4}',
    repair: '{"action":"escalate","target":"legal","confidence":0.88}',
  },
  {
    id: "stubborn",
    label: "repair still invalid",
    first: '{"action":"route","target":42,"confidence":0.5}',
    repair: '{"action":"route","target":42,"confidence":0.5}',
  },
];

function mockComplete(fixture: Fixture, pass: "first" | "repair", _errors?: Issue[]): string {
  if (pass === "first") return fixture.first;
  return fixture.repair ?? fixture.first;
}
Enter fullscreen mode Exit fullscreen mode

These are the shapes that show up in logs even when the provider promises schema adherence: a truncated stream, a string where a number should be, an enum the model invented, an extra field someone forgot to ban, a missing required key, a confidence above 1, and a repair that repeats the same mistake.

The last one matters. A repair loop without a reject path is an infinite apology.

Step 4: One Repair, Then Decide

type Outcome =
  | { status: "accept"; pass: "first" | "repair"; value: Decision }
  | { status: "reject"; pass: "first" | "repair"; issues: Issue[] };

function gate(fixture: Fixture): Outcome {
  const firstText = mockComplete(fixture, "first");
  const firstParsed = parseJson(firstText);
  if (!firstParsed.ok) {
    const issues: Issue[] = [{ path: "$", message: `JSON parse error: ${firstParsed.error}` }];
    return repairOnce(fixture, issues);
  }
  const firstIssues = validate(firstParsed.value, SCHEMA);
  if (firstIssues.length === 0) {
    return { status: "accept", pass: "first", value: firstParsed.value as Decision };
  }
  return repairOnce(fixture, firstIssues);
}

function repairOnce(fixture: Fixture, firstIssues: Issue[]): Outcome {
  const repairText = mockComplete(fixture, "repair", firstIssues);
  const repaired = parseJson(repairText);
  if (!repaired.ok) {
    return {
      status: "reject",
      pass: "repair",
      issues: [{ path: "$", message: `JSON parse error: ${repaired.error}` }],
    };
  }
  const repairIssues = validate(repaired.value, SCHEMA);
  if (repairIssues.length === 0) {
    return { status: "accept", pass: "repair", value: repaired.value as Decision };
  }
  return { status: "reject", pass: "repair", issues: repairIssues };
}
Enter fullscreen mode Exit fullscreen mode

The sequence is short on purpose.

Parse. Validate. If anything is wrong, ask once with the issues. Validate again. Stop.

In a real stack, mockComplete becomes your provider call, and firstIssues becomes the error payload you send back. Anthropic's docs already list schema violations as a reason people retry. This is that retry, owned by your harness instead of hoped for in the prompt.

One repair is a strategy. Infinite repair is a hang.

Step 5: Metrics and a Ship Gate

type Row = {
  id: string;
  label: string;
  status: "accept" | "reject";
  pass: "first" | "repair";
  detail: string;
};

const rows: Row[] = FIXTURES.map((f) => {
  const out = gate(f);
  if (out.status === "accept") {
    return {
      id: f.id,
      label: f.label,
      status: "accept",
      pass: out.pass,
      detail: JSON.stringify(out.value),
    };
  }
  return {
    id: f.id,
    label: f.label,
    status: "reject",
    pass: out.pass,
    detail: out.issues.map((i) => `${i.path}: ${i.message}`).join("; "),
  };
});

const n = rows.length;
const passAt1 = rows.filter((r) => r.status === "accept" && r.pass === "first").length;
const passAtRepair = rows.filter((r) => r.status === "accept").length;
const rejected = rows.filter((r) => r.status === "reject").length;

console.log(`tiny-schema-gate  fixtures=${n} (example data)\n`);
console.log("id            label                      result   pass     detail");
console.log("-".repeat(96));
for (const r of rows) {
  const id = r.id.padEnd(13);
  const label = r.label.padEnd(25);
  const result = r.status.padEnd(8);
  const pass = r.pass.padEnd(8);
  const detail = r.detail.length > 42 ? r.detail.slice(0, 39) + "..." : r.detail;
  console.log(`${id} ${label} ${result} ${pass} ${detail}`);
}

console.log("\n== metrics ==");
console.log(`pass@1       ${(passAt1 / n).toFixed(2)}  (${passAt1}/${n})`);
console.log(`pass@repair  ${(passAtRepair / n).toFixed(2)}  (${passAtRepair}/${n})`);
console.log(`reject rate  ${(rejected / n).toFixed(2)}  (${rejected}/${n})`);

const shipFloor = 0.75; // example threshold
const ship = passAtRepair / n >= shipFloor;
console.log(`\ngate: ${ship ? "SHIP" : "HOLD"}  (pass@repair >= ${shipFloor} required)`);
Enter fullscreen mode Exit fullscreen mode

Three numbers.

pass@1 is how often the first reply was already schema-valid.

pass@repair is how often you had a usable object after one constrained retry.

reject rate is how often you should refuse to act.

The ship floor here is an example: 0.75 after one repair. Pick yours from traffic, not from a blog.

Run it:

npx tsx gate.ts
Enter fullscreen mode Exit fullscreen mode

Pass rates and ship gate from the CLI

You should see:

tiny-schema-gate  fixtures=8 (example data)

id            label                      result   pass     detail
------------------------------------------------------------------------------------------------
valid         valid JSON                accept   first    {"action":"route","target":"billing","c...
truncated     truncated JSON            accept   repair   {"action":"route","target":"billing","c...
wrong-type    wrong types               accept   repair   {"action":"route","target":"support","c...
bad-enum      invalid enum              accept   repair   {"action":"escalate","target":"sales","...
extra-field   extra fields              accept   repair   {"action":"ignore","target":"spam","con...
missing       missing required          accept   repair   {"action":"route","target":"onboarding"...
out-of-range  confidence out of range   accept   repair   {"action":"escalate","target":"legal","...
stubborn      repair still invalid      reject   repair   target: expected string, got number

== metrics ==
pass@1       0.13  (1/8)
pass@repair  0.88  (7/8)
reject rate  0.13  (1/8)

gate: SHIP  (pass@repair >= 0.75 required)
Enter fullscreen mode Exit fullscreen mode

Only one fixture passed on the first try.

Six more became valid after the repair.

One stayed wrong. The gate rejected it instead of inventing a target.

That is the whole point.

Accept schema-valid output. Reject the rest. Count both.

Where It Breaks Down

Provider Structured Outputs Are Not Your Gate

Constrained decoding lowers the odds of a bad shape. It does not own your retry budget, your logging, or your decision to stop. OpenAI's guide still lists refusals and incomplete responses. Anthropic still documents schema violations as something applications handle. Keep the provider feature. Keep the gate too.

A Tiny Validator Is Not a Full JSON Schema Engine

This checker covers objects, required keys, enums, number bounds, and additionalProperties: false. It does not do $ref, anyOf, arrays of objects, or regex patterns. Production schemas usually need a real library. The pattern stays the same: issues in, one repair, accept or reject.

Repair Can Hallucinate a "Valid" Lie

Feeding errors back can make the model invent a target that never existed in the ticket. Schema validity is not semantic truth. If the field is high risk, pair the gate with a allowlist or a human step.

Eight Fixtures Are a Smoke Test

pass@1 of 0.13 here is an artifact of how the fixtures were written. On a constrained provider path it should be much higher. Build your matrix from production failures: truncated streams, enum drift, type coercion, extra keys your parser used to ignore.

The Bigger Idea

Structured output is not a prompt trick.

It is a contract between a probabilistic sampler and a deterministic system.

Model reply
  ↓
Parse JSON
  ↓
Validate schema ──→ issues
  ↓
One repair (errors fed back)
  ↓
Validate again
  ↓
Accept, or reject and stop
Enter fullscreen mode Exit fullscreen mode

The schema provides the shape.

The validator provides the evidence.

The repair provides one cheap second chance.

The reject path provides the stop.

The metrics provide the ship decision.

Your structured output passed once. Make it earn the next call.


Software should explain itself.

I'm building Helix around this idea: every change should come with the why. What changed, what got better, what got worse, and the evidence behind it.

What critical engineering knowledge is your team losing right now?

See what Helix reveals →

Top comments (0)