DEV Community

Geminate Solutions
Geminate Solutions

Posted on

Stop Trusting JSON From Your LLM: Zod Validation and Repair Prompts

Sooner or later your LLM will return JSON that parses but is wrong, or text that does not parse at all. The fix: check every response against a Zod schema, send exactly one targeted repair prompt when the check fails, and fall back to a safe path if the repair also fails. That way nothing malformed ever reaches your database.

This post walks through the pattern in TypeScript, from start to finish. The code works with any provider: you pass in a function that calls whatever model you use.

Why "JSON mode" is not enough

Structured output features help a lot, but you still need to validate. Common ways it goes wrong:

  • The JSON comes wrapped in a markdown fence, or has a friendly sentence in front of it.
  • A field exists but has the wrong type, like "quantity": "3" instead of a number.
  • The model invents an enum value: "priority": "critical" when your app only knows four priorities.
  • Required fields are missing because the response hit the token limit and got cut off.
  • You upgrade the model version and its formatting habits change overnight.

JSON.parse tells you the string is JSON. Only a schema tells you it is your JSON.

Step 1: Define the contract once

Say you use a model to triage incoming support tickets. The schema is the contract between the model and the rest of your system:

import { z } from 'zod';

export const TicketTriage = z.object({
  category: z.enum(['billing', 'bug', 'feature_request', 'account', 'other']),
  priority: z.enum(['low', 'medium', 'high', 'urgent']),
  summary: z.string().min(10).max(280),
  customerEmail: z.string().email().nullable(),
  tags: z.array(z.string().min(1)).max(5),
  needsHuman: z.boolean(),
});

export type TicketTriage = z.infer<typeof TicketTriage>;
Enter fullscreen mode Exit fullscreen mode

Two rules for writing these schemas:

  1. Be strict on anything your code branches on. Enums and booleans drive routing, so they get exact values.
  2. Use .nullable() for things the model may not know. If the ticket has no email, you want null, not a made-up address.

Step 2: Extract before you parse

Models love to add prose and fences. Strip them before parsing, and don't try to be cleverer than that:

export function extractJson(raw: string): unknown {
  const fenced = raw.match(/`{3}(?:json)?\s*([\s\S]*?)`{3}/i);
  const candidate = fenced ? fenced[1] : raw;
  const start = candidate.indexOf('{');
  const end = candidate.lastIndexOf('}');
  if (start === -1 || end <= start) {
    throw new Error('No JSON object found in model output');
  }
  return JSON.parse(candidate.slice(start, end + 1));
}
Enter fullscreen mode Exit fullscreen mode

If this throws, that is fine. The repair step handles it.

Step 3: Validate, repair once, then give up cleanly

First, a helper that turns any failure into a short, specific list of problems:

import { z } from 'zod';

type Message = { role: 'system' | 'user' | 'assistant'; content: string };
type CallModel = (messages: Message[]) => Promise<string>;

type Check<T> = { ok: true; data: T } | { ok: false; issues: string };

function check<T>(schema: z.ZodType<T>, raw: string): Check<T> {
  let parsed: unknown;
  try {
    parsed = extractJson(raw);
  } catch (err) {
    return { ok: false, issues: '- Output was not valid JSON: ' + (err as Error).message };
  }
  const result = schema.safeParse(parsed);
  if (result.success) return { ok: true, data: result.data };
  const issues = result.error.issues
    .map((i) => '- ' + (i.path.join('.') || '(root)') + ': ' + i.message)
    .join('\n');
  return { ok: false, issues };
}
Enter fullscreen mode Exit fullscreen mode

Then the main function:

type Result<T> =
  | { ok: true; data: T; repaired: boolean }
  | { ok: false; reason: string; raw: string };

export async function generateValidated<T>(opts: {
  schema: z.ZodType<T>;
  messages: Message[];
  callModel: CallModel;
}): Promise<Result<T>> {
  const { schema, messages, callModel } = opts;

  const first = await callModel(messages);
  const firstCheck = check(schema, first);
  if (firstCheck.ok) return { ok: true, data: firstCheck.data, repaired: false };

  const repairMessages: Message[] = [
    ...messages,
    { role: 'assistant', content: first },
    {
      role: 'user',
      content: [
        'Your previous reply failed validation:',
        firstCheck.issues,
        '',
        'Return the corrected JSON object only. No prose, no markdown fences.',
        'Keep every field that was already valid unchanged.',
      ].join('\n'),
    },
  ];

  const second = await callModel(repairMessages);
  const secondCheck = check(schema, second);
  if (secondCheck.ok) return { ok: true, data: secondCheck.data, repaired: true };

  return { ok: false, reason: secondCheck.issues, raw: second };
}
Enter fullscreen mode Exit fullscreen mode

Why the repair prompt is targeted

A generic "please return valid JSON" retry is a coin flip. Here the model gets the exact path and the exact complaint, for example priority: Invalid enum value. Expected 'low' | 'medium' | 'high' | 'urgent', received 'critical'. The exact wording depends on your Zod version. Either way, the model is asked for a small, concrete edit, and models handle those well.

Why only one repair

If the second attempt also fails, it is rarely bad luck. The cause is usually the prompt, an unusual input, or a schema that asks for something the model cannot know. More retries only add latency and token usage, and they hide the real problem. One repair fixes the formatting slips. Anything else should come to the surface.

Step 4: Make the fallback boring

The failure branch is where most teams cut corners. Decide it up front:

const result = await generateValidated({ schema: TicketTriage, messages, callModel });

if (result.ok) {
  await db.ticketTriage.insert({ ticketId, ...result.data, repaired: result.repaired });
} else {
  logger.warn({ ticketId, reason: result.reason }, 'triage validation failed');
  await db.ticketTriage.insert({
    ticketId,
    category: 'other',
    priority: 'medium',
    summary: 'Automatic triage unavailable. Needs manual review.',
    customerEmail: null,
    tags: [],
    needsHuman: true,
    repaired: false,
  });
  await reviewQueue.add({ ticketId, raw: result.raw });
}
Enter fullscreen mode Exit fullscreen mode

Use a simple rule to pick the fallback:

  • Safe default when a wrong guess is cheap and a person will see the record anyway, as with ticket triage.
  • Human review queue when the output touches payments, user-facing content, or anything with legal weight.
  • Fail the request with a clear error when a user is waiting and can retry.
  • Never write the raw model output to the database "to clean up later". Later never comes.

One more thing: run your fallback object through TicketTriage.parse() in a unit test. A fallback that breaks the schema is just a slower bug.

Production checklist

  • [ ] Log every repair with the prompt version. A rising repair rate is often the first sign that a model update changed behavior.
  • [ ] Store raw failed outputs with a retention limit. They can contain user data.
  • [ ] Put a timeout on each model call (for example 30 seconds), so a repair cannot double the wait on a hung request.
  • [ ] Unit test check() with fixtures: fenced output, truncated output, wrong enum, extra keys, empty string.
  • [ ] Version the schema and the prompt together. If a field changes, both change in the same commit.
  • [ ] Generate the provider's JSON schema from Zod (z.toJSONSchema in Zod 4, or the zod-to-json-schema package) so the two never drift apart.
  • [ ] Validate on the server, right before the write, not only in the client.

Trade-offs to keep in mind

  • Latency on failure. The repair adds a second round trip, but only when the first response fails. If repairs happen often, fix the prompt instead of tuning retries.
  • Strictness vs. acceptance. Strict schemas reject more responses. Stay strict on fields you branch on and loose on free text.
  • Context size. Sending the failed output back costs tokens. For large outputs, send the issues plus the original input and leave the failed output out.

Geminate Solutions treats this wrapper as the minimum for any model call that writes to a database. It is a small amount of code, and it turns "the model said something weird" from a data incident into a log line.

For the wider picture, including queues, streaming and provider choice, see the Node.js AI integration guide.

Top comments (1)

Collapse
 
vladzoff profile image
Vlad Zoff •

Zod is useful here, but I’d also keep the original model failure separate from the repaired result. Otherwise a repair loop can make an unreliable generation path look more deterministic than it actually is. Recording “invalid output → repair → accepted output” gives you a much better signal when these calls start driving database writes or other side effects.