DEV Community

Veaceslav
Veaceslav

Posted on

Stop prompting for valid JSON: an LLM output layer that does not break

I'm building JobCopilot: a small Python agent that reads a job description, compares it to my resume, and tells me what skills I'm missing. Like a skeptical recruiter in a terminal.

I want the agent's brain to be an LLM call. But my program needs an answer in a defined shape — like a form with three fields — not a wall of prose. And there's the problem.

The problem in one picture

I asked the model: return only raw JSON, no explanation.

It agreed politely. Then it sent:

Sure! Here's the breakdown in raw JSON:

{
  "verdict": "partial",
  "match_score": 68,
  "missing_skills": ["FastAPI", "Alemdic"]
}

See it? The JSON is wrapped in a hello + a code fence (the JSON-fence block). Python's json.loads() says there's no JSON there at all — it sees the words "Sure! Here's..." and dies. If I ask the model again, it does the same thing slightly differently. You can write "I really mean it this time!!!" in the prompt. It does not help. I tried for an hour.

And when the model does send clean JSON one night, there are three more ways it betrays you:

  1. Wrong key names — it sends skills_missing today and missing_skills tomorrow. Your code searched for one of them.
  2. Wrong types — the score arrives as 87.5, or even 105, when your program expects an integer up to 100.
  3. Correct format, wrong content — {"verdict": "no", "missing_skills": []}: shape is perfect, logic is nonsense.

"A prompt is a suggestion. Only code is a contract." So instead of fighting the prompt, I built a small validation layer — a wall between the model and my program. Whatever comes out of the model, my program either receives a valid form or a precise error report. Here it is, in four steps any beginner can follow.

Step 1: Unwrap the JSON (8 lines of Python)

The model keeps wrapping JSON in JSON-fence fences no matter what I write in the prompt. Fine — let the code unwrap it:

import json
import re

# match (.*?) between ```
{% endraw %}
 pairs, WITH newlines — that's DOTALL
_FENCE_RE = re.compile(r"
{% raw %}
```(?:json)?\s*(.*?)\s*```

", re.DOTALL)


def extract_json(raw: str) -> dict:
    """Takes the model's raw output. Returns clean JSON data."""
    try:
        return json.loads(raw)      # best case: it was already clean JSON
    except json.JSONDecodeError:
        pass                        # no? look for the fence below...

    match = _FENCE_RE.search(raw)
    if match is None:
        raise ValueError(f"no JSON found in: {raw[:200]}")

    loaded = json.loads(match.group(1))
    if not isinstance(loaded, dict):
        raise ValueError(f"expected a JSON object, got {type(loaded).__name__}")
    return loaded


Enter fullscreen mode Exit fullscreen mode

This function tries the honest path first (json.loads(raw)), and if the text has chatty greetings, it grabs whatever sits between the triple-backtick marks and parses THAT.

Three small details that each cost me an evening of confusion:

  • re.DOTALL — without it, the dot in my regex stops at the first line break, and the fenced JSON spans many lines. My first version matched nothing, silently.
  • (?:json)? — the model writes the language hint or a bare triple-backtick fence. This part makes the hint optional, so both work.
  • search, not fullmatch — because there's prose around the JSON: "Sure!", intro sentences, sometimes a closing joke.

One bonus detail I did not know: in Python, JSONDecodeError is a subclass of ValueError. That means except ValueError catches both — I'll use this in Step 3.

Step 2: Check the form field by field (pydantic)

The JSON parses now. But parsing is not validating: {"match_score": 105} parses perfectly fine. So the next layer checks the shape of the data field by field:


python
from typing import Literal
from pydantic import BaseModel, Field, field_validator, model_validator


class JobMatch(BaseModel):
    verdict: Literal["strong", "partial", "no"]   # ONLY these three values
    match_score: int = Field(ge=0, le=100)         # int, 0..100, nothing else
    missing_skills: list[str]                      # a list of strings

    @field_validator("missing_skills")
    @classmethod
    def clean_skills(cls, skills):
        # model sent ["React", "react", "  CSS ", ""] once. Clean it:
        seen = set()
        out = []
        for s in skills:
            key = s.strip().lower()
            if not key or key in seen:
                continue
            seen.add(key)
            out.append(s.strip())
        return out

    @model_validator(mode="after")
    def skeptical_check(self):
        # "no" verdict with empty list = the model gave up
        if self.verdict == "no" and not self.missing_skills:
            raise ValueError("verdict 'no' but no missing skills — model gave up")
        return self


Enter fullscreen mode Exit fullscreen mode

Read this as a form with three fields, each field guarded:

  • The verdict must be exactly "strong", "partial" or "no" — the model once answered "maybe"; with Literal that becomes impossible.
  • The score must be an integer between 0 and 100. The model sent 105 once. Now that is impossible too.
  • The skills list gets cleaned: duplicated "React" and "react" collapse into one, empty strings are dropped.
  • The final check catches nonsense logic: if the verdict says "no match" but lists zero missing skills, the model gave up instead of analyzing.

When a value breaks a rule, pydantic produces an error that is specific:


text
match_score: Input should be a valid integer, got a number with a fractional
             part [type=int_from_float, input_value=87.5]
verdict:     Input should be 'strong', 'partial' or 'no'
             [type=literal_error, input_value='maybe']


Enter fullscreen mode Exit fullscreen mode

Notice we threw out about thirty lines of handwritten if "verdict" not in data: ... code. The declarative model does the same job, and the errors are not "value error" mush — they name the field, the problem and the offending value.

One honest limit: this layer guarantees the shape, not the truth. It will not catch "Alemdic" — a misspelling of Alembic that is perfectly well-formed JSON. Catching wrong-but-valid content is the job of an eval set, and that is the next article in this series.

Step 3: Return an error as a result, not a crash

Now the top layer. My rule: the model can fail, and the program must handle that failure like any other value. If a function raises an exception, the caller has to catch it — and if the caller forgets, the whole program dies. If a function returns the error, the caller just decides what to do.

If you write TypeScript, you know zod.safeParse() — it returns a result object with a status. We can build the same in Python with two dataclasses:


python
from dataclasses import dataclass


@dataclass
class ParseOk:
    match: JobMatch                       # valid data


@dataclass
class ParseError:
    stage: str      # "json" — couldn't extract, or "schema" — invalid
    message: str    # exact reason, for me AND for the next LLM call


def parse_match(raw: str) -> ParseOk | ParseError:
    try:
        data = extract_json(raw)          # step 1
    except ValueError as exc:             # JSONDecodeError IS a ValueError
        return ParseError(stage="json", message=str(exc))

    try:
        return ParseOk(match=JobMatch.model_validate(data))   # step 2
    except ValidationError as exc:
        return ParseError(stage="schema", message=str(exc))


Enter fullscreen mode Exit fullscreen mode

The type ParseOk | ParseError works like a union type from TypeScript: every result is definitely one or the other, and the caller checks with isinstance(result, ParseOk).

Step 4: The retry — give the exact error back to the model

There is one more trick with ParseError.message. The model can repair itself if you tell it exactly what failed. "It failed" teaches nothing. But:

match_score: Input should be a valid integer, got a number with a fractional part [type=int_from_float, input_value=87.5]

...the model fixes almost every time. So the retry loop feeds the exact pydantic error back into the next call:


python
FIX_PROMPT = """Below was your previous output and the validation error.
Return corrected JSON only.

Previous output:
{previous}

Validation error:
{error}"""


def structured_match(chat, jd_text: str, max_retries: int = 2):
    prompt = MATCH_PROMPT.format(job=jd_text)

    result = None
    for attempt in range(1 + max_retries):
        raw = chat(prompt)
        result = parse_match(raw)
        if isinstance(result, ParseOk):
            return result
        prompt = FIX_PROMPT.format(previous=raw, error=result.message)

    return result   # after max retries: ParseError as data — never None


Enter fullscreen mode Exit fullscreen mode

Retry at most twice, then give up — but give up as a result, not a crash. The caller decides: log it, message the user, or end the run.

(A quick note about frameworks: many LLM providers have native JSON mode. If yours does, use it — it kills most fence problems at the root. But keep the pydantic layer anyway: native mode guarantees shape, not logic, and match_score <= 100 survives on every provider.)

The whole picture

Four small pieces, each doing one job:

  1. Unwrap the code fence — extract_json
  2. Validate field by field — the JobMatch pydantic model
  3. Convert errors into typed data — ParseOk / ParseError
  4. Retry with the exact error fed back — structured_match

In one sentence: code, not the prompt, decides when the model's output is good.

Two things I took from building it:

  • A prompt is a suggestion; a pydantic model is a contract.
  • "Errors as data" sounds fancy but is 15 lines: two dataclasses and an isinstance.

Next Tuesday, part 2: retrieval — embeddings with fastembed, cosine similarity computed by hand with numpy, and why my top-8 retrieval stopped discriminating on a resume of 20 chunks.

Follow if you want to see the rest of the copilot get built — including the parts that will break. Build log + daily code: t.me/novamind_hub.

How do you harden LLM JSON in your projects — native JSON mode, validators, or retries? I'm still collecting failure modes.

Top comments (0)