DEV Community

sharpenlee
sharpenlee

Posted on Originally published at json-tool.com

Why LLMs return invalid JSON — a language model has no parser

Ask a model for JSON and you often get something that is almost JSON: a value wrapped in a markdown fence, a trailing comma, a Python None where null belongs. The reflex is to blame the prompt, or to switch on "JSON mode". Both miss the mechanism. A language model predicts the next token; it never builds a document, and it cannot tell you that a document is broken, because nothing in it represents a document. The useful question is therefore not "how do I make it stop" but "where does the parser go, given there is not one upstream".

What comes back

The twelve source strings below are the shapes that actually break things — eight of them are repairs that the jsonrepair library documents by name. I ran all twelve through four parsers and recorded accept-or-reject. They were typed as literals rather than captured from a live call: there is no API key on the machine this was written on, so the shapes are the documented ones and what is measured here is how the parsers respond to them.

Source Node Python Ruby PHP
markdown fence ERR ERR ERR ERR
trailing comma {"a":1,} ERR ERR ERR ERR
single quotes {'a':1} ERR ERR ERR ERR
Python literals None / True ERR ERR ERR ERR
block comment /* … */ ERR ERR OK ERR
raw newline inside a string ERR ERR ERR ERR
truncated {"a":1,"b":[1,2 ERR ERR ERR ERR
NaN ERR OK ERR ERR
Infinity ERR OK ERR ERR
unquoted key {a:1} ERR ERR ERR ERR
trailing prose after the value ERR ERR ERR ERR
UTF-8 BOM before the object ERR ERR ERR ERR

Four parsers, twelve documents, and they agree on almost all of it. Nine of the twelve are rejected by every parser in the table, which is the reassuring part: for a fence, a trailing comma, a stray quote or a truncation, none of the mainstream parsers you are likely to be running will accept the text as JSON. Three rows break the unanimity, in two ways that matter more than they look.

Ruby accepts block comments

Ruby 2.6.10 — the build measured here, with its bundled json 2.1.0 — parses {"a":1 /* note */} and hands back {"a"=>1}. Every other parser in the table refuses it. So the exact bytes that are a hard 400 from a Node or PHP service are a clean 200 from a Ruby one. This is obscure, and that is the point: a strict parse is not a shared contract between two services in the same pipeline unless they happen to run the same parser.

Python accepts NaN and Infinity

JSON's number grammar has no NaN and no Infinity — RFC 8259 builds every number from an optional minus, digits, an optional fraction and an optional exponent, and those tokens are not in it. Node, Ruby and PHP reject {"x": NaN}. Python's json.loads returns {'x': nan}. Hand that value straight back to json.dumps and, with the default settings, you get {"x": NaN} — a document that is not valid JSON and that the other three parsers reject on sight; only allow_nan=False turns the silent round trip into an exception. This is the same divergence the companion article on number precision found from the write side; here it arrives from the read side, in data a model just produced.

There is no guarantee upstream, by construction

The reason models get this wrong is not that they are bad at JSON. It is that the mechanism cannot make the promise. Sampling emits one token at a time from a probability distribution; nothing in the loop represents "the text so far is a complete, well-formed document", so nothing in the loop can refuse to finish it. A constrained decoder does better — it can track whether the text so far is a valid prefix — but a valid prefix is not a valid document, and that gap is the whole story.

"JSON mode" is the vendor term for pushing that guarantee as far upstream as it will go, and the documentation is unusually direct about what it actually buys you. Microsoft's Azure OpenAI page states that JSON mode "guarantees valid JSON output, but it doesn't guarantee the output matches a specific schema". The braces will close; the fields inside them can still be wrong. The same page shows how thin even that guarantee is at the edges — it also requires the word "JSON" to appear somewhere in the messages, and when it does not, the API rejects the request outright:

Error code: 400 - {'error': {'message': "'messages' must contain the word
'json' in some form, to use 'response_format' of type 'json_object'."}}
Enter fullscreen mode Exit fullscreen mode

That error is the honest shape of the whole feature. When the instruction is missing, the documentation warns (quoting OpenAI) that the model can "generate an unending stream of whitespace and the request could run continually until it reaches the token limit" — a failure that looks, from your side, exactly like a hang.

The hole the guarantee does not cover

Validity per token is not validity per document, and the place they diverge is truncation. If the reply is longer than max_tokens, the stream simply stops mid-structure and finish_reason comes back as length. Azure's guidance is explicit on both counts: "You should check finish_reason for the value length before parsing the response. The model might generate partial JSON", and, in troubleshooting, "Don't parse partial JSON."

This is the case the guarantee cannot reach, and the case a repair layer is most likely to paper over. Take the truncated object from the table and let a repair pass close its brackets:

model (finish_reason: "length")
    {"a":1,"b":[1,2

after the repair pass closes the brackets
    {"a":1,"b":[1,2]}
Enter fullscreen mode Exit fullscreen mode

The result parses cleanly, and it is a lie. Nothing in the text says whether the list was going to end at two items or fifty; the repair chose two. Repair restores syntax, never data.

The repair layer, measured

Since the parser upstream does not exist, one has to exist downstream, and in practice it is a repair pass that runs before JSON.parse. I wrote the smallest honest version of that pass and ran it over the same twelve documents. A first stage of plain regex substitutions — strip the fence and the byte-order mark, drop comments and trailing commas, swap the single quotes, map the Python and JavaScript non-values, close unbalanced brackets — recovered nine of the twelve. The three it missed were simply rules that first pass did not contain, and two of them are cases a regex cannot handle safely on its own: a newline must be escaped only when it is inside a string, and trailing prose must be cut only once the value has ended. Adding a boundary-tracking pass, and the key-quoting rule, took it to twelve out of twelve:

Source pattern After repair What the repair did
markdown fence parses stripped the fence
trailing comma parses removed it
single quotes parses rewrote as double quotes
None / True parses mapped to null / true
block comment parses removed the comment
raw newline in a string parses escaped the control character
unquoted key parses added the quotes
trailing prose parses discarded everything after the value
UTF-8 BOM parses stripped the mark
NaN / Infinity parses substituted null — a value the model never wrote
truncated parses closed the brackets — and invented the data above

All twelve parse afterwards. The truncated one parses by being falsified. And some of those rows are decisions rather than restorations: the comment row silently dropped content the model took the trouble to write, and the NaN row is worse — the substituted null is indistinguishable from a real one, so a repair pass can convert "the model produced an undefined number" into "the model said there is nothing here" without leaving a trace.

None of this is improvisation on my part. The jsonrepair library publishes a list of what it fixes, and eight of the twelve shapes above are on it by name: add missing quotes around keys, add missing escape characters, repair truncated JSON, replace single quotes with double quotes, replace Python constants, strip trailing commas, strip comments and strip fenced code blocks. Three of the shapes it does not cover are the ones that matter most — NaN, Infinity and prose trailing after the value. Those are exactly the cases that survive a repair pass silently: NaN parses in Python and nowhere else, so a pipeline can carry it through the repair stage and lose it at whichever service happens to disagree. That a dedicated repair library and a model's failure modes overlap this heavily is the clearest evidence that the problem is structural rather than incidental.

What to do instead

  1. Move the guarantee upstream when you can. Prefer a schema-constrained mode over asking politely in the prompt: the documented difference is that JSON mode promises syntax while structured outputs constrain the model to a schema. Neither was exercised here (no API key), so treat the exact semantics as documented rather than measured.
  2. Treat the reply as untrusted text. The pipeline is repair → parse → validate against a schema, in that order, and the last step is not optional, because a syntactically perfect document with the wrong fields passes the first two.
  3. Check finish_reason before you parse, and never repair a truncated reply into acceptance. Raise and retry with a larger budget instead. Repair makes a cut-off document look finished; only you can decide that it is not finished.
  4. Do not assume "valid JSON" is a shared contract. Two services in one pipeline can disagree about the same bytes — Ruby takes the block comment, Python takes the NaN, and the third service rejects both. Validate at the boundary rather than trusting that the upstream parser agreed with yours.
  5. Look at the output before it reaches code. A repair layer hides exactly the cases you most need to see.

Check your own payload

Paste a model reply into the validator to get the line and column of the first real error rather than a repair that is invisibly wrong, and into the formatter to see the structure once it parses. For output that is valid JSON but the wrong shape, the JSON Schema validator is the missing test: "parses" and "correct" are different questions, and JSON mode only answers the first. The syntax cheat sheet covers what the grammar actually allows, including why NaN and Infinity are not in it.

Provenance

The parser table and the repair table were produced by running the documents shown through Node 24.18.0 on V8 15.0.245.15, CPython 3.10.11, Ruby 2.6.10 (bundled json 2.1.0) and PHP 8.3.35 on a 64-bit build. The Node runtime here is the one embedded in an Electron host rather than a stock Node distribution — its V8 build string carries an -electron suffix — and the parser exercised in either case is V8's own JSON.parse. The repair pass was written and executed for this article rather than taken from a library, and the Python NaN round trip was measured with the default json.dumps and again with allow_nan=False. The twelve source strings are the documented failure shapes rather than captures from a live call: there is no API key on the machine this was written on, so nothing here was produced by a model. The claims about JSON mode, structured outputs and finish_reason are quoted from Microsoft's Azure OpenAI documentation, which mirrors OpenAI's, and are labelled documented rather than measured for exactly that reason; the jsonrepair list is quoted from that library's repository README, which enumerates its repairs without naming models as their source.

Related

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.