Ask a model for JSON and you often get something that is almost JSON: a value wrapped in a markdown fence, a trailing comma, a Python None where null belongs. The reflex is to blame the prompt, or to switch on "JSON mode". Both miss the mechanism. A language model predicts the next token; it never builds a document, and it cannot tell you that a document is broken, because nothing in it represents a document. The useful question is therefore not "how do I make it stop" but "where does the parser go, given there is not one upstream".
What comes back
The twelve source strings below are the shapes that actually break things — eight of them are repairs that the jsonrepair library documents by name. I ran all twelve through four parsers and recorded accept-or-reject. They were typed as literals rather than captured from a live call: there is no API key on the machine this was written on, so the shapes are the documented ones and what is measured here is how the parsers respond to them.
| Source | Node | Python | Ruby | PHP |
|---|---|---|---|---|
| markdown fence | ERR | ERR | ERR | ERR |
trailing comma {"a":1,}
|
ERR | ERR | ERR | ERR |
single quotes {'a':1}
|
ERR | ERR | ERR | ERR |
Python literals None / True
|
ERR | ERR | ERR | ERR |
block comment /* … */
|
ERR | ERR | OK | ERR |
| raw newline inside a string | ERR | ERR | ERR | ERR |
truncated {"a":1,"b":[1,2
|
ERR | ERR | ERR | ERR |
NaN |
ERR | OK | ERR | ERR |
Infinity |
ERR | OK | ERR | ERR |
unquoted key {a:1}
|
ERR | ERR | ERR | ERR |
| trailing prose after the value | ERR | ERR | ERR | ERR |
| UTF-8 BOM before the object | ERR | ERR | ERR | ERR |
Four parsers, twelve documents, and they agree on almost all of it. Nine of the twelve are rejected by every parser in the table, which is the reassuring part: for a fence, a trailing comma, a stray quote or a truncation, none of the mainstream parsers you are likely to be running will accept the text as JSON. Three rows break the unanimity, in two ways that matter more than they look.
Ruby accepts block comments
Ruby 2.6.10 — the build measured here, with its bundled json 2.1.0 — parses {"a":1 /* note */} and hands back {"a"=>1}. Every other parser in the table refuses it. So the exact bytes that are a hard 400 from a Node or PHP service are a clean 200 from a Ruby one. This is obscure, and that is the point: a strict parse is not a shared contract between two services in the same pipeline unless they happen to run the same parser.
Python accepts NaN and Infinity
JSON's number grammar has no NaN and no Infinity — RFC 8259 builds every number from an optional minus, digits, an optional fraction and an optional exponent, and those tokens are not in it. Node, Ruby and PHP reject {"x": NaN}. Python's json.loads returns {'x': nan}. Hand that value straight back to json.dumps and, with the default settings, you get {"x": NaN} — a document that is not valid JSON and that the other three parsers reject on sight; only allow_nan=False turns the silent round trip into an exception. This is the same divergence the companion article on number precision found from the write side; here it arrives from the read side, in data a model just produced.
There is no guarantee upstream, by construction
The reason models get this wrong is not that they are bad at JSON. It is that the mechanism cannot make the promise. Sampling emits one token at a time from a probability distribution; nothing in the loop represents "the text so far is a complete, well-formed document", so nothing in the loop can refuse to finish it. A constrained decoder does better — it can track whether the text so far is a valid prefix — but a valid prefix is not a valid document, and that gap is the whole story.
"JSON mode" is the vendor term for pushing that guarantee as far upstream as it will go, and the documentation is unusually direct about what it actually buys you. Microsoft's Azure OpenAI page states that JSON mode "guarantees valid JSON output, but it doesn't guarantee the output matches a specific schema". The braces will close; the fields inside them can still be wrong. The same page shows how thin even that guarantee is at the edges — it also requires the word "JSON" to appear somewhere in the messages, and when it does not, the API rejects the request outright:
Error code: 400 - {'error': {'message': "'messages' must contain the word
'json' in some form, to use 'response_format' of type 'json_object'."}}
That error is the honest shape of the whole feature. When the instruction is missing, the documentation warns (quoting OpenAI) that the model can "generate an unending stream of whitespace and the request could run continually until it reaches the token limit" — a failure that looks, from your side, exactly like a hang.
The hole the guarantee does not cover
Validity per token is not validity per document, and the place they diverge is truncation. If the reply is longer than max_tokens, the stream simply stops mid-structure and finish_reason comes back as length. Azure's guidance is explicit on both counts: "You should check finish_reason for the value length before parsing the response. The model might generate partial JSON", and, in troubleshooting, "Don't parse partial JSON."
This is the case the guarantee cannot reach, and the case a repair layer is most likely to paper over. Take the truncated object from the table and let a repair pass close its brackets:
model (finish_reason: "length")
{"a":1,"b":[1,2
after the repair pass closes the brackets
{"a":1,"b":[1,2]}
The result parses cleanly, and it is a lie. Nothing in the text says whether the list was going to end at two items or fifty; the repair chose two. Repair restores syntax, never data.
The repair layer, measured
Since the parser upstream does not exist, one has to exist downstream, and in practice it is a repair pass that runs before JSON.parse. I wrote the smallest honest version of that pass and ran it over the same twelve documents. A first stage of plain regex substitutions — strip the fence and the byte-order mark, drop comments and trailing commas, swap the single quotes, map the Python and JavaScript non-values, close unbalanced brackets — recovered nine of the twelve. The three it missed were simply rules that first pass did not contain, and two of them are cases a regex cannot handle safely on its own: a newline must be escaped only when it is inside a string, and trailing prose must be cut only once the value has ended. Adding a boundary-tracking pass, and the key-quoting rule, took it to twelve out of twelve:
| Source pattern | After repair | What the repair did |
|---|---|---|
| markdown fence | parses | stripped the fence |
| trailing comma | parses | removed it |
| single quotes | parses | rewrote as double quotes |
None / True
|
parses | mapped to null / true
|
| block comment | parses | removed the comment |
| raw newline in a string | parses | escaped the control character |
| unquoted key | parses | added the quotes |
| trailing prose | parses | discarded everything after the value |
| UTF-8 BOM | parses | stripped the mark |
NaN / Infinity
|
parses | substituted null — a value the model never wrote |
| truncated | parses | closed the brackets — and invented the data above |
All twelve parse afterwards. The truncated one parses by being falsified. And some of those rows are decisions rather than restorations: the comment row silently dropped content the model took the trouble to write, and the NaN row is worse — the substituted null is indistinguishable from a real one, so a repair pass can convert "the model produced an undefined number" into "the model said there is nothing here" without leaving a trace.
None of this is improvisation on my part. The jsonrepair library publishes a list of what it fixes, and eight of the twelve shapes above are on it by name: add missing quotes around keys, add missing escape characters, repair truncated JSON, replace single quotes with double quotes, replace Python constants, strip trailing commas, strip comments and strip fenced code blocks. Three of the shapes it does not cover are the ones that matter most — NaN, Infinity and prose trailing after the value. Those are exactly the cases that survive a repair pass silently: NaN parses in Python and nowhere else, so a pipeline can carry it through the repair stage and lose it at whichever service happens to disagree. That a dedicated repair library and a model's failure modes overlap this heavily is the clearest evidence that the problem is structural rather than incidental.
What to do instead
- Move the guarantee upstream when you can. Prefer a schema-constrained mode over asking politely in the prompt: the documented difference is that JSON mode promises syntax while structured outputs constrain the model to a schema. Neither was exercised here (no API key), so treat the exact semantics as documented rather than measured.
- Treat the reply as untrusted text. The pipeline is repair → parse → validate against a schema, in that order, and the last step is not optional, because a syntactically perfect document with the wrong fields passes the first two.
-
Check
finish_reasonbefore you parse, and never repair a truncated reply into acceptance. Raise and retry with a larger budget instead. Repair makes a cut-off document look finished; only you can decide that it is not finished. -
Do not assume "valid JSON" is a shared contract. Two services in one pipeline can disagree about the same bytes — Ruby takes the block comment, Python takes the
NaN, and the third service rejects both. Validate at the boundary rather than trusting that the upstream parser agreed with yours. - Look at the output before it reaches code. A repair layer hides exactly the cases you most need to see.
Check your own payload
Paste a model reply into the validator to get the line and column of the first real error rather than a repair that is invisibly wrong, and into the formatter to see the structure once it parses. For output that is valid JSON but the wrong shape, the JSON Schema validator is the missing test: "parses" and "correct" are different questions, and JSON mode only answers the first. The syntax cheat sheet covers what the grammar actually allows, including why NaN and Infinity are not in it.
Provenance
The parser table and the repair table were produced by running the documents shown through Node 24.18.0 on V8 15.0.245.15, CPython 3.10.11, Ruby 2.6.10 (bundled json 2.1.0) and PHP 8.3.35 on a 64-bit build. The Node runtime here is the one embedded in an Electron host rather than a stock Node distribution — its V8 build string carries an -electron suffix — and the parser exercised in either case is V8's own JSON.parse. The repair pass was written and executed for this article rather than taken from a library, and the Python NaN round trip was measured with the default json.dumps and again with allow_nan=False. The twelve source strings are the documented failure shapes rather than captures from a live call: there is no API key on the machine this was written on, so nothing here was produced by a model. The claims about JSON mode, structured outputs and finish_reason are quoted from Microsoft's Azure OpenAI documentation, which mirrors OpenAI's, and are labelled documented rather than measured for exactly that reason; the jsonrepair list is quoted from that library's repository README, which enumerates its repairs without naming models as their source.
Related
- JSON number precision — why 19-digit IDs come back wrong — the same "the parsers disagree" problem, measured from the write side.
- JSON Schema validator — the check that catches a reply which parses but has the wrong shape.
- JSON guides — everything published so far.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.