DEV Community

robzepdev
robzepdev

Posted on

Structured output guarantees the shape, not the lengths

I run the first real batch of my release tracker on 2026-09-28: 126 releases, one Message Batches
call, done in 17 minutes. The API reported 126 successes out of 126. My pipeline stored 122.

Four summaries were thrown away by my own validation. Not because the model hallucinated, not
because the JSON was malformed, not because a field was missing. Because the title was six
characters too long.

IMAGE 1: the Axios 1.19.0 page in production, showing "error: summary unavailable"

The four

Here they are, with their length against a limit of 120:

Release Title Characters
Axios 1.19.0 Axios 1.19.0 raises the form-data floor to fix a CRLF injection advisory and fixes NO_PROXY, interceptor and progress bugs 124
Node.js 26.8.0 Node.js 26.8.0 adds ZIP APIs in zlib, AES-SIV cipher modes, REPL syntax highlighting and new sqlite statement methods 121
Turborepo Workspace task orchestration replaces topological chunking, plus remote shared build artifacts and sideEffectsCache rework. 125
pnpm 11.27.0 pnpm 11.27.0 isolates registry metadata caches per URL path, adds global nodeDownloadMirrors, and fixes a store symlink flaw 126

Read them. They are good titles. Every other field was correct: the key points were within the
allowed count, every citation resolved to a real line of the release notes, the link between the
breaking changes and the "action required" verdict held, and stop_reason was end_turn on all
four. These were paid, correct summaries, discarded over punctuation-sized overshoot.

Why structured output did not stop it

My request carries the Zod schema as output_config.format. Constrained decoding is doing real work
there: I have never once received malformed JSON, a missing required field, or a string where an
array belongs. That part is a guarantee, and it is worth having.

But the guarantee is about the shape of the JSON, not about the values inside it. A
maxLength on a string is not a structural constraint that a decoder can enforce token by token,
the way it can enforce that an object closes or that a key is one of five allowed names. By the time
the model is 118 characters into a title, nothing in the machinery is counting, and nothing is going
to stop it at 120.

So counting characters stays something the model does by eye. And by eye, it misses by a little.

Not for lack of asking. My prompt is explicit:

Hard limit: 120 characters, spaces included. Count them before answering.

It aimed well: 122 of 126 landed under the limit, and the four that missed were 1 to 6 characters
over. That is not a model that ignored the instruction. That is a model doing arithmetic on a string
it is in the middle of writing.

The part that cannot be expressed at all

Worse than a maxLength that is ignored is a rule that never reaches the schema in the first place.

My summary schema has a .refine() on it: a summary that lists breaking changes cannot also report
"safe to upgrade". That is the whole product in one line, and it is the one rule I would least like
to be decorative.

It does not survive the trip to JSON Schema. A cross-field invariant is not expressible there at
all, so it exists only in my own parsing step, after the response comes back. On this batch it held
on all 126, and I believe it mostly holds because the prompt makes the relationship obvious. But
"it held 126 times" is an observation, not a guarantee, and I should not confuse the two. The
.refine() stays, and it stays as the last word: anything that fails it is not shown.

The fix is a gap, not a number

The obvious move is to raise the limit. I did, but the interesting part is what I did not raise.

The schema maximum went from 120 to 140. The prompt still asks for 120.

That distance is the entire fix. The model keeps aiming at 120, which is the length I actually want,
and the few characters it overshoots land in a margin instead of destroying a paid summary. If I had
raised the prompt to 140 as well, I would have moved the same problem to 141 and learned nothing.

The prompt text is untouched, so PROMPT_VERSION stays at v4 and my evaluation set stays
comparable across the change. A schema widening is not a prompt change, and it was useful that the
versioning made me say which one this was.

I also considered truncating at the limit. I decided against it: a title cut mid-word is worse than
a title that runs slightly long, and it would hide the problem instead of measuring it. I checked
the layout with a real 140-character title instead: two lines at 1280px, four lines at 390px, no
overflow. The title lives in exactly one place, the <h2> of the release page, which made that
cheap to verify.

IMAGE 2: React 19.3.0 in production, a title that fit, with its five key points and line citations

Do not pay twice for the same answer

The four summaries existed. They were generated, they were billed, and they were correct. Resending
them would have cost about 12 cents, which is nothing, and would also have been the wrong habit.

So the repair reads the saved batch results from disk, revalidates them through exactly the same
parser the pages use, and writes UPDATE statements to a .sql file. All four pass at 140. The
script touches no database: applying the file stays a deliberate gesture in the SQL editor of the
environment I mean.

There is a trap hiding in that sentence, and it nearly cost me the option. The expires_at of a
batch is 24 hours from creation, not the result retention period of 29 days that the
documentation makes easy to assume. I had saved the 126 results to a local file the same day, out of
habit rather than foresight, and that habit is the only reason this repair was free. Save the
results file. It costs nothing and it buys you every repair you have not thought of yet.

What it cost to learn

The batch billed $4.16: a mean of $0.033 per summary, a median of $0.0245. I had estimated $2.94,
measured on 14 releases I had picked by hand. The median was nearly right. The tail was not: real
release notes are longer than the ones I chose to test with, and the ten most expensive summaries
were 21% of the spend. That is a separate lesson and it deserves its own post.

The 3% discard is what I actually came away with. It is not a bug I fixed, it is a number I did not
have before: the measured distance between what structured output promises and what I had assumed it
promised.

Honest status at the time of writing: that SQL has not been applied yet, so those four releases
still read "summary unavailable" in production. The screenshot above is the live page, not a staged
one. The repair is written and verified; running it is a manual step still sitting in my queue.


The tracker is a side project: it watches releases of languages, frameworks and libraries and turns
their release notes into summaries that say whether you need to act before upgrading. Everything
above is from the real run, not a reconstruction.

Top comments (0)