The moment an LLM hands you a wall of prose and you reach for a regex, you've moved the problem — you traded a model that knows the answer for a string you have to re-parse to get it back.
curl https://aibridge-api.com/v1/chat/completions \
-H "Authorization: Bearer mb-xxxxxxxx" \
-d '{
"model": "deepseek-v4-pro",
"messages": [
{"role": "system", "content": "You return ONLY valid JSON. No markdown fences, no prose, no commentary."},
{"role": "user", "content": "Extract name, date, and amount from this invoice text: \"...\""}
]
}'
The response lands as clean, machine-readable JSON you can feed straight into your pipeline — no .split(), no regex, no "it worked until the model added an emoji."
The prose-parsing trap
LLMs default to prose because prose is how they were trained to talk. But your code doesn't want prose — it wants fields. So developers end up doing one of three fragile things:
- Regex surgery — brittle, breaks the moment the model rephrases
- String splitting — "find the first colon, take everything after it" (works until it doesn't)
- "Just ask nicely" — prompt the model to return JSON, but never verify it actually did
Every one of those fails in production, usually at 2 AM, usually on a customer's most important request.
The fix is a contract, not a wish
Structured output works when you treat the format as a requirement, not a request:
- Say it in the system prompt — "return ONLY valid JSON, no fences, no prose"
- Give it a schema — describe the exact keys and types you expect, or better, paste an example
- Verify before you trust — parse the result with json.loads / JSON.parse; if it throws, retry with a harder prompt
The model is good at following a clear contract. It's bad at guessing what you meant. Spell out the shape and it will usually deliver it.
Where structured output pays for itself
- Extraction — pull name/date/amount from messy text into clean fields
- Classification — return {"label": "refund", "confidence": 0.97} instead of "This appears to be a refund request"
- API responses — build an endpoint that returns LLM-shaped JSON, not prose you have to clean
- Chained pipelines — one model's JSON output becomes the next model's typed input, no parsing in between The moment your output is typed and parseable, your LLM call becomes a normal function in your codebase — testable, composable, boring in the best way.
One more trick: it pairs with the model router
You don't need a flagship model to extract fields from an invoice. Route the cheap, fast models to the structured-output work and save the heavy reasoning for the hard stuff:
def extract(messages) -> dict:
# glm-4-flash is fast, cheap, and plenty for JSON extraction
resp = client.chat.completions.create(model="glm-4-flash", messages=messages)
return json.loads(resp.choices[0].message.content) # typed, or it raises
- Fast tier — glm-4-flash, deepseek-v4-flash, moonshot-v1-8k
- Balanced tier — glm-4-air, deepseek-chat, qwen-plus
- Flagship tier — deepseek-v4-pro, qwen3-235b-a22b, glm-4-plus
- Deep reasoning — deepseek-reasoner, kimi-k3 (1M context) 15 models, 4 vendors, one OpenAI-compatible endpoint.
Pricing that keeps extraction cheap
- Free tier: 500K tokens/month (weighted)
- Pro: $9.90/month for 5M tokens
- Top-ups: 1M / $2.99 · 5M / $9.90 · 20M / $29.90 (never expire)
The takeaway
An LLM that returns JSON is a function. An LLM that returns prose is a chore. Put the format in the contract and let your code trust the output.
Try structured output on all 15 models, free.
→ aibridge-api.com · support@aibridge-api.com





Top comments (0)