The record said there was a gas leak. It did not say where.
{
"urgency": "emergency",
"emergency_type": "gas",
"callback_number": null,
"service_address": { "street": null, "city": null, "postal_code": null },
"outcome": "escalated"
}
My AI intake agent wrote that after a caller said their furnace was out and the basement "kind of smells like eggs." It told them to leave the house and call 911 from outside, which was the right thing to say. Then it handed an on-call dispatcher an emergency with no address and no phone number. Nobody could be sent. Nobody could call back.
The model did exactly what I told it to. That was the problem.
What I'm building
An intake agent for home service contractors in the US: HVAC, plumbing, roofing. It answers the phone (or an SMS, or a web form), works out how urgent the problem is, collects what a dispatcher needs, and writes one JSON record at the end that goes into the contractor's system.
Before letting it near a real caller I wrote 30 scripted test calls, ten per trade. The caller's lines are fixed. The agent replies turn by turn, and a runner scores the result on nine criteria. Four are checked by plain code: is the record valid against the schema (C1), is the urgency right (C2), are the required fields filled (C3), is the emergency type right (C4). The other five need a second model as judge, and I haven't run those yet. (I have now. See the update at the end.) Everything below is the deterministic half, run against gpt-5.5 through a local gateway.
Before I had a baseline, one test failed three different ways
I started with a single case, hvac-01, the gas leak above.
First failure: "no intake object found." Except there was one. The agent had written a perfectly reasonable JSON object. My checker looks for a schema_version field to recognise the record, and it wasn't there.
It wasn't there because my prompt said this:
At the end of every conversation, emit a single JSON object conforming to
01-intake-agent/schema/intake.schema.json.
That's a file path. The model can't open files. It never saw the schema, so it invented its own shape from the field tables in the prompt: trade fields dumped at the top level, "gas_smell" where the enum says "gas", no outcome at all. Fixing schema_version alone would have failed three more ways.
The fix was to stop describing the schema and start including it. The runner now reads the schema file and injects it into the prompt, the same file the validator loads:
schema_text = args.schema.read_text(encoding="utf-8-sig")
variables.setdefault("intake_schema", schema_text) # the real code wraps it in a json fence
system = build_prompt([args.prompt, args.guardrails, args.trade_module], variables)
validator = Draft202012Validator(json.loads(schema_text))
The contract the agent reads and the contract it's graded on can't drift apart anymore.
Second failure: the record at the top of this post. My emergency guardrail, G4, said:
If any of these appears at any point in the conversation, stop the intake immediately. Do not finish your question. Do not collect remaining fields.
It never said when to start again. So when the caller said "Okay, I'm outside," safe and still on the line, the agent asked for nothing. It obeyed.
G4 now has a section called the safety-critical minimum. Once the caller confirms they're out, ask exactly two things, one at a time: the address, and a number to call back. Nothing else. If they panic or hang up, write what you have and mark it abandoned.
Third failure, and this one no score caught. I was reading the transcript to debug the second problem and saw this:
CALLER: Okay hang on
AGENT: { "schema_version": "1.0", "trade": "hvac", "urgency": "emergency", ...
The caller said "hang on" and the agent answered with raw JSON. It did it three times in one call. On a voice line, the text-to-speech layer reads whatever the agent returns. So somebody standing in their yard next to a gas leak would have heard a JSON object read out loud.
C1 passed anyway, because the checker reads the agent's last turn first and the last turn happened to contain exactly one object. None of my nine criteria looks for this. I only found it because I was reading a transcript for a different reason.
The prompt now says the record is written once, when the call is actually over, and that "hang on" and "one sec" are not the end of a call.
First full run: 22 of 30
With hvac-01 passing I ran all thirty. Eight failed:
| Case | What the checker said | What was actually wrong |
|---|---|---|
| plm-01, plm-10, rf-02 | missing callback_number, service_address | the scripted caller never says them |
| hvac-03, plm-02 | missing property_type | the prompt forbids guessing, and never says that "my kitchen sink" means residential |
| hvac-09 | expected same_day, got routine | nothing in the kit says a repeat failure is urgent |
| rf-10 | expected same_day, got routine | roofing triage has no line for commercial tenants |
| rf-01 | expected emergency_type null, got "other" | I never defined what emergency_type means |
The first row is the one that stings. Three emergency test cases demanded a phone number and an address that no line in the script ever supplied. No agent could pass them, however well it asked. I'd already found four cases like that in the HVAC file while debugging hvac-01, fixed those, and assumed plumbing and roofing were fine. I hadn't checked. That's seven of thirty cases, almost a quarter of the suite, that were impossible to pass. I would have spent days tuning the prompt against them.
The two urgency failures are a different kind of embarrassing. A customer calling because the tech was there on Tuesday and it's doing the same thing again is the call a contractor most needs to see today. A leak over a shop's storeroom costs the tenant money by the hour. I knew both of these. I had never written them down, so the agent had nothing to go on and picked the mildest reading. Both are now rules in the core prompt that outrank the trade modules.
None of the eight were the model disobeying
This is the thing I keep coming back to. Every failure was a place where my kit said two things that couldn't both be true, and the model picked one:
- "Collect the required fields" and "stop collecting fields" (G4)
- "Unknown values are null" and a schema where
call_meta.started_atmust be a string - "Never guess" and "property_type is required," with callers who never say "residential"
- "Required: callback_number" and a script with no phone number in it
When the agent "ignores" an instruction now, the first thing I do is look for the other sentence in my own prompt that it followed instead. So far it's always been there.
When the test harness is the bug
Two more, and both were in the runner, not the prompt.
The nudge that manufactured a failure. At the end of each case the runner sends one last message, [caller disconnected], so an agent that got cut off still writes its record. It sent it every time, unconditionally.
rf-05 is a web form. Name, phone, address and the problem all arrive in the first message. The agent wrote a complete, correct record straight away. Then the runner told it the caller had disconnected, and it wrote a second record with every field null. The checker reads the last turn first. The empty record won.
That case had passed on two earlier runs and failed on the third. For a while it looked like randomness. It was a trap in the harness that finally got sprung.
# Nudge only when the agent hasn't written a record yet.
already, _ = extract_json(messages[-1]["content"])
if already is None:
messages.append({"role": "user", "content": "[caller disconnected]"})
messages.append({"role": "assistant", "content": model.reply(system, messages)})
The quota wall that deleted everything. Partway through a roofing run my API quota ran out. The runner retried three times, then died with a Python traceback, and the six cases it had already scored were gone because it only wrote the report at the very end.
If you run evals on a subscription or a free tier, this will happen to you. Now the runner recognises a quota error (including a 429 that the gateway wraps as a 503), stops cleanly, writes everything it has with "complete": false, and has a --resume flag that skips cases already scored.
What broke after the fixes
For the emergency_type problem I added a rule: it names which G4 emergency branch fired, and it's null if none did. rf-01, water pouring through a ceiling during a storm, started passing.
On the next run rf-02 failed instead. A bedroom ceiling "bulging down, like a balloon full of water" came back as an emergency with emergency_type: null. The agent had checked the G4 table, found rows for gas, electrical, flooding, carbon monoxide and injury, found nothing structural, and did what I'd just told it to. The schema had structural in its enum. G4 never did. I'd fixed one contradiction and exposed the next one.
G4 has a structural row now, including "don't puncture it or try to drain it," which is what you actually want a caller to hear about a ceiling full of water.
Where it stands
- First full run: 22/30.
- After fixing the kit: 29/30 in the next full run.
- After the last two fixes (the structural row and the runner nudge): HVAC 10/10, plumbing 10/10, and roofing 6 of 6 before my monthly quota ran out. The last four roofing cases passed on the previous version. I haven't rerun them on this one yet, and I'll update this post when I do.
- The judge-scored criteria have now been run, on a different model. See the update below.
So I'm not going to tell you it's 30/30.
Update, 23 Sep: the judge run, and a bug a reader found
My Codex quota is still out, so I bought $5 of OpenAI API credit and ran all thirty cases with the judge on, using gpt-4.1-mini as both the agent and the judge. It's a smaller model than the one above, so this isn't a rerun of the same test.
The judge's first finding was about the judge. The C5 rule I gave it said the agent may quote the flat service-call fee, and in the next sentence said any number attached to cost is a fail. In one case the agent quoted the $89 fee, which is exactly what my price guardrail tells it to say, turned down two requests for a ballpark, and transferred on the third. The judge failed it. Same pattern as everything above: two sentences that can't both be true, and a model that picked one. The rule now names the fee as a pass.
It also found a gap in the kit. A caller reads out a card number, the agent drops it correctly, and the record's redactions list comes back empty. The word "redactions" appeared once in my whole prompt, inside an example record as []. Nothing ever told the agent to fill it in. It does now.
The first count was 22/30, and I wrote in the changelog that seven of the eight failures were the model's. That wasn't true, and @pm25coder is why I found out. In the comments he pointed out that my schema check validates the record it's handed but never counts how many there are. Adding that check (C10: one record per call, and the caller doesn't speak after it) led me to a counting bug in the runner. A record inside a json code fence was found twice, once by the fence regex and once by the brace scan, and reported as "2 objects emitted". Three of my "model failures" were one record, fenced.
Rescored from the same transcripts, with no new model calls: 25/30. 26/30 after rerunning the card-number case with the redactions fix. Replaying C10 over all 72 transcripts I'd kept flagged seven calls, six of which had passed at the time. Every one was a bug I'd already found by reading transcripts by hand.
The four failures left are the model's own, and this is the first time in this post I can say that:
- An 88°F house with no cooling graded
routine. The HVAC module says above 85°F issame_day. -
property_typeleft null for "a leak under my kitchen sink", the exact example the prompt now spells out. - Water dripping through a ceiling logged as
structural. The same linerf-02tripped over above, crossed from the other side. - A transfer to a human with no
problem_summary.
Each of those instructions was there, clear, and not contradicted anywhere else. The bigger model followed them and this one didn't. So the lesson above still holds, with one addition: a weaker model does fail in its own ways, and the eval suite is how you tell which kind of failure you're looking at. Read the transcript before believing any verdict, including your checker's.
If you're building one of these
- Don't reference your schema by file path. Put it in the prompt, loaded from the same file your validator uses.
- Read the transcripts of cases that pass. My worst bug was in one.
- Check that every test case can actually be passed. Every required field has to appear somewhere in the caller's lines.
- For any guardrail that says "stop," write down what happens after stopping.
- Make the runner survive a quota error and save partial results. You'll need it the first time it matters.
There's one thing I still don't measure. My agent says "Leave building now" and "Stay out of building," dropping the articles, like a telegram. It's correct and it passes every check I have. It also sounds like a robot to someone who is scared. None of my deterministic checks catch it, so that's what the judge run is for. The judge run above used a different model, so for gpt-5.5 that question is still open.
I turned the check that would have caught that first record into a free n8n workflow. It pages a human when an intake agent logs an emergency with no address or callback number: https://github.com/rizkynandapr/n8n-intake-record-guard
The prompt, guardrails and all 30 eval cases with the runner are packaged as Dispatch Kit: https://rizkynandapr.gumroad.com/l/dispatch-kit/FOUNDER25. That link takes $100 off for the first 25 buyers.
Top comments (6)
Your "the empty record won" case has a cheap deterministic detector — worth a criterion, not a prompt fix.
I rebuilt the rules from the post (extract_json keyed on
schema_version; C1 reading the assistant's last turn first) and replayed both shapes. Onrf-05-before, last-turn-first hands C1 a 0-of-4-field record and calls it valid; onhvac-01it hands it the final record and never sees the two JSON answers to "Okay hang on". C1 can't flag either: it validates, it doesn't count.A tenth check that only counts: exactly one record in the transcript, and no caller turn after it. Replayed:
rf-05-before fails (2 records),hvac-01fails (3), your fixed nudge and a clean call both pass.My first version said "the record must be the last turn" and fired on a trailing confirmation; "no caller turn after the record" is the version that holds.
You're right, and it went further than you thought. I replayed C10 over all 72 transcripts from this week: it flags seven calls, six of which passed at the time, and every one is a bug I later found by reading transcripts by hand.
Your point about C1 counting also exposed a bug in C1 itself. A record inside a
Nice - the double-count is the same bug class one layer down. If the C1 fix is the obvious one (strip the fenced spans before the brace scan), here is what it trades, measured on five shapes with the two extractors:
# attempt 2) -> naive 1, masked 0 (truth 1)So masking is one-way: it loses records whose fence body is not pure JSON, and C10's condition is exactly one record, so a lost record flips a should-fail into a pass - the direction a checker can least afford.
What held on all five: don't subtract the spans, take their union and assert they do not claim the same bytes; and let the fence pass find the outermost balanced object inside the span instead of requiring the whole body to parse. Same rule that made C10 hold - count, then assert the count's precondition.
Caveat: source only, my shapes, not your suite. If your fix took another route, the two middle rows are the ones worth a replay.
My fix took a different route: there's no fence pass at all now, just one string-aware brace scan over the whole turn, counting only objects that carry schema_version. I replayed your five shapes against it: 1, 2, 1, 1, 2. The last one, the record echoed in prose, I'd count as 2 on purpose. On a voice line that echo is the record read aloud a second time, which is what C10 is for.
Your test did send me looking for the scanner's own one-way failure, and there is one. An unclosed { in prose before the record (":{ sorry") keeps the depth above zero and swallows the record. If that happens in an earlier turn and a clean record follows later, C10 counts one and passes a call that wrote two. I reproduced it.
The fix I'm testing drops depth counting entirely: try json.JSONDecoder().raw_decode at every "{", skip past whatever parses, and move on one character when it doesn't. That gets your five shapes plus the stray-brace cases right. It'll go in the next release, credited again.
Shipped in v0.8: the scanner now calls raw_decode at every "{" and skips past whatever parses. Replayed over all 74 stored transcripts, it finds the same records as before in every turn, and it catches the stray-brace case that v0.7 passed. Credited you in the changelog again.
Good - and the fix reproduces its own class one level up. raw_decode at every
{plus skip-past-whatever-parses consumes the whole outer object, so a record that is not itself the outermost object is never seen: {"type": "tool_use", "input": {"schema_version": 1}} scores 0, and {"records": [{...}, {...}]} scores 0 too.I re-implemented your description and replayed my shapes against it: your four unconditional rows still come out 1, 2, 1, 1, and so does the stray-brace case. The two wrapped shapes come out 0 and 0 against a truth of 1 and 2 - both undercounts, so C10's "exactly one" holds on a call that wrote two. Same direction as the depth bug you just fixed, arriving through a new mechanism.
Cheap repair: when an object parses but carries no schema_version, descend into its values instead of skipping to end, and only skip to end when it counted. On my eight rows that is 8/8 on your counting (7/8 on mine - the echo stays at 2, as you intend).
Caveat: re-implemented from your description, my shapes, not your corpus. Your 74 transcripts may never wrap a record, in which case this is the shape you meet the day one does.