HTTP 200 doesn't mean the client got a good tool call. A nested integer can arrive as a string, a streamed call can lose a fragment, or a later assistant turn can ignore the tool result. The traces for these protocol mistakes can be awkward.
I built toolcall-check, a Python CLI that checks these behaviors on a chat completions compatible endpoint. It runs five probes: a forced echo call, a forced nested JSON call, streaming versions of both, and a two turn echo round trip. The round trip validates the first call, returns a random sentinel as a local tool result, then checks that the next assistant message matches it exactly. Generated tools aren't executed.
What the report checks
The nested payload includes an integer, a boolean, a decimal number, a list, and exact object keys. Comparison checks both value and JSON type, so true, 1, and 1.0 remain different. For streamed responses, the parser handles fragmented UTF-8, CRLF, comments, multiline data fields, and tool call deltas keyed by choice and call index. It requires a tool_calls finish reason and a [DONE] marker. Duplicate JSON keys and nonfinite values fail validation.
Each probe has its own outcome and bounded request trace. The HTTP worker runs in a disposable subprocess, and a parent wall time limit covers slow DNS, slow headers, and trickle responses. Redirects are rejected, response bodies have a byte cap, and authorization headers aren't saved in the trace. The report is self-contained HTML with escaped evidence and no scripts or external requests. Secrets are redacted on a best effort basis, so inspect the report before sharing it.
A synthetic demo is only a report demo
The built-in fixture exercises the real HTTP client and writes the same artifacts as an endpoint run, but its responses are synthetic. A passing fixture means the checker and report path work. It doesn't show that a remote model or serving stack passes these checks.
CompatCanary already documents a broader API compatibility scan with forced calls, streaming, and structured output. I kept this project focused on nested argument integrity, streamed fragment reconstruction, the fixed result second turn, and inspectable traces. I haven't validated the probes against remote models yet.
The project has no runtime dependencies. It writes report.html, results.json, and trace.json into a new private directory. I'd like feedback on the strict stream markers, response shapes that should be supported, and local endpoints worth checking.

Top comments (0)