A coding agent lives and dies by tool calls. With a frontier model you barely think about them. With a 7B model running on your laptop, the JSON is where things fall apart. This is what we ran into in oxi, what is safe to fix on the client, and why everything else should go back to the model instead of being hidden.
How a tool call gets to your code
In an OpenAI-compatible API, the model doesn't run anything. It returns a message with a tool_calls array: a tool name and an arguments field. That field is a string, and the string is supposed to contain a JSON object that matches the tool's schema:
{
"type": "function",
"function": {
"name": "read",
"arguments": "{\"path\": \"src/main.rs\"}"
}
}
With local models there is one more step. The model writes plain text in whatever format it was trained on, and the server (llama-server, Ollama, LM Studio) uses the model's chat template to recognise a tool call in that text and turn it into the structure above. oxi relies on that native tool calling rather than inventing its own text protocol, because the model was trained on that format and follows it best.
The weak spot is the arguments string. Nothing guarantees it is valid JSON. The server extracts what the model wrote, and a small model doesn't always write what it meant.
What broken arguments look like
The failures fall into a few recognisable shapes:
-
Markdown fences. The model is used to showing JSON to humans, so it wraps it:
```json {"path": "src/main.rs"} ```. -
Prose around the object.
Let me read the file: {"path": "src/main.rs"}. -
Double encoding. The object arrives as a JSON string containing JSON:
"{\"path\": \"src/main.rs\"}". -
Truncation. Generation stops mid-object, often on long
writeoreditpayloads:{"path": "src/main.rs. -
Not JSON at all. Unquoted strings, single quotes, trailing commas:
{"path": src/main.rs}.
The first three are formatting mistakes: the model knows what it wants to call, and the intent is all there. The last two are real errors, and no amount of cleverness on the client can recover the missing half of a file.
The worst fix: pretend it was empty
The tempting code, and the code oxi had until this week, is one line:
let args: Value = serde_json::from_str(&tc.arguments).unwrap_or(json!({}));
It never crashes, which is why it looks fine. But think about what the model sees next. It asked to read src/main.rs, the arguments didn't parse, the tool ran with {}, and the result that came back was:
missing path
That is a lie about what happened. The model did send a path. A strong model shrugs and tries again. A small one reads it literally: maybe the file doesn't exist, maybe it should list the directory first, maybe it should apologise and stop. It burns rounds solving a problem it doesn't have, and from the outside it looks like the model is dumb, when the client threw away the evidence.
Repair what is unambiguous, report the rest
The rule we settled on: only repair a call when there is exactly one reasonable reading of it. Fences, surrounding prose and double encoding all qualify, because the intended object is sitting right there. Everything else is sent back to the model as a tool error that says what went wrong and quotes what it sent, and the tool doesn't run at all.
This is the whole function in oxi (Rust, serde_json):
pub(crate) fn parse_tool_args(name: &str, raw: &str) -> Result<Value, String> {
let trimmed = raw.trim();
if trimmed.is_empty() {
return Ok(Value::Object(Default::default()));
}
let first_err = match serde_json::from_str::<Value>(trimmed) {
Ok(v @ Value::Object(_)) => return Ok(v),
// Double-encoded: a JSON string that itself holds the object.
Ok(Value::String(inner)) => match serde_json::from_str::<Value>(inner.trim()) {
Ok(v @ Value::Object(_)) => return Ok(v),
_ => "expected a JSON object, got a string".to_string(),
},
Ok(_) => "expected a JSON object".to_string(),
Err(e) => e.to_string(),
};
// Fences or prose: try the span from the first `{` to the last `}`.
if let (Some(start), Some(end)) = (trimmed.find('{'), trimmed.rfind('}'))
&& start < end
&& let Ok(v @ Value::Object(_)) = serde_json::from_str::<Value>(&trimmed[start..=end])
{
return Ok(v);
}
const MAX_ECHO: usize = 500;
let echo: String = trimmed.chars().take(MAX_ECHO).collect();
let ellipsis = if trimmed.chars().count() > MAX_ECHO { "…" } else { "" };
Err(format!(
"The arguments for `{name}` were not valid JSON ({first_err}). Received: {echo}{ellipsis}\n\
Call `{name}` again with a single JSON object that matches its parameters."
))
}
A few details matter more than they look. The repair only accepts a result that is a JSON object, since that's the only thing a tool's arguments can be. The error quotes the raw input, capped at 500 characters, so a truncated 20 KB write doesn't flood the context. And it names the tool, because the model may have made several calls in the same turn.
Here is what the model now gets back for a truncated call:
The arguments for `read` were not valid JSON (EOF while parsing a string at line 1 column 21).
Received: {"path": "src/main.rs
Call `read` again with a single JSON object that matches its parameters.
That message is true, specific and actionable. The model can see its own broken output and exactly where parsing stopped, which is everything it needs to fix the call on the next try. It's the same reason compilers quote the offending line.
Where the check runs matters
Parsing happens before anything else touches the call. A call with bad arguments:
-
never asks for approval. There's nothing meaningful to approve, and an approval prompt for
edit {}would only confuse the person at the keyboard. - is never batched with parallel read-only calls. It is handled on its own, so one bad call can't hold up or mix into the others.
- shows up as a failed tool in the UI, with the same message the model got, so when you're watching a run you can see exactly why it retried.
oxi talks to three wire formats (OpenAI-style chat completions, Anthropic messages and the Codex responses API), and all three loops go through the same function, even though it's the small local models that need it most.
Why not constrain the output instead?
The stronger fix is constrained decoding: llama.cpp can apply a grammar or a JSON schema at the sampler, so the model physically cannot emit invalid JSON. It's a good tool, and worth using when you control the server. We didn't start there for two reasons. oxi also runs against Ollama, LM Studio, remote boxes over SSH and hosted APIs, so the client needs a sensible answer for malformed output regardless of the backend. And constraints fix the syntax, not the intent: a model forced into valid JSON can still send the wrong field names, and that also has to come back as a clear error.
The two approaches stack well: constrain where you can, and report clearly everywhere.
If you're building on small models
-
Never swallow a parse error.
unwrap_or_default()on model output is a bug that just hasn't shown up yet. - Repair only what is unambiguous. Fences, prose and double encoding, yes. Guessing at truncated or malformed content, no.
- Quote the input back. Models correct what they can see.
- Keep the tool surface small. Fewer tools and shorter schemas leave a small model more room to get each call right.
- Test the sad path. oxi has an end-to-end test that streams a cut-off tool call from a mock server and checks that the model receives the error and the tool never runs.
The change shipped in oxi v1.11.2. It came from a question on X about small models drifting off the tool-call JSON, which is the kind of feedback we're always happy to get.
This post was drafted with help from Claude and reviewed by the oxi maintainer.
oxi is a native, open-source coding agent for any model. Run GGUF models on your own machine (oxi sets up llama-server for you), use Ollama or LM Studio, or bring your Claude Code, Codex or Cursor subscription. One Rust binary, around 110 MB of RAM, MIT licensed. Install it or star it on GitHub.
Top comments (0)