I maintain an MCP server that checks backtests: give it a strategy's daily returns and it tells you, among other things, how many years of track record you'd need before the Sharpe ratio means anything.
The obvious way to use it is to paste the returns into the chat and let the agent pass them to the tool. That turned out to be the worst way. Here is the measurement.
The setup
Three questions, each needing the whole audit chain. For example: "how many years of track record does this return series need before its Sharpe beats 1.0 at 95% confidence?" The series had 400, 504 and 756 daily returns. Each question was asked 3 times of the same model (gpt-5.4-mini), so 9 runs per arm. The ground truth is the same local computation the tool runs, so an answer is either right (within 2%) or it isn't.
Three arms:
- Old tools, returns pasted. The server as it was before the one-call audit tool existed, so the agent had to chain several validators itself.
- One-call audit tool, returns pasted. The series sits in the prompt as a JSON array; the agent copies it into the tool call.
- One-call audit tool, file path. The series is in a CSV; the prompt names the file and the tool reads it.
The result
| arm | correct | median tokens per task | median time |
|---|---|---|---|
| old tools, pasted | 0 of 9 | 11,180 | 2.4 s |
| audit tool, pasted | 4 of 9 | 27,036 | 10.0 s |
| audit tool, file path | 8 of 9 | 6,524 | 2.4 s |
Same model, same tool, same numbers. Moving the data out of the prompt doubled the accuracy and cut the tokens by about three quarters.
What went wrong when it pasted
The interesting part is how the pasted runs failed. On the 756-return question, all three pasted runs gave the same wrong answer: 0.1718 years instead of 0.1452. Not three different mistakes, one identical mistake, three times.
That's what you get when values go missing on the way into the call. The model is re-typing 756 numbers as output tokens. Somewhere in there it skipped some, the tool dutifully computed the right answer for the wrong series, and the agent reported it with full confidence. Nothing errored. If I hadn't had the true value, I'd have believed it.
Pasting also costs. Every number is paid for twice: once as input in the prompt, once as output when the model writes the tool call. One pasted run used 76,930 tokens for a single question.
The one miss in the file arm was different: the agent answered 388, which is the right answer in observations (about 1.54 years × 252 trading days) reported where years were asked. A unit slip, not lost data.
What I changed
- The audit tool takes
returns_file(a path to a CSV) on the local server. It reads numbers only, caps the file size, and reports errors by row position, never by echoing file content back into the context. - Tool descriptions say plainly that a long series should go in as a file.
- The hosted version can't read your disk, so it refuses file input instead of pretending.
Limits of this test
It's one small model, 9 runs per arm, on synthetic series. I didn't run Claude models for this one. So read it as "this failure is real and easy to hit", not "pasting loses exactly 55% accuracy". The run records and the task code are in the repo, if you'd like to rerun it with a different model.
The general lesson I took: if a tool needs a lot of data, don't make the model carry it. Pass a reference (a path, an ID, a URL) and let the tool fetch it. The model is good at deciding what to compute and bad at being a copy machine.
Have you seen the same thing with other kinds of data, like long JSON, CSV rows or logs? I'm curious where the length threshold sits for bigger models.
Top comments (0)