Agents rarely pass the exact ID a data source wants. They pass words: "10 year treasury yield", "core inflation", "silver price". So most tools that wrap a data source have a small resolver that turns words into an ID. Mine had a bug that never raised an error.
The bug
In canli-markets-mcp 0.1.0 (now one of the packs in canli-mcp), economic_series took a FRED series ID or plain words. The plain-word path scored each of 47 common series by the share of the request's words found in its description, and took the best one if at least half matched:
const words = s.toLowerCase().match(/[a-z0-9&]+/g) ?? [];
const scored = COMMON_SERIES
.map(([id, text]) => [id, words.filter((w) => text.split(" ").includes(w)).length / words.length])
.sort((a, b) => b[1] - a[1]);
if (!scored.length || scored[0][1] < 0.5) throw new NotFound(/* ... */);
return { id: scored[0][0] };
A two-word request needs only one shared word to reach half. CPI's description is "cpi consumer price index inflation all items", so "price" alone was enough. Here is the 0.1.0 resolver next to the current one, on the same requests:
| Request | 0.1.0 returned | Now |
|---|---|---|
| silver price | CPIAUCSL (US CPI) | refused, pointing to price_history
|
| copper price | CPIAUCSL | refused |
| bitcoin price | CPIAUCSL | refused |
| japan inflation | CPIAUCSL | refused, pointing to world_economy
|
| uk unemployment rate | UNRATE (US unemployment) | refused |
| wage growth | A191RL1Q225SBEA (real GDP growth) | CES0500000003 (average hourly earnings) |
| house prices | refused | CSUSHPINSA (Case-Shiller) |
| inflation | CPIAUCSL | CPIAUCSL |
| 10 year treasury yield | DGS10 | DGS10 |
Every wrong answer came back as a real series with real numbers. The result did carry the series title, so the mistake was there for anyone who read it. Nothing made an agent read it.
The pattern is the same in every row. The words that decide the question ("silver", "japan", "uk") are the ones that matched nothing. The words that did match ("price", "inflation", "rate") say what kind of number you want, not which one.
The rule that replaced it
One function now serves three resolvers: FRED series, World Bank indicators and OECD outlook measures. The two newer ones had the same flaw before they were released: "wage growth" gave GDP growth, "house prices" gave CPI, and the OECD "exchange rate" gave the unemployment rate.
Each word of a request falls into one of three groups:
- Filler is ignored: "what is", "show me", "since", "this year", and years such as 2020.
- Kind words say what sort of number: rate, growth, index, level, daily, monthly. They count toward the best match but are never required.
- Specific words are everything else, and every one of them must appear in the entry.
export function matchWords(entries, textOf, request, vocab) {
const asked = split(request).filter((w) => !vocab.filler.has(w) && !/^(1[89]|20)\d\d$/.test(w) && !/^[a-z]$/.test(w));
const specific = asked.filter((w) => !vocab.kind.has(w));
if (!specific.length) return null;
let best = null, most = 0;
for (const entry of entries) {
const has = new Set(split(textOf(entry)));
if (!specific.every((w) => has.has(w))) continue;
const shared = asked.filter((w) => has.has(w)).length;
if (shared > most) { best = entry; most = shared; }
}
return best;
}
split lowercases, keeps letters, digits and &, and folds plurals ("wages" becomes "wage"). Ties go to list order, so the more common series sits first. Each resolver adds its own words: the FRED series are all US series, so "us" and "fred" are filler there, and "interest" is a kind word.
When nothing qualifies, the tool says so and says where to look:
No common series matches "silver price". These are US series: for another country's economy use world_economy or economic_outlook, and for the price of a commodity, currency, crypto or stock use price_history ("gold" is COMEX gold futures). Or pass a FRED series id (find it at fred.stlouisfed.org).
An agent that gets this refusal can call price_history next. An agent that gets CPI has nothing telling it to.
Test the refusals first
The old tests checked that good requests found their series, and they all passed. None of them asked whether a request for something we don't have gets refused. The test file now starts with exactly that:
test("economic_series: a request naming something no series is refuses, with where to look instead", () => {
for (const q of ["silver price", "copper price", "bitcoin price", "gold price", "japan inflation",
"uk unemployment rate", "euro area gdp", "stock prices", "interest rates"]) {
assert.throws(() => resolveSeries(q),
(e) => e instanceof NotFound && /price_history/.test(e.message) && /world_economy/.test(e.message), q);
}
});
After it come 42 plain requests that must each resolve to one exact series (with filler, plurals and IDs inside the words), and the same two kinds of test for the World Bank and OECD resolvers.
"interest rates" is refused on purpose. There are dozens of interest rates, and picking one would be the same bug in a nicer coat.
What I'd take from it
- A resolver from words to IDs is a search engine that returns one result. It needs a way to return none.
- Score on the words that tell entries apart, not on the share of words that match. Generic words are the most likely to match and the least likely to matter.
- Write the refusal tests before the positive ones. The positive ones pass on day one.
Try it
canli-mcp is free and MIT-licensed, runs on your machine, and needs no API key for this (FRED, World Bank and OECD data are public):
npx -y canli-mcp
# Claude Code
claude mcp add canli -- npx -y canli-mcp
The code is on GitHub:
-
arhancanli/canlicapital: every MCP server, including
mcp-markets/src/words.mjsandtest/words.test.mjs, the benchmarks and the site. - arhancanli/alphac: the strategy research engine.
If you find a request it resolves wrong, open an issue. That is exactly the bug I want to hear about.
Top comments (0)