DEV Community

Arhan Canli
Arhan Canli

Posted on Originally published at canlicapital.com

"silver price" returned US CPI: plain-word lookups that answer the wrong question

Agents rarely pass the exact ID a data source wants. They pass words: "10 year treasury yield", "core inflation", "silver price". So most tools that wrap a data source have a small resolver that turns words into an ID. Mine had a bug that never raised an error.

The bug

In canli-markets-mcp 0.1.0 (now one of the packs in canli-mcp), economic_series took a FRED series ID or plain words. The plain-word path scored each of 47 common series by the share of the request's words found in its description, and took the best one if at least half matched:

const words = s.toLowerCase().match(/[a-z0-9&]+/g) ?? [];
const scored = COMMON_SERIES
  .map(([id, text]) => [id, words.filter((w) => text.split(" ").includes(w)).length / words.length])
  .sort((a, b) => b[1] - a[1]);
if (!scored.length || scored[0][1] < 0.5) throw new NotFound(/* ... */);
return { id: scored[0][0] };
Enter fullscreen mode Exit fullscreen mode

A two-word request needs only one shared word to reach half. CPI's description is "cpi consumer price index inflation all items", so "price" alone was enough. Here is the 0.1.0 resolver next to the current one, on the same requests:

Request 0.1.0 returned Now
silver price CPIAUCSL (US CPI) refused, pointing to price_history
copper price CPIAUCSL refused
bitcoin price CPIAUCSL refused
japan inflation CPIAUCSL refused, pointing to world_economy
uk unemployment rate UNRATE (US unemployment) refused
wage growth A191RL1Q225SBEA (real GDP growth) CES0500000003 (average hourly earnings)
house prices refused CSUSHPINSA (Case-Shiller)
inflation CPIAUCSL CPIAUCSL
10 year treasury yield DGS10 DGS10

Every wrong answer came back as a real series with real numbers. The result did carry the series title, so the mistake was there for anyone who read it. Nothing made an agent read it.

The pattern is the same in every row. The words that decide the question ("silver", "japan", "uk") are the ones that matched nothing. The words that did match ("price", "inflation", "rate") say what kind of number you want, not which one.

The rule that replaced it

One function now serves three resolvers: FRED series, World Bank indicators and OECD outlook measures. The two newer ones had the same flaw before they were released: "wage growth" gave GDP growth, "house prices" gave CPI, and the OECD "exchange rate" gave the unemployment rate.

Each word of a request falls into one of three groups:

  • Filler is ignored: "what is", "show me", "since", "this year", and years such as 2020.
  • Kind words say what sort of number: rate, growth, index, level, daily, monthly. They count toward the best match but are never required.
  • Specific words are everything else, and every one of them must appear in the entry.
export function matchWords(entries, textOf, request, vocab) {
  const asked = split(request).filter((w) => !vocab.filler.has(w) && !/^(1[89]|20)\d\d$/.test(w) && !/^[a-z]$/.test(w));
  const specific = asked.filter((w) => !vocab.kind.has(w));
  if (!specific.length) return null;
  let best = null, most = 0;
  for (const entry of entries) {
    const has = new Set(split(textOf(entry)));
    if (!specific.every((w) => has.has(w))) continue;
    const shared = asked.filter((w) => has.has(w)).length;
    if (shared > most) { best = entry; most = shared; }
  }
  return best;
}
Enter fullscreen mode Exit fullscreen mode

split lowercases, keeps letters, digits and &, and folds plurals ("wages" becomes "wage"). Ties go to list order, so the more common series sits first. Each resolver adds its own words: the FRED series are all US series, so "us" and "fred" are filler there, and "interest" is a kind word.

When nothing qualifies, the tool says so and says where to look:

No common series matches "silver price". These are US series: for another country's economy use world_economy or economic_outlook, and for the price of a commodity, currency, crypto or stock use price_history ("gold" is COMEX gold futures). Or pass a FRED series id (find it at fred.stlouisfed.org).

An agent that gets this refusal can call price_history next. An agent that gets CPI has nothing telling it to.

Test the refusals first

The old tests checked that good requests found their series, and they all passed. None of them asked whether a request for something we don't have gets refused. The test file now starts with exactly that:

test("economic_series: a request naming something no series is refuses, with where to look instead", () => {
  for (const q of ["silver price", "copper price", "bitcoin price", "gold price", "japan inflation",
                   "uk unemployment rate", "euro area gdp", "stock prices", "interest rates"]) {
    assert.throws(() => resolveSeries(q),
      (e) => e instanceof NotFound && /price_history/.test(e.message) && /world_economy/.test(e.message), q);
  }
});
Enter fullscreen mode Exit fullscreen mode

After it come 42 plain requests that must each resolve to one exact series (with filler, plurals and IDs inside the words), and the same two kinds of test for the World Bank and OECD resolvers.

"interest rates" is refused on purpose. There are dozens of interest rates, and picking one would be the same bug in a nicer coat.

What I'd take from it

  • A resolver from words to IDs is a search engine that returns one result. It needs a way to return none.
  • Score on the words that tell entries apart, not on the share of words that match. Generic words are the most likely to match and the least likely to matter.
  • Write the refusal tests before the positive ones. The positive ones pass on day one.

Try it

canli-mcp is free and MIT-licensed, runs on your machine, and needs no API key for this (FRED, World Bank and OECD data are public):

npx -y canli-mcp

# Claude Code
claude mcp add canli -- npx -y canli-mcp
Enter fullscreen mode Exit fullscreen mode

The code is on GitHub:

If you find a request it resolves wrong, open an issue. That is exactly the bug I want to hear about.

Top comments (0)