DEV Community

Arhan Canli
Arhan Canli

Posted on

canli-mcp: 306 free finance tools for Claude and any MCP client, measured against OpenBB, EdgarTools and Yahoo

I've been building canli-mcp: one MCP server that puts 306 finance tools behind three. It is free, MIT-licensed, needs no API key for almost everything, and runs on your own machine.

npx -y canli-mcp

# Claude Code
claude mcp add canli -- npx -y canli-mcp
Enter fullscreen mode Exit fullscreen mode

For Claude Desktop, Cursor or any other MCP client: { "mcpServers": { "canli": { "command": "npx", "args": ["-y", "canli-mcp"] } } }.

What it does

The model sees three tools: find_tool searches all 306 by what you ask for, describe_tool reads one tool's arguments, and run_tool runs any tool, or up to 25 in one round trip. Behind them are seven packs:

  • Markets (34 tools), read from the sources themselves:
    • SEC filings and their sections, and full-text search across every filing;
    • ownership: insider trades (Form 4), 13F holdings, 5% holders (13D/13G), fund and ETF portfolios (N-PORT);
    • trading: short interest and off-exchange volume (FINRA), fails to deliver, futures positioning (CFTC);
    • company detail: executive pay against performance, revenue by segment;
    • options: chains with IV and greeks, volatility surfaces and the VIX futures curve (Cboe's delayed data);
    • the Fed and macro: decisions, votes and dot plots, the US economic calendar, Treasury yields and auctions, the federal debt, FRED, World Bank and OECD data;
    • prices for stocks, currencies, futures, indices and crypto.
  • Quant (235 tools): performance and risk, options and exotics, fixed income, portfolios, econometrics, indicators, and a library of 399 strategy sleeves. They are checked against QuantLib, statsmodels, arch, TA-Lib, scipy and pandas in 461 reference cases.
  • Backtest validation (17 tools): deflated Sharpe, the probability of backtest overfitting, data-snooping tests, leakage checks and placebo tests, all run locally.
  • Also: SEC fundamentals as first reported (point in time), factor backtests that refuse lookahead, Alpaca paper trading behind pre-trade checks, and an open research record.

Data results carry their source URLs, and market data carries hashes of the exact bytes received, so a figure can be checked later.

How it compares

canli-mcp OpenBB EdgarTools Yahoo Finance MCP
Finance question types with a keyless tool (of 53) 49 39 11 13
Tool definitions sent with every request 859 tokens 432,001 (2,171 in its discovery mode) 3,801 2,205

The coverage map checks each server's own tool list (bench/rivals/COVERAGE.md in the repo). Having a tool doesn't mean answering correctly, so there is also a head-to-head:

  • Setup: the same model (gpt-5.4-mini), the same prompt and the same limits for every arm; only the MCP server differs.
  • Result: on the questions every server finished, canli-mcp answered 22 of 28 runs correctly (79%). The best rival setup, EdgarTools plus Yahoo, answered 15 of 28 (54%).
  • Caveats: we wrote the 24 questions, in areas our servers were built for. It is one small model with two runs per question. A held-out set is still to run.

The full report, with every answer, is bench/rivals/REPORT.md.

Context is the other cost. Installed separately, the seven Canli servers send 34,508 tokens of tool definitions and instructions with every request. canli-mcp sends 1,109, measured with bench/context.py. The 859 in the table counts tools in the shape OpenAI-style clients send them, the same way for every server.

Finding the right tool

With 306 tools, search is the product. I wrote 48 everyday requests after the last change to the search and never tuned on them: "buy spy in my alpaca paper account", "cot report for euro futures", "executive salaries at microsoft". find_tool puts the right tool first for 46 of them. The previous release managed 25.

The biggest single fix was embarrassing: the stemmer turned "rates" into "rat", so "mortgage rates" never matched anything that said "rate". Prices, shares and trades broke the same way.

Wrong answers we caught before you did

The most useful work this week was hunting answers that were wrong without warning:

  • "silver price" returned US CPI. The plain-word lookup for economic series took any series sharing half the words, and "price" is in "consumer price index". "japan inflation" returned US inflation the same way. Now every specific word must match, and a request it can't answer is refused, with where to look instead.
  • "facebook" resolved to an unrelated "Denny Hill/Facebook LLC" from EDGAR's entity search, and "jp morgan" to "JP Morgan AG". Names now go to the listed company first, so Google is Alphabet and Facebook is Meta. Fund managers go to the filer of their 13Fs.
  • Hostile inputs. Every tool was called with hostile arguments: "constructor", empty and 5,000-character strings, edge numbers, 20,000-item arrays. Of the 2,980 calls that passed validation, a few showed bugs: find_tool crashed on the word "constructor", and some tools returned NaN where they should refuse. All are fixed, and the sweep (bench/robustness.mjs) runs before every release now.

Privacy

Everything runs on your machine:

  • Where it connects: only the public sources (SEC, Treasury, FRED, FINRA, CFTC, the Fed, BLS, BEA, Census, OECD, World Bank, Cboe, and Yahoo's public chart data for prices).
  • What it keeps: it sends nothing to us and logs nothing.
  • Offline: CANLI_OFFLINE=1 keeps only the packs that use no network.
  • Provenance: packages are published from GitHub Actions with npm provenance.

Not investment advice. Data comes as the publishers release it, Cboe's quotes are delayed, and Yahoo's chart data is unofficial.

Try it

npx -y canli-mcp
Enter fullscreen mode Exit fullscreen mode

Then ask your assistant things like:

  • "who owns more than 5% of coca cola"
  • "did the fed raise rates and who dissented"
  • "short interest in GME"
  • "is my backtest overfit"

Everything is on GitHub:

If something is missing or wrong, open an issue. I would much rather hear about a wrong answer than ship it quietly.

Coming in the next release: earnings dates from companies' own 8-K filings, press-release headlines as news, and ready-made workflows as MCP prompts.

Top comments (4)

Collapse
 
deanlee profile image
Dean Lee •

Collapsing 306 tools behind a three-tool discovery layer (find_tool, describe_tool, run_tool) addresses the hidden token tax in tool-heavy workflows. Sending 34,508 tokens of tool schemas on every turn burns context budget before the model even reads the financial data, and attention degrades over long schema lists regardless of context window marketing. Trimming that overhead down to 859 tokens changes both the per-query economics and the routing accuracy.

The silent resolution bugs you caught in the audit are where most financial agent benchmarks fail in practice. In quantitative workflows, a hard crash on a malformed ticker is cheap because the harness catches it immediately. An entity resolver that quietly maps "Facebook" to an unrelated private LLC or returns US CPI when asked for silver prices introduces unhedged data corruption into the downstream calculation. Combining strict lexical matching with byte-level response hashes gives the caller an auditable trail without relying on the model to notice that a macro series looks off.

Collapse
 
arhancanli profile image
Arhan Canli •

Small correction on the number, since it's easy to mix up: 859 is the tool-definition size in the OpenAI function shape. In the MCP shape the three front tools come to 1,109 tokens per request, against 34,508 for the seven servers installed separately. Both counts cover only the definitions. A discovery layer adds calls (find_tool, then describe_tool for the one tool you pick), so the fair comparison is definitions plus those extra round trips, and I haven't priced that second part.

On routing: find_tool put the right tool first for 46 of 48 everyday requests it was never tuned on, up from 25 of 48 in the earlier version. One of the bugs fixed along the way was a stemmer that turned "rates" into "rat", so part of the accuracy came from repairing lookups like that, not from the smaller schema.

It's free (MIT, no key, runs locally) if you want to try it: npx -y canli-mcp, repo at github.com/arhancanli/canlicapital. I built it.

Collapse
 
jaroslav_marda_9f25d4d9c profile image
Jaroslav Šmarda •

Holding the model, prompt and limits fixed and swapping only the server is the comparison I haven’t seen anyone else run — it shows answer accuracy depends on the server, not just the model. Turning your own Part 34 point back on you: 22 of 28 has a Wilson 95% interval of roughly 60–90%, and 15 of 28 roughly 36–71%, so the gap is promising but not yet settled on 28 runs. The held-out set should tell — curious what it shows.

Collapse
 
arhancanli profile image
Arhan Canli •

Fair, and the interval is the honest way to read it. One more reason to be careful: the two runs of a question are not independent draws, since an easy question tends to be right twice and a hard one wrong twice. So the 28 runs behave more like 14 questions' worth of information than 28, and the real interval is wider than the Wilson one. A paired comparison per question (both arms saw the same questions) is the better test than comparing two proportions, and that is how I will report the held-out set, with the questions written before anything is tuned. I will post the result either way.