Most LLM dashboards show you the bill after the damage is done.
You run a coding agent for an afternoon, ship a feature, and two days later your
OpenAI dashboard says you spent $47. On what? Which model? Which request? You
have no idea — and by then it's too late to do anything about it.
This is the problem we kept hitting, so we built tokeneff — an open-source CLI that puts a real-time electricity meter on your LLM API spend.
⚡ tokeneff 电表 (CNY)
今日花费 ¥0.0284
本月累计 ¥0.2524
月终预测 ~¥0.93 (100% 置信)
累计节省 ¥0.0421
今日模型花费分布
deepseek-v4-flash ¥0.0192 15,797 tok
glm-4-flash ¥0.0092 8,273 tok
How it works
tokeneff runs a tiny local proxy on localhost:7860. You point your LLM client's base_url at it, and it:
- Routes your request to the upstream provider (OpenAI, DeepSeek, GLM, Kimi, MiniMax, Anthropic) using your own API key — BYOK, the key never leaves your machine
- Adapts request formats automatically (OpenAI ↔ Anthropic conversion)
- Meters the token usage from the response
- Calculates cost locally using a bundled pricing table — official price vs what you actually pay
- Stores the record in a local SQLite DB (token counts only — never your prompts)
your client → tokeneff proxy (localhost:7860) → LLM upstream
↓
local SQLite meter
It's a transparent pipe. Your requests still go direct to the provider; tokeneff just reads the usage counts from the response and keeps a running tab.
What makes it different
There are a few LLM cost trackers on GitHub already. Here's where tokeneff fits:
| Feature | tokeneff | toktrack | tokencost | LLM-Cost-Guardian |
|---|---|---|---|---|
| Capture method | local proxy | reads CLI logs | manual / proxy | local proxy |
| Real-time token count | ✅ | ❌ | ✅ | ✅ |
| Month-end forecast | ✅ | ❌ | ❌ | ❌ |
| Dual region / currency | ✅ | ❌ | ❌ | ❌ |
| BYOK + platform dual-mode | ✅ | ❌ | ❌ | ❌ |
| Budget alerts | ✅ | ❌ | ✅ | ✅ |
Three things nobody else does:
1. Month-end forecast
The killer feature. Based on your usage trend (weighted: last 7 days × 70% + linear extrapolation × 30%), it predicts what you'll spend by month-end — with a confidence score. Seven days of data and it locks in. "At this rate, I'll hit $31 by the 31st" is the kind of number that actually changes behavior.
2. Dual region / currency
If you're a Chinese developer using DeepSeek/GLM/Kimi, your spend is in CNY ¥. If you're using OpenAI/Claude, it's in USD $. Most tools mash these together into one meaningless number. tokeneff tracks them separately — CNY spend and USD spend are never mixed.
3. BYOK + platform dual-mode
BYOK mode is 100% free and local — your keys, your upstream, zero markup. But there's also a platform mode that routes through the TokenEff gateway (one key, all models, wholesale pricing). The meter shows the savings: official price vs platform price, side by side. You switch when it makes sense, not when you're forced to.
The privacy angle
This matters more than people admit. A cost tracker that sends your request metadata to a cloud service to compute your bill is... not great. tokeneff:
- Stores only token counts, model name, cost, timestamp
- Never stores prompt content or completions
- Runs entirely locally —
~/.tokeneff/meter.dbis a SQLite file you canrmanytime - API keys live in the OS keyring, never plaintext on disk
Try it
pip install tokeneff
tokeneff setup # pick provider, paste key
tokeneff start # starts the proxy on :7860
tokeneff stats # see the meter
tokeneff dashboard # live TUI (refreshes every 0.5s)
Point your client at http://localhost:7860/v1 and start building. Watch the meter while you work — not after.
Ubuntu 22.04 gotcha: the default
setuptools(59.6.0) is too old for PEP 621 and installs asUNKNOWN-0.0.0. Runpip install --user --upgrade "setuptools>=70"first. (Yes, this is in the README too.)
What's next
- Richer TUI — per-model trend charts, request-level drill-down
- More providers — community-contributable registry
- Cost comparison reports — "run this task on 5 models, here's what each costs"
Star ⭐ the repo if this solves a problem for you. Issues and PRs welcome — especially new provider additions.
Building tokeneff as part of a broader open-source LLM gateway stack. BYOK meter first, gateway later. The meter is and will remain free + local.
Top comments (3)
The month-end forecast is the feature I'd steal first. I run agents with several model providers and the real problem was never the provider dashboard — it's attributing spend to a specific run or a specific agent. A proxy that meters per request, with token counts read from the response itself, closes exactly that gap. One thing I'd ask about: streaming. My clients use SSE and the usage object only arrives in the final chunk — do you buffer the response or read usage from the trailing chunk without delaying the stream? And do you dedupe retries? On my side a timed-out request that the agent retries shows up twice in the meter, which inflates the real cost of a run by a factor I had to learn the hard way.
The after-the-fact bill is the default pain of agent work: the dashboard tells you $47 two days later with no attribution, and by then the loop that spent it is long gone. Metering at a local proxy is the right seam - one redirect of base_url and every request gets counted at the moment it happens, which is when you can still act on it. Token counts only, never prompts, is the correct privacy posture for a meter; a cost tool that stored prompts would be a leak wearing a utility's clothes. The month-end forecast line is a nice touch - the number that actually changes behavior is the one that predicts the damage before the invoice confirms it.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.