DEV Community

Phinn
Phinn

Posted on

Cost data surprised me

稿子写好了:drafts/devto_kinetaios_engines.md,直接可用(带 front-matter,published: false 默认草稿)。

英文全文如下:


I ran 4 AI coding agents on the same 400k-line project. The cost data surprised me.

This is a build-log, not a landing page. I built a multi-engine agent dashboard (KinetAios, GPLv3, macOS + Windows) and benchmarked four agent engines on the same real task: a 400k-line medical-device CRM cross-analysis job. Same prompt, zero human intervention, full token accounting. My own engine lost the first round of architectural arguments. Numbers below — including the ones that embarrassed me.

The setup

The task: cross-analyze a real 400k-line CRM dataset (multi-sheet Excel, hospital tiers × time windows), produce an interactive ECharts report. This is the kind of messy, multi-hour job people actually hand to coding agents — not a LeetCode benchmark.

Four engines, one prompt each, no babysitting:

  • Direct — my in-house ReAct loop with native tools
  • Claude Code — Anthropic's CLI, spawned as an engine
  • Codex — OpenAI's CLI
  • PEVJ V2 — my older Plan-Execute-Verify-Judge architecture

The scoreboard

Engine Score /10 Tokens What killed it (or didn't)
Direct (mine) 9.2 1.31M Native SAX streaming — big files never enter context
Claude Code 7.0 ~2M Best context management, but no plugin tools → detours through Bash+python every step
Codex 5.5 ~2.4M Sandbox-first design fights interactive data tasks
PEVJ V2 (also mine) 3.5 3.8M+ Verify stage with no guardrails = a token amplifier

Yes, my newer engine beat my older engine and Claude Code. Yes, the older one is in the table anyway. That's the point of publishing trace data.

The finding that actually matters

7.5% of API calls burned 63% of the total budget.

I dug into the traces expecting a model problem. It wasn't. The worst offender read the same big file sequentially, 16 times — and every read resent the full conversation history. The model wasn't dumb; the tool design was.

The fix wasn't a better model. It was a streaming reader (SAX-style, so the file never enters context at all). That line item went to zero.

If you're building anything that lets an LLM touch large files, instrument per-tool-call cost before you tune prompts. Your "expensive model" problem is probably a tool granularity problem.

What a real multi-engine workday costs

The benchmark is one data point. Here's an actual day: four parallel sessions, mixed engines, isolated cost accounting per session.

Session Engine Model Job Cost that day
A Claude Code Sonnet Refactor an 800-line module $2.41
B Codex o4-mini Fix 17 failing tests $0.87
C Direct GLM Code retrieval + doc Q&A $0.12
D Direct GLM Long-term memory housekeeping $0.03

$3.43 total. All in one window, all running concurrently. The boring economics: expensive models on hard tasks, cheap models on mechanical ones — and that only works if sessions can run side by side with per-session model choice. One terminal per CLI doesn't compose.

Three engineering lessons from making multi-engine actually work

1. Total state isolation, or sessions will eat each other

We shipped a bug where session A was waiting on a shell-command confirmation — and session B consumed it. One modal queue shared across sessions, in a multi-agent app, is a race condition wearing a UI. Every session now owns its confirm queue and its entire tool context. Isolation isn't a feature here; it's the architecture.

2. SQLite WAL is the right substrate for desktop agents

Readers never block the writer. Browsing session history while the current session streams writes: zero SQLITE_BUSY in months of daily use. Plus the file is the export format — copy it, it opens anywhere. And because it's a plain local file, other tools (scripts, other agents) can read your session history directly. Cloud apps lock that behind APIs; local SQLite hands it to you.

One adjacent gotcha: FTS5's default unicode61 tokenizer is a disaster for Chinese text. We went bigram + a custom index to pull recall up. If your agent serves non-English users, check this before launch, not after.

3. The local-first part is a threat model, not a vibe

The agent executes on your machine — shell, files, SQLite history, local vector index. The only cloud is your own API key for inference, and the free tier runs entirely on local Ollama models. No relay server, no account, no telemetry. Your code never transits anyone's middleman.

I keep this framing because "local-first" gets used as marketing glitter. The test is simple: can you point at the exact bytes that leave your machine? If the answer is "just the inference requests to your own key," that's a defensible design.

What's still hard

  • Interactive approvals across concurrent sessions (see lesson 1) — the modal-confirm pattern gets weird at 3+ parallel agents. Headless timeout-with-auto-deny is the next step.
  • Session transcript size grows unbounded; recall works, but "summarize my last 40 hours of agent work" is still an open problem for us.
  • Windows build parity (Electron) always trails macOS by a week or two. Solo-dev tax.

Links

  • KinetAios — the dashboard from this post: github.com/phinn/KinetAios (GPLv3, macOS + Windows, free tier with Ollama, BYO key). ~1,000+ users reached so far; if this post saved you time, a star genuinely helps an indie project get seen.
  • Full benchmark report with per-step traces (including my own engine's 3.5/10 failure details): excel-cross-analysis-engines.html

Ask me anything about the benchmark methodology, the token accounting, or multi-engine session isolation in the comments — happy to share traces.

Top comments (0)