DEV Community

Cover image for Baize v0.4.0: Tool Matching Upgraded — Optional Vector Retrieval, Lexical Baseline Unchanged
rebornace
rebornace

Posted on

Baize v0.4.0: Tool Matching Upgraded — Optional Vector Retrieval, Lexical Baseline Unchanged

What this release is about

Baize is a team-facing AI assistant runtime: you point it at an existing service, upload an OpenAPI / Swagger-style doc, and the endpoints become tools the assistant can call. A real internal backend, once wired up, usually brings hundreds of tools into the catalog.

That scale creates a problem: stuffing hundreds of tool schemas into every prompt is expensive, and the model has to pick from a very long list. Baize already solved the first part with a decision layer — before the main model call, a cheap deterministic step narrows the candidate tools. Published numbers: against 37 real read-only requests across 3 backends (390 tools total, DeepSeek-Flash), narrowing to the default 16 candidates cut turn-0 prompt tokens by about 34%, and 185 runs across 5 loops finished at 99.5% success.

The lexical baseline, however, has a known ceiling: it ranks by tokens. Same meaning, different wording, and a tool may simply not surface. v0.4.0 enhances that step without touching the baseline.

What v0.4.0 adds

Baize's tool retrieval has always been Tool-RAG in spirit: each tool is indexed as a small document (name, description, HTTP method and path, parameter names).

The lexical baseline stays exactly as it was. Every tool document gets a BM25 index plus a sparse TF-IDF index; at query time the two scores are fused with RRF. It ships on by default, needs zero dependencies, and stays the zero-friction path.

What's new is the dense channel. Once an embedder is configured, each tool document also gets a vector. At query time, the query is embedded and cosine similarity is computed — and in the fusion, the dense channel is the primary one, with BM25 demoted to an auxiliary channel. The two are merged again with RRF before being handed to the model. Division of labor: BM25 catches exact terms — proper nouns and fixed phrases in your business domain; the dense channel catches paraphrases, so a user asking the same thing with different words still surfaces the right tool. The mixed mode and the pure lexical mode are distinct internal states, observable and reversible.

You can pick either embedding source:

  • Local Ollama. For setups that don't want to send text to a cloud embedder. The settings page walks through the whole flow: one-click Ollama install , pulling the matching model, showing install and model paths, a custom models directory, and a clean uninstall when you don't need it. No command line required.
  • Any OpenAI-compatible Embedding API. If you already have a compatible endpoint, fill in the base URL and key — the same habit Baize uses for chat models.

Design tradeoffs

Making this an optional enhancement rather than the default was deliberate:

  • Zero-cost onboarding. Fresh install, lexical matching just works. You only opt into the dense channel when you feel lexical recall isn't enough.
  • Fail open to the lexical path. If the dense channel misbehaves — Ollama not ready, embedding API timeout — retrieval automatically falls back to the pure lexical index and the conversation continues. Consistent with the runtime's fail-open philosophy: speedups should never become the new failure point.
  • Narrowing never disables tools. Pruning only decides which schemas land in this step's prompt; any registered tool still runs if the model names it directly. When nothing matches, it falls back to the full set, and system / login tools are always retained.
  • Easier to find. The runtime settings page is now organized into tabs: common, speedups, memory & compaction, security. Tool matching lives under "speedups".

Reproducibility

The release also ships an evidence evaluation script: it replays and scores tool-call traces. After a change to matching, one run tells you whether recall regressed, instead of relying on gut feel. Corpus and script are in the repo.

(The previous v0.3.x line already brought the web workspace, long-horizon projection — model-facing hints only, saved chats never rewritten — and a built-in self-help skill. This post focuses on the v0.4.0 matching upgrade; the README covers everything else.)

Try it

You'll need Go 1.25+, or grab a prebuilt binary from GitHub Releases (which also ships the optional WeChat channel adapter). Unpack, fill in your .env file, and start — the console is at /ui. Out of the box you get lexical matching; to try the dense channel, open the tool matching page in settings and pick Ollama or an Embedding endpoint.

Wrapping up

The direction is straightforward: the more tools you wire in, the more retrieval at this step matters. Baize keeps it "zero-cost by default, enhancement on demand, fail open" — users who want no new dependencies see no change, and users willing to spend a little on embeddings get more reliable tool recall.

v0.4.0 is out: https://github.com/rebornace/baize

Feedback on matching quality, or on the Ollama one-click flow, is very welcome — I'm following the issues and will keep shipping.

Top comments (0)