<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: rebornace</title>
    <description>The latest articles on DEV Community by rebornace (@rebornace).</description>
    <link>https://dev.to/rebornace</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4127161%2F90390741-aa85-460d-a5fd-8e6e7f58bcb0.jpg</url>
      <title>DEV Community: rebornace</title>
      <link>https://dev.to/rebornace</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rebornace"/>
    <language>en</language>
    <item>
      <title>Baize v0.4.0: Tool Matching Upgraded — Optional Vector Retrieval, Lexical Baseline Unchanged</title>
      <dc:creator>rebornace</dc:creator>
      <pubDate>Thu, 08 Oct 2026 02:27:30 +0000</pubDate>
      <link>https://dev.to/rebornace/baize-v040-tool-matching-upgraded-optional-vector-retrieval-lexical-baseline-unchanged-20h4</link>
      <guid>https://dev.to/rebornace/baize-v040-tool-matching-upgraded-optional-vector-retrieval-lexical-baseline-unchanged-20h4</guid>
      <description>&lt;h2&gt;
  
  
  What this release is about
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/rebornace/baize" rel="noopener noreferrer"&gt;Baize&lt;/a&gt; is a team-facing AI assistant runtime: you point it at an existing service, upload an OpenAPI / Swagger-style doc, and the endpoints become tools the assistant can call. A real internal backend, once wired up, usually brings hundreds of tools into the catalog.&lt;/p&gt;

&lt;p&gt;That scale creates a problem: stuffing hundreds of tool schemas into every prompt is expensive, and the model has to pick from a very long list. Baize already solved the first part with a decision layer — before the main model call, a cheap deterministic step narrows the candidate tools. Published numbers: against 37 real read-only requests across 3 backends (390 tools total, DeepSeek-Flash), narrowing to the default 16 candidates cut turn-0 prompt tokens by about 34%, and 185 runs across 5 loops finished at 99.5% success.&lt;/p&gt;

&lt;p&gt;The lexical baseline, however, has a known ceiling: it ranks by tokens. Same meaning, different wording, and a tool may simply not surface. v0.4.0 enhances that step without touching the baseline.&lt;/p&gt;

&lt;h2&gt;
  
  
  What v0.4.0 adds
&lt;/h2&gt;

&lt;p&gt;Baize's tool retrieval has always been Tool-RAG in spirit: each tool is indexed as a small document (name, description, HTTP method and path, parameter names).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The lexical baseline stays exactly as it was.&lt;/strong&gt; Every tool document gets a BM25 index plus a sparse TF-IDF index; at query time the two scores are fused with RRF. It ships on by default, needs zero dependencies, and stays the zero-friction path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's new is the dense channel.&lt;/strong&gt; Once an embedder is configured, each tool document also gets a vector. At query time, the query is embedded and cosine similarity is computed — and in the fusion, the dense channel is the primary one, with BM25 demoted to an auxiliary channel. The two are merged again with RRF before being handed to the model. Division of labor: BM25 catches exact terms — proper nouns and fixed phrases in your business domain; the dense channel catches paraphrases, so a user asking the same thing with different words still surfaces the right tool. The mixed mode and the pure lexical mode are distinct internal states, observable and reversible.&lt;/p&gt;

&lt;p&gt;You can pick either embedding source:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Local Ollama.&lt;/strong&gt; For setups that don't want to send text to a cloud embedder. The settings page walks through the whole flow: one-click Ollama install , pulling the matching model, showing install and model paths, a custom models directory, and a clean uninstall when you don't need it. No command line required.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Any OpenAI-compatible Embedding API.&lt;/strong&gt; If you already have a compatible endpoint, fill in the base URL and key — the same habit Baize uses for chat models.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Design tradeoffs
&lt;/h2&gt;

&lt;p&gt;Making this an &lt;em&gt;optional&lt;/em&gt; enhancement rather than the default was deliberate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Zero-cost onboarding.&lt;/strong&gt; Fresh install, lexical matching just works. You only opt into the dense channel when you feel lexical recall isn't enough.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fail open to the lexical path.&lt;/strong&gt; If the dense channel misbehaves — Ollama not ready, embedding API timeout — retrieval automatically falls back to the pure lexical index and the conversation continues. Consistent with the runtime's fail-open philosophy: speedups should never become the new failure point.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Narrowing never disables tools.&lt;/strong&gt; Pruning only decides which schemas land in this step's prompt; any registered tool still runs if the model names it directly. When nothing matches, it falls back to the full set, and system / login tools are always retained.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Easier to find.&lt;/strong&gt; The runtime settings page is now organized into tabs: common, speedups, memory &amp;amp; compaction, security. Tool matching lives under "speedups".&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Reproducibility
&lt;/h2&gt;

&lt;p&gt;The release also ships an evidence evaluation script: it replays and scores tool-call traces. After a change to matching, one run tells you whether recall regressed, instead of relying on gut feel. Corpus and script are in the repo.&lt;/p&gt;

&lt;p&gt;(The previous v0.3.x line already brought the web workspace, long-horizon projection — model-facing hints only, saved chats never rewritten — and a built-in self-help skill. This post focuses on the v0.4.0 matching upgrade; the README covers everything else.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;You'll need Go 1.25+, or grab a prebuilt binary from GitHub Releases (which also ships the optional WeChat channel adapter). Unpack, fill in your .env file, and start — the console is at /ui. Out of the box you get lexical matching; to try the dense channel, open the tool matching page in settings and pick Ollama or an Embedding endpoint.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrapping up
&lt;/h2&gt;

&lt;p&gt;The direction is straightforward: the more tools you wire in, the more retrieval at this step matters. Baize keeps it "zero-cost by default, enhancement on demand, fail open" — users who want no new dependencies see no change, and users willing to spend a little on embeddings get more reliable tool recall.&lt;/p&gt;

&lt;p&gt;v0.4.0 is out: &lt;a href="https://github.com/rebornace/baize" rel="noopener noreferrer"&gt;https://github.com/rebornace/baize&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Feedback on matching quality, or on the Ollama one-click flow, is very welcome — I'm following the issues and will keep shipping.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>agents</category>
      <category>go</category>
    </item>
    <item>
      <title>Context as a File: Baize's CLM-Inspired Long-Conversation Projection</title>
      <dc:creator>rebornace</dc:creator>
      <pubDate>Sun, 04 Oct 2026 02:40:48 +0000</pubDate>
      <link>https://dev.to/rebornace/context-as-a-file-baizes-clm-inspired-long-conversation-projection-50n9</link>
      <guid>https://dev.to/rebornace/context-as-a-file-baizes-clm-inspired-long-conversation-projection-50n9</guid>
      <description>

&lt;blockquote&gt;
&lt;p&gt;Project: Baize — an AI assistant runtime for your team (Go 1.25+, MIT)&lt;br&gt;
Repo:&lt;br&gt;
&lt;a href="https://github.com/rebornace/baize" rel="noopener noreferrer"&gt;https://github.com/rebornace/baize&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  1. The short version
&lt;/h2&gt;

&lt;p&gt;Baize v0.3.x takes inspiration from the recent Context Language Models (CLM) papers and ships a constrained take: &lt;strong&gt;keep important facts in long chats&lt;/strong&gt;. It treats the context that reaches the model as something that can be curated — pin the essentials, roll up rolling summaries, drop oversized tool results — but it &lt;strong&gt;never rewrites the saved chat messages&lt;/strong&gt;, and on failure it automatically falls back to the existing compaction logic.&lt;/p&gt;

&lt;p&gt;One line for the tradeoff: CLM hands editing rights over the context to the model; Baize puts the editing in a system-side projection layer, trading some freedom for reliability.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Why long conversations are a hard problem for agents
&lt;/h2&gt;

&lt;p&gt;An agent's context is the classic append-only structure: every new message, tool result and reasoning trace gets appended to the history. Once a conversation gets long, the problems show up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Attention dilution&lt;/strong&gt;: with thousands of lines of history, one critical configuration item can get lost; the model may not reliably pick it up anymore.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Truncation is irreversible&lt;/strong&gt;: cutting old messages silently drops early agreements and context the user never notices.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Compaction is irreversible&lt;/strong&gt;: rolling summaries save space, but a summary is a lossy, one-way transformation; key details can simply disappear.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The traditional answer is to hand context management to an external framework: compress on a threshold, offload, retrieve. Those approaches work, but they are "the framework defines the actions, the model can only pick one".&lt;/p&gt;

&lt;h2&gt;
  
  
  3. What CLM brings
&lt;/h2&gt;

&lt;p&gt;CLM is a paper from late September 2026 (arXiv:2609.37725, by the University of Washington, Meta Superintelligence Labs, MIT and others). It turns the problem around: &lt;strong&gt;context is not an append-only log — it is a file the model has write permission on&lt;/strong&gt;. The model can append, but it can also edit, delete and rewrite that file, and every change is synced into the next turn's context.&lt;/p&gt;

&lt;p&gt;The observations in the paper are direct:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Models can learn on their own what is most worth keeping in context, and behaviors show up that previous frameworks never had — maintaining a tracker table for multi-agent collaboration, or defining reusable context-management functions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;On the deep-research benchmark BrowseComp-Plus, out-of-the-box CLM beats the strongest baseline by 11.4% in accuracy while using 21.5% fewer inference FLOPs; on a 24-hour, multi-repo agent-swarm task it achieves 65% higher end-to-end speedup at the same compute.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Because context management becomes model behavior rather than an external harness policy, the strategy itself can keep improving through reinforcement learning.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The appeal is that context management stops being a hard-coded rule set and becomes a product the model can understand and learn. But handing the model full editing rights has real production gates: is the editing reliable, does the archive stay trustworthy, and what happens when it fails?&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Baize's constrained take: keep important facts in long chats
&lt;/h2&gt;

&lt;p&gt;Baize did not copy CLM wholesale. It extracted the core idea — &lt;strong&gt;context is curatable output, not an append-only log&lt;/strong&gt; — and added three constraints:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The projection only lives on the model input side; saved chats are never rewritten.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When a thread gets long, the assistant keeps a model-facing projection of the facts it still needs: it pins what actually matters, summarizes the secondary stuff, and drops oversized tool results. The chat the user sees stays untouched, ready to review, fork or roll back at any time. This matters for team use: the archive is the basis of audit and trust. The model input can be curated aggressively; the archive must stay true.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Fail-open.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If the projection fails, the assistant falls back to the original compaction and the conversation keeps going — context curation must never become an availability bottleneck. This matches the philosophy Baize's decision layer already follows: optimization must not affect whether the tools themselves still work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. No promise of lower cloud token usage.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The projection itself costs tokens; the curated result is not necessarily cheaper than plain compaction. The docs say it plainly. What this feature solves is &lt;strong&gt;key-fact retention and context usability&lt;/strong&gt; — not the bill.&lt;/p&gt;

&lt;p&gt;The knob is optional (default off), hot-reloadable, and costs nothing when unused.&lt;/p&gt;

&lt;p&gt;Why "constrained"? Baize talks to OpenAI-compatible APIs: it cannot read logits, and there is no native model-side context-editing capability. Rather than letting the model rewrite history in an uncontrolled place, the editing lives in the system-side projection layer, where it can be made reliable and reversible. CLM is the radical route of native model editing; Baize's projection is the pragmatic route of framework-style editing. Same goal, different cost and reliability.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. A little design philosophy
&lt;/h2&gt;

&lt;p&gt;Looking back at this upgrade, a few decisions are consistent:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Optional, rather than deciding for the user by default&lt;/strong&gt;: the projection and the decision layer are both runtime knobs, off by default, taking effect via hot reload when enabled.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Fail-open first&lt;/strong&gt;: projection falls back to compaction, routing falls back to full tools — availability is never harmed by an optimization.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are also working on long-conversation handling for agents, come discuss your tradeoffs at the repo: &lt;a href="https://github.com/rebornace/baize" rel="noopener noreferrer"&gt;https://github.com/rebornace/baize&lt;/a&gt;. Everything mentioned here (the projection and workspaces) already exists in the public repository; docs live under docs/developers in both English and Chinese. I'll keep following up on issues and suggestions.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Reference: Context Language Models (arXiv:2609.37725): &lt;/em&gt;&lt;a href="https://arxiv.org/abs/2609.37725" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2609.37725&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>go</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>After Jev Blew Up, I Brought Its "System One Judgment" Idea Into My Open-Source AI Agent</title>
      <dc:creator>rebornace</dc:creator>
      <pubDate>Wed, 23 Sep 2026 15:22:04 +0000</pubDate>
      <link>https://dev.to/rebornace/after-jev-blew-up-i-brought-its-system-one-judgment-idea-into-my-open-source-ai-agent-584e</link>
      <guid>https://dev.to/rebornace/after-jev-blew-up-i-brought-its-system-one-judgment-idea-into-my-open-source-ai-agent-584e</guid>
      <description>

&lt;blockquote&gt;
&lt;p&gt;Project: Baize — an AI assistant runtime for your team (Go 1.25+, MIT)&lt;br&gt;
Repo: &lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/rebornace/baize" rel="noopener noreferrer"&gt;https://github.com/rebornace/baize&lt;/a&gt;&lt;br&gt;
This post is not a tutorial for "integrating Jev". It shares one judgment: &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jev is worth learning for more than its model — the engineering idea that high-frequency decisions shouldn't ask a generative model to write an essay.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And that idea lands without depending on Jev's service at all.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  1. Decision layer design: three principles + one red line
&lt;/h2&gt;

&lt;p&gt;Baize is a sidecar AI assistant runtime: a single process deployed next to your business systems, turning OpenAPI / MCP / HTTP plugins into callable tools, with human approval required for write operations. It calls models through OpenAI-compatible APIs, so it cannot read logits — and therefore cannot get calibrated probabilities. The decision layer was designed from the start with no probability-dependent logic.&lt;/p&gt;

&lt;p&gt;Three hard principles:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Pluggable&lt;/strong&gt; — one decision interface; local rules, a local small model, or a remote decision service are all implementations. Default to the cheapest. No dependency on any external service.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Degradable&lt;/strong&gt; — every decision point fails open to the existing rules path; the decision layer is never a single point of failure for a run.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Observable&lt;/strong&gt; — every judgment records its source and any degradation (decide.degraded, decide.tool_narrow events).&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One red line: &lt;strong&gt;no confidence numbers of any kind in judgment results.&lt;/strong&gt; Without calibrated logits, a number that looks like a probability but is not calibrated is more dangerous than no number at all — it invites branches like "auto-execute when p &amp;gt; 0.9".&lt;/p&gt;

&lt;h3&gt;
  
  
  Interface shape: verdicts only
&lt;/h3&gt;

&lt;p&gt;The return shape is deliberately minimal: verdicts are yes / no enums; multi-select returns the chosen items; every answer carries its source (rules, remote, or fallback) and a degraded flag. There is no float confidence field in the interface — a type-level guarantee against treating uncalibrated numbers as confidence.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Chain: errors must be swallowed by degradation
&lt;/h3&gt;

&lt;p&gt;All implementations are chained and tried in order; the first success returns. If all fail, the chain returns the fallback verdict hard-coded by the call site, marked degraded. The chain itself never returns an error — errors are consumed by degradation, never propagate into the main flow. Callers always get a legal enum value; a decision-layer failure looks like "old behavior", never like a crash.&lt;/p&gt;

&lt;p&gt;Easy to miss: &lt;strong&gt;the fallback verdict is hard-coded at the call site, not a config item.&lt;/strong&gt; Failure directions differ per decision point — memory extraction fails open to Yes (spend a call rather than lose a memory), tool narrowing fails open to the full set (spend prefill rather than miss a tool). As a config knob, someone would eventually flip it under cost pressure and silently lose data. Failure direction is a safety property, not an ops parameter.&lt;/p&gt;

&lt;h3&gt;
  
  
  Implementations today
&lt;/h3&gt;

&lt;p&gt;Two tiers on the chain, remote first: when a dedicated decision-model profile is configured, a small model answers under a prompt that forces enum-only output, validated by regex / JSON parsing; unparseable replies abstain and let the chain degrade. Multi-select answers accept only names from the offered set; out-of-set values are dropped. Because structured-output support varies across OpenAI-compatible endpoints, the contract is enforced by prompt + parsing, not by endpoint features.&lt;/p&gt;

&lt;p&gt;With no profile configured (or remote unavailable), the zero-latency rules backstop handles memory extraction only — it abstains everywhere else: short chitchat with no fact cue returns No; explicit fact cues ("remember that", "my number", passwords, addresses, phone numbers), a run of three or more digits, or substantial length returns Yes.&lt;/p&gt;

&lt;p&gt;Want to plug in a self-hosted inference endpoint that can read logits later? Add one more tier to the chain; callers don't change.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Four decision points: moving "judgment" out of "essay-writing"
&lt;/h2&gt;

&lt;p&gt;Four places in Baize were classic waste: paying a generative model to write a paragraph just to get a "yes / no / which one" answer. The decision layer moves each one down.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.1 DP-1: memory-extraction pre-check (the best cost/benefit)
&lt;/h3&gt;

&lt;p&gt;Baize used to fire an LLM call after every successful conversation to extract memory. Most conversations contain nothing worth remembering and return an empty array — but each one paid a full call plus a JSON-parse risk.&lt;/p&gt;

&lt;p&gt;Now, before paying for that call, the engine asks the layer "is this turn worth extracting?": a No skips extraction and records a skip event (marked as decided by the layer) together with the estimated input tokens saved, so the win is measurable; a Yes — or an unavailable layer — extracts exactly as before.&lt;/p&gt;

&lt;p&gt;The probe itself is capped at 1500 characters: the probe must never cost more than the call it tries to save.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.2 DP-2: two-level tool-candidate narrowing (the structural one)
&lt;/h3&gt;

&lt;p&gt;The most worthwhile change in Baize. Its positioning is "give it OpenAPI, get tools automatically"; the more connectors you attach, the more full tool schemas go to the main model every turn — prefill cost grows linearly with connectors. That is a structural cost that scales with your users.&lt;/p&gt;

&lt;p&gt;Narrowing has two levels:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Level 1: system routing (System Targets).&lt;/strong&gt; The decision model picks which backend systems — a handful of connector ids — this turn needs; it may pick several. Choosing a tool among hundreds is hard; choosing among three to five systems is easy. System descriptions are not hand-written: they are auto-derived from each system's tool names and descriptions via within-system term frequency times cross-system IDF, so adding a connector needs no routing-corpus maintenance.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Level 2: in-system keyword prefilter.&lt;/strong&gt; Within the chosen systems, deterministic IDF keyword scoring (tool names weighted 2x, descriptions 1x, stopwords removed, CJK tokenized as bigrams), merged round-robin across systems, narrowed to 16 candidates by default.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Narrowing is not one-size-fits-all; three floors prevent over-pruning:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Category-tool floor:&lt;/strong&gt; for "list the users", a generic pagination tool ranks poorly under IDF — but a list request needs exactly that tool. List / detail intent forces the matching admin read tool into the set.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Auth floor:&lt;/strong&gt; once a system is chosen, its login / current-session primitives are always kept — on a 401 the model must recover the session on the spot, and the user will never say "login".&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Hybrid routing:&lt;/strong&gt; discriminative domain words in the query (e.g. pets, orders) force their systems and can never be removed by the model — the small model can only add on top, so it cannot vote away the only correct system.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Also, main-model calls support a choice constraint (tool_choice=required) via optional interfaces, without breaking existing provider signatures; providers that don't support it fail open to an ordinary call. A constraint is an optimization, not a hard gate.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.3 DP-3: tool-result pruning
&lt;/h3&gt;

&lt;p&gt;In long conversations, a big JSON returned by an earlier tool call keeps bloating the context. When compaction triggers, the layer judges each of the bulkiest tool results "worth keeping verbatim?"; unworthy ones become a short placeholder that preserves the message / tool-call pairing — the model can re-invoke the tool, and the original result stays in the run log.&lt;/p&gt;

&lt;p&gt;Three bounds keep the judgment cheap: only tool results are judged (not ordinary turns), only those over roughly 500 estimated tokens, at most 8 per turn. All failures mean keep everything.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.4 DP-4: tier arbitration
&lt;/h3&gt;

&lt;p&gt;Baize used a pure heuristic (image present? code fences?) to route each turn to light / standard / power tiers. The problem: "help me refactor this function" contains no code fences, gets classified as light, and the strong model never runs.&lt;/p&gt;

&lt;p&gt;DP-4 does not replace the heuristic. It only consults the layer when the heuristic lands on the ambiguous standard tier, the turn is long enough (at least 400 characters), and routing is Auto — asking for a light / power choice. Short turns are not worth asking: "hello" barely differs across tiers, and every consult costs latency.&lt;/p&gt;

&lt;p&gt;DP-4's judgment goes through the same chain: a dedicated decision tier makes the two-way pick when available; on degradation, the heuristic's standard tier stands — the layer advises, never rewrites the routing outcome.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Results and validation
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Reproducible benchmark:&lt;/strong&gt; 37 real read-only business requests across 3 connected backends (390 tools total), model DeepSeek-Flash, stability run of 5 rounds (185 requests). With the default top-16: run success rate &lt;strong&gt;99.5% (184/185)&lt;/strong&gt;, average turn-0 prompt &lt;strong&gt;~3,090 tokens&lt;/strong&gt;, about &lt;strong&gt;34% lower&lt;/strong&gt; than top-32; sending all 390 tools directly measured &lt;strong&gt;~85k&lt;/strong&gt; turn-0 prompt tokens. The dataset and scripts live in scripts/tool-routing-eval — reproducible.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Default narrowing:&lt;/strong&gt; convergence kicks in above the threshold (default 12 tools); 16 candidates is the default cap — the measured cost / success sweet spot.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Regression harness:&lt;/strong&gt; sweep (width), stability, and analyze (offline) scripts that restore your previous configuration when they finish. An optimization that changes behavior should only go live with shadow data first, then flip shadow off.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything ships with the decision layer off by default and enabled per point: the master switch (decide_enabled) defaults to off, each decision point has its own knob, all controlled through the existing hot-reload settings — no wiring changes.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Honest boundaries
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;No calibrated probabilities.&lt;/strong&gt; Baize calls OpenAI-compatible APIs and cannot read logits. The layer only ever returns enums, and there is no "auto-execute above a probability threshold" feature. Write approval stays deterministic rules + human approval. That red line does not bend for Jev.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;This is not "integrating Jev".&lt;/strong&gt; The decision layer depends on no external decision service. If you later point your inference at a local vLLM running a decision plugin, the remote implementation works as-is — a free win, not an architectural prerequisite.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  5. Closing
&lt;/h2&gt;

&lt;p&gt;Jev made one thing click for me: &lt;strong&gt;models matter, of course — but plenty of high-frequency little decisions shouldn't ask a generative model to write an essay every time.&lt;/strong&gt; Pulling "judgment" out of "generation" into a pluggable, degradable, observable layer is engineering any agent project can do — it does not depend on any company's window.&lt;/p&gt;

&lt;p&gt;Baize's decision layer is live in the main branch (internal/decide + four decision points). Come take a look, run it, or open an issue:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;GitHub: &lt;a href="https://github.com/rebornace/baize" rel="noopener noreferrer"&gt;https://github.com/rebornace/baize&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;More: the README's decision-layer section, and the bilingual developer docs under docs/developers&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you have hit the same wall — tool schemas going to the model every turn while prefill gets more expensive — I'd love to hear your solution. I'll keep following up on feedback and issues.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
      <category>jev</category>
    </item>
    <item>
      <title>Building a Lightweight AI Agent in Go: Baize's Architecture and Trade-offs</title>
      <dc:creator>rebornace</dc:creator>
      <pubDate>Thu, 17 Sep 2026 02:06:02 +0000</pubDate>
      <link>https://dev.to/rebornace/building-a-lightweight-ai-agent-in-go-baizes-architecture-and-trade-offs-4c0k</link>
      <guid>https://dev.to/rebornace/building-a-lightweight-ai-agent-in-go-baizes-architecture-and-trade-offs-4c0k</guid>
      <description>

&lt;h2&gt;
  
  
  Preface
&lt;/h2&gt;

&lt;p&gt;Baize is a "sidecar" AI assistant runtime: a single process that sits beside services you already run, turns API documentation into tools the assistant can call, and pauses important writes until a person approves them. Stop it and it leaves almost nothing behind.&lt;/p&gt;

&lt;p&gt;One sentence positioning: &lt;strong&gt;it's a runtime, not a framework.&lt;/strong&gt; That positioning drives every architecture decision below.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Why Go
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Single binary, zero dependencies
&lt;/h3&gt;

&lt;p&gt;The core requirement of sidecar deployment is "one process, copy it over, run it." A Go binary needs no interpreter, no dependency installation, no virtualenv on the target machine. For an assistant meant to live in enterprise environments, that's the lowest-cost delivery form.&lt;/p&gt;

&lt;h3&gt;
  
  
  Goroutine concurrency
&lt;/h3&gt;

&lt;p&gt;One assistant process serves multiple entry points at once: the web console, signed alert/ticket ingress, and IM channels. Goroutines make "one process handling many sessions concurrently" straightforward — and parallel tool calls later become almost free.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cross-compilation
&lt;/h3&gt;

&lt;p&gt;Enterprise environments run everything: Windows, Linux, macOS, ARM. A single &lt;code&gt;GOOS=linux GOARCH=arm64 go build&lt;/code&gt; produces a binary for the target platform, with no toolchain setup on the destination machine.&lt;/p&gt;

&lt;h3&gt;
  
  
  Static typing and tool contracts
&lt;/h3&gt;

&lt;p&gt;Tool input schemas come from OpenAPI documents; mapping them onto Go's strong types catches many errors at compile time. For a long-running daemon, that saves a lot of operational pain.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Architecture Overview
&lt;/h2&gt;

&lt;p&gt;A clean three-layer structure: &lt;strong&gt;core loop → tool router → executors&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Users / Channels ──► Agent Core Loop (think → pick tool → run → report)
                           │
                           ▼
                     Tool Router (Registry)
                           │
          ┌────────────────┼────────────────┐
    OpenAPI connector    HTTP plugin      MCP connector
          └────────────────┼────────────────┘
                           ▼
              Invoker closures (registered in Registry)
                           │
                     [ HITL approval gate ]
                           │
             ┌─────────────┴─────────────┐
       Direct execution          HTTP callback executor
     (plugin / proxy)            (callback to your side)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Core loop&lt;/strong&gt; (&lt;code&gt;internal/run&lt;/code&gt;): LLM thinks → picks a tool → executes → reports. The event stream (&lt;code&gt;llm.thinking&lt;/code&gt;, &lt;code&gt;llm.tool_call&lt;/code&gt;, &lt;code&gt;tool.result&lt;/code&gt;) is persisted, so the console can follow along in real time.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Tool router&lt;/strong&gt; (&lt;code&gt;internal/tool&lt;/code&gt;): a mutex-protected map that registers not functions but "tool contracts + invoker closures".&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Executors&lt;/strong&gt; (&lt;code&gt;internal/connector&lt;/code&gt;): tools enter through three sources — OpenAPI docs, HTTP plugins, MCP tool servers — plus a "callback execution" mode that hands execution back to your side.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  3. Key Design Decisions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Why HTTP callback instead of a plugin protocol
&lt;/h3&gt;

&lt;p&gt;This is Baize's most important trade-off.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The problem with plugin protocols&lt;/strong&gt;: in-process plugins (Go plugins, shared libraries, language-binding SDKs) require the plugin to compile with the host process — language, version, and ABI must all align. Most enterprise systems aren't written in Go: legacy systems, Java/.NET/Python services can't load an in-process plugin at all. And even when they could, upgrading a plugin means restarting the process — "sidecar in, clean out" is gone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What Baize does instead&lt;/strong&gt;: it doesn't execute the tool itself; it POSTs the invocation to your own endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tool"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"create_ticket"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"arguments"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"run_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"run_xxx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"agent_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"agent_xxx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"idempotency_key"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"uuid-xxx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"callback_urls"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"event"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://your-service/baize-events"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your side executes it and returns the result. Benefits: language-agnostic, process-isolated, auditable; &lt;code&gt;idempotency_key&lt;/code&gt; makes retries safe (no duplicate execution); &lt;code&gt;callback_urls&lt;/code&gt; lets your side keep driving follow-up actions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The cost&lt;/strong&gt;: one extra network round-trip, the callback endpoint must be reachable, and to prevent forged callbacks you need signed requests (Baize uses callback signing with a TTL against replay).&lt;/p&gt;

&lt;h3&gt;
  
  
  Dynamic tool registration and discovery
&lt;/h3&gt;

&lt;p&gt;The registry (&lt;code&gt;tool.Registry&lt;/code&gt;) is the core data structure: a &lt;code&gt;sync.RWMutex&lt;/code&gt; guarding a map, with runtime register/unregister and per-connector bulk unregister — adding a tool or disabling a connector never requires a restart.&lt;/p&gt;

&lt;p&gt;Three tool sources share one registration path:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;OpenAPI docs&lt;/strong&gt;: import Swagger/OpenAPI/Postman-style docs, each operation becomes a tool;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;HTTP plugins&lt;/strong&gt;: a small companion service declares "what tools exist and how to invoke them";&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;MCP tool servers&lt;/strong&gt;: connect to external tool ecosystems as an MCP client.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Security policy is baked into each entry at registration time: &lt;code&gt;require_approval&lt;/code&gt; (needs a human), &lt;code&gt;require_login&lt;/code&gt; (needs a session), &lt;code&gt;security_schemes&lt;/code&gt; (which auth scheme to use). &lt;strong&gt;Security policy is decided at registration, not asked at execution time&lt;/strong&gt; — this is the precondition for letting the assistant actually act.&lt;/p&gt;

&lt;p&gt;Discovery is trivial: &lt;code&gt;Registry.List()&lt;/code&gt; / &lt;code&gt;Registry.Specs()&lt;/code&gt; feed the model's tool list, visible live in the console.&lt;/p&gt;

&lt;h3&gt;
  
  
  Graceful degradation on failure
&lt;/h3&gt;

&lt;p&gt;Failures are the norm in AI agents, so degradation design matters more than the happy path:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Timeout guardrail&lt;/strong&gt;: every tool invocation is bound to &lt;code&gt;context.WithTimeout&lt;/code&gt; (default 60s, configurable);&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Failure is content&lt;/strong&gt;: &lt;code&gt;Invoker&lt;/code&gt; returns &lt;code&gt;(content, isError, err)&lt;/code&gt; — &lt;code&gt;err&lt;/code&gt; is an infrastructure failure (timeout, network), &lt;code&gt;isError&lt;/code&gt; is a business-side failure. Both flow back to the model as structured content, so the model can retry, switch tools, or explain to the user — instead of crashing the whole session;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Approval rejection is not a crash&lt;/strong&gt;: when a human rejects a write, the run settles into an explicit "rejected" terminal state with a trail, no panic;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Everything is observable&lt;/strong&gt;: the &lt;code&gt;llm.tool_call → tool.result&lt;/code&gt; event stream is persisted, so any problem can be traced step by step;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Context compaction&lt;/strong&gt;: long sessions get rolling summaries so quality doesn't degrade as the thread grows.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  4. Go vs Python / Node.js for Agent Scenarios
&lt;/h2&gt;

&lt;p&gt;Let's be honest first: &lt;strong&gt;Python is the best choice in the AI/Agent ecosystem&lt;/strong&gt;. LangChain, LlamaIndex and most reference implementations live there. If your goal is fast experimentation and deep reuse of the LLM ecosystem, Python has no rival.&lt;/p&gt;

&lt;p&gt;Baize chose Go because its positioning is different:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Go&lt;/th&gt;
&lt;th&gt;Python&lt;/th&gt;
&lt;th&gt;Node.js&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Deployment&lt;/td&gt;
&lt;td&gt;Single binary, zero deps&lt;/td&gt;
&lt;td&gt;Interpreter + deps/venv&lt;/td&gt;
&lt;td&gt;Node runtime + node_modules&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resource footprint&lt;/td&gt;
&lt;td&gt;Low; one resident process is cheap&lt;/td&gt;
&lt;td&gt;Higher; resident processes need care&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Concurrency&lt;/td&gt;
&lt;td&gt;Native goroutines&lt;/td&gt;
&lt;td&gt;GIL-limited; multi-process/async&lt;/td&gt;
&lt;td&gt;Event loop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Type safety&lt;/td&gt;
&lt;td&gt;Static, compile-time checks&lt;/td&gt;
&lt;td&gt;Dynamic, found at runtime&lt;/td&gt;
&lt;td&gt;Dynamic / TypeScript&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LLM ecosystem&lt;/td&gt;
&lt;td&gt;Newer, catching up fast&lt;/td&gt;
&lt;td&gt;Richest&lt;/td&gt;
&lt;td&gt;Rich&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cross-platform&lt;/td&gt;
&lt;td&gt;Cross-compile to all platforms&lt;/td&gt;
&lt;td&gt;Needs interpreter on target&lt;/td&gt;
&lt;td&gt;Needs Node on target&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The conclusion isn't "Go is better than Python" — it's &lt;strong&gt;positioning decides the language&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Goal: framework / fast experimentation → Python;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Goal: sidecar, resident, one-command deployment to enterprise environments, running on modest hardware for a long time → Go's advantages in deployment and resource usage are hard to replace.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  5. Core Code Snippets (Go)
&lt;/h2&gt;

&lt;p&gt;All snippets are from the project source, lightly trimmed. Each comes with one line on what problem it solves.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. A tool = contract + invoker closure
&lt;/h3&gt;

&lt;p&gt;Modeling a "tool" as "a contract the model sees (Spec) + an invoker closure injected by the connector" fully decouples routing from execution:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;Invoker&lt;/span&gt; &lt;span class="k"&gt;func&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt; &lt;span class="k"&gt;map&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="n"&gt;any&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="k"&gt;map&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="n"&gt;any&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;isError&lt;/span&gt; &lt;span class="kt"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;err&lt;/span&gt; &lt;span class="kt"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;type&lt;/span&gt; &lt;span class="n"&gt;Meta&lt;/span&gt; &lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;Spec&lt;/span&gt;            &lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ToolSpec&lt;/span&gt;
    &lt;span class="n"&gt;ConnectorID&lt;/span&gt;     &lt;span class="kt"&gt;string&lt;/span&gt;
    &lt;span class="n"&gt;Method&lt;/span&gt;          &lt;span class="kt"&gt;string&lt;/span&gt;
    &lt;span class="n"&gt;Path&lt;/span&gt;            &lt;span class="kt"&gt;string&lt;/span&gt;
    &lt;span class="n"&gt;RequireLogin&lt;/span&gt;    &lt;span class="kt"&gt;bool&lt;/span&gt;
    &lt;span class="n"&gt;SecuritySchemes&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. Dynamic registration with policy baked in
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;require_approval&lt;/code&gt; / &lt;code&gt;require_login&lt;/code&gt; are written into the entry at registration; the tool list is "hot" — adding/removing connectors never requires a restart:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;Registry&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="n"&gt;RegisterMeta&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;meta&lt;/span&gt; &lt;span class="n"&gt;Meta&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;inv&lt;/span&gt; &lt;span class="n"&gt;Invoker&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;requireApproval&lt;/span&gt; &lt;span class="kt"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;mu&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Lock&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;defer&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;mu&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Unlock&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Spec&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Name&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;spec&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;            &lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Spec&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;invoker&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;         &lt;span class="n"&gt;inv&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;requireApproval&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;requireApproval&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;requireLogin&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;    &lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;RequireLogin&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;connectorID&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;     &lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ConnectorID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;method&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;          &lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Method&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;            &lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3. HTTP callback executor
&lt;/h3&gt;

&lt;p&gt;Hands execution back to your side; the idempotency key makes network retries safe:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="k"&gt;map&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="n"&gt;any&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="s"&gt;"tool"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;            &lt;span class="n"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s"&gt;"arguments"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;       &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s"&gt;"run_id"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;          &lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;RunID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s"&gt;"agent_id"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;        &lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;AgentID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="s"&gt;"idempotency_key"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;IdempotencyKey&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;strings&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;TrimSpace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CallbackEventURL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="s"&gt;""&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"callback_urls"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;map&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="kt"&gt;string&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="n"&gt;any&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="s"&gt;"event"&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;meta&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CallbackEventURL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="n"&gt;rawPayload&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Marshal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;http&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;NewRequestWithContext&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;http&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MethodPost&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bytes&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;NewReader&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rawPayload&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  4. Writes automatically enter the approval gate
&lt;/h3&gt;

&lt;p&gt;Non-GET/HEAD/OPTIONS operations are automatically flagged "needs approval" at registration, and only run after a human clicks approve/reject in the console:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="n"&gt;needApproval&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;RequireApproval&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;requireApprovalMutating&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;isMutatingMethod&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Method&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Source&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;store&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ToolSourceSpec&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;needApproval&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="no"&gt;true&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  5. Timeouts and failure degradation
&lt;/h3&gt;

&lt;p&gt;A timeout guardrail plus "failure is content" semantics keeps one bad tool call from blowing up the whole session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="n"&gt;toolCtx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;toolCancel&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;WithTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;toolTimeout&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;span class="k"&gt;defer&lt;/span&gt; &lt;span class="n"&gt;toolCancel&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;isError&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;invErr&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Tools&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Invoke&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;toolCtx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ToolName&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Arguments&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;invErr&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c"&gt;// Infrastructure failure (timeout/network): persist and close the round&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;finalizeFailedRun&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;runID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;invErr&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="c"&gt;// When isError is true, the failure flows back to the model as content;&lt;/span&gt;
&lt;span class="c"&gt;// the model decides whether to retry or explain.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;Baize is still early. The trade-offs above are far from "optimal" — especially the approval UX, channel adapters, and executor extensibility. If you have real-world scenarios, I'd love to hear them.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Repo: &lt;a href="https://github.com/rebornace/baize" rel="noopener noreferrer"&gt;https://github.com/rebornace/baize&lt;/a&gt; (MIT)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Issues: open a discussion with your scenario — I'll follow up.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




</description>
      <category>ai</category>
      <category>agents</category>
      <category>go</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Add AI to Your Legacy System in 5 Minutes: A Baize Hands-On Tutorial (Spring Boot Example)</title>
      <dc:creator>rebornace</dc:creator>
      <pubDate>Wed, 16 Sep 2026 04:34:42 +0000</pubDate>
      <link>https://dev.to/rebornace/add-ai-to-your-legacy-system-in-5-minutes-a-baize-hands-on-tutorial-spring-boot-example-4i0p</link>
      <guid>https://dev.to/rebornace/add-ai-to-your-legacy-system-in-5-minutes-a-baize-hands-on-tutorial-spring-boot-example-4i0p</guid>
      <description>&lt;p&gt;Got a backend system that's been running for years — piles of APIs, complex business logic, and the thought of adding AI feels like a massive undertaking? This tutorial shows you how to hook up an AI assistant to your existing system in 5 minutes using &lt;strong&gt;Baize&lt;/strong&gt; — no refactoring, no core code changes, at most one new endpoint.&lt;/p&gt;

&lt;p&gt;Baize is an open-source AI Agent runtime written in Go. Its core ideas: &lt;strong&gt;sidecar deployment, APIs-as-tools, and human-in-the-loop approval for critical operations&lt;/strong&gt;. It connects to business systems over HTTP, so it doesn't matter what your backend is — Spring Boot, Django, Express, Gin, ThinkPHP — if it has APIs, it can connect. The AI understands user intent, calls the right tools automatically, and pauses for human confirmation on sensitive operations.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Language-agnostic, but we'll use Spring Boot as the main example.&lt;/strong&gt; The principles are fully universal — for other frameworks, just swap the Controller code in Step 3 for your language of choice. The Baize config file stays exactly the same.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;p&gt;To follow along, you'll need two things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A backend service&lt;/strong&gt; (we'll use a Spring Boot ticket system as an example, but feel free to use your own project)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Baize binary&lt;/strong&gt; (a single executable, zero dependencies)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;You can clone the repo which includes a demo ticket system and Baize configs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/rebornace/baize.git
&lt;span class="nb"&gt;cd &lt;/span&gt;baize
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;The Spring Boot code examples in this tutorial use Java 17 + Spring Boot 3.x. Readers using other languages or frameworks can focus on the configuration and protocol sections, and map the code examples to their own stack.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Step 1: Download Baize and Start It with One Command
&lt;/h2&gt;

&lt;p&gt;Baize is a single static binary written in Go — no runtime to install.&lt;/p&gt;

&lt;h3&gt;
  
  
  Option 1: Download the Binary (Recommended)
&lt;/h3&gt;

&lt;p&gt;Grab the latest release for your platform from &lt;a href="https://github.com/rebornace/baize/releases" rel="noopener noreferrer"&gt;GitHub Releases&lt;/a&gt;, unzip it, and you're ready to go.&lt;/p&gt;

&lt;h3&gt;
  
  
  Option 2: Build from Source
&lt;/h3&gt;

&lt;p&gt;If you have Go 1.25+ locally, you can build it yourself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/rebornace/baize.git
&lt;span class="nb"&gt;cd &lt;/span&gt;baize
go build &lt;span class="nt"&gt;-o&lt;/span&gt; baize ./cmd/baize
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Start in Demo Mode
&lt;/h3&gt;

&lt;p&gt;Baize comes with a &lt;strong&gt;demo mode&lt;/strong&gt; that uses a Mock LLM (no API key needed) and a built-in demo ticket system — perfect for getting the full flow working first.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Windows&lt;/span&gt;
.&lt;span class="se"&gt;\b&lt;/span&gt;aize demo

&lt;span class="c"&gt;# macOS / Linux&lt;/span&gt;
./baize demo
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When you see output like this, it's running:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;baize demo listening on :8080
mock ticket server listening on :18080
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Open your browser and go to &lt;strong&gt;&lt;a href="http://localhost:8080/ui" rel="noopener noreferrer"&gt;http://localhost:8080/ui&lt;/a&gt;&lt;/strong&gt; — you'll see Baize's web dashboard.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;For production, use &lt;code&gt;baize start&lt;/code&gt; with a config file and a real LLM API key. We'll cover that later.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  No Config File? No Problem — Configure via the Web UI
&lt;/h3&gt;

&lt;p&gt;Baize supports &lt;strong&gt;zero-config startup&lt;/strong&gt; — just run &lt;code&gt;baize start&lt;/code&gt; without any config file, then open the Web UI and set everything up interactively: models, connectors, tools, approval rules — all of it. Great if you prefer a graphical interface over editing YAML by hand.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Windows PowerShell&lt;/span&gt;
&lt;span class="nv"&gt;$env&lt;/span&gt;:BAIZE_API_KEY &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"your-api-key"&lt;/span&gt;
.&lt;span class="se"&gt;\b&lt;/span&gt;aize start

&lt;span class="c"&gt;# macOS / Linux&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;BAIZE_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-api-key"&lt;/span&gt;
./baize start
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once it's running, go to &lt;a href="http://localhost:8080/ui" rel="noopener noreferrer"&gt;http://localhost:8080/ui&lt;/a&gt; → &lt;strong&gt;Settings&lt;/strong&gt; in the sidebar. You can configure your LLM provider, add connectors, and manage tools directly from the page. Changes take effect immediately — no restart needed.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This tutorial uses YAML config files for clarity in the step-by-step flow. Both approaches produce the same result — pick whichever you prefer.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Step 2: Declare Tools in a Config File
&lt;/h2&gt;

&lt;p&gt;The core of how Baize connects to business systems is the &lt;strong&gt;Connector&lt;/strong&gt;. The most common approach is the &lt;strong&gt;OpenAPI Connector&lt;/strong&gt; — if you have an OpenAPI/Swagger spec, Baize automatically turns every endpoint into a tool the AI can use.&lt;/p&gt;

&lt;h3&gt;
  
  
  2.1 Prepare the OpenAPI Spec
&lt;/h3&gt;

&lt;p&gt;Let's say your Spring Boot ticket system has these endpoints:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Path&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;th&gt;operationId&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GET&lt;/td&gt;
&lt;td&gt;/tickets&lt;/td&gt;
&lt;td&gt;List all tickets&lt;/td&gt;
&lt;td&gt;&lt;code&gt;list_tickets&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;POST&lt;/td&gt;
&lt;td&gt;/tickets&lt;/td&gt;
&lt;td&gt;Create a ticket&lt;/td&gt;
&lt;td&gt;&lt;code&gt;create_ticket&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GET&lt;/td&gt;
&lt;td&gt;/tickets/{id}&lt;/td&gt;
&lt;td&gt;Get ticket details&lt;/td&gt;
&lt;td&gt;&lt;code&gt;get_ticket&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PATCH&lt;/td&gt;
&lt;td&gt;/tickets/{id}&lt;/td&gt;
&lt;td&gt;Update ticket status&lt;/td&gt;
&lt;td&gt;&lt;code&gt;update_ticket_status&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Here's the corresponding OpenAPI spec (&lt;code&gt;ticket-api.yaml&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;openapi&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;3.0.3&lt;/span&gt;
&lt;span class="na"&gt;info&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Ticket System API&lt;/span&gt;
  &lt;span class="na"&gt;version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;1.0.0&lt;/span&gt;
&lt;span class="na"&gt;paths&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;/tickets&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;operationId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;list_tickets&lt;/span&gt;
      &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;List all tickets&lt;/span&gt;
      &lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;200"&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;List of tickets&lt;/span&gt;
    &lt;span class="na"&gt;post&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;operationId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;create_ticket&lt;/span&gt;
      &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Create a ticket&lt;/span&gt;
      &lt;span class="na"&gt;requestBody&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
        &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;application/json&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;object&lt;/span&gt;
              &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
              &lt;span class="na"&gt;properties&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;string&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;Ticket title&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
                &lt;span class="na"&gt;priority&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;string&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;Priority level&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
      &lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;201"&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Created successfully&lt;/span&gt;
  &lt;span class="s"&gt;/tickets/{id}&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;operationId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;get_ticket&lt;/span&gt;
      &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Get ticket details&lt;/span&gt;
      &lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;id&lt;/span&gt;
          &lt;span class="na"&gt;in&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;path&lt;/span&gt;
          &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
          &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;string&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
      &lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;200"&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Ticket details&lt;/span&gt;
    &lt;span class="na"&gt;patch&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;operationId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;update_ticket_status&lt;/span&gt;
      &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Update ticket status&lt;/span&gt;
      &lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;id&lt;/span&gt;
          &lt;span class="na"&gt;in&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;path&lt;/span&gt;
          &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
          &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;string&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
      &lt;span class="na"&gt;requestBody&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
        &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;application/json&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;object&lt;/span&gt;
              &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;status&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
              &lt;span class="na"&gt;properties&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;string&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;New status&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
      &lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;200"&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Updated successfully&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;If your Spring Boot project already has springdoc-openapi or Swagger integrated, you can export the JSON directly from &lt;code&gt;/v3/api-docs&lt;/code&gt; — Baize understands both YAML and JSON.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  2.2 Write the Baize Config File
&lt;/h3&gt;

&lt;p&gt;Create &lt;code&gt;baize-config.yaml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;listen&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;:8080"&lt;/span&gt;
&lt;span class="na"&gt;store&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;driver&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sqlite&lt;/span&gt;
  &lt;span class="na"&gt;sqlite_path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;./data/baize.db&lt;/span&gt;
&lt;span class="na"&gt;ui&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;enabled&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;

&lt;span class="na"&gt;llm&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openai_compatible&lt;/span&gt;
  &lt;span class="na"&gt;base_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://api.deepseek.com&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deepseek-chat&lt;/span&gt;
  &lt;span class="na"&gt;api_key_env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;BAIZE_API_KEY&lt;/span&gt;

&lt;span class="na"&gt;agent&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ticket-agent&lt;/span&gt;
  &lt;span class="na"&gt;system&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;are&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ticket&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;assistant.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;You&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;can&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;only&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;access&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ticket&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;through&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;tools."&lt;/span&gt;

&lt;span class="na"&gt;connector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ticket-api&lt;/span&gt;
  &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openapi&lt;/span&gt;
  &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;./ticket-api.yaml&lt;/span&gt;
  &lt;span class="na"&gt;base_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://localhost:8081&lt;/span&gt;   &lt;span class="c1"&gt;# Your Spring Boot service URL&lt;/span&gt;
  &lt;span class="na"&gt;require_approval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;create_ticket&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;update_ticket_status&lt;/span&gt;

&lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;max_steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;16&lt;/span&gt;
  &lt;span class="na"&gt;tool_timeout_sec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;60&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Key configuration points:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;connector.type: openapi&lt;/code&gt;&lt;/strong&gt; — Auto-discovers tools from the OpenAPI spec&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;connector.spec&lt;/code&gt;&lt;/strong&gt; — Path to your OpenAPI spec file&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;connector.base_url&lt;/code&gt;&lt;/strong&gt; — Your backend service URL; the AI calls this address when using tools&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;require_approval&lt;/code&gt;&lt;/strong&gt; — Tools listed here require human confirmation before execution (recommended for all write operations)&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2.3 Start with the Config File
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Set your LLM API Key (DeepSeek example)&lt;/span&gt;
&lt;span class="c"&gt;# Windows PowerShell&lt;/span&gt;
&lt;span class="nv"&gt;$env&lt;/span&gt;:BAIZE_API_KEY &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"your-api-key"&lt;/span&gt;

&lt;span class="c"&gt;# macOS / Linux&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;BAIZE_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-api-key"&lt;/span&gt;

&lt;span class="c"&gt;# Start&lt;/span&gt;
./baize start &lt;span class="nt"&gt;-c&lt;/span&gt; baize-config.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Open &lt;a href="http://localhost:8080/ui" rel="noopener noreferrer"&gt;http://localhost:8080/ui&lt;/a&gt; → click "Tools" in the sidebar, and you'll see all 4 ticket endpoints listed as AI-usable tools.&lt;/p&gt;

&lt;p&gt;At this step, &lt;strong&gt;you haven't changed a single line of your business code&lt;/strong&gt; — Baize is just another caller hitting your APIs, no different from a frontend app.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 3: Add a Callback Endpoint for Full Control (Spring Boot Example)
&lt;/h2&gt;

&lt;p&gt;The OpenAPI approach above works great when "the APIs already exist and the AI can call them directly." But sometimes you need more control:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You want to add custom logic before/after AI calls (audit logging, rate limiting, parameter validation)&lt;/li&gt;
&lt;li&gt;The tool logic is complex and not suitable for exposing as a direct API&lt;/li&gt;
&lt;li&gt;You don't want the AI touching internal services directly&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is where &lt;strong&gt;Execution Callback&lt;/strong&gt; mode comes in: instead of calling your business APIs directly, Baize sends a message saying "which tool to call and with what arguments" to an endpoint you specify. You decide how to execute it. The protocol is dead simple — any language that can serve HTTP can implement it.&lt;/p&gt;

&lt;p&gt;Below is a Spring Boot example; the same pattern applies to Django, Express, Gin, and any other framework.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.1 Add an &lt;code&gt;/execute&lt;/code&gt; Endpoint (Spring Boot Version)
&lt;/h3&gt;

&lt;p&gt;Add one Controller with one method to your Spring Boot project:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="kn"&gt;package&lt;/span&gt; &lt;span class="nn"&gt;com.example.ticket.controller&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;com.example.ticket.service.TicketService&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;org.springframework.web.bind.annotation.*&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;java.util.Map&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="nd"&gt;@RestController&lt;/span&gt;
&lt;span class="nd"&gt;@RequestMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/baize"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;BaizeCallbackController&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;

    &lt;span class="kd"&gt;private&lt;/span&gt; &lt;span class="kd"&gt;final&lt;/span&gt; &lt;span class="nc"&gt;TicketService&lt;/span&gt; &lt;span class="n"&gt;ticketService&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nf"&gt;BaizeCallbackController&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;TicketService&lt;/span&gt; &lt;span class="n"&gt;ticketService&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;ticketService&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ticketService&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;

    &lt;span class="nd"&gt;@PostMapping&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/execute"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
    &lt;span class="kd"&gt;public&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;Object&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;
            &lt;span class="nd"&gt;@RequestHeader&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"X-Baize-Protocol"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;protocol&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;
            &lt;span class="nd"&gt;@RequestBody&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;Object&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;

        &lt;span class="c1"&gt;// Protocol check&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;(!&lt;/span&gt;&lt;span class="s"&gt;"v0"&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;equals&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;protocol&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;of&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;
                &lt;span class="s"&gt;"content"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;of&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"error"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"unsupported protocol"&lt;/span&gt;&lt;span class="o"&gt;),&lt;/span&gt;
                &lt;span class="s"&gt;"is_error"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
            &lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;

        &lt;span class="nc"&gt;String&lt;/span&gt; &lt;span class="n"&gt;tool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"tool"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="nd"&gt;@SuppressWarnings&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"unchecked"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
        &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;Object&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;Object&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;)&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"arguments"&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;

        &lt;span class="c1"&gt;// Dispatch to your business logic based on tool name&lt;/span&gt;
        &lt;span class="nc"&gt;Object&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;switch&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tool&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="s"&gt;"list_tickets"&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;ticketService&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;listTickets&lt;/span&gt;&lt;span class="o"&gt;();&lt;/span&gt;
            &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="s"&gt;"create_ticket"&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;ticketService&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;createTicket&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;
                &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"title"&lt;/span&gt;&lt;span class="o"&gt;),&lt;/span&gt;
                &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"priority"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;);&lt;/span&gt;
            &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="s"&gt;"get_ticket"&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;ticketService&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getTicket&lt;/span&gt;&lt;span class="o"&gt;((&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"id"&lt;/span&gt;&lt;span class="o"&gt;));&lt;/span&gt;
            &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="s"&gt;"update_ticket_status"&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;ticketService&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;updateStatus&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;
                &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"id"&lt;/span&gt;&lt;span class="o"&gt;),&lt;/span&gt;
                &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"status"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
            &lt;span class="o"&gt;);&lt;/span&gt;
            &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;
        &lt;span class="o"&gt;};&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;of&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;
                &lt;span class="s"&gt;"content"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;of&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"error"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"unknown tool: "&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;tool&lt;/span&gt;&lt;span class="o"&gt;),&lt;/span&gt;
                &lt;span class="s"&gt;"is_error"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
            &lt;span class="o"&gt;);&lt;/span&gt;
        &lt;span class="o"&gt;}&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Map&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;of&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;
            &lt;span class="s"&gt;"content"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;
            &lt;span class="s"&gt;"is_error"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
        &lt;span class="o"&gt;);&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it — one Controller, one method, one switch statement for dispatch. You're in full control of the execution logic. Add logging, add authorization checks, add caching — whatever you need.&lt;/p&gt;

&lt;h3&gt;
  
  
  3.2 Update Baize Config to Point to the Callback URL
&lt;/h3&gt;

&lt;p&gt;Change the connector in &lt;code&gt;baize-config.yaml&lt;/code&gt; to callback mode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;connector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ticket-api&lt;/span&gt;
  &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openapi&lt;/span&gt;
  &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;./ticket-api.yaml&lt;/span&gt;
  &lt;span class="na"&gt;base_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;                           &lt;span class="c1"&gt;# Can be empty in callback mode&lt;/span&gt;
  &lt;span class="na"&gt;execution_callback_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://localhost:8081/baize/execute&lt;/span&gt;
  &lt;span class="na"&gt;require_approval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;create_ticket&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;update_ticket_status&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key addition is &lt;strong&gt;&lt;code&gt;execution_callback_url&lt;/code&gt;&lt;/strong&gt; — with this field set, Baize won't call &lt;code&gt;base_url&lt;/code&gt; directly. Instead, it sends tool invocation requests to this callback address.&lt;/p&gt;

&lt;p&gt;You still need the OpenAPI spec file because Baize uses it to discover what tools exist and what parameters they take — that's &lt;strong&gt;tool discovery&lt;/strong&gt;. The actual &lt;strong&gt;tool execution&lt;/strong&gt; is handled by your Spring Boot service.&lt;/p&gt;

&lt;p&gt;After restarting Baize, the user experience looks the same, but the execution path is completely different — every tool call now goes through your &lt;code&gt;/baize/execute&lt;/code&gt; endpoint.&lt;/p&gt;




&lt;h2&gt;
  
  
  Step 4: Send a Message and Watch the AI Call Tools
&lt;/h2&gt;

&lt;p&gt;Once Baize is running, you can talk to the AI in two ways: via the Web UI or the HTTP API.&lt;/p&gt;

&lt;h3&gt;
  
  
  Option 1: Web UI (Recommended for Getting Started)
&lt;/h3&gt;

&lt;p&gt;Open &lt;strong&gt;&lt;a href="http://localhost:8080/ui" rel="noopener noreferrer"&gt;http://localhost:8080/ui&lt;/a&gt;&lt;/strong&gt; and type in the chat box:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Show me all the tickets
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You'll see the AI's reasoning process step by step, including the full &lt;code&gt;list_tickets&lt;/code&gt; tool call — both the request parameters and the returned result.&lt;/p&gt;

&lt;p&gt;Try this next:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Create a ticket with title "Login page loads slowly" and high priority
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because &lt;code&gt;create_ticket&lt;/code&gt; is in the &lt;code&gt;require_approval&lt;/code&gt; list, the AI will pause and wait for your confirmation. Click &lt;strong&gt;Approve&lt;/strong&gt; on the approval card that pops up, and only then will it actually execute the create operation. This is &lt;strong&gt;HITL (Human-in-the-Loop)&lt;/strong&gt; — sensitive operations are never executed blindly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Option 2: HTTP API (Great for Integration)
&lt;/h3&gt;

&lt;p&gt;Baize exposes a full REST control plane that you can call from scripts or other systems.&lt;/p&gt;

&lt;p&gt;Start a new conversation run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST http://localhost:8080/v0/runs &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "agent_id": "ticket-agent",
    "messages": [
      {
        "role": "user",
        "content": "Show me all the tickets"
      }
    ]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response includes a &lt;code&gt;run_id&lt;/code&gt;. Use it to check status and stream events:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Get run result&lt;/span&gt;
curl http://localhost:8080/v0/runs/&lt;span class="o"&gt;{&lt;/span&gt;run_id&lt;span class="o"&gt;}&lt;/span&gt;

&lt;span class="c"&gt;# Stream events (SSE)&lt;/span&gt;
curl http://localhost:8080/v0/runs/&lt;span class="o"&gt;{&lt;/span&gt;run_id&lt;span class="o"&gt;}&lt;/span&gt;/stream
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Going Further
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Adding More Tools
&lt;/h3&gt;

&lt;p&gt;Tools come from the OpenAPI spec. Adding a tool = adding an endpoint to the spec, then restarting Baize (or hot-loading through the settings page).&lt;/p&gt;

&lt;p&gt;For example, if you want to add a "close ticket" tool, add this path to &lt;code&gt;ticket-api.yaml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;  &lt;span class="s"&gt;/tickets/{id}/close&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;post&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;operationId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;close_ticket&lt;/span&gt;
      &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Close a ticket&lt;/span&gt;
      &lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;id&lt;/span&gt;
          &lt;span class="na"&gt;in&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;path&lt;/span&gt;
          &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
          &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;string&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
      &lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;200"&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Closed successfully&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then add &lt;code&gt;close_ticket&lt;/code&gt; to &lt;code&gt;require_approval&lt;/code&gt;, and you're done.&lt;/p&gt;

&lt;h3&gt;
  
  
  Timeouts and Retries
&lt;/h3&gt;

&lt;p&gt;Baize has a default timeout for each tool call, which you can adjust in the config:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;max_steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;16&lt;/span&gt;
  &lt;span class="na"&gt;tool_timeout_sec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;120&lt;/span&gt;    &lt;span class="c1"&gt;# Timeout per tool call, default is 60 seconds&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your business APIs are occasionally flaky, you can handle retries inside your Spring Boot callback endpoint — Spring Retry or Resilience4j both work great. Baize just "sends the request and waits for the result" — retry strategy is up to you.&lt;/p&gt;

&lt;h3&gt;
  
  
  Handling Login / Auth
&lt;/h3&gt;

&lt;p&gt;If your system requires login before calling APIs, Baize has you covered. Configure &lt;code&gt;auth.capture&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;connector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ticket-api&lt;/span&gt;
  &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openapi&lt;/span&gt;
  &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;./ticket-api.yaml&lt;/span&gt;
  &lt;span class="na"&gt;execution_callback_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://localhost:8081/baize/execute&lt;/span&gt;
  &lt;span class="na"&gt;auth&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;mode&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;passthrough&lt;/span&gt;
    &lt;span class="na"&gt;capture&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;tool_name_glob&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*login*"&lt;/span&gt;
      &lt;span class="na"&gt;token_json_paths&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;accessToken"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data.accessToken"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;label_json_paths&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;email"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;header_template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;{{token}}"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the AI detects that login is required, it calls the login tool first, automatically extracts the token from the response, and attaches it to subsequent calls. Login info is stored in the session identity and persists across conversations.&lt;/p&gt;




&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Q: Does Baize have to run in Kubernetes?
&lt;/h3&gt;

&lt;p&gt;No. Baize is a single binary — it runs on bare metal, VMs, Docker, and Kubernetes alike. Put it on the same internal network as your legacy system, or even deploy it on the public network and call back over VPN (security will be weaker, of course).&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: What if my OpenAPI spec is outdated or doesn't cover all endpoints?
&lt;/h3&gt;

&lt;p&gt;Two options:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Use &lt;strong&gt;Execution Callback&lt;/strong&gt; mode — the OpenAPI spec just declares "what tools exist," and you write the actual execution logic yourself&lt;/li&gt;
&lt;li&gt;Use the &lt;strong&gt;HTTP Plugin v0&lt;/strong&gt; protocol — write a small standalone service to add capabilities your original system doesn't have&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Q: Could the AI call APIs randomly and mess up my data?
&lt;/h3&gt;

&lt;p&gt;No. Two layers of protection:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Human approval&lt;/strong&gt; — tools in the &lt;code&gt;require_approval&lt;/code&gt; list pause before execution and wait for someone to click approve in the dashboard&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read-only first&lt;/strong&gt; — it's good practice to start with read-only endpoints open and put all write operations behind approval, then gradually open things up as you gain confidence&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Q: Which LLMs are supported?
&lt;/h3&gt;

&lt;p&gt;Any model compatible with the OpenAI API format works — including OpenAI, DeepSeek, Anthropic (via compatible gateways), local Ollama, and more. Just change &lt;code&gt;llm.base_url&lt;/code&gt; and &lt;code&gt;llm.model&lt;/code&gt; in the config.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: Can it connect to messaging platforms like Slack or Discord?
&lt;/h3&gt;

&lt;p&gt;Yes. Baize's channels are designed as an extensible capability. You can use the Webhook channel to adapt any platform. One assistant, one set of tools, one approval flow — reachable from multiple channels.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: What's the difference between demo mode and production mode?
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Demo mode&lt;/strong&gt; (&lt;code&gt;baize demo&lt;/code&gt;): Uses Mock LLM, no API key needed, includes a built-in demo ticket system — great for trying it out quickly&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Production mode&lt;/strong&gt; (&lt;code&gt;baize start&lt;/code&gt;): Connects to a real LLM, works with real business systems — full feature set&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Wrapping Up
&lt;/h2&gt;

&lt;p&gt;Let's recap the whole flow — really just four steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Download Baize&lt;/strong&gt; → single binary, zero dependencies&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write an OpenAPI spec + configure the connector&lt;/strong&gt; → APIs become tools automatically&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;(Optional) Add an &lt;code&gt;/execute&lt;/code&gt; endpoint&lt;/strong&gt; → full control over execution logic&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chat&lt;/strong&gt; → the AI calls tools automatically, with human approval for sensitive operations&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Baize's design philosophy is "sidecar deployment, clean removal" — it doesn't intrude on your business system, it just adds an intelligent caller. Turn it on when you want to try it, turn it off when you don't. Your legacy system stays untouched.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Open Source Info&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Baize&lt;/strong&gt; · MIT License · Go 1.25 · Zero C dependencies&lt;/p&gt;

&lt;p&gt;GitHub: &lt;a href="https://github.com/rebornace/baize" rel="noopener noreferrer"&gt;https://github.com/rebornace/baize&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Core features: ReAct Agent · OpenAPI Connector · HTTP Plugin v0 · Execution Callback · MCP Bridge/Export · HITL (Human-in-the-Loop) · Built-in Web UI&lt;/p&gt;

&lt;p&gt;Supported LLMs: OpenAI / DeepSeek / any OpenAI-compatible API&lt;/p&gt;

&lt;p&gt;Stars, issues, and PRs are all welcome.&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>go</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
