<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: SAI RAM</title>
    <description>The latest articles on DEV Community by SAI RAM (@sai_ram_0000).</description>
    <link>https://dev.to/sai_ram_0000</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2169650%2F5f4ceeb8-5c63-4c17-85a8-52beb60125a5.jpeg</url>
      <title>DEV Community: SAI RAM</title>
      <link>https://dev.to/sai_ram_0000</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sai_ram_0000"/>
    <language>en</language>
    <item>
      <title>trelix v3.2.2 to v3.3.8: A GitHub App Already Hardened and Running in Production, and a Connector That Never Touches a Pixel</title>
      <dc:creator>SAI RAM</dc:creator>
      <pubDate>Sat, 26 Sep 2026 11:08:08 +0000</pubDate>
      <link>https://dev.to/sai_ram_0000/trelix-v322-to-v338-a-github-app-already-hardened-and-running-in-production-and-a-connector-3hng</link>
      <guid>https://dev.to/sai_ram_0000/trelix-v322-to-v338-a-github-app-already-hardened-and-running-in-production-and-a-connector-3hng</guid>
      <description>&lt;p&gt;&lt;code&gt;docker run --entrypoint trelix-mcp ghcr.io/sairam0424/trelix:3.2.1 --help&lt;/code&gt; returns exit code 127: command not found. Not a typo, not a stale tag — the published 3.2.1 image never contained &lt;code&gt;trelix-mcp&lt;/code&gt; at all. The builder stage copied core &lt;code&gt;trelix&lt;/code&gt; and nothing else, and &lt;code&gt;.dockerignore&lt;/code&gt;'s blanket &lt;code&gt;packages/&lt;/code&gt; exclusion would have blocked the MCP package even if someone had remembered the &lt;code&gt;COPY&lt;/code&gt; line. A second, unrelated defect shipped in the same release: the published &lt;code&gt;trelix-mcp&lt;/code&gt; console script ignored every flag you gave it — &lt;code&gt;--help&lt;/code&gt;, &lt;code&gt;-h&lt;/code&gt;, &lt;code&gt;--version&lt;/code&gt;, anything — and silently launched the real MCP stdio server instead. Both were found the only way they could have been found: by installing the actual published artifacts and running them, not by reading a source tree that had never been wrong.&lt;/p&gt;

&lt;p&gt;That's v3.2.2, and it's where this span starts. This article covers the thirteen tagged releases from v3.2.2 through v3.3.8 — 2026-08-26 through 2026-09-25 — picking up the day after the last article in this series ended at v3.2.1. v3.3.8 is the current release: tagged, pushed, marked "Latest," and — I checked this directly rather than trusting the changelog — live on PyPI right now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Thirteen releases, one throughline
&lt;/h2&gt;

&lt;p&gt;The prior span's throughline was a green test suite that never touched the bug it claimed to cover — a &lt;code&gt;MagicMock&lt;/code&gt; standing in for an embedder, an all-ones attention mask making masked and unmasked math identical, a unit test that asserted a bug as its own specification. That defect class is still present in this span (more on it below), but it's no longer the dominant story. Two things sit ahead of it now.&lt;/p&gt;

&lt;p&gt;The first is that the response to catching defects graduated from "write a regression test" to "build a system that makes this class of defect structurally unable to reach a tag again." The second is quieter and, I think, more interesting: almost everything genuinely new in this span — a diagram connector, an MCP protocol upgrade, a retriever-caching fix, a compression provider, an audit-log pruning mechanism — turned out to be sound at the code level when I went and checked it against source. What kept drifting was the layer just above the code: a renamed class the CHANGELOG never caught up to, a performance number with no artifact behind it, a guarantee that overstates what the code actually does. If the last article's angle was "the code lied to the tests," this one's second angle is closer to "the code is right, and the notes about the code keep being slightly wrong" — which is a more mature failure mode, and one that only shows up once the cruder failure mode has already been closed off.&lt;/p&gt;

&lt;h2&gt;
  
  
  The origin: a manual audit that automated itself
&lt;/h2&gt;

&lt;p&gt;v3.2.2's fixes were correct but ordinary. What matters is what happened next. v3.2.3, two releases later, turned the exact manual "install the real published artifacts and run them" pass that had just caught two live defects into permanent CI: a real-subprocess E2E suite that spawns &lt;code&gt;trelix-mcp&lt;/code&gt; as an actual OS process talking real stdio JSON-RPC, fresh-venv installs of all four published packages, a &lt;code&gt;smoke-test-built-artifacts&lt;/code&gt; job that &lt;code&gt;release.yml&lt;/code&gt;'s &lt;code&gt;publish&lt;/code&gt; job is now hard-gated on, a Docker check that &lt;code&gt;trelix-mcp&lt;/code&gt; is actually present in the built image, and a Helm-lint check that the rendered image tag matches &lt;code&gt;Chart.yaml&lt;/code&gt;. The same release fixed a real, separate bug along the way — a &lt;code&gt;LIKE&lt;/code&gt;-wildcard escaping gap in &lt;code&gt;path_filter&lt;/code&gt;-scoped BM25/grep queries, where unescaped &lt;code&gt;_&lt;/code&gt; and &lt;code&gt;%&lt;/code&gt; in ordinary directory names were being read as SQL wildcards. Not injection (the queries were already parameterized), just a distinct semantics gap, closed with a shared &lt;code&gt;escape_like_pattern()&lt;/code&gt; helper and an &lt;code&gt;ESCAPE '\'&lt;/code&gt; clause.&lt;/p&gt;

&lt;p&gt;The centerpiece is &lt;code&gt;scripts/verify_release.py&lt;/code&gt;, a 742-line script with five independent check functions that each return a list of failures and never raise — so one category's bug can't hide another's results. &lt;code&gt;check_pypi_installs&lt;/code&gt; does a fresh venv per package and a real &lt;code&gt;pip install &amp;lt;pkg&amp;gt;==&amp;lt;version&amp;gt;&lt;/code&gt;, then a real console-script and import smoke test against the version string. &lt;code&gt;check_docker_images&lt;/code&gt; pulls both the plain and &lt;code&gt;-local&lt;/code&gt; tags and runs the exact &lt;code&gt;--entrypoint trelix-mcp --version&lt;/code&gt; command that returned 127 on the unpatched 3.2.1 image — this check exists specifically because that command failed once, silently, in production. &lt;code&gt;check_helm_chart&lt;/code&gt; checks out &lt;code&gt;v&amp;lt;version&amp;gt;&lt;/code&gt; into a throwaway &lt;code&gt;git worktree&lt;/code&gt;, never the caller's own working tree, and runs &lt;code&gt;helm lint&lt;/code&gt;/&lt;code&gt;helm template&lt;/code&gt; across all three store backends. &lt;code&gt;check_github_release_binaries&lt;/code&gt; downloads and &lt;code&gt;file(1)&lt;/code&gt;-checks all four platform assets and executes whichever one matches the current host, printing rather than silently skipping when it can't run a foreign-arch binary. &lt;code&gt;check_security_audit&lt;/code&gt; runs &lt;code&gt;pip-audit --format json&lt;/code&gt; on all four packages in a fresh venv and separately scans every wheel for &lt;code&gt;.env*&lt;/code&gt;, &lt;code&gt;.git&lt;/code&gt;-path, or &lt;code&gt;credential&lt;/code&gt;/&lt;code&gt;secret&lt;/code&gt;-named files that should never ship.&lt;/p&gt;

&lt;p&gt;v3.2.4 wired this permanently into CI via &lt;code&gt;.github/workflows/verify-release.yml&lt;/code&gt;: it waits for both &lt;code&gt;release.yml&lt;/code&gt; and &lt;code&gt;docker-publish.yml&lt;/code&gt; to go green for a pushed tag, then runs &lt;code&gt;verify_release.py&lt;/code&gt; automatically and posts a PASS/FAIL comment, with the manual command still available for ad-hoc re-runs. I read the 99-line workflow directly — it uses &lt;code&gt;env:&lt;/code&gt; passthrough rather than &lt;code&gt;${{ }}&lt;/code&gt; shell interpolation, matching &lt;code&gt;docker-publish.yml&lt;/code&gt;'s own injection-avoidance convention. v3.2.5, the same day, fixed a smaller instance of the identical class of gap: the frozen PyInstaller binary's embedder error message told users to &lt;code&gt;pip install 'trelix[local]'&lt;/code&gt;, which does nothing at all for a binary that never touches the host's Python or pip.&lt;/p&gt;

&lt;p&gt;I ran the one check in this whole dossier that didn't rely on reading source: &lt;code&gt;curl https://pypi.org/pypi/trelix/json&lt;/code&gt;. &lt;code&gt;info.version&lt;/code&gt; is &lt;code&gt;3.3.8&lt;/code&gt;. The release history includes 3.2.2 through 3.2.5, then 3.3.0, 3.3.5, 3.3.6, 3.3.7, 3.3.8 — 3.3.1 through 3.3.4 are conspicuously absent from PyPI's own list, which I take to mean they were likely never tagged as standalone PyPI releases rather than any kind of cover-up; I didn't dig further and I'm flagging it rather than resolving it. &lt;code&gt;trelix-mcp&lt;/code&gt; is confirmed in lockstep at 3.3.8 too, wheel and sdist both uploaded 2026-09-25T09:54 UTC. That's a directly executed, non-paraphrased result — the single strongest claim in this whole article, because it's the one number here that isn't source-reading, it's a live check against the real index.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7sc5vnrk9dto1b27h1yz.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7sc5vnrk9dto1b27h1yz.gif" alt="A person holding a magnifying glass up to their eye, peering closely through it" width="270" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  v3.3.0: the migration that broke things on purpose
&lt;/h2&gt;

&lt;p&gt;Every other release in this span is a patch. v3.3.0 is a deliberate minor bump with four real breaking changes, which is the honest way to signal "this will require you to change something," and I want to give it credit for that framing even where its own migration doc undersells it.&lt;/p&gt;

&lt;p&gt;The four changes: the Python floor moved to &lt;code&gt;&amp;gt;=3.12&lt;/code&gt; across all three packages; the &lt;code&gt;TRELIX_RETRIEVAL_FLARE_MAX_ITER&lt;/code&gt; env alias was removed outright, with zero remaining references anywhere outside docs; &lt;code&gt;AWS_REGION&lt;/code&gt; is now hard-required for the Bedrock backend, raising &lt;code&gt;ValueError&lt;/code&gt; where it used to silently default to &lt;code&gt;us-east-1&lt;/code&gt;; and the Anthropic backend now silently no-ops on a &lt;code&gt;temperature&lt;/code&gt; argument instead of forwarding it, logging a one-time warning instead. That last pair — a required env var and a silently dropped parameter — are exactly the two changes most likely to actually break a deployment on upgrade, and they are, by my own read of &lt;code&gt;docs/migration/v3.2-to-v3.3.md&lt;/code&gt;, the two the migration guide doesn't mention at all. The doc covers three of the six real changes and its own closing section states the LLM-provider and MCP changes are "not yet released" — they shipped the same day as the doc, under the same version tag. If you're upgrading past v3.3.0, read the CHANGELOG, not the migration guide.&lt;/p&gt;

&lt;p&gt;The release also introduces a reasoning-content abstraction across both cloud backends — and here's a naming error worth stating plainly because it'll cost you time if you don't know about it: the CHANGELOG calls this type &lt;code&gt;ReasoningBlock&lt;/code&gt;. The actual class, which I confirmed directly in &lt;code&gt;src/trelix/llm/client.py&lt;/code&gt; (line 29), is named &lt;code&gt;ThinkingBlock&lt;/code&gt;. A repo-wide grep for &lt;code&gt;ReasoningBlock&lt;/code&gt; across &lt;code&gt;src/&lt;/code&gt;, &lt;code&gt;packages/&lt;/code&gt;, and &lt;code&gt;tests/&lt;/code&gt; returns zero matches. &lt;code&gt;ChatResponse.thinking_blocks: list[ThinkingBlock]&lt;/code&gt; plus a back-compat &lt;code&gt;ChatResponse.thinking: str | None&lt;/code&gt; is what actually ships, and both the Anthropic and Bedrock backends build &lt;code&gt;ThinkingBlock&lt;/code&gt; instances correctly, including the previously-dropped &lt;code&gt;redacted_thinking&lt;/code&gt; variant. If you grep the codebase for the CHANGELOG's own name, you'll find nothing — use &lt;code&gt;ThinkingBlock&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Separately, MCP gained SEP-2322 support: &lt;code&gt;AgentResult.needs_input&lt;/code&gt;, set only on an explicit &lt;code&gt;CLARIFY&lt;/code&gt; action from the agent loop, lets the &lt;code&gt;ask_agent&lt;/code&gt; MCP tool wrap an &lt;code&gt;ElicitRequest&lt;/code&gt; inside an &lt;code&gt;InputRequiredResult&lt;/code&gt; when the agent genuinely needs clarification, resuming later via a dedicated answer-extraction path. Non-elicitation callers are unaffected — it's additive. This is the most solidly verified claim in the release, and also the one place the CHANGELOG's own claim doesn't hold all the way through: Bedrock's &lt;code&gt;stream()&lt;/code&gt; method got a fix for dropped reasoning deltas, and the changelog describes it as fixing dropped reasoning &lt;em&gt;and&lt;/em&gt; tool-use deltas together — "now logged at debug level instead of vanishing untraced." Reading the actual diff, only the &lt;code&gt;reasoningContent&lt;/code&gt; half shipped; there is no branch at all for &lt;code&gt;toolUse&lt;/code&gt; deltas, and a &lt;code&gt;contentBlockDelta&lt;/code&gt; carrying incremental tool-call argument fragments still falls through with zero logging, exactly as it did before the fix. It's a small gap, but it's the kind of gap that only shows up when you read the code instead of the note describing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The parser-correctness class, one level deeper
&lt;/h2&gt;

&lt;p&gt;The prior article's four-instance spine — a green suite that never exercised the defect — recurs here too, but this time it's the product of a deliberate audit rather than an accident someone happened to trip over: a correctness sweep hunting one specific bug pattern, qualified-name collisions, that pulled 19 independently confirmed defects across seven subsystems in v3.2.4 alone.&lt;/p&gt;

&lt;p&gt;Two of them are worth walking through because they're the same shape at different layers. The installed tree-sitter-java grammar emits &lt;code&gt;formal_parameters&lt;/code&gt;/&lt;code&gt;formal_parameter&lt;/code&gt; for a Java record's component list, not the &lt;code&gt;record_parameters&lt;/code&gt;/&lt;code&gt;record_component&lt;/code&gt; node types the extractor had assumed — every record's component list had been rendering empty. The fix shares a single &lt;code&gt;_extract_type_list_edges()&lt;/code&gt; helper across six call sites rather than patching the one spot that got noticed. Rust had the mirror problem: &lt;code&gt;use foo::{Alpha, Beta}&lt;/code&gt; — the dominant real-world Rust import form — fell through to a generic single-name fallback because the code checked for &lt;code&gt;use_tree_list&lt;/code&gt;/&lt;code&gt;use_tree&lt;/code&gt; node types tree-sitter-rust doesn't actually emit; the real types are &lt;code&gt;scoped_use_list&lt;/code&gt;/&lt;code&gt;use_list&lt;/code&gt;. Pre-fix, that path would have produced the literal string &lt;code&gt;"{Alpha, Beta}"&lt;/code&gt; as a bogus import name.&lt;/p&gt;

&lt;p&gt;One layer up, in the federation retriever, &lt;code&gt;make_scip_symbol_id()&lt;/code&gt; hashed only &lt;code&gt;(package, version, qualified_name)&lt;/code&gt; — so two files sharing a qualified name, like two &lt;code&gt;def main()&lt;/code&gt; entry points in different modules, collided, and the second was silently dropped via &lt;code&gt;INSERT OR IGNORE&lt;/code&gt;. The fix folds &lt;code&gt;file_path&lt;/code&gt; into the hash. And a fix in v3.3.4's chunker — where oversized symbols had their excess body silently discarded and replaced with a truncation marker, unrecoverable by retrieval — created a &lt;em&gt;new&lt;/em&gt; collision surface one release later: multiple chunks could now share one &lt;code&gt;symbol_id&lt;/code&gt;, and the fusion/dedup layer's key didn't yet account for that. v3.3.5's own fix docstring is candid about this being the same defect class recurring one abstraction level down from a prior bug, in code the same engineers had just touched days earlier. That's the honest version of "we fixed it" — not zero recurrence, but recognizing the recurrence immediately and naming it as such rather than treating each instance as a surprise.&lt;/p&gt;

&lt;p&gt;A separate 94-agent chunking-strategy study, read in full, found trelix's one-symbol-per-chunk strategy underperforming alternative strategies by 3.57–5.64 percentage points of exact-match accuracy against a real published benchmark — and then rated its own recommendation "LOW-MEDIUM confidence a full redesign is worth building right now," because a redesign would break the 1-chunk-to-1-symbol cardinality assumption baked into the reranker and graph-context code — "the exact class of bug this session already found three separate times in adjacent code this same week." That's a research doc citing its own team's bug history as the reason to be cautious about its own conclusion, which is a genuinely unusual thing to see in a changelog-adjacent artifact, and I think it's a better outcome than either shipping the redesign on a hunch or ignoring the study.&lt;/p&gt;

&lt;h2&gt;
  
  
  trelix Code Review: not a Marketplace app, and one open question
&lt;/h2&gt;

&lt;p&gt;There are two real, shipped ways to get trelix's review on a PR, and both post a check literally named "trelix Code Review." The first is a GitHub Actions workflow that trusts nothing beyond the repo's own &lt;code&gt;GITHUB_TOKEN&lt;/code&gt;. The second is a standalone GitHub App — TypeScript/Express, webhook-driven, zero workflow YAML required in the installing repo, confirmed live right now on Railway (production environment, deployment from commit &lt;code&gt;ac94e1c&lt;/code&gt; on &lt;code&gt;main&lt;/code&gt;, instance status RUNNING). The tradeoff between the two is exactly what it looks like: the Actions path trusts nobody, the App path is more convenient but grants a real third party real permissions on your repo.&lt;/p&gt;

&lt;p&gt;It is not, and should not be described as, a Marketplace app. &lt;code&gt;infra/github-app/README.md&lt;/code&gt; says this outright, in words I'm quoting rather than paraphrasing: "Not claimed: GitHub Marketplace listing. This App is installable and hardened, not Marketplace-verified." &lt;code&gt;docs/ROADMAP.md&lt;/code&gt; lists the Marketplace listing as an explicit open backlog item, gated on needing roughly 100 installations before GitHub will even review it for the listing — an adoption gate, not an engineering one.&lt;/p&gt;

&lt;p&gt;The security posture underneath, which I read at the file level rather than taking on faith: webhook signature verification uses &lt;code&gt;timingSafeEqual&lt;/code&gt; via &lt;code&gt;@octokit/webhooks-methods&lt;/code&gt; and rejects unauthorized requests before the payload body is even read; installation tokens are minted and cached per-config in a &lt;code&gt;WeakMap&lt;/code&gt;, with JWT-mode auth correctly &lt;em&gt;not&lt;/em&gt; cached; PR checkouts clone into a temp workspace and fetch &lt;code&gt;refs/pull/&amp;lt;n&amp;gt;/head&lt;/code&gt;, which is the correct ref for fork PRs, with the token passed only via a &lt;code&gt;GIT_ASKPASS&lt;/code&gt; environment variable — never in a URL, argv, or git config, where it could leak into logs or process listings; and there's a 25MB payload cap applied &lt;em&gt;before&lt;/em&gt; signature verification, which is the only defense against an oversized unsigned body. As of v3.3.8, a review that times out or crashes now always posts a completed Check run instead of leaving the PR with no status at all — &lt;code&gt;timed_out&lt;/code&gt; conclusion when the process was SIGTERM'd, &lt;code&gt;neutral&lt;/code&gt; otherwise. That's a real, shipped fix for a real silence-on-failure gap, and it's a good instinct for a system whose entire value proposition is a status check people trust.&lt;/p&gt;

&lt;p&gt;What's still open, and what I want to be careful not to resolve in either direction just because it would make a tidier ending: there's a documented, not-yet-confirmed-resolved concern that the App's own review of trelix's own large codebase has been getting SIGKILLed before indexing finishes, suspected to trace back to &lt;code&gt;review-runner.ts&lt;/code&gt;'s hardcoded 5-minute &lt;code&gt;INDEX_TIMEOUT_MS&lt;/code&gt;/&lt;code&gt;REVIEW_TIMEOUT_MS&lt;/code&gt;. I checked those constants directly — they're unchanged at 5 minutes, in this span and since. But I also checked every CHANGELOG line, every ROADMAP line, and the git log for any commit connecting those two constants to a SIGKILL-during-self-review incident, and found nothing. The only &lt;code&gt;SIGKILL&lt;/code&gt; mentions anywhere in the App's code concern a completely different scenario — OOM-killed processes leaving stale temp workspaces behind, which a &lt;code&gt;sweepStaleWorkspaces()&lt;/code&gt; cleanup at boot handles. So: the timeout constants are exactly as short as the concern says, but nothing in the repository substantiates that they're actually the cause of a self-review failure, or that the failure is still happening at all. I'm stating this as an open question because it is one, not because I want to end this section on a cliffhanger. It should not be confused with a different, already-closed issue in the same App: a missing &lt;code&gt;TRELIX_EMBEDDER_PROVIDER=azure&lt;/code&gt; env var on Railway, which caused indexing to silently fall back to a never-installed local embedder and fail — that was fixed on 2026-09-22 via a deployment config change, not a code change, and it's fully resolved. These are two unrelated bugs at two unrelated layers; only one of them is closed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The .drawio connector: diagrams as text
&lt;/h2&gt;

&lt;p&gt;v3.3.6 shipped a fifth connector alongside Jira, TestRail, Xray, and Linear: &lt;code&gt;trelix connector sync &amp;lt;repo&amp;gt; diagram&lt;/code&gt;, for local &lt;code&gt;.drawio&lt;/code&gt; (diagrams.net) files. The mechanism is more mundane than it sounds, in the best way — every shape in a drawio file's mxGraph XML carries its label in a plain &lt;code&gt;value="..."&lt;/code&gt; attribute, and every edge is an explicit &lt;code&gt;&amp;lt;mxCell edge="1" source=... target=...&amp;gt;&lt;/code&gt;, so the whole diagram is fully text-describable without any vision model. The connector truncates the XML to 20,000 characters and sends it through the same text-only chat client every other LLM call site in trelix already uses. Raster images — PNG, JPG — are explicitly out of scope; there's no glob or handling for them anywhere.&lt;/p&gt;

&lt;p&gt;v3.3.7 fixed a real bug in the fallback path: when captioning failed, the connector used to discard the diagram's real content entirely and report just a character count. Now it regex-extracts actual &lt;code&gt;mxCell value="..."&lt;/code&gt; shape labels, HTML-unescapes before stripping inline tags (the order matters, since drawio escapes attribute values, so entities have to decode before tags can be safely stripped), de-duplicates in first-seen order, caps at 50 labels, and only falls back to the old length-only message when literally zero labels exist. It's a small fix but a real one — the difference between "the diagram failed to caption, here's 4,200 characters" and "the diagram failed to caption, here are the 12 real shape labels it contains" is the difference between a useless fallback and a usable one. There's no diagram-specific config or env var; it rides whichever LLM provider is already configured, and the artifact linker treats a diagram exactly like a Jira ticket — same edge shape, same matching logic, no special-cased code path anywhere.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2u1uedem5aktwuco0nw1.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2u1uedem5aktwuco0nw1.gif" alt="An animated isometric architectural line diagram assembling itself piece by piece" width="360" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Performance: real numbers, and one number that isn't
&lt;/h2&gt;

&lt;p&gt;v3.3.8 fixed a real inefficiency: &lt;code&gt;trelix-mcp&lt;/code&gt;, the LangChain retriever, and the LlamaIndex retriever all used to rebuild a fresh &lt;code&gt;Retriever&lt;/code&gt; object on every single call. They're cached now — module-level with a lock in the MCP server, invalidated correctly on reindex, per-instance in both framework adapters. The mechanism is sound and I'd recommend the fix without reservation. What I can't recommend repeating is the specific number attached to it: the CHANGELOG cites a "24.5x–52.7x measured speedup," and that figure appears nowhere except the CHANGELOG's own prose — not in the commit message I pulled directly, not in the PR, not in any test. The commit itself says only that three consecutive calls in one session took 13-16 seconds, then 5, then 5, with zero warm-up benefit — no multiplier, and no recorded post-fix warm baseline anywhere to derive one from. The fix is real; the multiplier isn't backed by anything outside the CHANGELOG's own unsupported prose, so I'm not repeating it as measured.&lt;/p&gt;

&lt;p&gt;The rest of this theme holds up better. &lt;code&gt;trelix-mcp&lt;/code&gt;'s and the graph community-detection code's use of &lt;code&gt;nx.pagerank&lt;/code&gt; now catches &lt;code&gt;ImportError&lt;/code&gt; and degrades to a uniform-score fallback instead of crashing, because scipy is only imported lazily on first call and the frozen binary deliberately excludes it — a real crash class closed with two small, defensive catches. Qdrant gained int8/binary quantization support, with one small flag: the commit claims verification against &lt;code&gt;qdrant-client==1.19.0&lt;/code&gt;, but the lockfile currently pins 1.18.0. Not a contradiction — a dev-venv version never captured in the lockfile is the more likely explanation than anything adversarial — but not independently confirmable either, so treat "verified against 1.19.0" as unconfirmed rather than false. A direct Cohere embedder landed using &lt;code&gt;cohere.ClientV2&lt;/code&gt; rather than routing through Bedrock, with a 96-item batch limit. And an abstractive compression provider shipped with a real accuracy gap between what the CHANGELOG claims and what the code guarantees: the CHANGELOG says the signature &lt;em&gt;and&lt;/em&gt; docstring line are "always kept byte-for-byte." The actual &lt;code&gt;_must_keep()&lt;/code&gt; logic only locates and preserves the declaration line itself — the docstring on subsequent lines is forwarded to the LLM as context, with no verbatim guarantee, and can come back paraphrased. The test suite's own docstring gets this right where the CHANGELOG doesn't: "the declaration line is ALWAYS kept verbatim," singular. If you're relying on this feature to preserve docstrings exactly, don't — that's what the separate extractive compressor genuinely does.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcdxl9z0iuli7fy3h93pj.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcdxl9z0iuli7fy3h93pj.gif" alt="A rocket igniting and lifting off from the launch pad in a cloud of exhaust" width="480" height="412"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Security: two CRITICAL CVEs and an audit trail that undersells itself
&lt;/h2&gt;

&lt;p&gt;The litellm floor moved to &lt;code&gt;&amp;gt;=1.90.2&lt;/code&gt;, closing two CVEs the project's own pins cite as CVSS 9.8 — a Host-header authentication bypass and an unauthenticated SSTI leading to RCE — plus a third SSRF issue via nested &lt;code&gt;user_config&lt;/code&gt;. I can confirm these citations are in the repo's own dependency comments; I could not independently cross-check the CVE IDs against GitHub's live advisory database from this sandbox, so I'm attributing the "two CRITICAL, CVSS 9.8" framing to trelix's own release notes rather than an outside source I verified myself. Separately, &lt;code&gt;pydantic-settings&lt;/code&gt; CVE-2026-58203 was patched by a floor bump, and I confirmed the vulnerable feature — &lt;code&gt;secrets_nested_subdir&lt;/code&gt; — was never actually called anywhere in the codebase, so this one was zero real exposure even before the bump; a clean case of patching a dependency that was never actually exploitable through this code.&lt;/p&gt;

&lt;p&gt;The most substantial security addition is &lt;code&gt;trelix audit prune&lt;/code&gt; — a real hash-chain watermarking implementation for pruning old audit-log rows without breaking chain verification, with an &lt;code&gt;--dry-run&lt;/code&gt; mode that's confirmed not wired into any in-process scheduler. What struck me reading it is that the module's own docstring is more conservative than the CHANGELOG's framing: it states plainly that this is "tamper-evident, not tamper-proof," and it enumerates exactly which tamper shapes remain undetectable — a total wipe that also clears SQLite's own internal sequence table, for instance. A security feature whose own documentation undersells its guarantees relative to the marketing copy around it is the rare direction I'm happy to see the drift run in.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fam3zqeop3svu0tmjeuqr.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fam3zqeop3svu0tmjeuqr.gif" alt="An animated shield illustration with a glowing lava-lamp pattern moving inside it" width="480" height="440"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The smaller robustness releases
&lt;/h2&gt;

&lt;p&gt;Two releases of CLI/API/extension hardening are worth a fast pass. v3.3.1 fixed four real VS Code extension bugs: a JSON.parse crash whenever the MCP client returned a plain-text error instead of JSON, a path escape hatch for GUI-launched editors whose PATH excludes pip/uv, inline symbol links in chat responses, and a CodeLens replacing a global QuickPick for blast-radius visualization. v3.3.7 fixed a &lt;code&gt;.vscodeignore&lt;/code&gt; gap that had been shipping a local vector database file and gigabyte-plus of sourcemaps inside the packaged extension — &lt;code&gt;vsce&lt;/code&gt; ignores &lt;code&gt;.gitignore&lt;/code&gt; entirely once &lt;code&gt;.vscodeignore&lt;/code&gt; exists, so the two files having drifted apart had real consequences. The same release moved four CLI/REST failure modes from raw tracebacks or bare 500s to labeled, structured errors: &lt;code&gt;eval&lt;/code&gt;/&lt;code&gt;eval-synthesis&lt;/code&gt;/&lt;code&gt;migrate-vectors&lt;/code&gt; now print one-line labeled errors instead of stack traces; a missing golden file for synthesis eval now raises &lt;code&gt;FileNotFoundError&lt;/code&gt; instead of silently returning zero-filled metrics; the REST API's &lt;code&gt;DimensionMismatchError&lt;/code&gt; and &lt;code&gt;ImportError&lt;/code&gt; now return structured JSON 500s instead of Starlette's plain-text default; and static-token auth now checks &lt;code&gt;Authorization: Bearer&lt;/code&gt; independently of &lt;code&gt;X-Trelix-Api-Key&lt;/code&gt;, via &lt;code&gt;hmac.compare_digest&lt;/code&gt;, closing a path where the Bearer header was silently ignored.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern behind the corrections
&lt;/h2&gt;

&lt;p&gt;I want to name the pattern across this article's corrections rather than let them read as scattered nitpicks, because I think they're one thing: &lt;code&gt;ReasoningBlock&lt;/code&gt; vs. &lt;code&gt;ThinkingBlock&lt;/code&gt;, the half-shipped &lt;code&gt;toolUse&lt;/code&gt;-delta logging claim, the unaudited 24.5x-52.7x figure, the docstring-verbatim overstatement, a migration guide that's stale relative to its own release date, a &lt;code&gt;docs/WHY_TRELIX.md&lt;/code&gt; version-scope stamp that still reads v3.3.7 while the package is at v3.3.8 — none of these are code bugs. Every one of them is the layer just above the code: the sentence describing what the code does, written slightly faster or slightly more optimistically than the code itself justifies. That's a different failure mode than a &lt;code&gt;MagicMock&lt;/code&gt; hiding a broken embedder, and it's arguably a sign of health rather than decline — it only becomes the dominant story once the release-verification discipline from earlier in this span has already closed off the cruder failure mode of a defect reaching a tag at all. It's also, not coincidentally, exactly the discipline this article itself needed applied to it: this dossier exists because the same standard — check the sentence against the source, not against the intent — got turned on the release notes themselves, and turned up a correction on nearly every theme it touched.&lt;/p&gt;

&lt;h2&gt;
  
  
  What v3.3.8 actually means for someone upgrading today
&lt;/h2&gt;

&lt;p&gt;If you're upgrading from anywhere before v3.3.0, read the CHANGELOG for that release specifically rather than the migration guide — the guide is missing the &lt;code&gt;AWS_REGION&lt;/code&gt; requirement and the silently-dropped &lt;code&gt;temperature&lt;/code&gt; parameter, which are the two changes most likely to actually break your deployment. If you're grepping the codebase for the reasoning-content type, look for &lt;code&gt;ThinkingBlock&lt;/code&gt;, not &lt;code&gt;ReasoningBlock&lt;/code&gt;. If you're evaluating the retriever-caching fix, trust the mechanism and not the specific multiplier. If you're relying on abstractive compression to preserve a function's docstring exactly, it doesn't — only the declaration line is guaranteed verbatim. And if you're deciding whether to install the GitHub App or rely on the Actions workflow for PR review, both post the same "trelix Code Review" check, but only one of them asks you to trust a third party — and if you run the App against a codebase the size of trelix's own, know that a timeout-related self-review concern from a week ago hasn't been confirmed fixed, or confirmed to still be happening either.&lt;/p&gt;

&lt;p&gt;None of that is a reason not to upgrade. The live PyPI check I ran directly — not read, not paraphrased, executed — confirms v3.3.8 across all four packages, published in lockstep on 2026-09-25. The release-verification infrastructure built in v3.2.3 and v3.2.4 means that check itself, along with a Docker &lt;code&gt;--version&lt;/code&gt; smoke test and a Helm chart lint, now runs automatically against every tag before it's allowed to publish, which is a meaningfully different guarantee than the source tree alone ever gave you. What changed across these thirteen releases isn't that trelix stopped shipping defects — the parser sweep alone found 19 of them in one release. What changed is that the mechanism for catching what a green test suite misses stopped being a person doing it once, by hand, out of frustration, and started being a system that does it on every tag, whether anyone remembers to ask or not. The part that hasn't caught up yet is the sentence describing the fix, which is a smaller problem to have than the one this project started with nine weeks ago, and a more honest place to end an audit than pretending it isn't there at all.&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>python</category>
      <category>testing</category>
      <category>security</category>
    </item>
    <item>
      <title>MindForge v12.0.0: An Agentic Framework for Claude Code — What It Ships, How to Install It, and What's Actually Enforced</title>
      <dc:creator>SAI RAM</dc:creator>
      <pubDate>Thu, 24 Sep 2026 05:42:32 +0000</pubDate>
      <link>https://dev.to/sai_ram_0000/mindforge-v1200-an-agentic-framework-for-claude-code-what-it-ships-how-to-install-it-and-3ncb</link>
      <guid>https://dev.to/sai_ram_0000/mindforge-v1200-an-agentic-framework-for-claude-code-what-it-ships-how-to-install-it-and-3ncb</guid>
      <description>&lt;h2&gt;
  
  
  What MindForge Is (And Isn't)
&lt;/h2&gt;

&lt;p&gt;Agentic frameworks for Claude Code have exploded over the last year — swarms of subagents, skill libraries, protocol layers stacked on top of a single model. Most of them describe what they ship in one flattened register: "AI capabilities." MindForge (&lt;code&gt;mindforge-cc&lt;/code&gt; on npm) is a governance and orchestration layer in that same category, sitting on top of Claude Code — it is not a replacement for it, and it does not run without it — but it's one of the more mature, more self-critical entries: it treats its own surface area as five mechanically distinct pieces rather than one undifferentiated pile, and it draws an explicit line between the parts that are genuinely enforced and the parts that are just well-organized advice. As of this writing, &lt;code&gt;v12.0.0&lt;/code&gt; is tagged and live on npm — both &lt;code&gt;latest&lt;/code&gt; and &lt;code&gt;stable&lt;/code&gt; point to it, with a real GitHub Release at the matching tag — and its CHANGELOG entry says, verbatim, this is "MindForge's first release deliberately cut for real external users." That framing, and how it holds up against the actual source tree, is what this piece checks.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0og1xwe971r8idpo3jmy.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0og1xwe971r8idpo3jmy.gif" alt="An animated description of what is happening in the GIF" width="480" height="339"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Concretely, MindForge ships as three things: an npm package that installs slash commands, skills, personas, subagents, and hooks into a &lt;code&gt;.claude/&lt;/code&gt; (and mirrored &lt;code&gt;.agent/&lt;/code&gt;) directory; a set of Claude Code plugin-marketplace entries; and a separate MCP server plus TypeScript SDK. Almost everything it ships — the commands, the skills, the personas, the workflow scripts, the governance docs — works by being loaded into the model's context. None of that is code the model is forced to obey. The one exception, and the one piece of this system that is a real technical enforcement mechanism, is a small set of Claude Code hooks. That distinction — advisory context vs. enforced blocking — is the single most important thing to understand about this project, and it's the project's own README that draws the line, not outside auditors.&lt;/p&gt;

&lt;p&gt;This matters because MindForge's own CHANGELOG and its repo-local &lt;code&gt;.claude/CLAUDE.md&lt;/code&gt; spend real effort disowning parts of their own marketing language. The "Unified Protocol Engine" framing that shows up in some MindForge-derived global configs — SwarmController, PersonaFactory, WaveExecutor, &lt;code&gt;soul-engine.js&lt;/code&gt;, &lt;code&gt;shard-controller.js&lt;/code&gt; — is explicitly called out in the repo's own &lt;code&gt;.claude/CLAUDE.md&lt;/code&gt; as aspirational naming with no backing code: those are "role names in the specs under &lt;code&gt;.mindforge/engine/&lt;/code&gt;, not importable code." That kind of self-correction, verified directly against the source tree rather than taken on faith, is the throughline of this whole release.&lt;/p&gt;

&lt;p&gt;The real working loop that MindForge scaffolds is plan-phase → execute-phase → verify-phase → ship, and it is detailed rather than just a name: plan-phase spawns a research subagent and writes atomic XML plan files with explicit &lt;code&gt;&amp;lt;verify&amp;gt;&lt;/code&gt; steps; execute-phase runs dependency-aware waves against a five-level escalating validation ladder (static, unit, build, integration, edge) and writes a Deviation Report per task; verify-phase walks the human through REQUIREMENTS.md deliverables and writes UAT.md, spawning a debug subagent on failure; ship gates on UAT.md reading "All passed," runs &lt;code&gt;tsc --noEmit&lt;/code&gt;, &lt;code&gt;eslint&lt;/code&gt;, &lt;code&gt;npm test&lt;/code&gt;, and &lt;code&gt;npm audit&lt;/code&gt;, and generates the PR description. It's real machinery, not command-name theater — though even here there's a live wrinkle worth knowing about: &lt;code&gt;ship.md&lt;/code&gt;'s Step 4 currently hardcodes a literal branch name in its &lt;code&gt;git push&lt;/code&gt; rather than deriving the active branch, so anyone not literally on that named branch could push to the wrong place.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fag00hfotfnoot0i87sgf.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fag00hfotfnoot0i87sgf.gif" alt="An animated description of what is happening in the second GIF" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Building Blocks: Five Different Mechanisms, Five Different Levels of Trust
&lt;/h2&gt;

&lt;p&gt;MindForge's surface area breaks into five genuinely distinct mechanisms, and the differences between them are mechanical, not just naming conventions. Here's what it actually gives you, at a glance, verified count by verified count:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Count&lt;/th&gt;
&lt;th&gt;Invoked via&lt;/th&gt;
&lt;th&gt;What it mechanically is&lt;/th&gt;
&lt;th&gt;Enforced or advisory&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Slash commands&lt;/td&gt;
&lt;td&gt;221&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/mindforge:&amp;lt;name&amp;gt;&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;.md&lt;/code&gt; prompt specs the model reads and can choose to follow&lt;/td&gt;
&lt;td&gt;Advisory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Skills — engine tier&lt;/td&gt;
&lt;td&gt;232&lt;/td&gt;
&lt;td&gt;Auto-triggered by keyword match&lt;/td&gt;
&lt;td&gt;An LLM-followed protocol spec (&lt;code&gt;.mindforge/engine/skills/loader.md&lt;/code&gt;) does the matching — non-deterministic, not a parser&lt;/td&gt;
&lt;td&gt;Advisory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Skills — extended tier&lt;/td&gt;
&lt;td&gt;122&lt;/td&gt;
&lt;td&gt;Invoked explicitly by name&lt;/td&gt;
&lt;td&gt;Same skill mechanism, lenient schema (only a name is required)&lt;/td&gt;
&lt;td&gt;Advisory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Personas&lt;/td&gt;
&lt;td&gt;216&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/mindforge:agent &amp;lt;name&amp;gt;&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;An in-session role overlay — same context, not a new agent&lt;/td&gt;
&lt;td&gt;Advisory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Subagents&lt;/td&gt;
&lt;td&gt;164 (154 adapted from VoltAgent's MIT-licensed library + 10 original)&lt;/td&gt;
&lt;td&gt;Claude Code's native subagent mechanism, via the plugin marketplace&lt;/td&gt;
&lt;td&gt;A genuinely isolated execution context — a separate mechanism from personas&lt;/td&gt;
&lt;td&gt;Advisory (the definition; the isolation itself is Claude Code's)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dynamic workflows&lt;/td&gt;
&lt;td&gt;35, across 5 tiers (Research 5 / Dev 14 / Ops 6 / Intelligence 7 / Beast 3)&lt;/td&gt;
&lt;td&gt;Claude Code's own host-level &lt;code&gt;Workflow&lt;/code&gt; tool&lt;/td&gt;
&lt;td&gt;Curated multi-agent orchestration scripts targeting a host capability MindForge doesn't itself implement&lt;/td&gt;
&lt;td&gt;Advisory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hooks&lt;/td&gt;
&lt;td&gt;8 registered (7 executed by preflight, 3 "deny-class")&lt;/td&gt;
&lt;td&gt;Automatic, on tool-call boundaries&lt;/td&gt;
&lt;td&gt;Real shell scripts wired into Claude Code's own hook system&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Enforced&lt;/strong&gt; — the one row on this table that actually is&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;"Enforced" in that last row only ever means on Claude Code, via the npx installer's &lt;code&gt;--local&lt;/code&gt; or &lt;code&gt;--global&lt;/code&gt; target — see the next section for exactly what that does and doesn't cover.&lt;/p&gt;

&lt;p&gt;Commands are 221 slash-command &lt;code&gt;.md&lt;/code&gt; files under &lt;code&gt;.claude/commands/mindforge/&lt;/code&gt;, mirrored exactly in &lt;code&gt;.agent/mindforge/&lt;/code&gt; — verified by direct file count, not just README's say-so. They're prompt specs the model reads and follows; they don't execute independently of the model choosing to follow them.&lt;/p&gt;

&lt;p&gt;Skills split into two tiers totaling 354 files: 232 "engine tier" skills under &lt;code&gt;.mindforge/skills/&lt;/code&gt;, auto-triggered by keyword match, and 122 "extended tier" skills under &lt;code&gt;.agent/skills/&lt;/code&gt;, explicitly invoked by name. The engine tier enforces a strict schema via &lt;code&gt;tests/skills-platform.test.js&lt;/code&gt; — semver version, a status field, at least ten trigger terms, a "Mandatory actions" section — while the extended tier only requires a name. Here's the part worth being honest about: the file that's supposed to do the actual trigger-matching, &lt;code&gt;bin/engine/skill-loader.js&lt;/code&gt;, is dead code — a four-line stub (&lt;code&gt;loadSkill: () =&amp;gt; null, matchTriggers: () =&amp;gt; []&lt;/code&gt;) with zero real callers, and the CHANGELOG says so directly. The real trigger-matching mechanism is &lt;code&gt;.mindforge/engine/skills/loader.md&lt;/code&gt;, an LLM-followed protocol spec, not deterministic code. That means "skill triggering" in MindForge is non-deterministic by design — the model decides what counts as a match, not a parser.&lt;/p&gt;

&lt;p&gt;Personas are 216 named role-overlay files under &lt;code&gt;.mindforge/personas/&lt;/code&gt;, loaded via &lt;code&gt;/mindforge:agent &amp;lt;name&amp;gt;&lt;/code&gt; into the &lt;em&gt;same&lt;/em&gt; session context — a role swap, not a new agent. This is a meaningfully different mechanism from subagents, and MindForge's own docs are explicit about the distinction rather than blurring it for effect.&lt;/p&gt;

&lt;p&gt;Subagents are 164 real Claude-Code-native subagent definition files under &lt;code&gt;subagents/categories/&lt;/code&gt;, each with proper frontmatter (name/description/tools/model) designed for Claude Code's isolated-context subagent mechanism — genuinely separate execution contexts, unlike personas. 152 of those files are attributed to VoltAgent's MIT-licensed &lt;code&gt;awesome-claude-code-subagents&lt;/code&gt;, with the full license text reproduced; that attribution checks out against the vendored upstream repo. One correction worth flagging: the README's claimed "152 adapted / 12 original" split undercounts the adapted set by two — &lt;code&gt;dotnet-framework-48-expert.md&lt;/code&gt; and &lt;code&gt;powershell-51-expert.md&lt;/code&gt; are near-byte-identical VoltAgent copies with dots stripped from the filename for path safety, so the real split is closer to 154 adapted / 10 original. Small, but it's the kind of thing worth naming rather than repeating uncritically.&lt;/p&gt;

&lt;p&gt;Dynamic workflows are 35 JavaScript scripts across five tiers (Research 5, Dev 14, Ops 6, Intelligence 7, Beast 3) under &lt;code&gt;.mindforge/dynamic-workflows/scripts/&lt;/code&gt;, and the ones that were read in full are genuinely substantial multi-agent orchestration code — &lt;code&gt;security-hardening.js&lt;/code&gt;, for instance, runs five parallel OWASP scout agents, then for every critical/high finding spins up a nested three-vote adversarial verification round (refute / challenge-exploitability / assess-impact) requiring a real two-of-three quorum before accepting a finding. That's not a cosmetic label. But here's the structural fact that matters most: the actual execution primitives these scripts call — &lt;code&gt;agent()&lt;/code&gt;, &lt;code&gt;parallel()&lt;/code&gt;, &lt;code&gt;phase()&lt;/code&gt; — are not implemented anywhere in the MindForge repository. They're supplied entirely by Claude Code's own host-level Workflow tool at runtime. MindForge ships the script library, the JSON registry, and the CLI browser; it does not ship a self-contained execution engine. Every one of these scripts also contains a top-level &lt;code&gt;return&lt;/code&gt; statement, so none of them are valid standalone Node.js files — they only run wrapped by that external harness. This is a real architectural fact, not a defect, but it's exactly the kind of thing a launch article shouldn't gloss over: "35 dynamic workflows" means 35 curated orchestration scripts that target a host capability, not 35 things MindForge alone can run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Advisory Context vs. Enforced Blocking
&lt;/h2&gt;

&lt;p&gt;This is the section that most agent-tooling projects don't write, and MindForge's README does, in a dedicated "What is actually enforced" section: hooks are the only real blocking mechanism in the entire framework, and they only block on Claude Code, via the npx installer's &lt;code&gt;--local&lt;/code&gt; target. They are explicitly not enforced on Cursor, Copilot, Gemini, or Antigravity; not on &lt;code&gt;--global&lt;/code&gt; installs; not on self-installs; and not on Windows.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftlorxmm0ttl79p661np8.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftlorxmm0ttl79p661np8.gif" alt="An animated description of the goalkeeper save" width="200" height="200"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fve9j7ho2on750evewhrn.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fve9j7ho2on750evewhrn.gif" alt="An animated description of the goal celebration" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The live, dogfooded numbers behind that claim hold up under direct verification: the repo's own &lt;code&gt;.claude/settings.json&lt;/code&gt; registers exactly 8 hooks (2 PostToolUse, 4 PreToolUse, 2 SessionStart). Of those, 7 are executed by the installer's preflight check (one, &lt;code&gt;mindforge-check-update&lt;/code&gt;, is deliberately skipped). Of the 7 executed, only 3 are "deny-class" — capable of returning a hard block: &lt;code&gt;trust-gate&lt;/code&gt;, &lt;code&gt;mindforge-block-no-verify&lt;/code&gt;, and &lt;code&gt;mindforge-config-protection&lt;/code&gt;. All three were confirmed live to return exit code 2 on trigger. Everything else in this system — the 221 commands, the 354 skills, the 216 personas, the 35 workflows, the governance and audit docs — can be ignored by the model at any time. A launch article shouldn't imply that installing MindForge "activates" or "enforces" its governance or security-scan language across arbitrary tools or channels. It doesn't. It narrows enforcement to three specific hooks on one specific install channel, and says so.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Audit Chain and Why This Release Is Different
&lt;/h2&gt;

&lt;p&gt;MindForge maintains a tamper-evident audit log at &lt;code&gt;.planning/AUDIT.jsonl&lt;/code&gt;, hash-chained with SHA-256 via a shared hasher (&lt;code&gt;bin/governance/audit-hash.js&lt;/code&gt;) used by both the writer and the verifier. Running the verifier live against this repo's own log returned "audit chain valid: 6396 entries." SECURITY.md scopes the claim precisely — the chain detects mid-file mutation and deletion, and explicitly does not detect tail truncation or replay — and a dedicated honesty test suite (&lt;code&gt;tests/audit-claims-honesty.test.js&lt;/code&gt;) checks that the documentation doesn't overclaim beyond that, including confirming the writer doesn't sign entries and the docs never call this a Merkle tree.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9nkxca0um8rykctygiyo.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9nkxca0um8rykctygiyo.gif" alt="An animated description of what is happening in the fourth GIF" width="480" height="300"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The more interesting story is what the project's own pre-ship audits actually catch, because it's a real mix of genuine bugs and genuine overclaiming, disclosed rather than buried. The v11.9.9 release describes an eight-agent audit workflow run as a release gate before shipping to real external users, and it's not a rubber stamp: that pass found 5 CRITICAL issues, including a jailbreak skill named "godmode" that had been shipping live, a security check with no real CLI entrypoint (so it could structurally never fail), and a &lt;code&gt;--minimal&lt;/code&gt; install flag that silently shipped the full persona set anyway — plus 9 HIGH findings. Verified directly: neither the godmode skill nor any of its assets exist anywhere in the shipped tree, and a packaging-allowlist test explicitly asserts its absence. That same audit pass also fixed a set of docs (&lt;code&gt;help.md&lt;/code&gt;, &lt;code&gt;status.md&lt;/code&gt;, &lt;code&gt;health.md&lt;/code&gt;, &lt;code&gt;security-scan.md&lt;/code&gt;) that had been overclaiming a post-quantum-signature feature (PQAS) and biometric gating as active; the real code is honest about this now — &lt;code&gt;bin/governance/quantum-crypto.js&lt;/code&gt; labels its Dilithium-5 implementation "SIMULATED... NOT real ML-DSA/FIPS-204," gated off by default in config. Earlier releases show the same pattern repeating: v11.9.8 found two real crashing bugs plus 12 pure documentation inaccuracies out of 113 claims checked; v11.9.6 found a crashing &lt;code&gt;/mindforge:learn&lt;/code&gt; command and an auth-token-leaking browser daemon.&lt;/p&gt;

&lt;p&gt;Two things are worth naming honestly rather than hiding. First, that PQAS-overclaim pattern had a residual instance the 11.9.9 audit missed: &lt;code&gt;.mindforge/governance/policies/sovereign-default.json&lt;/code&gt; still marked PQAS and biometric checks as "ENABLED," even though that file is dead config the real &lt;code&gt;PolicyEngine&lt;/code&gt; never loads. Second, the same "feature that isn't wired to what it implies" pattern shows up in a place the audits hadn't caught as of this research pass: MindForge's README describes a "two-model adversarial PR review" and a "4-voice consensus council," but as currently wired, both mechanisms route every voice and every round to the exact same single backing model — the cross-review engine's caller-supplied model names are used only as display labels, and the council's four voices all resolve through an unmapped persona to the same tier-2 default. That gap between naming and behavior was not in the CHANGELOG as of v11.9.9 — it's the kind of thing this project's own audit discipline should catch in its next pass, and naming it here is more useful to a reader than pretending it doesn't exist.&lt;/p&gt;

&lt;p&gt;Also worth knowing: several of the CLI/doc claims that have surfaced in past audits are still true today and still disclosed rather than fixed. &lt;code&gt;/mindforge:health --repair&lt;/code&gt; silently drops the &lt;code&gt;--repair&lt;/code&gt; flag (byte-identical output to plain &lt;code&gt;health&lt;/code&gt;); &lt;code&gt;spawn &amp;lt;persona&amp;gt;&lt;/code&gt; still exits 1 with "not implemented in v1.0"; and two doc inaccuracies that survived multiple audit passes are still live — &lt;code&gt;docs/commands-reference.md&lt;/code&gt; still lists a CLI command, &lt;code&gt;quantum-verify&lt;/code&gt;, that doesn't exist (confirmed by running it), and &lt;code&gt;docs/tutorial.md&lt;/code&gt; still documents an &lt;code&gt;--ads&lt;/code&gt; flag on &lt;code&gt;/mindforge:plan-phase&lt;/code&gt; that was never wired into the actual command file.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost-Aware Model Routing: The Honest Version
&lt;/h2&gt;

&lt;p&gt;MindForge advertises cost-aware, difficulty-aware model routing across five real providers — Anthropic, OpenAI, Gemini, Bedrock, and Ollama — and the provider layer is genuinely live: each provider makes real HTTPS calls and computes real per-call cost through a shared pricing registry. What's worth being precise about is what actually decides which model gets used in production: it's persona plus a caller-supplied security tier (0–3), not computed task difficulty. Tier 3 forces a dedicated security model; named personas map to fixed model settings; everything else falls to sensible defaults (Haiku for quick tasks, Sonnet as the general executor).&lt;/p&gt;

&lt;p&gt;There are, in fact, three separate difficulty-aware routing systems built into this codebase, and none of them are live. A real 1–10 difficulty scorer exists and is unit-tested, but its only caller runs it in shadow mode — it logs what model it would have picked and changes nothing. A second system, a full "multi-cloud arbitrage" broker with its own passing tests, is required nowhere except by its own test file (it even lists a fourth "azure" provider with no corresponding implementation). A third, MIR-based steering engine is explicitly gated behind a &lt;code&gt;shadow_mode&lt;/code&gt; config flag and is imported by the one module that's supposed to call it — but never actually invoked. This is disclosed territory: the CHANGELOG documents fixing a near-identical prior bug in the same routing module (a regex mismatch that made every config value resolve to &lt;code&gt;undefined&lt;/code&gt;), so the pattern of "wired but not connected" recurring in cost routing isn't new to this codebase — it's just not fully cleaned up yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Shipping Surface: MCP Server, SDK, and Three Real Channels
&lt;/h2&gt;

&lt;p&gt;MindForge ships three real, independently-verifiable channels, not one. &lt;code&gt;mindforge-mcp-server@12.0.0&lt;/code&gt; is genuinely published on npm right now, matches the repo source exactly (including SLSA provenance), and registers exactly 8 MCP tools — 6 read-only (health, status, memory query/stats/find-related, audit log), one guarded write (&lt;code&gt;memory_remember&lt;/code&gt;), and one guarded, open-world tool (&lt;code&gt;browse&lt;/code&gt;, which proxies to a separate pre-existing bearer-token-authenticated Playwright daemon and never spawns it itself). Worth flagging honestly, because the project's own README already does: the official MCP registry listing for this server lags one patch version behind what's actually published on npm, and README tells readers to check what that listing actually serves, or install from npm directly, rather than trusting the registry blindly.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;mindforge-sdk@12.0.0&lt;/code&gt; is also live on npm and matches source. Its &lt;code&gt;batchExecute&lt;/code&gt; spawns real child processes with SIGTERM-then-SIGKILL escalation — and there's a named regression test in the SDK's own test suite that exists specifically to guard against a previously-shipped fake stub that used to return &lt;code&gt;{executed:true}&lt;/code&gt; without doing anything. That's a good tell for how seriously this project treats its own past mistakes: it doesn't just fix a caught bug, it writes a test whose name records what the bug was.&lt;/p&gt;

&lt;p&gt;The underlying persistence layer for MindForge's local knowledge graph is a genuinely zero-native-dependency setup: &lt;code&gt;sql.js&lt;/code&gt;, a WASM-compiled SQLite, backing an 11.8MB file (&lt;code&gt;celestial.db&lt;/code&gt;) with real FTS4 full-text search, plus a separate SHA-256-checksummed, file-locked JSONL edge list for relationship data. The live dashboard at &lt;code&gt;localhost:7339&lt;/code&gt; is a real Express + SSE server, bearer-token-authenticated on mutating routes, rate-limited, localhost-only, with seven API-backed tabs. Worth being precise here too: the more elaborate dashboard sub-features described in some MindForge-adjacent marketing language — a Temporal Slider, Hindsight Steering Vectors, a $100/hr AgRevOps ROI hub — were not independently verified beyond confirming the dashboard's base HTTP response; treat those as unverified narrative, not confirmed capability.&lt;/p&gt;

&lt;h2&gt;
  
  
  Installing MindForge
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft1k3zyj60ijxbpejdlz5.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft1k3zyj60ijxbpejdlz5.gif" alt="An animated description of what is happening in the fifth GIF" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;MindForge ships five real, independently-checkable install channels — pick whichever matches how you actually work. As of this writing every one of them installs &lt;code&gt;v12.0.0&lt;/code&gt; directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;npx (recommended)&lt;/strong&gt; — writes &lt;code&gt;.mindforge/&lt;/code&gt; governance, memory, and planning into your project, and registers the 8 hooks described above:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx mindforge-cc@latest &lt;span class="nt"&gt;--claude&lt;/span&gt; &lt;span class="nt"&gt;--local&lt;/span&gt;      &lt;span class="c"&gt;# Claude Code, this project only&lt;/span&gt;
npx mindforge-cc@latest &lt;span class="nt"&gt;--antigravity&lt;/span&gt; &lt;span class="nt"&gt;--local&lt;/span&gt; &lt;span class="c"&gt;# Antigravity, this project only&lt;/span&gt;
npx mindforge-cc@latest                       &lt;span class="c"&gt;# interactive wizard, pre-selects a detected runtime&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run non-interactively (CI, piped, scripted) and it skips the wizard entirely and installs &lt;code&gt;--claude&lt;/code&gt; by default, regardless of what's actually on the machine. For a system-wide install: &lt;code&gt;npx mindforge-cc@latest --claude --global&lt;/code&gt; — note that &lt;code&gt;npm install -g mindforge-cc@latest&lt;/code&gt; on its own only puts the &lt;code&gt;mindforge&lt;/code&gt;/&lt;code&gt;mindforge-cc&lt;/code&gt; binaries on your PATH, it doesn't scaffold anything; you still need to run the &lt;code&gt;--global&lt;/code&gt; command above once. Other supported runtimes use the same flag pattern (&lt;code&gt;--cursor&lt;/code&gt;, &lt;code&gt;--copilot&lt;/code&gt;, &lt;code&gt;--gemini&lt;/code&gt;). A few advanced flags worth knowing: &lt;code&gt;--runtime claude,cursor&lt;/code&gt; (combined runtimes in one pass), &lt;code&gt;--minimal&lt;/code&gt; (essential scaffolding only, no persona library), &lt;code&gt;--force&lt;/code&gt; (rewrite an existing schema file with the current, stricter one).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude Code plugin marketplace&lt;/strong&gt; — writes no project files at all; only the plugin's hooks fire (see the enforcement section above for exactly what that does and doesn't cover):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/plugin marketplace add sairam0424/MindForge
/plugin &lt;span class="nb"&gt;install &lt;/span&gt;mindforge@mindforge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Want just a slice instead of everything? &lt;code&gt;mindforge-lang@mindforge&lt;/code&gt; and 9 other focused sub-packs exist for narrower installs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Standalone MCP server&lt;/strong&gt; — if all you want is the 8 read/write-guarded stdio tools described earlier, without any of the command/skill/persona tree:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add mindforge &lt;span class="nt"&gt;--&lt;/span&gt; npx &lt;span class="nt"&gt;-y&lt;/span&gt; mindforge-mcp-server
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Also listed on the official MCP Registry, though that listing is manually republished and can lag behind npm — install &lt;code&gt;mindforge-mcp-server&lt;/code&gt; directly if you want to pin an exact version.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Homebrew&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;brew &lt;span class="nb"&gt;install &lt;/span&gt;sairam0424/tap/mindforge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The formula pins to a specific npm tarball with a checked sha256 — currently pointed at &lt;code&gt;mindforge-cc-12.0.0.tgz&lt;/code&gt;, confirming this channel is already caught up for this release too, not a stale mirror.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SDK&lt;/strong&gt; — build on MindForge programmatically, without installing any of the rest:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm i mindforge-sdk
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once installed through any of the above, the natural entry points are &lt;code&gt;/mindforge:init-project&lt;/code&gt; to scaffold a new project under the framework, &lt;code&gt;/mindforge:next&lt;/code&gt; for auto-discovery of what to do given current state, and the plan-phase → execute-phase → verify-phase → ship loop described earlier for anything substantive. The subagent library is best installed through the plugin marketplace above rather than through &lt;code&gt;bin/spawn-agent.js&lt;/code&gt;, which — despite what its own README implies about installing subagents "by name" via that script — is a dry-run/verification tool only; without &lt;code&gt;--dry-run&lt;/code&gt; it always exits 1 with "not implemented in v1.0."&lt;/p&gt;

&lt;h2&gt;
  
  
  What v12.0.0 Actually Means
&lt;/h2&gt;

&lt;p&gt;By the time research for this piece began, the goalpost had already moved once — from v11.9.9's own framing as a "release-readiness milestone" to a live git branch, spotted mid-research, whose commit message named v12.0.0 as the actual "first-user-release cut." It moved again shortly after: merged to &lt;code&gt;main&lt;/code&gt; via PR #298, with a formal CHANGELOG entry titled, verbatim, "[12.0.0] — First release aimed at real external users," but not yet tagged or published. As of this update, it's moved a third and final time: &lt;code&gt;v12.0.0&lt;/code&gt; is tagged, published, and live — &lt;code&gt;latest&lt;/code&gt; and &lt;code&gt;stable&lt;/code&gt; on npm both point to it, and &lt;code&gt;mindforge-mcp-server&lt;/code&gt; and &lt;code&gt;mindforge-sdk&lt;/code&gt; shipped alongside it at the same version. The CHANGELOG entry says plainly what that title implies and nothing more: "MindForge's first release deliberately cut for real external users, not just internal iteration. The major-version bump is a milestone marker for that shift, not a signal of a breaking API/behavior change — there is none in this release; every item below is a fix."&lt;/p&gt;

&lt;p&gt;That's the honest way to read a major-version bump in this project: not a rewrite, a readiness gate. It shipped because a second, independent eight-agent audit — security and STRIDE threat modeling, staff-engineer code review, dependency and license review, a live production dry-run across all six supported runtimes, docs-accuracy checking, re-verification of the prior release's deferred backlog, test-coverage gap analysis, and a full trace of the CI/CD release pipeline — was run specifically to answer one question: was v11.9.9 actually ready for real external users? The answer was no, not quite. That audit surfaced 1 CRITICAL and 4 HIGH findings, plus 8 MEDIUM/LOW findings judged worth fixing before this exact cutover, and every one of them was independently adversarially re-verified — 10 out of 10 confirmed, zero refuted — before being fixed.&lt;/p&gt;

&lt;p&gt;The CRITICAL finding is worth naming, because it's the same species of bug this piece's own research caught mid-flight, independently, the same week: a &lt;code&gt;--global&lt;/code&gt; install printed a fabricated "PAYLOAD MANIFEST" claiming 216 personas, 122 skills, and more were now "active," while writing none of them. The success banner had been counting the source package tree instead of what a global install actually writes by design — an entry file, commands, and subagents only. It's fixed now, replaced with an honest message describing exactly what a global install does and doesn't include. The HIGH-severity fixes closed a related gap — an install banner still claiming a confirmed no-op feature was "active," unhedged, one line below a properly-hedged claim about a different feature — plus a crash fix in one of the dynamic-workflow scripts and two new regression tests written specifically to keep the godmode-removal and &lt;code&gt;--minimal&lt;/code&gt; fixes from 11.9.9 from silently regressing. The MEDIUM/LOW list is the unglamorous kind of hardening that doesn't make headlines but is exactly what "ready for real users" should mean in practice: a missed curl-chmod-run dropper-chain pattern in the install-time security gate, a prototype-pollution guard added to a shared config singleton, a CodeQL-flagged regex-escape bug, a non-LTS base image reverted, a persona-table doc regenerated wholesale from source instead of hand-maintained, and a release-pipeline step that had been silently failing on four consecutive prior releases for a root-caused, now-fixed reason.&lt;/p&gt;

&lt;p&gt;So, the honest final answer to what this release means: a mature, iteratively-hardened orchestration layer, now on its 77th npm publish, just finished a second full external-readiness audit, found and fixed one more real critical bug in the process, and shipped — tagged, published, live — as exactly what its own CHANGELOG calls itself: the first release cut deliberately for real external users, not just internal iteration. That's a more credible launch story than an inflated "v1.0" narrative would ever be, precisely because the discipline of catching real bugs before shipping to real users kept recurring right up to the moment this one actually went out the door.&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>claudecode</category>
      <category>agenttooling</category>
      <category>developertools</category>
    </item>
    <item>
      <title>tracehub-mcp: Giving AI Assistants a Real Query Interface Into Your LLM Traces</title>
      <dc:creator>SAI RAM</dc:creator>
      <pubDate>Sun, 13 Sep 2026 14:06:24 +0000</pubDate>
      <link>https://dev.to/sai_ram_0000/tracehub-mcp-giving-ai-assistants-a-real-query-interface-into-your-llm-traces-22ck</link>
      <guid>https://dev.to/sai_ram_0000/tracehub-mcp-giving-ai-assistants-a-real-query-interface-into-your-llm-traces-22ck</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe3vf1hjds2jnth2q1d2l.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe3vf1hjds2jnth2q1d2l.gif" alt="Debugging LLM traces manually" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Copy-Paste Problem
&lt;/h2&gt;

&lt;p&gt;Here's what debugging an LLM application looks like for most people today, including me until recently: something is slow, or wrong, or expensive, and the AI assistant sitting in your editor wants to help. It can read your code. It can read your logs if you paste them in. What it cannot do, natively, is ask your trace backend a question. So you open Jaeger or Datadog in a browser tab, find the trace, copy the JSON blob, and paste it into the chat window. Then you do it again for the next trace, because "compare the last five calls to this model" isn't a query you can run — it's five manual round trips.&lt;/p&gt;

&lt;p&gt;tracehub-mcp exists to delete that loop. It's an MCP (Model Context Protocol) server that gives Claude, Cursor, Windsurf, Gemini CLI, or any MCP client a direct query interface into your OpenTelemetry trace backend. Instead of pasting JSON, the assistant calls a tool — &lt;code&gt;search_traces&lt;/code&gt;, &lt;code&gt;get_llm_expensive_traces&lt;/code&gt;, &lt;code&gt;get_llm_model_stats&lt;/code&gt; — and gets back structured trace data it can actually reason about. Find expensive calls, debug errors, compare model performance across providers, track token usage over a time window: these become questions you ask, not spreadsheets you build by hand.&lt;/p&gt;

&lt;p&gt;The part that matters most for an LLM-observability tool specifically is that it doesn't treat a trace as an anonymous bag of span attributes. It speaks OpenTelemetry's &lt;code&gt;gen_ai.*&lt;/code&gt; semantic conventions natively: prompts and completions pulled from actual span events, a token-usage fallback chain for providers that don't emit a clean total, finish reasons as a queryable field rather than a string buried in an attributes dictionary. That's why this is worth building over pointing an assistant at a generic OTel viewer, and it's the throughline for the rest of this piece: the five backends, the hardening bar, and the scope decision about which parts of the code get held to it.&lt;/p&gt;

&lt;p&gt;Today is 2026-09-13. tracehub-mcp is at version 0.3.0, live on PyPI, tagged &lt;code&gt;v0.3.0&lt;/code&gt; — the shipping version I'm describing, not a changelog recap of getting from 0.2 to 0.3.&lt;/p&gt;

&lt;h2&gt;
  
  
  From Three Backends to Five, and What "Full Implementation" Means
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu5madpvn4cjzygsclh8a.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu5madpvn4cjzygsclh8a.gif" alt="Debugging LLM traces with an AI assistant" width="384" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;tracehub-mcp didn't start as a five-backend project. It inherited three — Jaeger, Grafana Tempo, and Traceloop — and the growth story worth telling is how it became a five-backend, security-hardened server without the two new backends being thin wrappers bolted on for coverage.&lt;/p&gt;

&lt;p&gt;Every backend implements the same &lt;code&gt;BaseBackend&lt;/code&gt; interface (&lt;code&gt;src/opentelemetry_mcp/backends/base.py&lt;/code&gt;), so every MCP tool works identically no matter which backend is configured — &lt;code&gt;search_traces&lt;/code&gt;, &lt;code&gt;search_spans&lt;/code&gt;, &lt;code&gt;get_trace&lt;/code&gt;, &lt;code&gt;list_services&lt;/code&gt;, &lt;code&gt;get_service_operations&lt;/code&gt;, &lt;code&gt;health_check&lt;/code&gt;, whether you're pointed at a local Jaeger container with no auth or a paid Datadog account with an API key and app key pair. Jaeger and Tempo cover local/self-hosted setups; Traceloop covers cloud LLM observability with API-key auth. Those three came from the upstream fork.&lt;/p&gt;

&lt;p&gt;Datadog and Sentry are new, and genuinely new: full implementations of all four MCP-facing operations for each, not stubs that return "not implemented" past a basic search. &lt;code&gt;src/opentelemetry_mcp/backends/datadog.py&lt;/code&gt; and &lt;code&gt;src/opentelemetry_mcp/backends/sentry.py&lt;/code&gt; each implement &lt;code&gt;search_traces&lt;/code&gt;, &lt;code&gt;search_spans&lt;/code&gt;, &lt;code&gt;get_trace&lt;/code&gt;, and &lt;code&gt;list_services&lt;/code&gt; as independent code paths with their own request building, pagination, and parsing logic. That's the bar I hold "full backend" to: every tool a client can call has to actually work, not just the one that made for a good demo.&lt;/p&gt;

&lt;p&gt;The more interesting decision was holding those two backends to a security bar that's specific rather than generic. "We hardened it" means nothing on its own. Here's what it concretely means:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;HTTPS-only enforcement.&lt;/strong&gt; Both refuse to construct at all if the configured URL isn't &lt;code&gt;https://&lt;/code&gt; — a &lt;code&gt;ValueError&lt;/code&gt; raised in &lt;code&gt;__init__&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Datadog backend requires an https:// URL - DD-API-KEY and &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DD-APPLICATION-KEY must not be sent over plain http&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Sentry has the identical check, worded for its single bearer token instead of an API-key pair. Datadog goes further and disables automatic redirect-following on its &lt;code&gt;httpx&lt;/code&gt; client, because its non-standard &lt;code&gt;DD-API-KEY&lt;/code&gt;/&lt;code&gt;DD-APPLICATION-KEY&lt;/code&gt; headers aren't stripped by &lt;code&gt;httpx&lt;/code&gt; on a cross-origin redirect the way a standard &lt;code&gt;Authorization&lt;/code&gt; header is — a subtlety Sentry's own module docstring calls out as the reason it doesn't need the same override. The hardening was reasoned about per-backend, not copy-pasted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Query-injection-safe escaping.&lt;/strong&gt; Every value spliced into a Datadog query goes through &lt;code&gt;_escape_dd_query_value&lt;/code&gt;, which escapes embedded backslashes and quotes and wraps the result in exact quotes. Field &lt;em&gt;names&lt;/em&gt; — the part you'd assume is safer because it isn't free text — go through a separate allowlist, &lt;code&gt;_VALID_DD_FIELD_RE = re.compile(r"^@?[A-Za-z0-9_.]+$")&lt;/code&gt;, before they're anywhere near the query string, because a filter's &lt;code&gt;field&lt;/code&gt; parameter is an unvalidated string reachable from any MCP tool call. Sentry has the equivalent pair, &lt;code&gt;_escape_sentry_query_value&lt;/code&gt; and &lt;code&gt;_VALID_SENTRY_FIELD_RE&lt;/code&gt;, adapted to Sentry's syntax.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bounded pagination.&lt;/strong&gt; Datadog's span search loops through cursor pages capped at &lt;code&gt;_MAX_SEARCH_PAGES = 10&lt;/code&gt;, stopping when the requested &lt;code&gt;limit&lt;/code&gt; is satisfied or the cursor (&lt;code&gt;meta.page.after&lt;/code&gt;) runs out, whichever comes first, and logs a warning if it hits the cap with more data available:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;_MAX_SEARCH_PAGES&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;remaining&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;limit&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;collected&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;remaining&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;
    &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="n"&gt;cursor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;page_meta&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;after&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;isinstance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;cursor&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;warning&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Search truncated at %d pages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_MAX_SEARCH_PAGES&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Sentry does the same shape of loop through HTTP &lt;code&gt;Link&lt;/code&gt;-header cursors instead, matching how its pagination actually works, also capped at 10 pages. A single tool call can't turn into an unbounded crawl against either backend, and the code says so out loud when it truncates rather than returning a partial result silently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Exact-ID re-verification.&lt;/strong&gt; This one has the sharpest story. Datadog has no "get trace by ID" endpoint, so &lt;code&gt;get_trace&lt;/code&gt; reconstructs a trace by searching for spans matching the trace ID and grouping the results — inherently a search, which means it could match more broadly than intended. So before any span leaves the function, it's re-checked against the exact ID that was asked for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Belt-and-suspenders: only keep spans that exactly match the
# requested trace_id, in case the query above ever matches more
# broadly than intended.
&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;span&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;trace_id&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;trace_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;spans&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;span&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Sentry's &lt;code&gt;get_trace&lt;/code&gt; does the same re-check after calling its native &lt;code&gt;/trace/{id}/&lt;/code&gt; endpoint, and its parser goes one step further: it rejects a trace item that's missing its own &lt;code&gt;trace_id&lt;/code&gt; rather than falling back to the requested ID, because substituting the requested ID would make the whole re-verification check a no-op.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No fabricated data on malformed responses.&lt;/strong&gt; Both backends reject a span outright — return &lt;code&gt;None&lt;/code&gt;, log a warning — rather than filling in a placeholder. Datadog's parser has explicit rejection blocks for missing timestamps and for missing service/operation identity, with comments stating why: a substituted &lt;code&gt;now()&lt;/code&gt; would silently corrupt trace ordering, a substituted &lt;code&gt;"unknown"&lt;/code&gt; would silently merge unrelated spans. Sentry mirrors this and adds a sanity bound on duration — anything past ten years in milliseconds is rejected as almost certainly a clock-skew artifact.&lt;/p&gt;

&lt;p&gt;None of this landed in one pass. Two rounds of adversarial CodeRabbit review hit the Datadog backend: the first cites eight actionable findings addressed (HTTPS enforcement, redirect-disabling, exact-ID re-verification, cursor pagination, trace-ID escaping, timestamp rejection, and others); the second, an explicit follow-up after CodeRabbit re-reviewed the code the first round had just introduced, addressed four more — query escaping, response-shape validation, and &lt;code&gt;get_trace&lt;/code&gt;'s pagination limit handling, the same wording in the commit subject line. A hardening pass can introduce its own gaps; it's worth re-reviewing the fix, not just the original code.&lt;/p&gt;

&lt;p&gt;One honest caveat on Sentry: its module docstring says plainly that several response-shape assumptions — exact column names in raw event rows, the shape of Sentry's &lt;code&gt;SerializedTraceItem&lt;/code&gt; type, the escaping convention its Discover query syntax expects — were built from published documentation and source reading, not verified against a live Sentry account. That's disclosed in the code, and I'm disclosing it here rather than implying this backend has been battle-tested in production.&lt;/p&gt;

&lt;p&gt;The test suites reflect the same weight: &lt;code&gt;tests/test_datadog.py&lt;/code&gt; runs 781 lines and &lt;code&gt;tests/test_sentry.py&lt;/code&gt; runs 1117, with named tests for exactly the things above — &lt;code&gt;test_datadog_backend_rejects_non_https_url&lt;/code&gt;, &lt;code&gt;test_equals_rejects_injection_field&lt;/code&gt;, &lt;code&gt;test_escapes_trace_id_in_query&lt;/code&gt;, &lt;code&gt;test_requests_full_pagination_capacity&lt;/code&gt;, &lt;code&gt;test_malformed_data_field_does_not_crash&lt;/code&gt;. That's the difference between a README bullet list and something a future contributor can actually run and watch fail if they break it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Decision Not to Touch Jaeger, Tempo, or Traceloop
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe0zzann7sjqqmj9lh1ye.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe0zzann7sjqqmj9lh1ye.gif" alt="AI assistant query interface and traces" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The part I think is actually the most interesting design decision in the project: I did not retrofit any of that discipline onto the three inherited backends.&lt;/p&gt;

&lt;p&gt;Open &lt;code&gt;jaeger.py&lt;/code&gt; next to &lt;code&gt;datadog.py&lt;/code&gt; and the gap is immediate. Four lines into Jaeger's span parser, the default service name is the literal string &lt;code&gt;"unknown"&lt;/code&gt; — if neither the processes map nor the span's own &lt;code&gt;process&lt;/code&gt; field resolves a real name, that placeholder ships as-is rather than the span being rejected. &lt;code&gt;traceloop.py&lt;/code&gt; has the same shape of problem: it falls back to an empty string for &lt;code&gt;service.name&lt;/code&gt; when the field is absent. Neither Jaeger nor Tempo has an HTTPS check anywhere — &lt;code&gt;BaseBackend.__init__&lt;/code&gt; just stores the URL, no scheme validation, so both will happily talk to a plain &lt;code&gt;http://&lt;/code&gt; endpoint. Tempo's TraceQL builder, &lt;code&gt;_filter_to_traceql&lt;/code&gt;, splices filter values directly into f-strings with no escaping function anywhere in the file — the exact pattern &lt;code&gt;_escape_dd_query_value&lt;/code&gt; and &lt;code&gt;_escape_sentry_query_value&lt;/code&gt; exist to close off.&lt;/p&gt;

&lt;p&gt;I could have quietly left that gap unaddressed and let the README imply the hardening covers "the backends." I could also have spent a week retrofitting Jaeger, Tempo, and Traceloop to match. I did neither. The README states the position directly:&lt;/p&gt;

&lt;p&gt;"The three backends inherited from upstream — Jaeger, Tempo, Traceloop — predate this discipline and haven't been retrofitted; that's deliberate scope discipline, not an oversight, mirroring this project's own precedent of not reaching into shared/inherited code without full regression coverage for it."&lt;/p&gt;

&lt;p&gt;There's a specific kind of dishonesty in open source that's more common than outright lying: a security-adjacent hardening pass described as covering "the project" when it actually covers the two files someone happened to be working in that week. The work that got done, got done — but the framing implies a blanket guarantee that was never earned. The honest alternative isn't silence about the gap, and it isn't burning a week retrofitting code you don't have the depth of intuition for well enough to trust the retrofit. It's saying plainly which parts got the bar and which didn't, and why. "These three predate this discipline" is a true, checkable statement — I just checked it against the actual source above. "Zero known issues" would not have been.&lt;/p&gt;

&lt;p&gt;That "full regression coverage" framing matters because Jaeger, Tempo, and Traceloop are code I didn't write and don't have the same depth of intuition for. Reaching in to add HTTPS checks or escaping without a test suite built to catch the ways I could break it quietly is how a regression goes unnoticed until a production trace query starts silently dropping spans — which is exactly what happened in a different corner of the codebase.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bugs That Taught Me the Hard Way
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fldejzm435wq2rxtfea3e.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fldejzm435wq2rxtfea3e.gif" alt="AI assistant querying trace data" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The hardening bar on Datadog and Sentry didn't come from first principles. It came partly from bugs I hit and fixed in the inherited code, which is why I trust the no-fabrication rule enough to enforce it strictly on new code.&lt;/p&gt;

&lt;p&gt;The sharpest one: any span carrying a &lt;code&gt;gen_ai.response.finish_reasons&lt;/code&gt; value — any LLM span representing a completed response — was silently disappearing from every query result on both Jaeger and Tempo. Not erroring, disappearing. I found it during a live end-to-end dry run against a real Jaeger container and the actual published PyPI package, not a unit test. The root cause was a type mismatch: &lt;code&gt;finish_reasons&lt;/code&gt; is typed &lt;code&gt;list[str]&lt;/code&gt; in &lt;code&gt;attributes.py&lt;/code&gt;'s &lt;code&gt;SpanAttributes&lt;/code&gt; model, but Jaeger hands the value over as a JSON-encoded string, and Tempo's OTLP parser was stringifying the raw &lt;code&gt;arrayValue&lt;/code&gt; structure into an unparseable repr instead of extracting the actual elements. Both shapes failed Pydantic validation, and both backends wrapped span construction in a broad &lt;code&gt;except&lt;/code&gt; that returned &lt;code&gt;None&lt;/code&gt; — so a parse failure and "no span here" were indistinguishable. The fix added a &lt;code&gt;field_validator&lt;/code&gt; that coerces any of those shapes into a real list without raising, and rewrote Tempo's &lt;code&gt;arrayValue&lt;/code&gt; handling to pull each element's scalar value by key presence rather than truthiness, so a real &lt;code&gt;0&lt;/code&gt; or &lt;code&gt;False&lt;/code&gt; doesn't get dropped along with it.&lt;/p&gt;

&lt;p&gt;That bug is why the no-fabrication rule on Datadog and Sentry isn't theoretical caution — it's a lesson learned on the backends that predate it, and it's a large part of why I trust the new backends' stricter behavior (reject the span, log why) over the older ones' looser one (silently return nothing).&lt;/p&gt;

&lt;p&gt;The rest of this arc is smaller but real. An earlier fix replaced sequential &lt;code&gt;get_trace&lt;/code&gt; round trips with a single &lt;code&gt;asyncio.gather&lt;/code&gt; call across the four tools that hydrate full traces after a search — &lt;code&gt;expensive_traces.py&lt;/code&gt;, &lt;code&gt;list_models.py&lt;/code&gt;, &lt;code&gt;model_stats.py&lt;/code&gt;, &lt;code&gt;slow_traces.py&lt;/code&gt; — turning N sequential calls into one concurrent batch, with a regression test asserting the maximum in-flight call count equals the batch size. The same commit deleted a dead &lt;code&gt;else 0&lt;/code&gt; branch in &lt;code&gt;model_stats.py&lt;/code&gt;'s success-rate calculation, left over from before an earlier guard made it unreachable. On the CI side, &lt;code&gt;v0.3.0&lt;/code&gt; closed off a pattern in &lt;code&gt;docker-publish.yml&lt;/code&gt; where &lt;code&gt;${{ github.ref }}&lt;/code&gt; and &lt;code&gt;${{ github.run_id }}&lt;/code&gt; were textually inlined into shell script text rather than passed through environment variables — not an exploited vulnerability here, since neither value is attacker-influenceable in this step, but a known-dangerous pattern worth closing regardless. The same release removed a step from that workflow that was supposed to flip the published GHCR image to public visibility but always reported success even though its &lt;code&gt;gh api&lt;/code&gt; calls 404'd under the default &lt;code&gt;GITHUB_TOKEN&lt;/code&gt;, chained with &lt;code&gt;||&lt;/code&gt; into a trailing &lt;code&gt;echo&lt;/code&gt; so the exit code was always zero. Green CI, private package, until someone checked; I removed the fake step and documented the one-time manual flip instead of fake-automating a fix. And &lt;code&gt;start_locally.sh&lt;/code&gt; had a live Traceloop API key hardcoded as the active default — inherited from before the fork point, still live on the upstream public repo, not mine to rotate — so the fix was scoped honestly: switch the default to the credential-free Jaeger backend, without claiming to have remediated an exposure that predates this repository. The Docker image now publishes to GHCR on every version tag, for &lt;code&gt;linux/amd64&lt;/code&gt; and &lt;code&gt;linux/arm64&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fork, the Pivot, and a Changelog Anomaly
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feqk3nfmxaxyagz7lnuvd.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feqk3nfmxaxyagz7lnuvd.gif" alt="A developer looking thoughtfully at trace logs" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;tracehub-mcp began as &lt;code&gt;traceloop/opentelemetry-mcp-server&lt;/code&gt;, Apache 2.0, with full attribution and fork history preserved in &lt;code&gt;NOTICE&lt;/code&gt; — created via &lt;code&gt;git clone&lt;/code&gt; plus &lt;code&gt;git remote&lt;/code&gt;, not by copying files into a fresh repository, so the original commit history stays visible instead of collapsing into one "initial commit." The first commit landed 2025-11-02, with a first release, &lt;code&gt;v0.2.0&lt;/code&gt;, on 2025-11-17. There was a long gap, then three commits on 2026-09-12 completed the rename and pivot to &lt;code&gt;mcpsmiths/tracehub-mcp&lt;/code&gt;: a fork-setup commit adding &lt;code&gt;NOTICE&lt;/code&gt;, a rename commit, and a follow-up, &lt;code&gt;fix: complete the tracehub-mcp rename, remove old-org CI dependencies&lt;/code&gt;, closing out roughly 90 stale references to the old name across the README, &lt;code&gt;CLAUDE.md&lt;/code&gt;, the Dockerfile, &lt;code&gt;start_locally.sh&lt;/code&gt;, &lt;code&gt;server.py&lt;/code&gt;, and CI, and dropping CI's dependency on the upstream org's private infrastructure — a self-hosted runner label, a private composite Trivy action, a GitHub App for version bumps — in favor of GitHub-hosted runners and the repo's own token and git identity.&lt;/p&gt;

&lt;p&gt;That pivot also explains something you'll notice scrolling &lt;code&gt;CHANGELOG.md&lt;/code&gt;: two separate &lt;code&gt;v0.2.0&lt;/code&gt;/&lt;code&gt;v0.2.1&lt;/code&gt;/&lt;code&gt;v0.2.2&lt;/code&gt; entries months apart, one set from November 2025 through February 2026, another from September 2026. Resetting &lt;code&gt;pyproject.toml&lt;/code&gt;'s version to &lt;code&gt;0.1.0&lt;/code&gt; during the rename left the Commitizen version files still pointing at the old &lt;code&gt;0.2.2&lt;/code&gt;, so the automated bumper started counting up again on the new codebase — a real, mechanical drift, confirmed by dry-running the actual bump logic rather than assumed, and fixed in a commit that also caught an unrelated bug where a separate "configure git identity" CI step did nothing, because the commitizen action's own entrypoint re-runs &lt;code&gt;git config&lt;/code&gt; with its own defaults right before committing. I'd rather the changelog look odd and be explained than look clean because I quietly rewrote history.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Actually Is
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F659asvqvy4hw468iwxh5.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F659asvqvy4hw468iwxh5.gif" alt="A developer troubleshooting an issue" width="216" height="216"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I ran the test suite directly against this checkout rather than trusting the README: 467 passed, 2 skipped, 87% overall coverage — the README currently says 458, which is stale, and I'm correcting it here instead of repeating it. The two skips are environment-dependent integration tests needing cassette fixture data not present in this run, not disabled tests. There are 5 open pull requests right now, every one a Dependabot bump to a GitHub Action's minor or patch version — &lt;code&gt;docker/login-action&lt;/code&gt;, &lt;code&gt;actions/cache&lt;/code&gt;, &lt;code&gt;sticky-pull-request-comment&lt;/code&gt;, &lt;code&gt;astral-sh/setup-uv&lt;/code&gt;, &lt;code&gt;trivy-action&lt;/code&gt; — routine and low-risk. It's listed on the official MCP registry as &lt;code&gt;io.github.mcpsmiths/tracehub-mcp&lt;/code&gt; and carries a glama.ai score badge. It has 0 GitHub stars as of today.&lt;/p&gt;

&lt;p&gt;What I'm actually claiming is narrower than "production-hardened observability platform," and more useful for being narrow: two new backends built to a specific, checkable security bar, three inherited backends explicitly not held to that bar yet, a real data-loss bug found and fixed the hard way, and a test suite I ran myself before writing any of these numbers down. If you're running an LLM application on Datadog or Sentry and want your AI assistant to query the actual trace data instead of you pasting it in by hand, that's what this is for. If you're on Jaeger, Tempo, or Traceloop, it works today, on the discipline the fork already had — just don't assume it has the newer bar until I've said it does.&lt;/p&gt;

</description>
      <category>aiengineering</category>
      <category>mcp</category>
      <category>observability</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Why I Built cost-guard-mcp: Pre-Flight Cost Guardrails for AI Agents Talking to Data Warehouses</title>
      <dc:creator>SAI RAM</dc:creator>
      <pubDate>Sun, 13 Sep 2026 14:02:20 +0000</pubDate>
      <link>https://dev.to/sai_ram_0000/why-i-built-cost-guard-mcp-pre-flight-cost-guardrails-for-ai-agents-talking-to-data-warehouses-25f2</link>
      <guid>https://dev.to/sai_ram_0000/why-i-built-cost-guard-mcp-pre-flight-cost-guardrails-for-ai-agents-talking-to-data-warehouses-25f2</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuqi2vqbnqf2nfmr0bddc.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuqi2vqbnqf2nfmr0bddc.gif" alt="Description of the GIF" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The query that shouldn't have run
&lt;/h2&gt;

&lt;p&gt;Give an AI agent a warehouse MCP server and, sooner or later, it will write a query that scans a full multi-terabyte table and quietly runs up a bill in the hundreds of dollars because it forgot a partition filter, or it will &lt;code&gt;SELECT *&lt;/code&gt; from something it thought was small and get back four million rows straight into its own context window. Neither of these requires a malicious agent or a bad prompt. It just requires an agent doing what agents do: writing plausible SQL against a schema it partially understands, then running it. A human DBA would eyeball the query, guess the table size, maybe run &lt;code&gt;EXPLAIN&lt;/code&gt; first out of habit. An agent given a generic "run this SQL" tool has no equivalent instinct, and as far as I can tell, no warehouse MCP server on the market tells the agent what a query will cost or how much data it will return before that query actually executes.&lt;/p&gt;

&lt;p&gt;That gap is the entire reason cost-guard-mcp exists. It is a Model Context Protocol server that sits in front of BigQuery and Snowflake and gives an agent three tools: &lt;code&gt;describe_engine_capabilities(engine)&lt;/code&gt; to learn what an engine can and can't tell you, &lt;code&gt;estimate_query_cost(engine, sql, warehouse)&lt;/code&gt; to get a cost and byte estimate before running anything, and &lt;code&gt;run_query_bounded(engine, sql, max_bytes_billed, max_rows, max_estimated_cost_usd)&lt;/code&gt; to actually execute a query with those caps enforced. It is a genuinely small project — 15 Python files, 715 lines under &lt;code&gt;src/cost_guard_mcp/&lt;/code&gt; — and it is genuinely new. The first commit landed 2026-09-12 at 22:23 IST, the feature-complete v1 landed a few hours later at 02:27, v0.1.0 shipped that same day, and v0.1.1 — the hardening pass I'll get to below — shipped today, 2026-09-13. The whole history of this project fits inside about eighteen hours. I'm not going to pretend otherwise; there's no multi-month backstory here, just a focused tool built fast and then immediately hardened.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj1tueyp5tptt6e423brs.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj1tueyp5tptt6e423brs.gif" alt="Alternative description of the GIF" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  An estimate that tells you how much to trust itself
&lt;/h2&gt;

&lt;p&gt;The design decision I care about most in this codebase isn't the warehouse integration — it's that &lt;code&gt;estimate_query_cost&lt;/code&gt; never returns a bare number. Every &lt;code&gt;CostEstimate&lt;/code&gt; it produces carries an &lt;code&gt;accuracy_tier&lt;/code&gt; field, and that field is not a comment or a README promise, it's a required field on the Pydantic model in &lt;code&gt;src/cost_guard_mcp/types.py&lt;/code&gt; with no default. You cannot construct a &lt;code&gt;CostEstimate&lt;/code&gt; without deciding what tier it belongs to. AGENTS.md in the repo states the intent behind this directly: it's "enforced structurally by the Pydantic model, not by convention." That distinction matters more than it sounds. A convention is something a future contributor forgets. A required field with no default is something the type checker refuses to let you skip.&lt;/p&gt;

&lt;p&gt;There are three tiers: &lt;code&gt;PRECISE&lt;/code&gt;, &lt;code&gt;UPPER_BOUND&lt;/code&gt;, and &lt;code&gt;HEURISTIC&lt;/code&gt;. BigQuery is where &lt;code&gt;PRECISE&lt;/code&gt; actually happens, and it happens because BigQuery's dry-run API is not a client-side guess — it's a real submission to BigQuery's own query planner. &lt;code&gt;dry_run()&lt;/code&gt; in &lt;code&gt;src/cost_guard_mcp/engines/bigquery.py&lt;/code&gt; builds a &lt;code&gt;bigquery.QueryJobConfig(dry_run=True, use_query_cache=False)&lt;/code&gt;, sends it through &lt;code&gt;client.query()&lt;/code&gt;, and reads back &lt;code&gt;total_bytes_processed&lt;/code&gt; from the real API response. BigQuery validates and fully plans the query server-side without executing or billing it, so the byte figure that comes back is the same figure the real execution would have produced. That's what earns the &lt;code&gt;PRECISE&lt;/code&gt; label.&lt;/p&gt;

&lt;p&gt;Snowflake gets a structurally different treatment, and the code is explicit about why. &lt;code&gt;explain_estimate()&lt;/code&gt; runs &lt;code&gt;EXPLAIN USING JSON {sql}&lt;/code&gt;, parses the returned plan, and reads &lt;code&gt;plan['GlobalStats']['bytesAssigned']&lt;/code&gt; as its byte figure — but that number describes what the query planner expects to scan, not what actually gets scanned, and Snowflake's own documentation (quoted directly in a code comment) says runtime plan optimizations "can reduce the number of partitions and bytes scanned." Because of that, &lt;code&gt;explain_estimate()&lt;/code&gt; doesn't have a downgrade table the way BigQuery does — it hardcodes &lt;code&gt;AccuracyTier.UPPER_BOUND&lt;/code&gt; on line 84, unconditionally, because there is no better tier available to fall from. BigQuery's dry-run and Snowflake's EXPLAIN are answering genuinely different questions: one is "what will this actually cost," the other is "what is the most this should cost."&lt;/p&gt;

&lt;p&gt;The place this stops being an abstract distinction and becomes something you can watch happen is inside BigQuery itself. BigQuery's dry-run response includes its own confidence signal — &lt;code&gt;totalBytesProcessedAccuracy&lt;/code&gt; — and &lt;code&gt;dry_run()&lt;/code&gt; reads it via &lt;code&gt;query_job._properties['statistics']['query']['totalBytesProcessedAccuracy']&lt;/code&gt;, defaulting to &lt;code&gt;'UNKNOWN'&lt;/code&gt; if it's missing, then maps it through a fixed table: &lt;code&gt;PRECISE&lt;/code&gt; stays &lt;code&gt;PRECISE&lt;/code&gt;; &lt;code&gt;LOWER_BOUND&lt;/code&gt;, &lt;code&gt;UPPER_BOUND&lt;/code&gt;, &lt;code&gt;UNKNOWN&lt;/code&gt;, and anything else all fall through to &lt;code&gt;AccuracyTier.UPPER_BOUND&lt;/code&gt; via the dict's own default. A unit test, &lt;code&gt;test_dry_run_treats_unrecognized_accuracy_value_as_upper_bound&lt;/code&gt;, feeds it a value BigQuery has never returned before — &lt;code&gt;'SOME_FUTURE_VALUE'&lt;/code&gt; — and asserts the tool still downgrades safely rather than defaulting to &lt;code&gt;PRECISE&lt;/code&gt;. The practical consequence: run the exact same SQL text against an ordinary settled table and you get &lt;code&gt;PRECISE&lt;/code&gt; with an empty &lt;code&gt;caveats&lt;/code&gt; list. Run that identical SQL against a table with a pending streaming buffer, or a wildcard, or a federated source, and the tool automatically flips to &lt;code&gt;UPPER_BOUND&lt;/code&gt; and attaches a caveat quoting BigQuery's own raw accuracy string verbatim — something like "BigQuery reported this estimate's own accuracy as 'LOWER_BOUND', not PRECISE — treating it conservatively as UPPER_BOUND." Nothing about the query changed. Only BigQuery's own confidence in its byte count changed, and the tool surfaces that shift instead of quietly reporting one undifferentiated number both times.&lt;/p&gt;

&lt;p&gt;The third tier, &lt;code&gt;HEURISTIC&lt;/code&gt;, is reserved for Databricks, and I want to be precise about its status: it exists in the type system — &lt;code&gt;types.py&lt;/code&gt; defines the enum value, and the agent-facing docstring in &lt;code&gt;server.py&lt;/code&gt; mentions it — but there is no &lt;code&gt;engines/databricks.py&lt;/code&gt;, no pricing module, nothing. &lt;code&gt;describe_engine_capabilities('databricks')&lt;/code&gt; raises &lt;code&gt;ValueError&lt;/code&gt;, and a test asserts exactly that. I deferred Databricks rather than ship it, because a Databricks estimate would necessarily be a heuristic guess with no dry-run and no EXPLAIN-equivalent bound behind it, and shipping a guess dressed up as a real number is precisely the failure mode this whole project exists to prevent. Better to leave the third tier as a load-bearing placeholder in the schema than half-honest in the field.&lt;/p&gt;

&lt;p&gt;The server itself bakes the sequencing into its own tool docstrings, not just into documentation: &lt;code&gt;describe_engine_capabilities&lt;/code&gt;'s docstring tells an agent to call it "before &lt;code&gt;estimate_query_cost&lt;/code&gt; or &lt;code&gt;run_query_bounded&lt;/code&gt; to understand how much to trust the &lt;code&gt;accuracy_tier&lt;/code&gt; on their responses for this engine," and &lt;code&gt;estimate_query_cost&lt;/code&gt;'s docstring spells out what each tier means in the response itself. That's the actual contract an agent operates against — not marketing copy, code the agent reads.&lt;/p&gt;

&lt;h2&gt;
  
  
  Refusing is safer than guessing
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdumxb8r17cxmf4i3xg4y.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdumxb8r17cxmf4i3xg4y.gif" alt="A short description of what is happening in the GIF" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The second thread running through this project is fail-closed, least-privilege design, and it shows up as a pattern rather than a single feature. The clearest instance is in &lt;code&gt;run_query_bounded&lt;/code&gt;. If a caller passes &lt;code&gt;max_estimated_cost_usd&lt;/code&gt; but the estimate that came back has &lt;code&gt;estimated_cost_usd=None&lt;/code&gt; — which happens for a BigQuery Editions/capacity-billed project, since capacity billing has no fixed dollar-per-byte rate to convert from — the tool doesn't run the query with that cap silently unenforced. It refuses. The check in &lt;code&gt;src/cost_guard_mcp/tools/run_query_bounded.py&lt;/code&gt; returns a normal &lt;code&gt;BoundedQueryResult&lt;/code&gt; with &lt;code&gt;status="refused"&lt;/code&gt; and &lt;code&gt;reason=RefusalReason.COST_CAP_EXCEEDED&lt;/code&gt;, plus a plain-English hint explaining exactly why: it couldn't produce a dollar estimate, so it's refusing rather than running uncapped. This is a normal, successful MCP tool result, not an exception and not the MCP error channel — that's a deliberate choice recorded in the project's own ADR #4, so a caller can distinguish "your query is too expensive" from "your MCP client is broken" cleanly.&lt;/p&gt;

&lt;p&gt;Two other examples of the same instinct: Snowflake connections require an explicit &lt;code&gt;SNOWFLAKE_ROLE&lt;/code&gt; environment variable with no default, and &lt;code&gt;config.py&lt;/code&gt; raises a &lt;code&gt;ConfigError&lt;/code&gt; if it's unset, specifically to prevent ever falling back to &lt;code&gt;ACCOUNTADMIN&lt;/code&gt;. And every function in both engine modules that touches &lt;code&gt;snowflake.connector&lt;/code&gt; or the BigQuery client libraries is wrapped by a &lt;code&gt;sanitize_exceptions&lt;/code&gt; decorator, which regex-redacts five categories of secret material — passwords, private keys and PEM blocks, tokens/API keys, and &lt;code&gt;user:password@host&lt;/code&gt; URIs — from any exception before it can reach a caller or a log line, because &lt;code&gt;snowflake-connector-python&lt;/code&gt; has a documented history of leaking credentials into raw exception text.&lt;/p&gt;

&lt;p&gt;The part of this story I find most worth being honest about is that two of the security fixes weren't found by a later audit — they were found and closed the same day the vulnerable code shipped. Git history has the receipts: a commit titled &lt;code&gt;fix(security): validate Snowflake warehouse name before USE WAREHOUSE&lt;/code&gt;, because the &lt;code&gt;warehouse&lt;/code&gt; parameter — caller-supplied, same trust level as the SQL itself — was being interpolated straight into &lt;code&gt;USE WAREHOUSE {warehouse}&lt;/code&gt;. That statement is executed directly against the live Snowflake session via &lt;code&gt;cur.execute()&lt;/code&gt;, so an unvalidated warehouse name wasn't just a cosmetic risk: a caller could have appended a semicolon and a second statement after the intended identifier, turning a parameter meant to just pick a compute resource into a general SQL execution primitive scoped by whatever the connection's role could already reach. The second commit, &lt;code&gt;fix(bigquery): validate project ID before interpolating into API filter&lt;/code&gt;, whose own message says it was "flagged by automated security review after the previous commit introduced the query filter," closes a related but structurally different gap: the GCP project ID isn't actually reachable from an MCP tool parameter today, but the code treats the identifier as if it could be — validated on the theory that anything string-interpolated into an API filter expression deserves the same allowlist treatment as untrusted input, not just the fields a caller happens to be able to reach right now. Both are closed with allowlist regex validators — &lt;code&gt;_validate_warehouse()&lt;/code&gt; and &lt;code&gt;_validate_project_id()&lt;/code&gt; — that fail loudly with a &lt;code&gt;ValueError&lt;/code&gt; (itself sanitized) rather than let a malformed identifier reach a query-language context. Both incidents are written up together as ADR #6 in &lt;code&gt;DECISIONS.md&lt;/code&gt;, which turns them into a standing rule for any engine this project adds in the future.&lt;/p&gt;

&lt;p&gt;One more mechanism worth naming because it's small and clever: row-bounded execution on both engines wraps the caller's SQL as &lt;code&gt;SELECT * FROM (&amp;lt;sql&amp;gt;) AS cost_guard_row_cap LIMIT max_rows + 1&lt;/code&gt; before fetching. The &lt;code&gt;+1&lt;/code&gt; is the whole trick — fetch exactly &lt;code&gt;max_rows&lt;/code&gt; and you can't tell "there were exactly that many rows" apart from "there were more, and we truncated." Fetch more than &lt;code&gt;max_rows + 1&lt;/code&gt; and the row cap stops doing its job of keeping the fetch itself bounded. This pattern trips Ruff's bandit-style SQL-injection ruleset (&lt;code&gt;S608&lt;/code&gt;) in exactly two places, once per engine, and both are suppressed inline with &lt;code&gt;# noqa: S608&lt;/code&gt; because the string interpolation is the caller's own already-validated SQL being wrapped, not an unvalidated identifier.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hardening the guardrail tool, immediately
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbyz4uwkyn2tm8ucygis7.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbyz4uwkyn2tm8ucygis7.gif" alt="A short description of what is happening in the third GIF" width="400" height="300"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I shipped v0.1.0 on 2026-09-12. One day later, before the project had a single GitHub star, I ran a security pass against my own new code and shipped v0.1.1 with PyPI Trusted Publishing (OIDC via &lt;code&gt;id-token: write&lt;/code&gt;, scoped to the &lt;code&gt;pypi&lt;/code&gt; environment, with zero stored API token anywhere in &lt;code&gt;release.yml&lt;/code&gt;), CodeQL scanning on a weekly cron, OpenSSF Scorecard with SARIF results uploaded to GitHub's Security tab, secret scanning, push protection, Dependabot security updates, and branch protection requiring CI to pass on main. All of that is free for a public repository and was simply off for the first day. Concretely: secret scanning and push protection are what would catch an accidentally committed API key or private key before it ever settles into git history, not after; Dependabot security updates are what would catch a CVE landing in a transitive dependency next month, long after I've stopped actively watching this repo. OpenSSF Scorecard in particular runs a battery of named checks against the repo's own supply-chain hygiene — &lt;code&gt;Dangerous-Workflow&lt;/code&gt; for GitHub Actions patterns that let untrusted input execute privileged code, &lt;code&gt;Branch-Protection&lt;/code&gt; for exactly the main-branch CI gate I just turned on, and &lt;code&gt;Pinned-Dependencies&lt;/code&gt; for whether third-party actions are pinned to a mutable tag or a fixed SHA. It felt right to treat the guardrail tool itself as something that needed guardrails, and to do it before the project attracted any real attention rather than after.&lt;/p&gt;

&lt;p&gt;The CI structure separates required from advisory checks rather than treating everything as one gate: &lt;code&gt;tests.yml&lt;/code&gt; runs Ruff, mypy, and the unit suite with &lt;code&gt;--cov-fail-under=80&lt;/code&gt;, and that's what actually blocks a merge to main. &lt;code&gt;audit.yml&lt;/code&gt; runs &lt;code&gt;pip-audit&lt;/code&gt; but is deliberately not a required check — its own header comment explains the reasoning: a new CVE in an already-approved, unrelated dependency shouldn't block an unrelated PR from merging. &lt;code&gt;integration.yml&lt;/code&gt; is separate again, triggered only by &lt;code&gt;workflow_dispatch&lt;/code&gt; or a weekly schedule, and runs the five tests under &lt;code&gt;tests/integration/&lt;/code&gt; against real BigQuery and Snowflake credentials in a &lt;code&gt;live-integration&lt;/code&gt; environment. Those five tests are not part of the 60 unit tests I count as the project's real coverage number — I ran &lt;code&gt;pytest tests/ --ignore=tests/integration&lt;/code&gt; directly against this checkout and got 60 passed, 0 failed, and one asyncio deprecation warning coming from an early bootstrap test in &lt;code&gt;tests/unit/test_server.py&lt;/code&gt; that still calls &lt;code&gt;asyncio.get_event_loop()&lt;/code&gt; and says in its own comment that it "documents current state" pending cleanup — a small piece of unfinished housekeeping from the first day, not noise from an external dependency. The integration tests need cloud credentials I'm not going to hand to a CI runner on every PR, so they stay gated and scheduled instead.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;DECISIONS.md&lt;/code&gt; also documents a bug that's worth naming precisely because it's the same failure class as the fail-closed cap-check I described above: the original &lt;code&gt;run_query_bounded&lt;/code&gt; logic used a flat boolean chain that fell through to "allow" whenever a cap was requested but its matching estimate field was &lt;code&gt;None&lt;/code&gt;. An automated review caught it. The fix — refuse rather than silently execute unenforced — is now the standing convention for the whole tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this doesn't do yet, on purpose
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvfj79pxltua7o17bif2t.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvfj79pxltua7o17bif2t.gif" alt="A short description of what is happening in the fourth GIF" width="250" height="250"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I'd rather state the gaps plainly than let anyone find them the hard way. Databricks isn't supported — the &lt;code&gt;HEURISTIC&lt;/code&gt; tier is a type-level placeholder with zero engine code behind it, deferred past v1 because I'm not willing to ship a guess dressed as a measurement. Snowflake's &lt;code&gt;UPPER_BOUND&lt;/code&gt; estimate excludes Cortex AI Function ("AI Credits") cost entirely — &lt;code&gt;EXPLAIN&lt;/code&gt;'s &lt;code&gt;bytesAssigned&lt;/code&gt; reflects warehouse compute and scan, not AI-inference calls embedded in the SQL, and both the static capability lookup and the live caveat on every Snowflake estimate say so. And BigQuery projects on Editions or capacity billing get a byte estimate with no dollar figure at all, because capacity billing charges for reserved slot-hours, not bytes scanned — there's no fixed dollar-per-byte rate to convert with, so the tool declines to fabricate one rather than print a number that would be fiction.&lt;/p&gt;

&lt;p&gt;None of these are bugs waiting for a fix. They're the same accuracy-tier discipline applied to the project's own documentation that the project applies to its output: say what you don't know instead of rounding it up to something confident-sounding. cost-guard-mcp is two days old, has one GitHub star, zero open issues, and I built it because I got tired of agents finding out a query was expensive after it already ran. An estimate that admits what it doesn't know is more useful, and safer, than one that guesses and calls itself precise — that's the whole idea, and it's the same idea whether you're reading it off a &lt;code&gt;CostEstimate&lt;/code&gt; object or off this changelog.&lt;/p&gt;

</description>
      <category>aiengineering</category>
      <category>mcp</category>
      <category>python</category>
      <category>opensource</category>
    </item>
    <item>
      <title>trelix v3.2.2 to v3.2.5: The Source Tree Was Fine. The Published Package Wasn't.</title>
      <dc:creator>SAI RAM</dc:creator>
      <pubDate>Sun, 06 Sep 2026 12:29:30 +0000</pubDate>
      <link>https://dev.to/sai_ram_0000/trelix-v322-to-v325-the-source-tree-was-fine-the-published-package-wasnt-55i</link>
      <guid>https://dev.to/sai_ram_0000/trelix-v322-to-v325-the-source-tree-was-fine-the-published-package-wasnt-55i</guid>
      <description>&lt;p&gt;Run this against the real, published image and watch it fail:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;--rm&lt;/span&gt; &lt;span class="nt"&gt;--entrypoint&lt;/span&gt; trelix-mcp ghcr.io/sairam0424/trelix:3.2.1 &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1we64cepe98ckh9mhy5x.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1we64cepe98ckh9mhy5x.gif" alt=" " width="384" height="216"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Exit code 127. Not a crash inside trelix-mcp, not a stack trace, not a permissions error — &lt;code&gt;127&lt;/code&gt; is the shell's own way of saying the binary you asked for does not exist. And it didn't. The console script trelix-mcp is supposed to install as part of every trelix package was simply absent from the image, on both the slim tag and the &lt;code&gt;-local&lt;/code&gt; tag, for the entire life of the 3.2.1 release. Every unit test in the suite was green. Every line of source that builds trelix-mcp was correct. The thing a user would actually get from &lt;code&gt;docker pull&lt;/code&gt; did not have the binary its own &lt;code&gt;--version&lt;/code&gt; flag implies exists.&lt;/p&gt;

&lt;p&gt;This article covers four releases — v3.2.2, v3.2.3, v3.2.4, and v3.2.5 — spanning 173 commits and 88 changed files since v3.2.1, which is where the last article in this series left off. That one was about tests that pass without exercising the code they claim to cover: a &lt;code&gt;MagicMock&lt;/code&gt; standing in for a real embedder, an all-ones attention mask that makes masked and unmasked math identical, a unit test that asserted a bug as its own specification. This one, on the heels of the mutation-testing push that closed out that arc, is about a different and in some ways more uncomfortable failure mode: tests that pass while exercising the wrong artifact entirely. A green pytest run against &lt;code&gt;src/&lt;/code&gt; says nothing about whether the wheel on PyPI, the image on GHCR, or the binary on the GitHub Releases page actually does what it claims. Those are three separate build products, built by three separate pipelines, and none of trelix's 4,353 collected unit tests had ever touched any of them directly. v3.2.2 through v3.2.4 is the story of finding that gap and closing it with an actual gate, not a promise to be more careful next time. v3.2.5 is a short postscript proving the discipline stuck.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Docker image that shipped without its own server
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzfhacx5m4swplsfraelm.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzfhacx5m4swplsfraelm.gif" alt=" " width="500" height="367"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The 127 above wasn't hypothetical or reconstructed after the fact — it's the literal command a human ran by hand against the real published v3.2.1 image, now baked verbatim into &lt;code&gt;scripts/verify_release.py&lt;/code&gt;'s Docker check with a comment explaining why: "this exact command returned exit 127 on the published 3.2.1 image before trelix-mcp was added to the Dockerfile." Root cause was doubly blocked. The Dockerfile's builder stage never had a &lt;code&gt;COPY packages/trelix-mcp/&lt;/code&gt; line — it only ever copied and installed core trelix. And even if someone had added that line, the repo's &lt;code&gt;.dockerignore&lt;/code&gt; had a single bare entry, &lt;code&gt;packages/&lt;/code&gt;, that would have silently excluded the whole directory from the build context anyway. Two independent gaps, each individually sufficient to explain the missing binary, both present at once — the kind of thing an ordinary Dockerfile-diff review would not catch, since the second gap lives in a different file entirely. The fix, across commits &lt;code&gt;aabaa14&lt;/code&gt; and &lt;code&gt;08233f8&lt;/code&gt;, rewrites &lt;code&gt;.dockerignore&lt;/code&gt; to &lt;code&gt;packages/*&lt;/code&gt; followed by &lt;code&gt;!packages/trelix-mcp&lt;/code&gt;, and adds the missing &lt;code&gt;COPY packages/trelix-mcp/ packages/trelix-mcp/&lt;/code&gt; plus its install into the same &lt;code&gt;pip install&lt;/code&gt; invocation as core. Bundling it costs roughly 84MB on the slim tag by the release's own measurement — trelix-mcp's dependencies are &lt;code&gt;mcp&lt;/code&gt;, &lt;code&gt;fastmcp&lt;/code&gt;, and &lt;code&gt;trelix&lt;/code&gt; itself, nothing that pulls in torch — but that 84MB figure is reported prose in the changelog, not a number any script in this repo computes or asserts. The mechanism that makes it plausible (no ML dependency in trelix-mcp's own &lt;code&gt;pyproject.toml&lt;/code&gt;) is verifiable; the exact delta isn't.&lt;/p&gt;

&lt;p&gt;The second 3.2.2 bug lived one layer up, in the console script itself. Before the fix, &lt;code&gt;trelix-mcp&lt;/code&gt;'s &lt;code&gt;main()&lt;/code&gt; never inspected &lt;code&gt;sys.argv&lt;/code&gt; at all — not incorrectly, not partially, not at all. There was no &lt;code&gt;import argparse&lt;/code&gt;, no reference to &lt;code&gt;sys.argv&lt;/code&gt;, nothing. Running &lt;code&gt;trelix-mcp --help&lt;/code&gt; from a real shell didn't print usage; it silently launched the actual MCP stdio server and sat waiting for JSON-RPC input on stdin, the exact opposite of every CLI convention a &lt;code&gt;--help&lt;/code&gt; flag implies. The fix is small enough to quote in full:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Entry point for the trelix-mcp server (stdio transport).

    Parses argv only for --help/--version/unknown-flag rejection — the normal path (no
    args, launched by an MCP client&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s server config) falls straight through to running
    the server, unchanged from before this parser existed.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;parser&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;argparse&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;ArgumentParser&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;prog&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;trelix-mcp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;MCP server for trelix — semantic code search over stdio.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;parser&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_argument&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--version&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;version&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;version&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;trelix-mcp &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;__version__&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;parser&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse_args&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both 3.2.2 bugs share the same shape: something was correct in principle and broken in the specific packaging or invocation path a real user takes. Neither one could have failed a unit test, because no unit test in the suite ever installed a wheel, pulled an image, or ran a console script as a subprocess against &lt;code&gt;sys.argv&lt;/code&gt;. They were found by a manual production-verification pass — installing and running the actual shipped PyPI packages and Docker images rather than the source tree — and that pass had, by this point, been run by hand twice: once for 3.2.1, once for 3.2.2. Running it a third time by hand was the thing the rest of this arc exists to stop doing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The wildcard leak and the response that mattered more than the fix
&lt;/h2&gt;

&lt;p&gt;v3.2.3's actual bug is narrow and specific: &lt;code&gt;Database.bm25_search&lt;/code&gt;'s &lt;code&gt;path_filter&lt;/code&gt; branch, plus three &lt;code&gt;path_filter&lt;/code&gt;-scoped queries in &lt;code&gt;src/trelix/retrieval/grep_search.py&lt;/code&gt;, built a SQL &lt;code&gt;LIKE&lt;/code&gt; pattern directly from a caller-supplied path without escaping it. &lt;code&gt;LIKE&lt;/code&gt; treats &lt;code&gt;%&lt;/code&gt; and &lt;code&gt;_&lt;/code&gt; as wildcards inside the pattern value itself, which has nothing to do with SQL injection — every one of these queries was already parameterized with &lt;code&gt;?&lt;/code&gt; — and everything to do with &lt;code&gt;LIKE&lt;/code&gt;'s own semantics. A &lt;code&gt;path_filter&lt;/code&gt; of &lt;code&gt;src_auth&lt;/code&gt; matches not just the directory &lt;code&gt;src_auth&lt;/code&gt; but also, say, &lt;code&gt;srcXauth&lt;/code&gt;, because the underscore in the filter is read as "any single character," not a literal underscore. Given how common underscores are in real directory and file names (&lt;code&gt;test_utils.py&lt;/code&gt;, &lt;code&gt;my_function&lt;/code&gt;), this wasn't an edge case; it was a routine collision waiting for the right two sibling paths to exist in the same repository. The fix lives in &lt;code&gt;escape_like_pattern()&lt;/code&gt; at &lt;code&gt;src/trelix/store/db.py:291&lt;/code&gt;, whose docstring is worth quoting because it states the distinction precisely:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;escape_like_pattern&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Escape `%`, `_`, and `&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="s"&gt;` in `value` so it is safe as a SQL LIKE pattern segment.

    LIKE treats `%` and `_` as wildcards in the PATTERN VALUE itself — this has nothing
    to do with SQL injection (the callers here already use `?` parameterization) and
    everything to do with LIKE&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s own semantics.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy7l12z6xnkz6uojfz6e6.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy7l12z6xnkz6uojfz6e6.gif" alt=" " width="245" height="176"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The sibling method &lt;code&gt;get_index_metadata_with_prefix&lt;/code&gt; had already been doing this correctly for a while, pairing &lt;code&gt;escape_like_pattern()&lt;/code&gt; with an explicit &lt;code&gt;ESCAPE '\\'&lt;/code&gt; clause — so the fix wasn't a novel idiom, it was applying an already-proven pattern to the four call sites that had been missed. It's also not a total fix: &lt;code&gt;grep_search.py&lt;/code&gt;'s own docstring flags a second, separate, still-open instance of the same defect class on the symbol-name prefix match in &lt;code&gt;_name_search&lt;/code&gt;, explicitly out of scope for this release. I'd rather state that than let the fix read as more complete than it is.&lt;/p&gt;

&lt;p&gt;What makes v3.2.3 matter for this article isn't the bug — it's what came after it. The team's own reasoning, written into the new test suite's module docstring, is that no existing test, including Click's &lt;code&gt;CliRunner&lt;/code&gt;-style in-process mocking and the MCP SDK's own official in-memory &lt;code&gt;Client&lt;/code&gt;, is structurally capable of seeing a defect at the real process boundary — the exact class that let trelix-mcp's &lt;code&gt;--help&lt;/code&gt; bug ship in 3.2.1. In-memory transports never cross an actual OS process boundary; they can't see a console script that isn't installed, or a subprocess that hangs instead of exiting. So &lt;code&gt;tests/e2e/test_mcp_stdio_e2e.py&lt;/code&gt; spawns trelix-mcp as a genuine child process, located via &lt;code&gt;shutil.which("trelix-mcp")&lt;/code&gt; — the test fails loudly if the binary isn't on &lt;code&gt;PATH&lt;/code&gt;, not just importable from &lt;code&gt;src/&lt;/code&gt; — and talks to it over real stdin/stdout JSON-RPC via the MCP SDK's &lt;code&gt;stdio_client&lt;/code&gt;. A companion test, &lt;code&gt;test_pypi_dist_install_e2e.py&lt;/code&gt;, builds real wheels for all four published packages, installs each into a fresh venv, and asserts the installed &lt;code&gt;__version__&lt;/code&gt; matches the working tree — never editable, never a live PyPI pull, by explicit design.&lt;/p&gt;

&lt;p&gt;That test suite only matters if something forces it to run before a release goes out. In &lt;code&gt;.github/workflows/release.yml&lt;/code&gt;, the &lt;code&gt;publish&lt;/code&gt; job's dependencies now include a new &lt;code&gt;smoke-test-built-artifacts&lt;/code&gt; job that downloads the just-built wheel artifacts, installs them with &lt;code&gt;pip install&lt;/code&gt; — never editable — and runs the e2e suite against that install before &lt;code&gt;publish&lt;/code&gt; is permitted to execute. It is a real edge in the workflow's dependency graph, not a comment or a convention. Two smaller CI extensions round it out: &lt;code&gt;ci.yml&lt;/code&gt;'s Docker job now runs &lt;code&gt;docker run --rm --entrypoint trelix-mcp trelix:ci-test --version&lt;/code&gt; against every built image, and &lt;code&gt;helm-lint.yml&lt;/code&gt; asserts the rendered Helm chart's image tag matches &lt;code&gt;Chart.yaml&lt;/code&gt;'s &lt;code&gt;appVersion&lt;/code&gt;. Both are the kind of cheap check that would have caught a 3.2.1-shaped regression on the next PR instead of the next production audit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Automating the verification instead of re-typing it
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwamhc2oun3mii2y70y13.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwamhc2oun3mii2y70y13.gif" alt=" " width="500" height="280"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;v3.2.4's capstone answers a question the last two releases raised without answering: once you've written the checks, who runs them, and when? Before this release, the answer was a person, watching the GitHub Actions tab for both the Release and Docker Publish workflows to finish, then running &lt;code&gt;scripts/verify_release.py&lt;/code&gt; by hand and reading its output. &lt;code&gt;.github/workflows/verify-release.yml&lt;/code&gt; replaces that with a workflow that listens for &lt;code&gt;workflow_run&lt;/code&gt; completion events from both &lt;code&gt;Release&lt;/code&gt; and &lt;code&gt;Docker Publish&lt;/code&gt;. Because that event fires twice — once per upstream workflow finishing — the job does a single, non-blocking check via the GitHub CLI (&lt;code&gt;gh run list --workflow "$wf" --branch "$tag"&lt;/code&gt;) to see whether the &lt;em&gt;other&lt;/em&gt; workflow has also gone green for the same tag; if not, it exits cleanly rather than polling, and whichever of the two triggering runs happens to fire second is the one that proceeds. Once both are confirmed green, it runs the same &lt;code&gt;scripts/verify_release.py&lt;/code&gt; script a human used to run by hand — PyPI installs into fresh venvs for all four packages, Docker images smoke-tested with the exact &lt;code&gt;--entrypoint trelix-mcp&lt;/code&gt; command that returned 127 on 3.2.1, a Helm chart checked out into an isolated worktree and rendered across all three backends, a GitHub Release binary actually executed, and a &lt;code&gt;pip-audit&lt;/code&gt; plus wheel-content secret scan — and posts its own pass/fail summary as a workflow run. &lt;code&gt;CONTRIBUTING.md&lt;/code&gt; still documents the manual command for ad-hoc re-verification, which matters, because automating the check doesn't mean giving up the ability to run it by hand when you're debugging the check itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The other audit: when a symbol's nickname collides with someone else's
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnf4ttud032r4y2yp4n2q.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnf4ttud032r4y2yp4n2q.gif" alt=" " width="512" height="288"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Running in parallel with the artifact-verification work, v3.2.4 also shipped a distinct methodology: an audit for silent symbol-collision, not artifact drift. The pattern, once named, shows up in four unrelated corners of the codebase: something that should carry a full, unique address instead answers to a bare nickname, and when two different things share that nickname, one of them silently disappears or gets misattributed.&lt;/p&gt;

&lt;p&gt;The clearest instance is in Java. Before the fix, every nested class got a bare &lt;code&gt;qualified_name&lt;/code&gt; — a class named &lt;code&gt;Config&lt;/code&gt; nested inside &lt;code&gt;ServerConfig&lt;/code&gt; and a different class also named &lt;code&gt;Config&lt;/code&gt; nested inside &lt;code&gt;ClientConfig&lt;/code&gt; both indexed under the identical name &lt;code&gt;Config&lt;/code&gt;, with &lt;code&gt;parent_id&lt;/code&gt; set to &lt;code&gt;None&lt;/code&gt; regardless of which outer class actually enclosed them. Re-indexing one silently left the other's stale row permanently in the database. The fix threads an &lt;code&gt;outer_qualified_name&lt;/code&gt; recursively through the class walk in &lt;code&gt;java.py&lt;/code&gt;, so a nested class now gets &lt;code&gt;f"{outer_qualified_name}.{name}"&lt;/code&gt; and a real &lt;code&gt;parent_id&lt;/code&gt;. The Java record-component bug is a different kind of mistake in the same file: the extractor checked for tree-sitter node types &lt;code&gt;record_parameters&lt;/code&gt; (the container) and &lt;code&gt;record_component&lt;/code&gt; (the item) that the installed grammar simply doesn't emit — it emits &lt;code&gt;formal_parameters&lt;/code&gt; and &lt;code&gt;formal_parameter&lt;/code&gt; — so every record in the corpus indexed zero fields, silently, because the code was written against a node vocabulary that never existed in the grammar it ran against.&lt;/p&gt;

&lt;p&gt;Rust has the identical collision shape one file over. A function defined inside &lt;code&gt;mod inner { }&lt;/code&gt; previously got the same bare &lt;code&gt;qualified_name&lt;/code&gt; as any other function of the same name elsewhere in the file — the fix in &lt;code&gt;rust.py&lt;/code&gt; threads a &lt;code&gt;module_path&lt;/code&gt; through the module walk and rewrites the qualified name to &lt;code&gt;inner::nested&lt;/code&gt; for anything defined inside it. A separate Rust fix corrected the brace-import flattener, which checked for node types &lt;code&gt;use_tree_list&lt;/code&gt; and &lt;code&gt;use_tree&lt;/code&gt; that don't exist in the installed grammar either — the real node types are &lt;code&gt;scoped_use_list&lt;/code&gt; and &lt;code&gt;use_list&lt;/code&gt; — so the dominant Rust import form, &lt;code&gt;use foo::{Alpha, Beta}&lt;/code&gt;, fell through to a generic fallback that stored the literal text &lt;code&gt;"{Alpha, Beta}"&lt;/code&gt; as a single bogus import name, making every brace-imported symbol invisible to import-graph queries.&lt;/p&gt;

&lt;p&gt;The same audit surfaced the collision shape outside the parsers entirely. &lt;code&gt;AgentLoop._do_get_symbol&lt;/code&gt;, when asked for an exact qualified name that didn't resolve, used to fall back to an arbitrary bare-name match — silently handing the agent a different symbol's body than the one it asked for, with no error signal. The fix requires exactly one exact match or reports "not found." Federation's &lt;code&gt;make_scip_symbol_id()&lt;/code&gt; used to hash a cross-repo symbol identity from just &lt;code&gt;(package, version, qualified_name)&lt;/code&gt;, omitting &lt;code&gt;file_path&lt;/code&gt; — meaning two different files in the same package defining &lt;code&gt;def main()&lt;/code&gt; hashed to the identical 16-character id, and &lt;code&gt;INSERT OR IGNORE&lt;/code&gt; against that id as a primary key silently dropped the second file's row rather than erroring. It's the most explicit statement of the whole pattern in the release: an identity hash that omitted one required scoping field, and a database that dropped the collision without complaint.&lt;/p&gt;

&lt;p&gt;Not every bug the same audit pass caught fits that collision shape, and I'd rather say so than force the frame. The chunker's &lt;code&gt;token_count&lt;/code&gt; bug was a stale cached value — computed against the pre-truncation text instead of the actual, truncated chunk that got stored. &lt;code&gt;bm25&lt;/code&gt;'s stop-word handling had a related but distinct problem: a query made entirely of stop words was supposed to hit the empty-query fallback and return the FTS5 sentinel path, but the stop-word filter ran after that check instead of before it, so a stop-words-only query was treated as a normal, non-empty query instead. The walker's fix normalizes &lt;code&gt;rel_path&lt;/code&gt; to NFC via &lt;code&gt;unicodedata.normalize&lt;/code&gt;, because a single real file could take two different Unicode-normalization forms across walks (NFD from an HFS+ filesystem versus NFC elsewhere), fracturing one file's identity into what change detection read as a phantom delete-and-add. The graph community-detection fix removed a force-cast to &lt;code&gt;int&lt;/code&gt; on node ids that crashed with &lt;code&gt;ValueError&lt;/code&gt; on any ticket-linked repo, because ticket node ids are strings by design. And the CLI's &lt;code&gt;update_index&lt;/code&gt; command now actually checks &lt;code&gt;result["status"]&lt;/code&gt; and exits 1 on failure, instead of printing the JSON error payload and returning exit 0 regardless. Different defect shapes, same audit pass, same release — worth naming honestly as a second act rather than folding into a pattern it doesn't share.&lt;/p&gt;

&lt;h2&gt;
  
  
  A short coda: the binary that gave advice it couldn't act on
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvzm1h901zxuz1ww110qo.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvzm1h901zxuz1ww110qo.gif" alt=" " width="480" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;v3.2.5 is one fix, and it's a fitting close because it was caught by the same discipline the verify-release capstone exists to encode: someone read the error message the real, shipped v3.2.4 binary actually prints, in production, rather than the source that built it. The standalone GitHub Release binary's local embedder — used when &lt;code&gt;sentence-transformers&lt;/code&gt; isn't available — told users to run &lt;code&gt;pip install 'trelix[local]'&lt;/code&gt;. That advice is correct for the pip-installed package. It is actively useless for the frozen PyInstaller binary, which is built with &lt;code&gt;sentence-transformers&lt;/code&gt;, &lt;code&gt;torch&lt;/code&gt;, and the rest of the ML stack deliberately excluded to keep it small, and which never consults the host's Python or pip environment at all — installing anything on the host has zero effect on a binary that doesn't read it. The fix, in &lt;code&gt;src/trelix/embedder/base.py&lt;/code&gt;, checks &lt;code&gt;getattr(sys, "frozen", False)&lt;/code&gt; — the standard attribute PyInstaller sets — and gives the frozen binary a message pointing at an API-backed provider or the Python package instead, leaving the pip-installed package's message byte-for-byte unchanged. It's covered by a test that asserts the frozen message never contains the string "pip install." One line of misdirection, found the same way the Docker gap and the console-script gap were found ten weeks and four releases earlier: by running the actual thing that ships, not the code that built it.&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>python</category>
      <category>devops</category>
      <category>testing</category>
    </item>
    <item>
      <title>trelix v3.1.2 to v3.2.1: The Tests That Passed Without Testing Anything</title>
      <dc:creator>SAI RAM</dc:creator>
      <pubDate>Wed, 26 Aug 2026 13:53:41 +0000</pubDate>
      <link>https://dev.to/sai_ram_0000/trelix-v312-to-v321-the-tests-that-passed-without-testing-anything-1gp4</link>
      <guid>https://dev.to/sai_ram_0000/trelix-v312-to-v321-the-tests-that-passed-without-testing-anything-1gp4</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg5107vncg56p649hyld.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg5107vncg56p649hyld.gif" alt="A developer watching tests pass while unaware of the underlying bug" width="243" height="243"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In FlagEmbedding 1.4.0, &lt;code&gt;from FlagEmbedding import FlagModel&lt;/code&gt; is an alias for &lt;code&gt;from .base import BaseEmbedder as FlagModel&lt;/code&gt; — the encoder-only base class, whose default &lt;code&gt;pooling_method&lt;/code&gt; is &lt;code&gt;"cls"&lt;/code&gt; and whose &lt;code&gt;pooling()&lt;/code&gt; for that method is one line: &lt;code&gt;return last_hidden_state[:, 0]&lt;/code&gt;. &lt;code&gt;BAAI/bge-code-v1&lt;/code&gt; is not an encoder. It is a causal Qwen2 decoder, published with &lt;code&gt;1_Pooling/config.json&lt;/code&gt; declaring &lt;code&gt;pooling_mode_lasttoken: true&lt;/code&gt; and &lt;code&gt;pooling_mode_cls_token: false&lt;/code&gt;. Position 0 of a causal decoder cannot attend forward. So the CLS vector &lt;code&gt;bge-code&lt;/code&gt; was computing depended on exactly one token and nothing else that followed it.&lt;/p&gt;

&lt;p&gt;I measured what that means with a randomly initialized &lt;code&gt;Qwen2Model&lt;/code&gt; pooled by FlagEmbedding's real &lt;code&gt;pooling()&lt;/code&gt;: two sequences identical at token 0 and different in every token after it came back with cosine similarity 1.0 and max absolute difference 0.0 — bitwise identical. The same hidden states through the real &lt;code&gt;last_token_pool&lt;/code&gt; gave cosine 0.10; through &lt;code&gt;mean&lt;/code&gt;, 0.65. Across 200 resamples of every token after position 0, position 0's output never moved once. And because &lt;code&gt;encode_queries&lt;/code&gt; prefixes every query with the same instruction, token 0 is that shared prefix for every query trelix ever sends — meaning every query embedding &lt;code&gt;bge-code&lt;/code&gt; produced was identical to every other one, full stop. Five distinct code chunks sharing a first token collapsed into a single vector, against five distinct vectors from &lt;code&gt;last_token_pool&lt;/code&gt; on the same input.&lt;/p&gt;

&lt;p&gt;This shipped in 28 tagged releases — every tag from v2.0.0 through v3.1.6 contains the commit that introduced the embedder — and stayed green through every one of them because &lt;code&gt;tests/unit/test_embedder_bge.py&lt;/code&gt; replaces &lt;code&gt;FlagModel&lt;/code&gt; with a &lt;code&gt;MagicMock&lt;/code&gt;. A &lt;code&gt;MagicMock&lt;/code&gt; answers to any attribute you ask of it. It never once called the real &lt;code&gt;pooling()&lt;/code&gt;, so a defect that destroys every query embedding a provider produces never had a chance to fail a test.&lt;/p&gt;

&lt;p&gt;v3.1.6 fixed a real, adjacent bug: &lt;code&gt;BGECodeEmbedder.dimension&lt;/code&gt; called a method FlagEmbedding has never had, raising &lt;code&gt;AttributeError&lt;/code&gt; before an index could even start, and the dimension it fell back to (768) was wrong anyway — the model emits 1536. That fix was correct and it shipped correctly. But removing the raise made a path reachable that had never been reachable before: the pooling defect. The v3.1.6 changelog entry read as though &lt;code&gt;bge-code&lt;/code&gt; now worked. It didn't, and one release later, v3.1.7 said so in public, in the same file: "Retracted: &lt;code&gt;bge-code&lt;/code&gt; does not 'work now'." I don't read that as an embarrassment. A team that ships a wrong claim and then corrects it in the next release, with the receipts, is doing exactly what a correctness process is for. The alternative — quietly letting the framing stand — is the actual failure mode.&lt;/p&gt;

&lt;h2&gt;
  
  
  Eight releases, one arc
&lt;/h2&gt;

&lt;p&gt;This is the story of eight tagged releases — v3.1.2, v3.1.3, v3.1.4, v3.1.5, v3.1.6, v3.1.7, v3.2.0, and v3.2.1 — spanning 190 commits and 324 changed files since v3.1.1, where the last article in this series left off. v3.2.1 is the current shipping release: tagged, dated 2026-08-26, live on PyPI, installable today with &lt;code&gt;pip install trelix==3.2.1&lt;/code&gt;. The unit suite sits at 4,376 tests collected on this exact checkout. trelix itself is about nine weeks old — the first commit landed 2026-06-25 — which matters for how to read what follows: this isn't a decade of debt being paid down, it's a project auditing itself hard and often, early, while the cost of finding a defect is still one release instead of ten.&lt;/p&gt;

&lt;p&gt;Unlike the v3.0.0 span this series already covered — audit trail, OIDC SSO, the VS Code extension, context compression, model-aware budgeting, extended thinking — this span has no comparable feature drop. It is, almost entirely, a correctness and audit arc. I'm not going to apologize for that framing or bury it under a feature list that doesn't fit the material. The interesting thing about these eight releases isn't what got built; it's what got caught, and specifically the &lt;em&gt;shape&lt;/em&gt; of what got caught, which repeats often enough across independent parts of the codebase that it deserves to be named as a pattern rather than four unrelated bug reports.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: a green suite that never touched the defect
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjjfuhedljqpqz29cw5k7.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjjfuhedljqpqz29cw5k7.gif" alt="A developer reacting to discovering a bug in tests" width="450" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here is the pattern, stated once so the four instances below don't have to keep re-deriving it: a test can pass without ever exercising the behavior it claims to cover. That's not the same claim as "the tests were bad" or "coverage was low." These tests ran. They asserted things. They turned green in CI, release after release. And in every one of the four cases below, the reason they turned green is diagnosable and specific — not vague test debt, but one of four concrete mechanisms.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;bge-code&lt;/code&gt;'s pooling defect survived because a &lt;code&gt;MagicMock&lt;/code&gt; answers any attribute asked of it, so no test ever called the real &lt;code&gt;pooling()&lt;/code&gt; method that carried the bug. The sparse-vector padding defect, below, survived because every existing fake for the tokenizer returned an all-ones attention mask by construction, which makes masked and unmasked aggregation produce identical numbers regardless of whether the masking code is even present. The &lt;code&gt;intent_hint&lt;/code&gt; dispatch bug survived because a unit test asserted the observed, buggy output as the &lt;em&gt;correct&lt;/em&gt; value — it had to be deleted, not fixed, once the real behavior was understood. And the federated &lt;code&gt;search-all&lt;/code&gt; deduplication bug survived because the existing test happened to use the one input distribution — globally unique row identifiers — that cannot trigger a collision, out of all the distributions that could.&lt;/p&gt;

&lt;p&gt;Four independent parts of the codebase — an embedder, a different embedder's aggregation math, a retrieval planner, and a federation layer — produced the same failure mode by four different routes. That repetition is why v3.2.1 exists, and why I'm leading with it rather than a version-by-version changelog walk.&lt;/p&gt;

&lt;h2&gt;
  
  
  The sparse leg: clean queries scored against contaminated documents
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmtnt0wfyup1ojkjvig5n.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmtnt0wfyup1ojkjvig5n.gif" alt="A developer looking shocked or surprised while reviewing test results" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;SparseEmbedder.embed&lt;/code&gt; tokenizes a batch with &lt;code&gt;padding=True&lt;/code&gt;, which pads every sequence in that batch out to the length of its longest member, and then aggregates with &lt;code&gt;torch.log(1 + torch.relu(logits)).max(dim=1))&lt;/code&gt; — a max taken over every sequence position, padding included. A masked-language-model head predicts a real, nonzero logit distribution at pad positions too; it doesn't know they're padding. Nothing in the aggregation multiplied by the attention mask to zero those positions back out. &lt;code&gt;grep -c attention_mask src/trelix/embedder/sparse.py&lt;/code&gt; returned 0. A tree-wide &lt;code&gt;grep -rln attention_mask src/trelix/&lt;/code&gt; matched no file in the codebase at all.&lt;/p&gt;

&lt;p&gt;I measured it against the real &lt;code&gt;naver/splade-v3-distilbert&lt;/code&gt; at &lt;code&gt;top_k=128&lt;/code&gt;. A 28-token chunk embedded alone, compared against the identical chunk embedded in a batch alongside a 185-token chunk, gained 28 phantom terms and lost 28 real ones at the cutoff — 22% of the stored vector — with 0.478 max weight drift on the terms that survived both runs. The phantom terms are ordinary English words the chunk never contained: &lt;code&gt;where&lt;/code&gt;, &lt;code&gt;store&lt;/code&gt;, &lt;code&gt;numbers&lt;/code&gt;, &lt;code&gt;text&lt;/code&gt;, &lt;code&gt;gene&lt;/code&gt;, &lt;code&gt;sequence&lt;/code&gt;, &lt;code&gt;phrase&lt;/code&gt;, &lt;code&gt;messages&lt;/code&gt;. The control that pins the cause is the one that matters most here: the same chunk batched with an &lt;em&gt;equal-length&lt;/em&gt; chunk, which needs no padding at all, came back bit-identical to the chunk embedded alone. Batching was never the problem. Padding was.&lt;/p&gt;

&lt;p&gt;Two things made this worth a minor release rather than a footnote. &lt;code&gt;sparse_embeddings&lt;/code&gt; is a persisted table in &lt;code&gt;store/db.py&lt;/code&gt;, so the corruption wasn't transient — reindexing the same repository in a different file order, or with a different &lt;code&gt;TRELIX_SPARSE_BATCH_SIZE&lt;/code&gt;, rewrote every row with a different, equally wrong answer, and no index was reproducible against itself. And &lt;code&gt;embed_query&lt;/code&gt; routes through &lt;code&gt;embed([text])&lt;/code&gt; — a batch of exactly one, which needs no padding — so every query trelix ever issued against the sparse leg was clean, scored against documents that weren't. No amount of query-side tuning could have surfaced or corrected that asymmetry, because the query side was never where the bug lived.&lt;/p&gt;

&lt;p&gt;The existing &lt;code&gt;test_sparse_embedder.py&lt;/code&gt; had ten tests covering this embedder, and every one of its fakes returned &lt;code&gt;attention_mask=torch.ones(...)&lt;/code&gt;. An all-ones mask makes the masked and the unmasked aggregation mathematically identical, so all ten tests passed whether or not the masking code existed at all — which it didn't. The replacement, &lt;code&gt;test_sparse_padding_contamination.py&lt;/code&gt;, is nine tests, every one mutation-verified: deliberately dropping the mask multiplication from the aggregation fails five of them while leaving the equal-length control passing, which is exactly the signature that localizes the cause to padding rather than to batching in general.&lt;/p&gt;

&lt;h2&gt;
  
  
  intent_hint: a test that asserted the bug as the spec
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fglcn59qj6xgkgubqzlt0.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fglcn59qj6xgkgubqzlt0.gif" alt="A developer looking stressed or overwhelmed while reviewing code" width="550" height="550"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;intent_hint&lt;/code&gt; is an optional parameter on both the REST &lt;code&gt;/search&lt;/code&gt; endpoint and the MCP &lt;code&gt;search_code&lt;/code&gt; tool, meant to let a caller skip trelix's own LLM-based intent classification and specify a retrieval strategy directly. The hint builder had a routing bug: it stamped the "direct answer" tier onto all eight recognized intent values while separately setting the correct leg strategy for each — but the executor checks the tier first, so the strategy it computed was never read. Every &lt;code&gt;intent_hint&lt;/code&gt; value, regardless of which of the eight intents you actually asked for, took the same direct-answer shortcut.&lt;/p&gt;

&lt;p&gt;I measured this against trelix's own codebase: all eight intent values returned byte-identical output — 40 README sections, zero code files, and no overlap at all with the actual correct result set for any of them. &lt;code&gt;intent_hint&lt;/code&gt; had been silently disabling retrieval since v2.10.0, over both protocols, for anyone who used it.&lt;/p&gt;

&lt;p&gt;The reason it took this long to catch is the plainest of the four: a unit test existed that asserted the broken output was the expected one. That's not a gap in coverage — it's a test that actively encodes the bug as the specification. It had to be deleted and replaced with assertions on the actual retrieval outcome, not patched.&lt;/p&gt;

&lt;h2&gt;
  
  
  search-all: a test that used the one distribution that can't fail
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F28oaa2kj3oi17r2bqsd9.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F28oaa2kj3oi17r2bqsd9.gif" alt="A developer looking frustrated or deep in thought while debugging" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Federated &lt;code&gt;search-all&lt;/code&gt; fans a query out across every registered repository and merges the results. The merge deduplicated on each result's per-database autoincrement row id — and every repository's ids start at 1. Two repositories' top hits collide on id 1 unavoidably; first-seen wins the merge, and "first" was whichever repository's thread finished first in the pool, which is nondeterministic. The consequence: an entire repository's results could vanish from a federated search, silently, while the response reported &lt;code&gt;repos_skipped: 0&lt;/code&gt; — nothing told you a repository had been dropped, because nothing had been skipped; its results had simply lost a dedup race.&lt;/p&gt;

&lt;p&gt;The existing test for this path passed because it constructed its fixture repositories with globally distinct identifiers — the one input distribution, out of every distribution the real world produces, under which two repositories' row ids cannot collide. The fix keys deduplication on a globally unique identity instead of the raw row id, and the new test deliberately reuses row ids across repositories to force the collision the old test structurally avoided.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do you know your check actually checks: the audit trail and the self-audit
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmioiv0jcrxle10b363v3.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmioiv0jcrxle10b363v3.gif" alt="A developer looking deep in thought while working on code" width="600" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two pieces of this arc aren't part of the four-instance spine above, but they belong in the same conversation because they ask the identical question from a different angle: not "does this test exercise the code," but "does this &lt;em&gt;check&lt;/em&gt; actually check the thing it claims to."&lt;/p&gt;

&lt;p&gt;v3.1.4 found that &lt;code&gt;trelix audit&lt;/code&gt;'s &lt;code&gt;list&lt;/code&gt;, &lt;code&gt;verify&lt;/code&gt;, and &lt;code&gt;export&lt;/code&gt; commands — the read path over the hash-chained audit trail introduced in v3.0.0 — used the same store constructor as the writer, which runs &lt;code&gt;CREATE TABLE IF NOT EXISTS&lt;/code&gt; on open. Point any of the three read commands at a file that wasn't an audit log, and they added the audit schema to it on the spot, then reported "Audit chain intact" at exit 0, because the chain they had just silently created was empty, and an empty chain is, trivially, consistent. One measured case took an 8 KB, one-table SQLite file to 32 KB and five tables just by being read. A CI integrity gate pointed at the wrong path would pass green while mutating the exact file it was supposed to verify. All three commands now open the file read-only (&lt;code&gt;file:&amp;lt;path&amp;gt;?mode=ro&lt;/code&gt;), run no DDL, and exit 2 unless the audit schema is already present. It is the same question as the mutation-testing thesis, asked of an integrity check instead of a unit test: a verification step that can succeed against input it was never supposed to accept isn't verifying anything.&lt;/p&gt;

&lt;p&gt;v3.1.2 is the plainest statement of the whole arc's thesis, predating the four-instance spine above by weeks. It's a self-audit: the team indexed trelix's own repository with trelix and checked whether every feature the configuration turned on was actually doing anything. The finding, in the release's own words, was "features that were on and doing nothing." File summaries were enabled and the index had zero of them. PageRank boosting had never once fired. &lt;code&gt;taint_flows&lt;/code&gt; was empty on a repository semgrep does find real flows in. The query planner had classified all 219 recorded queries as the same one of its eight intents. None of these failures were visible from outside, because each failure path either logged at DEBUG while the CLI runs at WARNING, or was swallowed by a bare &lt;code&gt;except&lt;/code&gt;, or reported a number nobody had reason to doubt. That release's own retrospective on its test suite reads almost like an early draft of this article's thesis: the taint parser was verified against a fixture invented to match its own misreading; every planner test supplied no credentials, so none could notice credentials being silently dropped; the eval metric tests used unique IDs only, so none could notice a repeated ID scoring twice.&lt;/p&gt;

&lt;p&gt;Briefly, because it doesn't compete for space with any of this: v3.1.3 also shipped real security fixes — a REST API containment check that validated a caller-supplied path against a root the same caller supplied, and a stored XSS in generated graph HTML from an unescaped symbol name or filename. Both are fixed and both matter, but they're a different kind of bug from everything above, and this article already has its through-line without them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: mutation testing that measures whether a test can fail
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa3i3qv1l70n45c1omev6.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa3i3qv1l70n45c1omev6.gif" alt="A developer looking thoughtfully at a computer screen" width="382" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;v3.2.1 is a test-infrastructure release. It changes no stored vectors, no retrieval behavior, and no CLI surface — its one production fix is that &lt;code&gt;retry.py&lt;/code&gt;'s status-code extractor imported every supported LLM provider SDK (&lt;code&gt;anthropic&lt;/code&gt;, &lt;code&gt;openai&lt;/code&gt;, &lt;code&gt;google.genai&lt;/code&gt;, &lt;code&gt;boto3&lt;/code&gt;, &lt;code&gt;azure&lt;/code&gt;) on every retry decision regardless of which backend actually raised, pulling &lt;code&gt;torch&lt;/code&gt; transitively into processes that use no LLM feature at all. Everything else in the release is instrumentation, and it exists because coverage percentage — the metric the previous eight releases had been implicitly relying on — cannot answer the question every instance above turned on: can this test actually fail?&lt;/p&gt;

&lt;p&gt;The core of it is a new scoped mutation-testing driver, &lt;code&gt;scripts/mutation.py&lt;/code&gt;, wrapping &lt;code&gt;mutmut&lt;/code&gt;. It runs against a throwaway git worktree rather than the live tree — unscoped mutant generation writes roughly 143 MB of Python that has no business near a commit — and it ratchets a per-module survivor &lt;em&gt;count&lt;/em&gt; in &lt;code&gt;scripts/mutation_baseline.json&lt;/code&gt;, deliberately never a survivor &lt;em&gt;ratio&lt;/em&gt;: a repo-wide ratio moves whenever an unrelated segfault appears or disappears in a part of the codebase mutation testing can't safely reach, and parts of trelix — anywhere a real torch model gets constructed — are exactly that. First real measurements landed for the parser's per-language extractors (split from one coarse scope key into 23 granular ones), &lt;code&gt;store.db&lt;/code&gt;, &lt;code&gt;store.vector&lt;/code&gt;, &lt;code&gt;indexing.chunker&lt;/code&gt;, &lt;code&gt;compression.extractive&lt;/code&gt;, and &lt;code&gt;graph&lt;/code&gt;. Closing the gaps mutation testing surfaced meant real new coverage: walker extension-map handling, &lt;code&gt;ContextualChunker&lt;/code&gt; boundary conditions, graph community-detection survivor patterns, and operator-environment-variable leak/scrub coverage.&lt;/p&gt;

&lt;p&gt;Branch coverage is now on (&lt;code&gt;--cov-branch&lt;/code&gt;), with per-package floors enforced separately from the unit run itself — and here the changelog for this release actually gets its own mechanism wrong, so it's worth stating plainly: the floors are not in a file called &lt;code&gt;tests/coverage-floors.json&lt;/code&gt;, which doesn't exist anywhere in this repository. They're a &lt;code&gt;FLOORS&lt;/code&gt; dict hardcoded directly inside &lt;code&gt;scripts/check_coverage_floors.py&lt;/code&gt;, checked in CI against a coverage report the unit job writes to &lt;code&gt;coverage-unit.json&lt;/code&gt; (&lt;code&gt;pytest tests/unit/ --cov-report=json:coverage-unit.json&lt;/code&gt;, followed by &lt;code&gt;python scripts/check_coverage_floors.py coverage-unit.json&lt;/code&gt;). The script's own docstring explains why it's a script and not a test: pytest-cov only writes its report at session finish, after every test has already resolved, so a test can't read its own run's coverage — and a check that can only skip when the report file is missing is a green identical to a check that never ran, which is precisely the defect class this whole release exists to close.&lt;/p&gt;

&lt;p&gt;The rest of the hardening follows the same logic. The suite is now hermetic: outbound sockets are banned via &lt;code&gt;pytest-socket&lt;/code&gt;, every job is timeout-bounded, and the live-LLM integration tests in &lt;code&gt;tests/integration/test_llm_e2e.py&lt;/code&gt; now require an explicit &lt;code&gt;TRELIX_LIVE_LLM_TESTS=1&lt;/code&gt; instead of running whenever a &lt;code&gt;.env&lt;/code&gt; file happened to be present — cutting roughly half the suite's wall clock along with the credential exposure that came bundled with it. Sixteen known Java and Rust extractor defects are pinned as &lt;code&gt;xfail(strict=True)&lt;/code&gt; rather than left silently uncovered, so each one XPASSes loudly, failing the build, if it's ever fixed without an accompanying changelog entry. A new marker taxonomy means a typo in a &lt;code&gt;-m&lt;/code&gt; selector can no longer silently collect and pass the entire suite. None of this catches every possible defect, and it doesn't claim to; the mutation baseline is a ratchet, not a finish line. What it does is make "the test passed but never touched the bug" structurally harder to repeat than it was for 28 releases of &lt;code&gt;bge-code&lt;/code&gt;, one release of the sparse leg, and however many releases &lt;code&gt;intent_hint&lt;/code&gt; and &lt;code&gt;search-all&lt;/code&gt; shipped broken before anyone measured directly instead of trusting the green check.&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>python</category>
      <category>testing</category>
      <category>aiengineering</category>
    </item>
    <item>
      <title>trelix v2.11.0 to v3.1.1: Six Feature Areas, Every One of Them Off By Default</title>
      <dc:creator>SAI RAM</dc:creator>
      <pubDate>Sat, 15 Aug 2026 15:33:10 +0000</pubDate>
      <link>https://dev.to/sai_ram_0000/trelix-v2110-to-v311-six-feature-areas-every-one-of-them-off-by-default-1go1</link>
      <guid>https://dev.to/sai_ram_0000/trelix-v2110-to-v311-six-feature-areas-every-one-of-them-off-by-default-1go1</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxkory274zmo39nf4zmww.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxkory274zmo39nf4zmww.gif" alt="A descriptive summary of what happens in the GIF" width="480" height="438"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Seed three events into an audit database, then reach past the application and change one row by hand:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;sqlite3 audit.db &lt;span class="s2"&gt;"UPDATE audit_log SET principal='attacker' WHERE id=2"&lt;/span&gt;
&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;trelix audit verify &lt;span class="nt"&gt;--db&lt;/span&gt; audit.db
&lt;span class="go"&gt;Audit chain TAMPERED — first divergent entry id: 2
&lt;/span&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="nv"&gt;$?&lt;/span&gt;
&lt;span class="go"&gt;1
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Delete the newest row instead and it still catches it, naming id 3, even though the surviving rows form a perfectly valid chain. Point it at something SQLite cannot open and it exits 2 rather than 0, because "I could not check" and "I checked and it is clean" must never collapse into the same green build.&lt;/p&gt;

&lt;p&gt;None of that existed six releases ago. &lt;code&gt;trelix audit verify&lt;/code&gt; is one command out of six feature areas that landed in trelix v3.0.0, and it is the one that most changes what the project is for.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the major bump actually is
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffphor9rlc57d1e35iy0x.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffphor9rlc57d1e35iy0x.gif" alt="A descriptive summary of what happens in this second GIF" width="450" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The span from v2.11.0 to v3.1.1 is six releases — v2.11.1, v2.12.0, v3.0.0, v3.0.1, v3.1.0 and v3.1.1, the last of them dated 2026-08-15 — 68 commits, 137 files changed, +19,829/-1,211 lines. v2.11.0 closed out the Jira and Linear connector work, which has its own story. Everything after it is a different kind of release.&lt;/p&gt;

&lt;p&gt;v3.0.0 carries six new feature areas: Anthropic extended thinking, a model-aware context budget, a VS Code extension that acts instead of merely displaying, a hash-chained append-only audit trail, OIDC SSO, and query-conditioned context compression. Alongside them, an opt-in FTS5 declaration boost for keyword ranking.&lt;/p&gt;

&lt;p&gt;It is a major bump because of scope, not breakage. Every one of those six is additive and off by default: &lt;code&gt;TRELIX_AUDIT_ENABLED=false&lt;/code&gt;, &lt;code&gt;TRELIX_OIDC_ENABLED=false&lt;/code&gt;, &lt;code&gt;TRELIX_LLM_THINKING_ENABLED=false&lt;/code&gt;, &lt;code&gt;TRELIX_RETRIEVAL_COMPRESSION=false&lt;/code&gt;, &lt;code&gt;declaration_boost_enabled&lt;/code&gt; False, and &lt;code&gt;context_token_budget&lt;/code&gt; still the exact &lt;code&gt;12_000&lt;/code&gt; integer it was in v2.12.0. A default v3.0.0 install assembles context byte-identically to a default v2.12.0 install, and there is a test that proves it rather than a release note that asserts it.&lt;/p&gt;

&lt;h2&gt;
  
  
  An audit trail you can hand to somebody else
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmc419u41qaqrcxye3ijm.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmc419u41qaqrcxye3ijm.gif" alt="A descriptive summary of what happens in this third GIF" width="400" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The trail is a separate SQLite file. &lt;code&gt;AuditConfig.resolved_db_path&lt;/code&gt; defaults to &lt;code&gt;&amp;lt;cwd&amp;gt;/.trelix/audit.db&lt;/code&gt;, never the index database, for one blunt reason: the index is disposable, and an audit trail that disappears on reindex is not an audit trail.&lt;/p&gt;

&lt;p&gt;Each row's &lt;code&gt;entry_hash&lt;/code&gt; is &lt;code&gt;sha256(prev_hash || canonical_json(content))&lt;/code&gt;, canonicalised in ten lines of &lt;code&gt;src/trelix/audit/store.py&lt;/code&gt; with &lt;code&gt;sort_keys=True&lt;/code&gt; and &lt;code&gt;separators=(",", ":")&lt;/code&gt;. Exactly eleven columns are hashed, in a fixed &lt;code&gt;_CONTENT_COLUMNS&lt;/code&gt; tuple, and the DB-assigned &lt;code&gt;id&lt;/code&gt; is deliberately not among them: a writer cannot hash a value it does not have until after the INSERT.&lt;/p&gt;

&lt;p&gt;That exclusion is why row ordering needs a second defence, and it explains the deleted-tail catch from the opening: a hash chain is structurally blind to truncation. An &lt;code&gt;audit_meta&lt;/code&gt; table carries a running &lt;code&gt;count&lt;/code&gt; and &lt;code&gt;head_hash&lt;/code&gt;, upserted inside the same &lt;code&gt;with self._conn:&lt;/code&gt; transaction as the insert, and &lt;code&gt;verify_chain&lt;/code&gt; checks both against the chain it just walked. &lt;code&gt;audit_log.id&lt;/code&gt; is &lt;code&gt;INTEGER PRIMARY KEY AUTOINCREMENT&lt;/code&gt; rather than a rowid alias, so a delete-then-refill cannot close the gap it made.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;export TRELIX_AUDIT_ENABLED=true&lt;/code&gt; turns the trail on and &lt;code&gt;trelix serve ./my-repo&lt;/code&gt; starts recording (&lt;code&gt;TRELIX_AUDIT_DB_PATH&lt;/code&gt; moves the file if you want it off &lt;code&gt;&amp;lt;cwd&amp;gt;/.trelix/audit.db&lt;/code&gt;), &lt;code&gt;trelix audit list -n 50&lt;/code&gt; reads it, and &lt;code&gt;trelix audit export --format ndjson&lt;/code&gt; writes one JSON object per line so a SIEM can be fed by Filebeat or Vector. All of it is stdlib &lt;code&gt;sqlite3&lt;/code&gt;, 54 tests cover the surface, and the store is fail-open by default (&lt;code&gt;AuditConfig.fail_closed: bool = False&lt;/code&gt;) so a full disk is not an outage.&lt;/p&gt;

&lt;p&gt;SSO is the identity half, and it is deliberately narrow: trelix is a resource server, not an OIDC client. No &lt;code&gt;/login&lt;/code&gt;, no &lt;code&gt;/callback&lt;/code&gt;, no code exchange, no refresh. It verifies tokens callers already hold. Three variables stand it up — &lt;code&gt;TRELIX_OIDC_ENABLED=true&lt;/code&gt;, &lt;code&gt;TRELIX_OIDC_ISSUER&lt;/code&gt; and &lt;code&gt;TRELIX_OIDC_AUDIENCE&lt;/code&gt; — on top of &lt;code&gt;pip install 'trelix[sso]'&lt;/code&gt;; &lt;code&gt;TRELIX_OIDC_ALGORITHMS&lt;/code&gt; defaults to &lt;code&gt;["RS256", "ES256"]&lt;/code&gt; and is asymmetric-only, enforced in &lt;code&gt;OidcVerifier.__init__&lt;/code&gt;, again in &lt;code&gt;authenticate&lt;/code&gt; against the unverified JWS header before any key resolution, and a third time in &lt;code&gt;jwt.decode&lt;/code&gt;. The test that matters skips trelix's own header gate entirely: it signs a token with HS256 using a real 2048-bit public key's PEM as the HMAC secret — the canonical algorithm-confusion forgery — and asserts &lt;code&gt;jwt.decode&lt;/code&gt; alone still refuses it. &lt;code&gt;Principal.principal_id&lt;/code&gt; is &lt;code&gt;f"{self.subject}@{self.issuer}"&lt;/code&gt;, never derived from email, because email-keyed identity is an account-takeover primitive, and the JWKS fetch is HTTPS-only, host-pinned and capped at 1 MiB — not hypothetical, since the pre-fix &lt;code&gt;response.read()&lt;/code&gt; let a hostile issuer drive RSS from 59 MB to about 662 MB in 0.2 seconds.&lt;/p&gt;

&lt;p&gt;The limits are in the docs rather than the marketing. The trail is tamper-evident, not tamper-proof: the chain and the anchor that checks it live in the same file, so anyone with write access can rewrite a row, recompute every subsequent hash and update the anchor in one transaction. There is sha256 and no key. Only the HTTP surface is audited, so an MCP-only deployment produces an empty trail and the agent loop's per-turn calls are invisible. &lt;code&gt;TRELIX_AUDIT_RETENTION_DAYS&lt;/code&gt; is declarative only — nothing prunes — and &lt;code&gt;/health&lt;/code&gt; is audited, so liveness probes will dominate the row count. This is authentication, not authorization — OIDC &lt;code&gt;groups&lt;/code&gt; are captured and stored but enforced nowhere — and there is no SAML, with no plan for it; put a broker in front.&lt;/p&gt;

&lt;h2&gt;
  
  
  Thinking, and a budget that knows which model you are talking to
&lt;/h2&gt;

&lt;p&gt;Extended thinking is a request parameter, not a model. &lt;code&gt;AnthropicBackend._thinking_kwargs()&lt;/code&gt; returns one dict, merged into both &lt;code&gt;messages.create&lt;/code&gt; and &lt;code&gt;messages.stream&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;thinking&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;enabled&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;budget_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;thinking_budget_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No model string in trelix changes. &lt;code&gt;TRELIX_LLM_THINKING_ENABLED&lt;/code&gt; defaults False and &lt;code&gt;TRELIX_LLM_THINKING_BUDGET_TOKENS&lt;/code&gt; to 4096.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy8w9um75t6ab5usynutf.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy8w9um75t6ab5usynutf.gif" alt="A descriptive summary of what happens in this fourth GIF" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Only one component opts in — &lt;code&gt;RetrievalSynthesizer&lt;/code&gt;, at two call sites. The index-time callers, &lt;code&gt;ContextualChunker&lt;/code&gt; and &lt;code&gt;FileSummarizer&lt;/code&gt;, do not pass the keyword at all; a single global flag would have turned an indexing run into roughly five to ten times the LLM cost to produce reasoning nobody reads.&lt;/p&gt;

&lt;p&gt;One consequence is worth stating plainly: enabling thinking forces &lt;code&gt;temperature&lt;/code&gt; to 1.0 in both &lt;code&gt;complete()&lt;/code&gt; and &lt;code&gt;stream()&lt;/code&gt;, overriding whatever the caller passed — and the synthesizer does pass 0.0 and 0.2, so a flag named &lt;code&gt;thinking_enabled&lt;/code&gt; silently ends near-deterministic synthesis. Thinking also bills as output tokens with no separate counter, so &lt;code&gt;ChatResponse.output_tokens&lt;/code&gt; mixes answer and reasoning.&lt;/p&gt;

&lt;p&gt;The model-aware budget is opt-in via a value rather than a flag. &lt;code&gt;context_token_budget: int | None&lt;/code&gt; still defaults to &lt;code&gt;12_000&lt;/code&gt;; set it to &lt;code&gt;null&lt;/code&gt;, &lt;code&gt;none&lt;/code&gt;, &lt;code&gt;auto&lt;/code&gt; or empty and &lt;code&gt;Retriever._resolve_effective_budget()&lt;/code&gt; returns &lt;code&gt;int(window * context_window_fraction)&lt;/code&gt;, fraction defaulting to 0.5. &lt;code&gt;resolve_window()&lt;/code&gt; is a first-match scan over a hand-ordered 37-entry table in &lt;code&gt;src/trelix/llm/context_windows.py&lt;/code&gt;, so &lt;code&gt;gpt-4o-2024-11-20&lt;/code&gt; resolves to 128000. Provider-prefixed Bedrock ids do not resolve: &lt;code&gt;us.anthropic.claude-sonnet-4-20250514-v1:0&lt;/code&gt; returns &lt;code&gt;None&lt;/code&gt; and falls back to 12,000 with a WARNING — exactly the default a Bedrock user setting &lt;code&gt;auto&lt;/code&gt; was trying to escape.&lt;/p&gt;

&lt;p&gt;Raising the budget alone changes very little, and the config docstring says so: &lt;code&gt;rerank_top_n&lt;/code&gt; defaults to 15 and &lt;code&gt;top_k_vector&lt;/code&gt;/&lt;code&gt;top_k_bm25&lt;/code&gt; to 20/20, which cap the candidate pool before the packer ever sees a budget. &lt;code&gt;TRELIX_RETRIEVAL_SCALE_TOP_K_TO_BUDGET&lt;/code&gt; is the explicit opt-in that multiplies &lt;code&gt;top_k_vector&lt;/code&gt; and &lt;code&gt;rerank_top_n&lt;/code&gt; — but not &lt;code&gt;top_k_bm25&lt;/code&gt; — by &lt;code&gt;effective_budget / 12_000&lt;/code&gt;, taking &lt;code&gt;top_k_vector&lt;/code&gt; from 20 to 106 on gpt-4o's 64,000. It stays off because it raises per-query cost on three axes at once.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compression that cannot cost you a result
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcbf09i29tdjcl9pr2can.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcbf09i29tdjcl9pr2can.gif" alt="A descriptive summary of what happens in this fifth GIF" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;pack_compressed&lt;/code&gt; runs two waves. Wave 1 is the exact pre-existing uncompressed pack, with the compressor never called. Wave 2 walks the same candidate pool and skips anything already selected, so it only ever gets a second look at candidates wave 1 could not fit — the ones the old packer silently dropped. Accepted entries are appended, never substituted, then re-sorted into ranking order. The compressed selection is a strict superset of the uncompressed one, so Recall, MRR and nDCG cannot regress by construction. That is a property you read; no ranking experiment required. It cannot blow the ceiling either: &lt;code&gt;remaining&lt;/code&gt; is recomputed per candidate, and anything that will not fit even at &lt;code&gt;_FLOOR_RATIO = 0.01&lt;/code&gt; lands in &lt;code&gt;stats["skipped"]&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Citation fidelity is structural. &lt;code&gt;format_compressed_blocks()&lt;/code&gt; emits one &lt;code&gt;[Lines a-b] &amp;lt;qualified_name&amp;gt;&lt;/code&gt; header per kept span with an explicit &lt;code&gt;# ... N lines elided ...&lt;/code&gt; marker in every gap, and each header's text is sliced from the body by the same arithmetic that produced the header. Spans are clamped where they are created and re-filtered again at render time — two independent gates on one invariant, which matters because roughly 35 extractor sites store a truncated body while keeping the full AST span. A header claiming lines its own text does not contain manufactures a confident, wrong citation in a prompt that will quote it verbatim.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ExtractiveCompressor&lt;/code&gt;, the default provider, does zero query-time inference — it reads sub-chunk vectors that already exist from index time. The tripwire that lets me claim the off path is unchanged is &lt;code&gt;tests/unit/test_assembler_backcompat_golden.py&lt;/code&gt;, which loads the pre-compression assembler out of &lt;code&gt;git show v2.12.0:src/trelix/retrieval/assembler.py&lt;/code&gt; and diffs the new one across 8 intents — 177 assertions, including &lt;code&gt;[id(r) for r in new_ctx.results] == [id(r) for r in old_ctx.results]&lt;/code&gt;, so no reordering hides behind equal text.&lt;/p&gt;

&lt;p&gt;The headline "query-conditioned" path is by default lexically conditioned: sub-chunk rows require &lt;code&gt;TRELIX_CHUNKER_MULTI_GRANULARITY=true&lt;/code&gt; (default false), are Python-only even then, and need the query embedding already resident in the embedder's LRU. The changelog's performance claim keeps its hedge: roughly 30-60% fewer synthesis input tokens and 15-35% lower latency on a network-API synthesis path, on the compressible intents. That is an expectation, not a benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  The editor surface, and one ranking flag
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffx2b5l6r1dcme9ng7l5h.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffx2b5l6r1dcme9ng7l5h.gif" alt="A descriptive summary of what happens in this sixth GIF" width="342" height="253"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The VS Code extension is the most visible feature of the release and it needed no server work, because the MCP surface already was the API: it spawns &lt;code&gt;trelix-mcp&lt;/code&gt; over stdio and builds all four features from &lt;code&gt;search_code&lt;/code&gt;, &lt;code&gt;get_symbol&lt;/code&gt;, &lt;code&gt;ask_agent&lt;/code&gt; and &lt;code&gt;blast_radius&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Code lenses re-fire on every keystroke, so the performance contract is enforced in the provider rather than documented. &lt;code&gt;provideCodeLenses&lt;/code&gt; makes zero MCP calls — its only external call is &lt;code&gt;vscode.executeDocumentSymbolProvider&lt;/code&gt; — and only the count-bearing lens reaches MCP, through &lt;code&gt;resolveCodeLens&lt;/code&gt;, for lenses VS Code actually paints, keyed on &lt;code&gt;${docUri}@${docVersion}::${symbolName}&lt;/code&gt;. The version component is the difference between a cache and a correctness bug: any edit must miss rather than show a stale dependent count.&lt;/p&gt;

&lt;p&gt;Chat became possible because of a bug fix that is the same change as the feature. &lt;code&gt;TrelixMcpClient.ask()&lt;/code&gt; used to call &lt;code&gt;getPrompt({ name: "trelix-search" })&lt;/code&gt; and render the joined result as the answer — but &lt;code&gt;getPrompt&lt;/code&gt; returns a filled-in prompt &lt;em&gt;template&lt;/em&gt;, so what users saw was scaffolding telling them to use the &lt;code&gt;search_code&lt;/code&gt; tool. No type checker could have caught it, because both calls return well-formed data and only one of them is an answer. The fix calls &lt;code&gt;ask_agent&lt;/code&gt;, which runs the multi-turn ReAct loop and returns &lt;code&gt;{answer, session_id, turn_count}&lt;/code&gt;; that &lt;code&gt;session_id&lt;/code&gt; is what makes follow-up questions work, so multi-turn chat was structurally impossible until &lt;code&gt;ask()&lt;/code&gt; moved. The regression test asserts &lt;code&gt;getPromptCalls === 0&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The limitations sit next to the constants that cause them. &lt;code&gt;const THREAD_KEY = "default"&lt;/code&gt; exists because the 1.95 chat API gives the handler no stable thread id, so two chat threads open at once share one agent session. The &lt;code&gt;/explain&lt;/code&gt;, &lt;code&gt;/search&lt;/code&gt; and &lt;code&gt;/impact&lt;/code&gt; slash commands are stateless one-shot calls carrying no session at all. The extension is not on the Marketplace: &lt;code&gt;cd workspace-vscode &amp;amp;&amp;amp; npm install &amp;amp;&amp;amp; npm run build &amp;amp;&amp;amp; npm run package&lt;/code&gt; gets you a &lt;code&gt;.vsix&lt;/code&gt;. Its manifest is still at 0.3.0, so the file name does not tell you which feature set it contains.&lt;/p&gt;

&lt;p&gt;The declaration boost is the smallest feature in the release and the easiest to demonstrate, because trelix's own search engine could not find its own search engine. Query the live self-index for &lt;code&gt;bm25_search&lt;/code&gt; and the method that implements it — &lt;code&gt;Database.bm25_search&lt;/code&gt; in &lt;code&gt;src/trelix/store/db.py&lt;/code&gt; — comes back at rank 34 of 90 matches under default unweighted FTS5, beaten by sixteen test functions and four documentation headings that merely mention the term. With &lt;code&gt;top_k_bm25&lt;/code&gt; defaulting to 20 that is a recall failure, not a ranking nuisance: it is out of the candidate pool before fusion or reranking runs. &lt;code&gt;declaration_boost_weight=5.0&lt;/code&gt; moves it to rank 9. The mechanism is a reweighted &lt;code&gt;bm25(symbols_fts, ?, ?, 1.0, 1.0, 1.0)&lt;/code&gt; call applying the weight to &lt;code&gt;name&lt;/code&gt; and &lt;code&gt;qualified_name&lt;/code&gt; only — no schema change, no reindex — and it is one-directional: rows with no name match score byte-identically at both weights. It is gated twice, &lt;code&gt;declaration_boost_enabled&lt;/code&gt; False and &lt;code&gt;declaration_boost_weight&lt;/code&gt; 1.0, with the default no-op pinned by an exact-equality test over the returned score pairs. To reproduce the rank move: &lt;code&gt;export TRELIX_RETRIEVAL_DECLARATION_BOOST=true&lt;/code&gt; and &lt;code&gt;export TRELIX_RETRIEVAL_DECLARATION_BOOST_WEIGHT=5.0&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then three releases proving the surface was true
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiifq2pzg1xssj7j2208d.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiifq2pzg1xssj7j2208d.gif" alt="A descriptive summary of what happens in this seventh GIF" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A big feature drop earns a specific obligation, and v3.0.1, v3.1.0 and v3.1.1 are what paying it looks like.&lt;/p&gt;

&lt;p&gt;v3.0.1 is the consequential one. &lt;code&gt;PythonParser.parse()&lt;/code&gt; recorded local indices into its &lt;code&gt;symbols&lt;/code&gt; list during the walk, then ran &lt;code&gt;symbols.insert(0, &amp;lt;module&amp;gt;)&lt;/code&gt; afterwards whenever the module had a docstring — shifting every recorded index by one. &lt;code&gt;Symbol.parent_id&lt;/code&gt;, &lt;code&gt;CallEdge.caller_id&lt;/code&gt; and &lt;code&gt;TypeEdge.from_symbol_id&lt;/code&gt; all pointed into that list, so one off-by-one corrupted the call graph, the symbol hierarchy and the type graph at once. It never raised and never produced a NULL, because &lt;code&gt;Indexer&lt;/code&gt; builds &lt;code&gt;local_to_db&lt;/code&gt; from the final symbols list, so a pre-insert index still resolved to a valid row. Just the wrong one. A validity check passes; only a correctness check catches this. The fix reserves &lt;code&gt;symbols[0]&lt;/code&gt; before the walk. The release-time measurement was 8,815 of 8,815 index references wrong across 139 source files, which is arithmetic rather than sampling: all 2,179 symbols lived in the 131 docstring-bearing files.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;v3.0.1 requires a reindex.&lt;/strong&gt; v3.0.0's release note said upgrading from v2.12.0 needed no reindex and no migration — true of the schema, false of the graph. The graph the old parser wrote is wrong on disk and no code path corrects it in place. If you upgraded from v2.12.0 or v3.0.0 and did not reindex, your call graph is still wrong: run &lt;code&gt;trelix index ./my-repo&lt;/code&gt; again. v3.1.0 and v3.1.1 need no reindex and no schema change.&lt;/p&gt;

&lt;p&gt;v3.1.0 is a rendering-correctness release, and its best proof needs no attacker at all: &lt;code&gt;src/trelix/indexing/parser/extractors/rust.py:1035&lt;/code&gt; contains &lt;code&gt;re.sub(r"^//[/!]?\s?", ...)&lt;/code&gt;, which Rich reads as an unmatched closing tag, so &lt;code&gt;trelix ask&lt;/code&gt; against trelix's own repository died with &lt;code&gt;MarkupError&lt;/code&gt; instead of showing results. The silent mode is worse: a balanced-looking pair is swallowed and the command exits 0 having dropped characters. A directory named &lt;code&gt;deep[&lt;/code&gt; then raised the same &lt;code&gt;MarkupError&lt;/code&gt; from inside the &lt;code&gt;except&lt;/code&gt; block meant to skip one file, discarding the whole index run and hiding its own cause.&lt;/p&gt;

&lt;p&gt;The machine-readable side needed the opposite fix: &lt;code&gt;_print_json()&lt;/code&gt; leaves the payload byte-identical and disables the renderer instead. &lt;code&gt;soft_wrap=True&lt;/code&gt; is the part that mattered, because Rich hard-wraps at console width and a wrap landing inside a JSON string injects a newline that &lt;code&gt;json.loads&lt;/code&gt; rejects — which the existing short-payload contract tests were too short to catch. Under &lt;code&gt;FORCE_COLOR=1&lt;/code&gt;, routinely set in CI, &lt;code&gt;trelix graph --json | jq&lt;/code&gt; produced garbage at exit 0 because the spinner wrote to stdout. 51 new regression tests, all but five demonstrated failing against v3.0.1 — the exceptions are negative and byte-identity controls that must pass on both sides — and the markup tests each assert the payload's literal characters appear in the output rather than merely that nothing raised.&lt;/p&gt;

&lt;p&gt;v3.1.1's headline is a safety guarantee that was never true. From v2.x through v3.1.0, &lt;code&gt;SECURITY.md&lt;/code&gt; said trelix "does not follow symlinks outside the repo boundary." &lt;code&gt;FileWalker&lt;/code&gt; had no symlink handling at all: &lt;code&gt;_iter_files&lt;/code&gt; used &lt;code&gt;entry.is_dir()&lt;/code&gt;/&lt;code&gt;entry.is_file()&lt;/code&gt;, both of which follow, and &lt;code&gt;rel_path&lt;/code&gt; is computed on the unresolved path, so an out-of-tree file was indexed &lt;em&gt;and&lt;/em&gt; reported as sitting inside the repo. &lt;code&gt;TRELIX_WALKER_FOLLOW_SYMLINKS=false&lt;/code&gt; now makes the boundary real by comparing resolved paths on both sides, because &lt;code&gt;Path.is_relative_to&lt;/code&gt; is lexical and an unresolved comparison would let &lt;code&gt;repo/link -&amp;gt; /etc&lt;/code&gt; straight through. It is opt-in for a non-security reason: confining by default would silently drop files from any repository that symlinks to vendored directories.&lt;/p&gt;

&lt;p&gt;The unit suite sits at 2,457 collected on this branch, from 2,341 at v3.0.0 and 2,445 at v3.1.0. One fact about that span is more telling than the count: &lt;code&gt;src/trelix/audit/&lt;/code&gt; and &lt;code&gt;src/trelix/auth/&lt;/code&gt; have not been touched since the commit that introduced them. The integrity core and the token verifier needed zero source changes across three releases; the code that renders them needed three. And the &lt;code&gt;audit verify&lt;/code&gt; guard that started this article shipped with a test that only ever passed a directory — a path SQLite genuinely cannot open — so it exercised the branch that already worked. Covering the working half of a two-branch guard buys confidence rather than earning it, which is exactly what three releases of looking for the other halves were for.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the 3.x line is for
&lt;/h2&gt;

&lt;p&gt;If you are deciding whether to build on trelix, the useful summary is not the feature list. It is that the 3.x line has a verifiable trail with a CI-gateable exit code, a real OIDC resource-server gate with a forged-token test rather than a config assertion, a context packer whose quality floor is provable by reading it, and an editor integration built entirely out of the MCP surface. Each is off until you turn it on, and each has its boundaries written next to the code that draws them: tamper-evident and HTTP-only, authentication without authorization, lexical compression by default, one agent session across two chat threads.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;pip install 'trelix[serve,sso]'&lt;/code&gt; at 3.0.0 or later gets you the audit and SSO surface. v3.1.1 is cut and dated but not yet tagged or published; when it lands it will change nothing at default settings. The one instruction that carries real consequences is the oldest in the span: if you came from v2.12.0 or v3.0.0, reindex.&lt;/p&gt;

</description>
      <category>python</category>
      <category>security</category>
      <category>systemdesign</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Trelix v2.11.0: Jira and Linear Now Live Inside the Code Graph</title>
      <dc:creator>SAI RAM</dc:creator>
      <pubDate>Sun, 02 Aug 2026 10:13:29 +0000</pubDate>
      <link>https://dev.to/sai_ram_0000/trelix-v2110-jira-and-linear-now-live-inside-the-code-graph-3e6l</link>
      <guid>https://dev.to/sai_ram_0000/trelix-v2110-jira-and-linear-now-live-inside-the-code-graph-3e6l</guid>
      <description>&lt;p&gt;I've had the same conversation with trelix a hundred times: "what does &lt;code&gt;_verify_auth&lt;/code&gt; do." It answers instantly, correctly, and completely uselessly for the question I actually had, which was why does this function exist. Git blame gets you partway there — a commit message, maybe a PR description if you're lucky. But the real answer, the actual requirement, almost always lives somewhere else entirely: a Jira ticket, a Linear issue, a bug report. Code and "the work" have always lived in two systems that never talk to each other. trelix v2.11.0, shipped 2026-08-02, is the release where that stops being true.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7me6f71d7naxxciqbuaa.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7me6f71d7naxxciqbuaa.gif" alt="A person typing excitedly on a laptop and nodding in approval" width="400" height="224"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The headline feature is two new connectors — Jira and Linear — joining the Xray and TestRail connectors trelix already had for test-management platforms, with all four now sharing one interface. Sync a project's tickets and by the time the command returns, those tickets are already wired into the code graph and already influencing search ranking. Not a CSV dump, not a webhook you have to build your own consumer for. Real typed integrations, with each platform's actual authentication quirks handled correctly, sitting behind one shared contract.&lt;/p&gt;

&lt;h2&gt;
  
  
  One interface, four connectors
&lt;/h2&gt;

&lt;p&gt;Every connector — &lt;code&gt;JiraConnector&lt;/code&gt;, &lt;code&gt;LinearConnector&lt;/code&gt;, &lt;code&gt;XrayConnector&lt;/code&gt;, &lt;code&gt;TestRailConnector&lt;/code&gt; — implements the same &lt;code&gt;ArtifactSource&lt;/code&gt; abstract base class (&lt;code&gt;src/trelix/indexing/connectors/base.py&lt;/code&gt;). Two methods: &lt;code&gt;validate_config()&lt;/code&gt; and &lt;code&gt;fetch()&lt;/code&gt;. That's it.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;validate_config()&lt;/code&gt; exists specifically so a missing credential fails loud and fast. Look at what &lt;code&gt;JiraConnector.validate_config()&lt;/code&gt; actually checks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;missing&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;val&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TRELIX_JIRA_BASE_URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TRELIX_JIRA_EMAIL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;email&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TRELIX_JIRA_API_TOKEN&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;api_token&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TRELIX_JIRA_PROJECT_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;_config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;project_key&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;val&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;missing&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;JiraConnector is missing required config: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;missing&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you forgot &lt;code&gt;TRELIX_JIRA_API_TOKEN&lt;/code&gt;, you get told that directly, before a single HTTP request goes out — not a 401 three requests deep into a paginated search, not a silent empty result you have to debug later. &lt;code&gt;fetch()&lt;/code&gt; is the other half of the contract: paginate internally, however the target platform does pagination, and hand back a fully materialized &lt;code&gt;list[Artifact]&lt;/code&gt;. Callers never see a page token or a cursor.&lt;/p&gt;

&lt;p&gt;That uniform shape is the whole reason adding Linear as the fourth &lt;code&gt;ArtifactSource&lt;/code&gt; implementation was tractable inside this release cycle instead of a future one. It didn't have to invent its own calling convention — it just had to satisfy the same two methods over a completely different platform reality underneath.&lt;/p&gt;

&lt;p&gt;The part that actually matters for what you get out of this, though, is &lt;code&gt;sync()&lt;/code&gt;, which lives once on the base class and every connector inherits unchanged:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpyg1sunkmxqdvgmhi2sm.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpyg1sunkmxqdvgmhi2sm.gif" alt="A person celebrating or reacting with excitement" width="500" height="281"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;sync&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;db_writer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ArtifactWriter&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;linker&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ArtifactLinker&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;ConnectorSyncResult&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;validate_config&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;fetched&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;written&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="n"&gt;errors&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="n"&gt;edges_linked&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;artifact&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;fetched&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;db_writer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;upsert_artifact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;artifact&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;written&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
        &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;errors&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
            &lt;span class="k"&gt;continue&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;linker&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;edges_linked&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;linker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;link_one&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;artifact&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;source_ref&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;pass&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;ConnectorSyncResult&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;artifacts_fetched&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;fetched&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;artifacts_written&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;written&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;edges_linked&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;edges_linked&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read that loop carefully: when a linker is supplied, every successfully-written artifact gets passed through &lt;code&gt;linker.link_one()&lt;/code&gt; in the same call. Not queued for a later batch job. Immediately, synchronously, before &lt;code&gt;sync()&lt;/code&gt; returns. Run &lt;code&gt;trelix connector sync ./repo jira&lt;/code&gt; and the output isn't "downloaded 84 tickets into a table" — it's &lt;code&gt;Synced jira: fetched 84, wrote 84, errors 0, linked 79 edge(s)&lt;/code&gt;. Those 79 edges already exist in &lt;code&gt;generic_edges&lt;/code&gt; and are already visible to PageRank the moment the process exits. There's a standalone &lt;code&gt;trelix link-artifacts&lt;/code&gt; command if you want a manual full re-link pass — useful after a &lt;code&gt;--no-link&lt;/code&gt; bulk sync, or after a schema change surfaces new matches against artifacts you already synced — but it's an option, not a requirement. The default path gives you a fully linked graph in one command.&lt;/p&gt;

&lt;h2&gt;
  
  
  Jira and Linear: same contract, opposite platforms
&lt;/h2&gt;

&lt;p&gt;This is the part I find genuinely interesting to build, because Jira and Linear satisfy the identical &lt;code&gt;ArtifactSource&lt;/code&gt; contract while being about as different as two ticket-tracker APIs can be.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjanvmbngqr15cfzkffxn.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjanvmbngqr15cfzkffxn.gif" alt="A humorous or abstract animation showing a visual concept" width="320" height="240"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Jira Cloud is REST, authenticated with HTTP Basic — email plus API token, Jira's own recommended scheme for this kind of integration, no OAuth dance required. Pagination is a cursor token (&lt;code&gt;nextPageToken&lt;/code&gt;), not offset/limit. The interesting wrinkle is the description field: Jira Cloud's v3 API returns &lt;code&gt;description&lt;/code&gt; as Atlassian Document Format, a nested JSON tree, on every single ticket, unconditionally — never a plain string. That's not a rare shape you might hit on a rich-text ticket; it's the only shape. So the connector ships a real renderer, &lt;code&gt;_adf_to_text()&lt;/code&gt;, that walks the tree — paragraphs, headings, bullet and ordered lists, code blocks, inline code, blockquotes, panels, expand sections, rules, links, embedded cards — and produces plain text good enough for &lt;code&gt;ArtifactLinker&lt;/code&gt;'s regex matching downstream. It's not a full ADF renderer (no tables, no user-mention name resolution), but it covers everything actually observed on a live Jira Cloud site.&lt;/p&gt;

&lt;p&gt;Linear is a different animal entirely. There's no REST surface at all — it's GraphQL-only, a single POST endpoint (&lt;code&gt;https://api.linear.app/graphql&lt;/code&gt;), and there's no official Python SDK, only a TypeScript one. Pagination is Relay-style: &lt;code&gt;first&lt;/code&gt;/&lt;code&gt;after&lt;/code&gt; arguments, &lt;code&gt;pageInfo { hasNextPage, endCursor }&lt;/code&gt;. The loop has to terminate on &lt;code&gt;hasNextPage&lt;/code&gt;, not on a short page — Relay pagination can legitimately hand back fewer results than requested while &lt;code&gt;hasNextPage&lt;/code&gt; is still true, so checking &lt;code&gt;len(results) &amp;lt; page_size&lt;/code&gt; the way TestRail and Xray do would end the sync early and silently drop issues.&lt;/p&gt;

&lt;p&gt;The auth detail is worth calling out by name, because it's the exact kind of thing that breaks silently if you're not paying attention: Linear's personal API key goes into the &lt;code&gt;Authorization&lt;/code&gt; header with no &lt;code&gt;Bearer&lt;/code&gt; prefix. Just the raw key. Every other connector in this codebase uses &lt;code&gt;Bearer &amp;lt;token&amp;gt;&lt;/code&gt; — Xray's connector literally does &lt;code&gt;f"Bearer {self._token}"&lt;/code&gt; — so copy-pasting that pattern for Linear produces a request that authenticates against nothing and fails with a 401 whose message doesn't obviously point at the missing prefix. The module docstring in &lt;code&gt;linear.py&lt;/code&gt; calls this out explicitly as "the one detail most likely to get corrected back to Bearer by someone skimming Xray's connector — don't."&lt;/p&gt;

&lt;p&gt;Linear also runs a real query-complexity budget: 10,000 points per request, where a connection multiplies its children's cost by its own &lt;code&gt;first&lt;/code&gt; argument. This connector's field selection costs roughly 7.7 points per issue, so at &lt;code&gt;page_size=100&lt;/code&gt; that's about 1 + 770 ≈ 771 points — comfortably under the cap, with real headroom for schema drift. And Linear signals rate-limiting differently from everything else in this codebase: HTTP 400, not 429, carrying a GraphQL body error with &lt;code&gt;extensions.code == "RATELIMITED"&lt;/code&gt;, and no &lt;code&gt;Retry-After&lt;/code&gt; header at all. Reset timing has to come from reading &lt;code&gt;X-RateLimit-Requests-Reset&lt;/code&gt;/&lt;code&gt;X-RateLimit-Complexity-Reset&lt;/code&gt; response headers directly, so &lt;code&gt;_fetch_issues_page()&lt;/code&gt; runs its own small, connector-local retry loop for exactly this case rather than teaching the shared retry contract a Linear-specific response shape.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually makes a synced ticket useful
&lt;/h2&gt;

&lt;p&gt;Downloading a ticket into a table doesn't do anything by itself. The piece that turns a synced artifact into something that changes search results is &lt;code&gt;ArtifactLinker&lt;/code&gt; (&lt;code&gt;src/trelix/indexing/artifact_linker.py&lt;/code&gt;), and it's worth walking through because the design has a real precision problem baked into it that's easy to get wrong.&lt;/p&gt;

&lt;p&gt;Stage one is regex extraction: &lt;code&gt;_IDENTIFIER_RE&lt;/code&gt; finds identifier-shaped tokens in an artifact's title and body — function/class/variable-looking strings — and checks each one, casefolded, against an index of every symbol name and qualified name in the codebase. A hit becomes a &lt;code&gt;GenericEdge&lt;/code&gt; with &lt;code&gt;edge_kind="references_artifact"&lt;/code&gt; and &lt;code&gt;weight=1.0&lt;/code&gt; — a real graph edge connecting that ticket directly to that code symbol.&lt;/p&gt;

&lt;p&gt;The obvious failure mode here is that plenty of function names are also ordinary English words. A ticket titled "the test suite failed to run and update the process" would, on pure token overlap, spuriously link to functions literally named &lt;code&gt;run&lt;/code&gt;, &lt;code&gt;test&lt;/code&gt;, &lt;code&gt;update&lt;/code&gt;, and &lt;code&gt;process&lt;/code&gt; — none of which have anything to do with that ticket. &lt;code&gt;ArtifactLinker&lt;/code&gt; handles this with &lt;code&gt;_COMMON_WORD_STOPLIST&lt;/code&gt;, roughly 40 common English/programming words (&lt;code&gt;get&lt;/code&gt;, &lt;code&gt;set&lt;/code&gt;, &lt;code&gt;run&lt;/code&gt;, &lt;code&gt;test&lt;/code&gt;, &lt;code&gt;process&lt;/code&gt;, &lt;code&gt;update&lt;/code&gt;, &lt;code&gt;data&lt;/code&gt;, &lt;code&gt;file&lt;/code&gt;, &lt;code&gt;main&lt;/code&gt;, and so on) excluded specifically from the bare-name side of the index. It's a real precision guard, not an afterthought, and it's scoped carefully: it only gates the bare &lt;code&gt;name&lt;/code&gt; key, not &lt;code&gt;qualified_name&lt;/code&gt;, so a genuinely qualified match like &lt;code&gt;auth.run&lt;/code&gt; still links correctly. Only the literal stoplisted string on its own gets excluded.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F39rwdmddpte6s4twh0sy.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F39rwdmddpte6s4twh0sy.gif" alt="A humorous animation showing garbage sorting and sifting" width="384" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Stage two is an opt-in embedding-similarity fallback, gated behind &lt;code&gt;ArtifactLinkerConfig.embedding_fallback_enabled&lt;/code&gt;, and it only runs for artifacts where stage one found zero matches. Those fallback edges get &lt;code&gt;weight=0.5&lt;/code&gt; — deliberately half the weight of a regex hit — so a fuzzy semantic match never outranks an exact identifier match for PageRank purposes. Exact reference wins; embedding similarity is a second-chance net, not a peer signal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Xray and TestRail: proving the abstraction generalizes
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fznmlspcmg8p0qun9ges8.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fznmlspcmg8p0qun9ges8.gif" alt="A satisfying line production animation showing items moving along a track" width="480" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If the &lt;code&gt;ArtifactSource&lt;/code&gt; contract only worked for issue trackers, I'd be less confident in it. Xray and TestRail are both test-management platforms, and they slot into the exact same shape.&lt;/p&gt;

&lt;p&gt;TestRail is the simpler of the two: HTTP Basic auth (username plus API key), offset/limit pagination capped at TestRail's own 250-per-page ceiling. Xray Cloud is odder — its tests are Jira issues under the hood, but test-specific content (steps, expected results) has no Jira-native equivalent, so a single GraphQL &lt;code&gt;getTests&lt;/code&gt; query pulls both the Jira fields and the Xray-specific step data in one call, with no separate REST round-trip needed. Auth is its own thing too: a &lt;code&gt;client_id&lt;/code&gt;/&lt;code&gt;client_secret&lt;/code&gt; pair, distinct from a personal Jira API token, exchanged via &lt;code&gt;POST /api/v2/authenticate&lt;/code&gt; for a short-lived bearer JWT attached to every subsequent request. Both connectors write &lt;code&gt;artifact_kind="test_case"&lt;/code&gt; records, and both flow through the exact same &lt;code&gt;sync()&lt;/code&gt; → &lt;code&gt;ArtifactLinker&lt;/code&gt; path as Jira and Linear's tickets. Same registry lookup, same CLI surface, same auto-link behavior. The abstraction holds.&lt;/p&gt;

&lt;p&gt;The CLI is identical across all four: &lt;code&gt;trelix connector sync ./repo &amp;lt;jira|testrail|xray|linear&amp;gt;&lt;/code&gt;, with a &lt;code&gt;--link/--no-link&lt;/code&gt; flag defaulting to on, resolved through one factory — &lt;code&gt;get_artifact_source("jira"|"testrail"|"xray"|"linear")&lt;/code&gt; in &lt;code&gt;registry.py&lt;/code&gt; — that mirrors the same match-statement pattern trelix already uses for embedder and vector-store selection.&lt;/p&gt;

&lt;h2&gt;
  
  
  The plumbing underneath
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fccntu831wn601o40rdr8.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fccntu831wn601o40rdr8.gif" alt="A blueprint or civil engineering drafting animation showing technical designs" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;None of the above would mean much if a sync silently swallowed failures or nobody could tell why a request hung. This release also landed a unified retry/backoff contract (&lt;code&gt;core/retry.py&lt;/code&gt;, built on tenacity) shared across every LLM backend, embedder, and connector — full-jitter exponential backoff on 429s and 5xx responses, honoring a server's &lt;code&gt;Retry-After&lt;/code&gt; header when present. Structured JSON logging shipped alongside it, with trace-ID correlation when OpenTelemetry is enabled. And Personalized PageRank landed as an opt-in flag (&lt;code&gt;TRELIX_RETRIEVAL_PAGERANK_PERSONALIZATION&lt;/code&gt;) that concentrates teleport mass on ticket-linked symbols instead of spreading it uniformly — its own docs flag a real interaction risk: combined with &lt;code&gt;pagerank_boost_enabled&lt;/code&gt; on a repo with few ticket-linked symbols, it can invert the ranking &lt;code&gt;get_top_central_symbols()&lt;/code&gt; produces. Opt-in for a reason.&lt;/p&gt;

&lt;p&gt;None of this was validated against mocks alone, either. Live-testing the Jira connector against a real production Jira Cloud site turned up two real bugs — ADF descriptions silently rendering as empty bodies on every ticket, and bad credentials silently reporting success because &lt;code&gt;/rest/api/3/search/jql&lt;/code&gt; returns an identical empty result set for a bad token, no token, and a genuinely empty project. Both got fixed this cycle: a real &lt;code&gt;_adf_to_text()&lt;/code&gt; renderer for the first, and a pre-flight auth check against &lt;code&gt;/rest/api/3/myself&lt;/code&gt; for the second, since that endpoint correctly returns a real 401 where the search endpoint doesn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this actually unlocks
&lt;/h2&gt;

&lt;p&gt;The full unit suite sits at 1,909 passing tests as of this release, up from 1,643 at v2.9.0 — not because coverage was thin before, but because four connectors, a linker, a retry contract, and Personalized PageRank all needed their own test surface.&lt;/p&gt;

&lt;p&gt;But the number that matters isn't the test count. It's that trelix can now sit in front of a question like "why does this function exist" and actually answer it, because the ticket that requested the change, or the bug report that prompted the fix, is sitting in the graph one edge away from the symbol itself. Ask "what's supposed to happen when I touch this code" and the answer isn't just call-graph traversal anymore — it's the test case that verifies it, linked in through the same mechanism, the same weight-1.0 edge, the same &lt;code&gt;ArtifactLinker&lt;/code&gt; pass. Before this release, trelix understood code. Now it understands why the code is there.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>developertools</category>
      <category>python</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>trelix v2.7 to v2.9: The Release Where the Pipeline Itself Became the Product</title>
      <dc:creator>SAI RAM</dc:creator>
      <pubDate>Fri, 24 Jul 2026 16:44:32 +0000</pubDate>
      <link>https://dev.to/sai_ram_0000/trelix-v27-to-v29-the-release-where-the-pipeline-itself-became-the-product-21kb</link>
      <guid>https://dev.to/sai_ram_0000/trelix-v27-to-v29-the-release-where-the-pipeline-itself-became-the-product-21kb</guid>
      <description>&lt;p&gt;On 2026-07-09 I shipped trelix v2.7.0. The architecture felt done — seven retrieval legs, a knowledge graph, an agentic loop. Then I opened the GitHub Release page and counted two binary assets where there should have been three. &lt;code&gt;release.yml&lt;/code&gt; built the macOS and Linux PyInstaller binaries both as &lt;code&gt;dist/trelix&lt;/code&gt;, and &lt;code&gt;softprops/action-gh-release&lt;/code&gt; uploads assets by basename — two files with the same name collide into one asset, with no metadata telling you which OS survived. I genuinely could not tell, from the published release, whether the surviving binary was macOS or Linux. That's not a bug in trelix's retrieval logic; it's a bug in the thing that puts trelix in front of users, and it shipped because nobody checked the page after the workflow went green.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4nlmhdy0bzqqyi9ydteo.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4nlmhdy0bzqqyi9ydteo.gif" alt="Description of the GIF" width="500" height="281"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That bug is why this article exists. v2.7.1 through v2.9.0 — six releases over two weeks — spend a surprising amount of energy on parts of the project that aren't retrieval quality at all: the release pipeline, the concurrency model, a language-parser migration, the deployment story, and a VS Code extension and GitHub App with real, embarrassing problems. I'm grouping by theme instead of walking the changelog top to bottom.&lt;/p&gt;

&lt;h2&gt;
  
  
  Shipping infra is its own product surface, and it broke
&lt;/h2&gt;

&lt;p&gt;The binary collision got fixed in v2.7.1 (2026-07-10) by renaming each binary uniquely before upload. Fresh eyes on the pipeline turned up company: the PR-time CI workflow had never built a Linux binary, even though the release workflow built one at tag time. I added the missing matrix entry and a verify step.&lt;/p&gt;

&lt;p&gt;The most embarrassing find: &lt;code&gt;trelix-mcp&lt;/code&gt;'s own test suite had never run in CI. It let a real regression sit undetected — a test asserted the MCP server had "exactly 6 tools" when it had actually registered 8 since two subscription tools shipped in v2.5.0. A wrong-but-passing test is worse than no test. I wired the tests into CI and fixed the assertion.&lt;/p&gt;

&lt;p&gt;Smaller mistakes, same release: three companion packages' dependency floors on trelix core had been bumped in v2.7.0 on the assumption they used new APIs. None did, so I reverted them. The changelog had also rotted: an entry was defined twice with conflicting URLs, silently resolving to the last definition, so the first link was dead. Rebuilt the footer from the actual git tags.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five concurrency bugs, found only once I wrote real stress tests
&lt;/h2&gt;

&lt;p&gt;v2.7.2 (2026-07-12) reads like the payoff of finally taking concurrency seriously instead of assuming &lt;code&gt;check_same_thread=False&lt;/code&gt; meant "safe." It shipped real scale features — Qdrant Cloud readiness, incremental per-symbol embedding on partial re-index, an opt-in parallel BM25 read pool — but the five bugs stress-testing surfaced matter more.&lt;/p&gt;

&lt;p&gt;A TOCTOU race in the sparse embedder's lazy-load checked whether the model was loaded before acquiring its lock, so two racing threads could both see "not loaded" and both start loading; fixed with double-checked locking. An MCP stdout write race let concurrent notification writes interleave partial JSON-RPC lines, corrupting a client's output; fixed with a lock around the write-and-flush pair. And unbounded subscription-registry growth let a misbehaving client grow the registry forever, fixed with a max-subscriber cap and a TTL sweep.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fndm7kuuc8wjmrkpih7hu.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fndm7kuuc8wjmrkpih7hu.gif" alt="A person juggling multiple items while working" width="500" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The one I find genuinely unsettling in hindsight: silent foreign-key corruption on partial re-index. The parent-symbol, call-callee, and type-edge columns are all set to null on delete, which sounds safe until deleting a changed symbol's old row silently nulls those links on every row that referenced it — including unchanged rows with nothing to do with the edit. No error, just quietly severed graph edges accumulating as files change. I added snapshot-and-repoint helpers so the indexer captures stale links before the delete and repoints them.&lt;/p&gt;

&lt;p&gt;And the BM25 lock was incomplete even after I thought I'd handled it. The shared database connection is opened with &lt;code&gt;check_same_thread=False&lt;/code&gt;, which I'd treated as a green light for concurrent use — it isn't; that flag disables SQLite's thread-affinity check, not concurrent-statement safety. Grep, sparse, and vector legs all hydrate through that connection from sibling worker threads. I added a real lock everywhere it's touched from a worker thread, verified with a 60-thread by 10-iteration by 3-leg stress test and zero errors — the only way I'd trust this, given the previous fix had looked equally trustworthy.&lt;/p&gt;

&lt;p&gt;Same release: a Qdrant client API migration to keep pace with an upstream deprecation, pinned so a future major bump can't break it again. One honest non-win: Windows ARM64 binaries were briefly added to the build matrices, then reverted — two core dependencies publish no wheel for that target. Linux ARM64 shipped; Windows ARM64 didn't.&lt;/p&gt;

&lt;p&gt;v2.7.3 was a pure documentation release. A full README audit fixed 15+ factual bugs: wrong env var names, fabricated pip extras, a broken Homebrew tap, a config value that crashes on use, wrong REST method and table names. The architecture diagram got redrawn to show all 7 retrieval legs instead of 3. I backfilled the changelog's empty &lt;code&gt;[2.2.0]&lt;/code&gt; entry, which had shipped 5 real features — agentic ReAct loop, data-flow analysis, taint analysis, sparse+dense hybrid retrieval, multi-granularity indexing — never documented anywhere. The README shrank by roughly a third by consolidating duplicated API sections into pointers. None of it changed behavior; it changed whether docs matched reality.&lt;/p&gt;

&lt;h2&gt;
  
  
  Becoming genuinely deployable, not just runnable
&lt;/h2&gt;

&lt;p&gt;v2.9.0 is where trelix stopped being "clone it, pip install, run the CLI" and became something you could put in front of other services. Four pieces went in, and the order was the point.&lt;/p&gt;

&lt;p&gt;Typed REST API response models came first — every route now declares a real Pydantic model instead of an untyped object in the OpenAPI schema, groundwork for the next item. Cursor pagination on the search endpoint came second, the one deliberately narrow breaking change in an otherwise additive release: the endpoint now returns a results-plus-cursor envelope instead of a bare list, matching the MCP tool's existing pagination contract, done so the new SDK wouldn't lock in against the old shape.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frthvnqawh3e15izdcfzq.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frthvnqawh3e15izdcfzq.gif" alt="A container ship or shipping container being discharged" width="480" height="176"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The TypeScript SDK came third — a hand-written HTTP client with generated types covering every route, with &lt;code&gt;/ask&lt;/code&gt;'s streaming response getting its own async generator because token-plus-terminator-plus-error-frame semantics don't fit the same shape as everything else. Fourth, OpenTelemetry tracing — opt-in, off by default, emitting one span per retrieval leg plus spans for each pipeline stage. Worth explaining: OTel's context propagation doesn't cross a thread-pool boundary automatically, and trelix's parallel sub-query execution runs inside exactly that kind of pool, so naive instrumentation produces disconnected, orphaned spans. The fix wraps the traced function to carry the current context across the pool boundary — the docs note the underlying span-naming conventions are still "Development," not "Stable."&lt;/p&gt;

&lt;p&gt;The deployment story rounds out with an official multi-arch Docker image in two variants — a slim, API-embedder-only build and a &lt;code&gt;-local&lt;/code&gt; build bundling the offline embedding stack — running as a non-root user, with the entrypoint overriding the CLI's loopback-only default to listen on all interfaces (a loopback bind is a silent dead-end in a container unless overridden). Plus a Helm chart modeling the server's real behavior: every route re-derives its config from the request's own repo parameter, so one Deployment is already multi-repo-capable — but the chart's persistent volume is a shared data directory across every repo you serve through it, called out loudly in the docs. Ingress defaults to disabled: the server ships with zero auth middleware, and I'd rather state that directly than bury it.&lt;/p&gt;

&lt;p&gt;Building this surfaced the same documentation-rot pattern again: docs referenced an embedder env var that was actually a silent no-op, plus a nonexistent CLI flag and Docker examples with the wrong port — docs drift silently unless something forces someone to run the commands in them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Python 3.13, and how one dependency swap turned into six bugs
&lt;/h2&gt;

&lt;p&gt;The Python 3.13 item sounds like the least interesting line in the whole changelog: bump the minimum version, done. The actual blocker was &lt;code&gt;tree-sitter-languages&lt;/code&gt;, abandoned upstream with no wheels for the new interpreter. The fix was a single swap to the actively-maintained &lt;code&gt;tree-sitter-language-pack&lt;/code&gt;, behind trelix's one chokepoint for grammar loading — that's the whole diff at the chokepoint. It is not the whole blast radius: the new library exposes different AST node names and shapes, which propagated silently into nearly every per-language extractor, because each had been written against the old grammar's node vocabulary without anyone realizing how much was implicit knowledge rather than a documented contract.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5rmeaehdafewf2xryq7s.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5rmeaehdafewf2xryq7s.gif" alt="Abstract animation representing internet technology" width="500" height="281"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Six separate bugs came out of this: C#'s grammar name changed, silently tolerated before this release only because the old library accepted both names; Kotlin's extractor had to be rewritten entirely after its field-lookup API stopped returning anything, breaking class, interface, enum, function, and property extraction at once; Python's docstring extraction broke because the new grammar drops a wrapper node; Go's interface-method node was renamed; TypeScript's interface-body node got a new name; and C#'s &lt;code&gt;using&lt;/code&gt; alias imports lost their wrapper node.&lt;/p&gt;

&lt;p&gt;Every one of these was catchable only because per-language extractor tests already existed before the migration — without them, most would have shipped as silent failures. A seventh casualty I almost missed: the PyInstaller build spec still imported the retired package, breaking every binary build with a flat import error, because that build runs in its own workflow, outside the test suite and linters. Dropped the stale entries, added the new package as a hidden import, verified with a local build plus a smoke test.&lt;/p&gt;

&lt;p&gt;One behavior change worth knowing: grammar loading is now network-on-first-use with local caching, not bundled in the wheel — fine on a laptop, mildly alarming for an air-gapped job. A prefetch function warms the cache during image builds; CI runs it automatically now.&lt;/p&gt;

&lt;h2&gt;
  
  
  The VS Code extension, the GitHub App, and an admission I'm not going to soften
&lt;/h2&gt;

&lt;p&gt;The search command became a debounced (250ms) search-as-you-type picker with live snippet preview, instead of the old one-shot input-then-static-list flow. The preview runs through a virtual document provider that gets real syntax highlighting for free by keeping the real file's extension on its virtual URI. The state machine lives in its own testable class, verified with a fake-timer harness proving the debounce actually debounces.&lt;/p&gt;

&lt;p&gt;The security fix here is direct: the ask-panel's Webview interpolated the raw, unescaped LLM answer string straight into HTML, with no content-security policy and no script restrictions. A crafted or adversarial answer — not a hypothetical when you're piping retrieval results through a model — could execute arbitrary script inside the Webview's context. Fixed with HTML-escaping plus disabled scripts and an explicit deny-all CSP; it's an XSS vulnerability, and it's fixed now. Separately: search results had been silently mis-parsed the entire time the extension existed — it read the wrong field names off each result, confirmed against the actual server source, so those fields were always empty strings and clicking a result opened a broken, empty file URI.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmtcgs92ejdsc13wcooye.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmtcgs92ejdsc13wcooye.gif" alt="A hacker typing quickly on a keyboard" width="700" height="394"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And then the line I'm quoting close to verbatim because paraphrasing softens it: the trelix Code Review Check has never posted a single real annotation since the workflow shipped. Status messages ran unconditionally to stdout even in JSON output mode, and combined with the workflow's output redirect, the annotation-posting step's JSON parse had been throwing on every run, silently swallowed since the day this workflow shipped. Every PR ever reviewed by this Check got nothing. Fixed by routing status output to stderr and narrowing the redirect to stdout only. Even after that fix, the mapping logic still wouldn't have worked — wrong response keys, lowercase severity strings compared against real uppercase values. New regression tests were verified against the pre-fix code first — most failed with the exact parse error this bug produces — before trusting they passed for the right reason. The GitHub App also reached GA-readiness: real installation-token minting behind an expiry-aware cache, webhook signature verification via constant-time comparison, a request-body size cap matching GitHub's own limit, and a subprocess timeout on the review shell-out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multi-repo federation, and the security fix found before it hit anyone
&lt;/h2&gt;

&lt;p&gt;v2.8.0 and v2.8.1 shipped the same day, 2026-07-20, and the second was a direct response to auditing the first before letting it sit. v2.8.0 exposed the existing federation infrastructure to MCP clients through four new tools, and added a CLI command apparently missing despite the underlying method already existing. Persistent agent memory landed for the agentic loop too — a follow-up call can now resume a prior conversation with full context, sessions auto-evicting after a week of inactivity.&lt;/p&gt;

&lt;p&gt;Building this surfaced two real, previously invisible bugs: federated search had silently lost repo provenance in an earlier refactor, so the search command's repo column had been blank with no test catching it; and a per-repo weighting setting — settable, stored, documented — had never actually been forwarded into the fusion math, silently doing nothing since it was added.&lt;/p&gt;

&lt;p&gt;v2.8.1 is where the real security finding lives. All four federation MCP tools passed a caller-supplied config path straight into the registry's load/save calls with zero validation, meaning an MCP client — including a prompt-injected agent, exactly the threat model MCP has to take seriously — could point registry I/O at an arbitrary filesystem path. I found this in a pre-push audit of v2.8.0, before it reached anyone running the released version. The fix confines the path to one of two known-safe directories via a proper containment check, not a naive string-prefix check, which would also incorrectly match a similarly-named sibling directory. Same release: repo-count and fan-out caps so a runaway add-repo loop can't scale every search linearly against an unbounded repo count, plus a pagination fix for a per-repo candidate pool that had been widening as the cursor grew, letting later pages get fused from a differently-shaped pool than earlier ones.&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>devops</category>
      <category>python</category>
      <category>observability</category>
    </item>
    <item>
      <title>Tombstone v1.3-v1.4: Resilience Was the Easy Layer</title>
      <dc:creator>SAI RAM</dc:creator>
      <pubDate>Thu, 09 Jul 2026 19:57:31 +0000</pubDate>
      <link>https://dev.to/sai_ram_0000/tombstone-v13-v14-resilience-was-the-easy-layer-370e</link>
      <guid>https://dev.to/sai_ram_0000/tombstone-v13-v14-resilience-was-the-easy-layer-370e</guid>
      <description>&lt;p&gt;I removed four &lt;code&gt;|| true&lt;/code&gt; statements from a GitHub Actions workflow on July 9th and watched CI go red in four different ways within the same run. Not one failure. Four — a Python pytest install that had never actually finished, a ruff lint violation nobody had looked at, a Ruby require path that resolved to nothing, and a Java Gradle wrapper that didn't exist in the repo. All four had been "passing" for who knows how long, because the test steps were configured to succeed no matter what came back.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwxpn2rf3idpkjmelfohg.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwxpn2rf3idpkjmelfohg.gif" alt="A massive chain reaction of colorful dominoes toppling over" width="200" height="200"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That's the theme of this release window. v1.2 was about making the running system survive failure — retries, circuit breakers, idempotency keys, DLQs. This one is about making the layers &lt;em&gt;above&lt;/em&gt; the running system — the Helm chart, the SDKs, the GitOps pipeline, the CI config — tell the truth about their own state. Resilience isn't a feature you ship, it's a property you discover you're missing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Helm chart was only deploying two of five services
&lt;/h2&gt;

&lt;p&gt;Tombstone has five application services — flag-api, gateway, evaluator, intelligence, marketplace — plus the operator. Until v1.3.0, the Helm chart (&lt;code&gt;infra/helm/flagmind&lt;/code&gt;) only had Deployment templates for two of them. This wasn't a secret; it was written down in &lt;code&gt;COMPATIBILITY.md&lt;/code&gt; under a section literally titled "Known Gap." Run &lt;code&gt;helm install&lt;/code&gt; in a fresh cluster and you got flag-api and gateway, nothing else. Anyone deploying evaluator, intelligence, or marketplace was hand-rolling manifests or copy-pasting the two existing templates and hoping the env vars lined up.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fougtjushud9jqfg1wgw5.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fougtjushud9jqfg1wgw5.gif" alt="An animation of a white puzzle where the last piece is placed into a missing gap" width="580" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;v1.3.0 closes that gap: &lt;code&gt;deployment-evaluator.yaml&lt;/code&gt;, &lt;code&gt;deployment-intelligence.yaml&lt;/code&gt;, and &lt;code&gt;deployment-marketplace.yaml&lt;/code&gt; now ship in the chart. evaluator gets an optional HPA behind &lt;code&gt;evaluator.autoscaling.*&lt;/code&gt;. intelligence exposes &lt;code&gt;IS_PRIMARY_REGION&lt;/code&gt; straight from &lt;code&gt;values.yaml&lt;/code&gt;, which matters for the multi-region setup where secondary regions run in read-only relay mode.&lt;/p&gt;

&lt;p&gt;Here's the footgun worth flagging, because it isn't obvious until it bites you: every one of those templates uses &lt;code&gt;tombstone.selectorLabels&lt;/code&gt; in &lt;code&gt;spec.selector.matchLabels&lt;/code&gt;, not the separate &lt;code&gt;tombstone.labels&lt;/code&gt; helper, which includes a version label. It's tempting to use the one "labels" helper everywhere for consistency. Don't. Kubernetes Deployment selectors are immutable once the object exists. If your selector helper includes a label that changes on every release — like a chart version — the second &lt;code&gt;helm upgrade&lt;/code&gt; you ever run will fail outright, because the new selector no longer matches the old one. &lt;code&gt;tombstone.selectorLabels&lt;/code&gt; is a narrower, stable subset — name and component, nothing that changes across releases — specifically so &lt;code&gt;matchLabels&lt;/code&gt; never drifts. One line in a template, and it's the difference between a chart that upgrades cleanly forever and one that works exactly once.&lt;/p&gt;

&lt;h2&gt;
  
  
  SDK parity is a correctness bug, not a feature request
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzo9d316xhqfueadbace3.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzo9d316xhqfueadbace3.gif" alt="A man in a suit doing a dramatic double-take and looking back at the camera" width="480" height="360"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The TypeScript SDK (&lt;code&gt;@tombstone/core&lt;/code&gt;) has had the full 5-step evaluation pipeline since v2.0 — preliminary checks, prerequisites, individual targeting, rule matching, fallthrough — with the complete operator set and semver comparisons. The Python SDK didn't. It had basic targeting but was missing large chunks of the operator surface and all of the prerequisite evaluation logic. That's not a nice-to-have gap: a flag with a &lt;code&gt;semver_gte&lt;/code&gt; rule or a prerequisite on another flag could evaluate to &lt;code&gt;true&lt;/code&gt; in a Node service and silently fall through to the default in a Python service, evaluating the same flag against the same user. Two SDKs, two answers, same input — the kind of bug that looks like a targeting mistake in the dashboard when it's actually a parity gap in the client.&lt;/p&gt;

&lt;p&gt;v1.3.0 closes it. &lt;code&gt;packages/sdks/flagmind-python/tombstone/matching.py&lt;/code&gt; now implements the full operator set — eq/neq/in/nin/contains/startsWith/endsWith, the four numeric comparisons, and all five semver operators (&lt;code&gt;semver_gt/gte/lt/lte/eq&lt;/code&gt;) — plus &lt;code&gt;date_before&lt;/code&gt;/&lt;code&gt;date_after&lt;/code&gt;. The semver comparison is hand-rolled: a &lt;code&gt;_padded_version()&lt;/code&gt; helper left-pads each numeric segment to 5 characters and appends a &lt;code&gt;~&lt;/code&gt; sentinel for 3-part releases, so &lt;code&gt;1.0.0-beta&lt;/code&gt; sorts below &lt;code&gt;1.0.0&lt;/code&gt; using pure string comparison. It's the same GrowthBook &lt;code&gt;paddedVersionString()&lt;/code&gt; pattern the TypeScript SDK already used, with zero new runtime dependencies — no &lt;code&gt;semver&lt;/code&gt; package, no extra install footprint.&lt;/p&gt;

&lt;p&gt;Prerequisite evaluation was the other half. &lt;code&gt;evaluation.py&lt;/code&gt; now threads an &lt;code&gt;evaluation_cache: dict[str, bool]&lt;/code&gt; through the recursive prerequisite check, so a flag with three prerequisites sharing a common ancestor doesn't re-evaluate that ancestor three times. Circular chains are rejected via a &lt;code&gt;_seen_keys&lt;/code&gt; tracking set rather than recursing forever. The SDK also distinguishes two failure modes with dedicated exception types: &lt;code&gt;InconclusiveMatchError&lt;/code&gt; means a targeting condition couldn't be evaluated locally — missing attribute, type mismatch — and the caller should move on to the next rule. &lt;code&gt;RequiresServerEvaluation&lt;/code&gt; means the evaluation genuinely needs data the local cache doesn't have, and the client falls back to a REST round-trip instead of silently returning a wrong default. Conflating the two used to mean callers couldn't tell "skip this rule" apart from "call the server."&lt;/p&gt;

&lt;h2&gt;
  
  
  Making the API impossible to not find
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvrwcsvkesp8choqacec4.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvrwcsvkesp8choqacec4.gif" alt="A bright spotlight shining down against a dark background" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The previous article's two worst bugs — the Slack kill switch sending &lt;code&gt;environment&lt;/code&gt; as a query param when the handler read it from the JSON body, and the four-eyes approval routes that were built, reviewed, and merged but never registered in &lt;code&gt;flag-api/cmd/main.go&lt;/code&gt; — share a root cause that has nothing to do with the code itself. Nobody had an easy way to look at what &lt;code&gt;flag-api&lt;/code&gt; actually exposed versus what the proto files said it should expose. The OpenAPI spec existed; nobody was looking at it, because there was nowhere convenient to look.&lt;/p&gt;

&lt;p&gt;v1.3.0 adds a Redoc explorer at &lt;code&gt;GET /api/v1/docs&lt;/code&gt;, embedded via &lt;code&gt;go-redoc&lt;/code&gt; rather than pulled from a CDN — it reads the existing grpc-gateway OpenAPI spec at &lt;code&gt;/api/v1/openapi.json&lt;/code&gt;, so there's no second source of truth to keep in sync. The implementation detail that made this take longer than expected: the plan referenced a chi adapter for go-redoc that doesn't exist in any published version — the library only ships gin/fiber/echo adapters. The fix was to call &lt;code&gt;goredoc.Redoc{SpecPath: specURL}.Body()&lt;/code&gt; directly for pre-rendered HTML and wrap it in a plain &lt;code&gt;http.HandlerFunc&lt;/code&gt;. Small feature, but it's the direct answer to how those two bugs happened in the first place: the API surface wasn't something anyone could casually glance at.&lt;/p&gt;

&lt;h2&gt;
  
  
  GitOps: ordering discipline, then a second controller on purpose
&lt;/h2&gt;

&lt;p&gt;v1.4.1 brought Flux CD v2.3+ into &lt;code&gt;gitops/&lt;/code&gt;, structured as three layered Kustomizations — infrastructure, apps, flags — each declaring &lt;code&gt;dependsOn&lt;/code&gt; against the one before it. The infrastructure layer also sets &lt;code&gt;healthChecks&lt;/code&gt; against the tombstone-operator HelmRelease, which matters more than it sounds: &lt;code&gt;dependsOn&lt;/code&gt; alone only blocks &lt;em&gt;applying&lt;/em&gt; a Kustomization until the upstream has been applied — it does not wait for the upstream to actually be healthy. Without &lt;code&gt;healthChecks&lt;/code&gt;, Flux considers the operator's CRDs "done" the moment the HelmRelease manifest hits the API server, before the CRDs are actually established, and the apps layer — which deploys FeatureFlag CRs via the flagmind chart — could start reconciling against CRDs that technically exist but aren't ready yet. Pairing &lt;code&gt;dependsOn&lt;/code&gt; with &lt;code&gt;healthChecks&lt;/code&gt; on the upstream is what actually guarantees ordering — this is exactly the kind of GitOps bug that works fine in every test until the day the operator pod is slow to come up.&lt;/p&gt;

&lt;p&gt;The more interesting decision is v1.4.2's addition of Argo CD v2.11 — not as a Flux replacement, as a second controller with a deliberately split job. Flux keeps infrastructure: the operator's CRDs, image update automation across the tombstone container images. Argo CD takes over the flagmind chart plus the FeatureFlag and FlagPolicy custom resources themselves. The reason to split: FeatureFlag CRs have a &lt;code&gt;rolloutPct&lt;/code&gt; field that the intelligence service's ML rollout recommendations mutate live in the cluster, independent of Git. A naive GitOps setup would see that mutation as drift and revert it on the next reconcile — the platform's own ML-driven rollout logic fought and undone by its own deployment tooling every sync interval.&lt;/p&gt;

&lt;p&gt;The fix lives in &lt;code&gt;gitops/clusters/production/argocd/apps.yaml&lt;/code&gt;: an &lt;code&gt;ignoreDifferences&lt;/code&gt; block excluding &lt;code&gt;/spec/environments/production/rolloutPct&lt;/code&gt; and &lt;code&gt;/spec/environments/staging/rolloutPct&lt;/code&gt; from diff detection, paired with &lt;code&gt;RespectIgnoreDifferences=true&lt;/code&gt; in &lt;code&gt;syncOptions&lt;/code&gt;. Both are required — &lt;code&gt;ignoreDifferences&lt;/code&gt; alone only suppresses the OutOfSync &lt;em&gt;display&lt;/em&gt;; without the sync-options flag, Argo CD still overwrites the field back to the Git value on every sync. Miss either half and the ML rollout percentage gets silently stomped on a fixed interval, a nasty class of bug to chase down because nothing in the intelligence service's logs would look wrong.&lt;/p&gt;

&lt;p&gt;Argo CD also needed Lua health checks for Tombstone's own CRDs — FeatureFlag (Pending to Progressing, Synced to Healthy, Error to Degraded) and FlagPolicy (Compliant to Healthy, Violation to Degraded) — because its generic health rollup doesn't understand custom CRD status fields out of the box. Without them, the root Application would just show every FeatureFlag as permanently "Unknown" rather than reflecting real reconciliation state.&lt;/p&gt;

&lt;h2&gt;
  
  
  Blast radius, but as a deployment gate now
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpl74a5ooww9tktsqz44k.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpl74a5ooww9tktsqz44k.gif" alt="A digital fingerprint scan with a glowing blue circuit interface signifying security" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the part I like best, because it closes a loop back to the v1.0 launch: blast-radius scoring — the LOW/MEDIUM/HIGH/BLOCKED classification the evaluator computes for flag changes — now gates Kubernetes deployments, not just flag rollouts. v1.4.1 adds an Argo Rollouts &lt;code&gt;AnalysisTemplate&lt;/code&gt; named &lt;code&gt;tombstone-blast-radius&lt;/code&gt; that polls &lt;code&gt;GET /api/v1/blast-radius?flag_key=&amp;lt;key&amp;gt;&lt;/code&gt; on the evaluator service during a canary step. &lt;code&gt;gitops/apps/production/tombstone/rollout-analysis.yaml&lt;/code&gt; wires this into flag-api's own Rollout: step to 20% traffic, run the analysis template, promote to 100% only if the result is LOW or MEDIUM, abort immediately on HIGH or BLOCKED (&lt;code&gt;failureLimit: 1&lt;/code&gt;, so the first bad reading stops it, no averaging across a window). The same scoring engine that decides whether a flag change is safe to ship now also decides whether a code deployment is safe to ship — the same signal doing double duty at two layers of the stack.&lt;/p&gt;

&lt;p&gt;Argo CD Notifications routes sync-failed events into the existing marketplace Slack endpoint (&lt;code&gt;marketplace.tombstone.svc:8086/api/v1/marketplace/slack/actions&lt;/code&gt;) rather than standing up a second webhook — one less integration surface to maintain.&lt;/p&gt;

&lt;h2&gt;
  
  
  Removing the safety net
&lt;/h2&gt;

&lt;p&gt;Back to where I started. The &lt;code&gt;|| true&lt;/code&gt; removal in &lt;code&gt;.github/workflows/ci.yml&lt;/code&gt; (commit &lt;code&gt;9c9bff9&lt;/code&gt;) touched test steps across the Python intelligence service, the TypeScript SDK build, the Python SDK, the Ruby SDK, and the Java SDK — every one configured to report green regardless of outcome. The very next commit (&lt;code&gt;031d041&lt;/code&gt;) is titled, accurately, "fix pre-existing test failures surfaced by removing || true," and it fixes four of them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The Python SDK's &lt;code&gt;pytest&lt;/code&gt; step installed &lt;code&gt;mmh3&lt;/code&gt; and the package itself but never installed &lt;code&gt;pytest&lt;/code&gt; — the runner didn't exist in the environment, so &lt;code&gt;python -m pytest&lt;/code&gt; had presumably been failing at the shell level the whole time.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ruff check . --ignore F401&lt;/code&gt; in the intelligence service flagged unused local variables (&lt;code&gt;now_ts&lt;/code&gt;, &lt;code&gt;model&lt;/code&gt;) a real lint gate should have caught the moment they were written.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;packages/sdks/flagmind-ruby/lib/tombstone.rb&lt;/code&gt; didn't exist at all — spec files &lt;code&gt;require "tombstone"&lt;/code&gt; but the gem's actual entry point is &lt;code&gt;flagmind.rb&lt;/code&gt;, a leftover from the gem's rename. I added a one-line alias file.&lt;/li&gt;
&lt;li&gt;The Java SDK's step ran &lt;code&gt;./gradlew test&lt;/code&gt;, but no &lt;code&gt;gradlew&lt;/code&gt; wrapper is committed to the repo — the step was failing to even start.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fixing the Java one turned into its own small saga across four more commits, because each fix uncovered the next problem: swapping to &lt;code&gt;gradle/actions/setup-gradle&lt;/code&gt; failed because the pinned SHA wasn't resolvable; falling back to &lt;code&gt;apt-get install gradle&lt;/code&gt; got Gradle installed, but the Ubuntu-runner package is old enough that it chokes on &lt;code&gt;useJUnitPlatform()&lt;/code&gt;; installing Gradle 8.7 directly from &lt;code&gt;services.gradle.org&lt;/code&gt; fixed that but exposed that &lt;code&gt;build.gradle&lt;/code&gt; declared &lt;code&gt;sourceCompatibility = JavaVersion.VERSION_21&lt;/code&gt;, an enum literal the older Gradle parser doesn't accept, which needed to become the plain string &lt;code&gt;'21'&lt;/code&gt;. Even after all that, the Java tests still fail, but for a real reason: the source files declare &lt;code&gt;package io.tombstone.*&lt;/code&gt; while living under &lt;code&gt;io/flagmind/&lt;/code&gt; directories, the same rename leftover that broke the Ruby require. That's now &lt;code&gt;continue-on-error: true&lt;/code&gt; with a comment pointing at v1.5.0, not a silent &lt;code&gt;|| true&lt;/code&gt; — the failure is visible in the CI UI and tracked, instead of invisible and untracked.&lt;/p&gt;

&lt;p&gt;Four bugs weren't introduced in this release. They'd been there for a while, hiding under a shell operator that made "it ran" indistinguishable from "it passed." The same commit mirrors a fix v1.2.1 made earlier, adding &lt;code&gt;pytest-asyncio&lt;/code&gt; back for a similar reason — a test suite that can fail silently isn't testing the thing it claims to test, it's testing that the CI runner can execute a command.&lt;/p&gt;

&lt;h2&gt;
  
  
  Supply chain and the honest caveat
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4kaam89jlsyuf73dmi4f.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4kaam89jlsyuf73dmi4f.gif" alt="A rocket launching into the sky with a large plume of smoke and fire" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Alongside the CI honesty pass, all remaining GitHub Actions workflow files got pinned to immutable commit SHAs instead of mutable tags like &lt;code&gt;@v4&lt;/code&gt; — nine files in one commit (&lt;code&gt;bfb167f&lt;/code&gt;), closing out the supply-chain hardening that earlier workflows (&lt;code&gt;flux-bootstrap.yml&lt;/code&gt;) had already established as the convention. Two other release blockers landed in the same commit: an &lt;code&gt;ImageUpdateAutomation&lt;/code&gt; resource still on the wrong beta API version, and a structural fix to the production rollouts kustomization that was silently missing a resource reference.&lt;/p&gt;

&lt;p&gt;Everything above — the layered Flux Kustomizations, the Argo CD split, the blast-radius AnalysisTemplate — is validated against a local k3d test cluster: every kustomize build passes, Flux bootstraps cleanly, Argo CD installs and reconciles. What it is not yet validated against is the actual production target, Oracle Cloud Kubernetes, blocked on an Oracle Cloud account signup that hasn't happened yet. The operator Helm chart is already published to &lt;code&gt;ghcr.io/sairam0424/charts/tombstone-operator&lt;/code&gt; at v0.1.0, ready for the day the cluster exists. Better to say that plainly than let "GitOps shipped" imply more than k3d has actually proven.&lt;/p&gt;

&lt;p&gt;The throughline across v1.2 through v1.4 is the same lesson at three different altitudes. v1.2 was runtime resilience — the system surviving its own dependencies failing. v1.3 was correctness at the API and SDK boundary — two clients agreeing on what a flag evaluates to. v1.4 is deployment and CI honesty — the pipeline that ships the system telling the truth about its own state, in the right order, without silently eating failures. None of these layers were broken in an obvious way. They were all quietly returning something other than the truth, and the only way to find that out was to stop letting them.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>gitops</category>
      <category>kubernetes</category>
      <category>go</category>
    </item>
    <item>
      <title>trelix v1.0 to v2.7: When "It Works" Meets "It Scales"</title>
      <dc:creator>SAI RAM</dc:creator>
      <pubDate>Thu, 09 Jul 2026 19:39:19 +0000</pubDate>
      <link>https://dev.to/sai_ram_0000/trelix-v10-to-v27-when-it-works-meets-it-scales-2f25</link>
      <guid>https://dev.to/sai_ram_0000/trelix-v10-to-v27-when-it-works-meets-it-scales-2f25</guid>
      <description>&lt;p&gt;Twelve days after I shipped trelix v1.0.0, I was staring at a &lt;code&gt;RetrievalConfig&lt;/code&gt; object with two conflicting sets of values and no idea which one was actually running. I'd built &lt;code&gt;AdaptiveRouter&lt;/code&gt; to accept a &lt;code&gt;retrieval_config&lt;/code&gt; parameter so callers could override the environment-variable defaults programmatically. Except it didn't. The constructor took the parameter, and then quietly ignored it and built its own instance from env vars anyway. Nobody had wired the plumbing from &lt;code&gt;Retriever&lt;/code&gt; through &lt;code&gt;QueryPlanner&lt;/code&gt; down to &lt;code&gt;AdaptiveRouter.__init__&lt;/code&gt;. It's the kind of bug that doesn't throw — it just makes your carefully-set config a decoy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqvnxct8ywk7qlhibrsxy.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqvnxct8ywk7qlhibrsxy.gif" alt="A computer monitor displaying a Blue Screen of Death (BSOD) error" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That fix landed in v2.7.0, PR #55, thirteen days and seven minor releases after launch. In between, trelix went from "search my repo well" to something closer to a platform: a knowledge graph, seven fused retrieval legs, an agentic loop, federated multi-repo search, and a GitHub Actions bot that reviews your PRs. Here's what actually shipped, grouped by what it was trying to solve rather than by version number.&lt;/p&gt;

&lt;h2&gt;
  
  
  From flat search to a knowledge graph
&lt;/h2&gt;

&lt;p&gt;v1.0 already had hybrid BM25 + vector + call-graph search. What it didn't have was any notion of the codebase as a &lt;em&gt;system&lt;/em&gt; — which files cluster into modules, which symbols sit at the center of the import graph, which concepts a human would use to describe an architecture. v2.0.0 (2026-06-28) and v2.1.0 (2026-06-30) fixed that.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5rmeaehdafewf2xryq7s.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5rmeaehdafewf2xryq7s.gif" alt="An animation showing data flowing through a network of connected nodes" width="500" height="281"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The new &lt;code&gt;trelix/graph/&lt;/code&gt; module builds a &lt;code&gt;CodeGraph&lt;/code&gt; as a NetworkX &lt;code&gt;MultiDiGraph&lt;/code&gt;, unifying call, import, and type edges into one traversable structure. On top of that, Louvain community detection clusters the graph into architectural modules — run &lt;code&gt;trelix graph ./repo&lt;/code&gt; and you get the top communities, not just a flat symbol list. &lt;code&gt;ConceptExtractor&lt;/code&gt; layers an LLM on top of symbol batches to name those communities in plain English, and it's built to fail quietly: any extraction error returns &lt;code&gt;[]&lt;/code&gt; rather than crashing the pipeline. &lt;code&gt;GraphVisualizer.export_html()&lt;/code&gt; renders the whole thing as an interactive Pyvis HTML page with community coloring, gated behind &lt;code&gt;pip install trelix[knowledge-graph]&lt;/code&gt; so the base install doesn't inherit the dependency weight.&lt;/p&gt;

&lt;p&gt;Graph search became a first-class retrieval leg — &lt;code&gt;graph_search_enabled=True&lt;/code&gt; runs a CodeGraph BFS as a fourth leg after RRF fusion — and &lt;code&gt;pagerank_boost_enabled&lt;/code&gt; uses import-graph centrality to boost symbols that sit at architectural chokepoints. None of this is static: &lt;code&gt;GraphUpdater.update_file()&lt;/code&gt; is wired into &lt;code&gt;trelix watch&lt;/code&gt;, so the graph and its communities update incrementally as files change, instead of requiring a full rebuild.&lt;/p&gt;

&lt;p&gt;This came with the release's one deliberate breaking change: &lt;code&gt;trelix graph&lt;/code&gt; — which used to mean "show me callers and callees of this symbol" — got renamed to &lt;code&gt;trelix call-graph&lt;/code&gt;. The name &lt;code&gt;trelix graph&lt;/code&gt; now means "build the knowledge graph." I made the call that a growing surface area needed the more intuitive name reserved for the bigger feature, and documented the rename explicitly in the changelog rather than let people discover it by trial and error.&lt;/p&gt;

&lt;h2&gt;
  
  
  Seven legs, one fusion function
&lt;/h2&gt;

&lt;p&gt;v1.0 had three retrieval legs. By v2.2.0 it had seven, and every one of them is grounded in a specific paper rather than a hunch.&lt;/p&gt;

&lt;p&gt;The fourth leg was the graph BFS above. The fifth is RAPTOR-style (arXiv:2401.18059) file-level summarization — &lt;code&gt;file_summary_leg_enabled&lt;/code&gt;, gated behind &lt;code&gt;TRELIX_FILE_SUMMARIES_ENABLED=true&lt;/code&gt; at index time — which lets trelix answer "explain this codebase" questions that no symbol-level chunk could answer alone. The sixth is HyDE (arXiv:2212.10496): instead of embedding your raw natural-language query, &lt;code&gt;hyde_fallback_enabled&lt;/code&gt; generates a synthetic code snippet and embeds &lt;em&gt;that&lt;/em&gt;, closing the semantic gap between "how do I validate a JWT" and the actual token-validation code. The seventh is multi-query expansion, which decomposes one query into N variants and RRF-fuses the independent retrievals for broader recall.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fifls44urcetin2dlrx0o.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fifls44urcetin2dlrx0o.gif" alt="A hand placing the final piece into a jigsaw puzzle, completing the picture" width="200" height="200"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Layered over all seven is FLARE (arXiv:2305.06983) — a confidence-gated re-retrieval loop that watches synthesis output for uncertainty phrases and triggers another retrieval pass when it finds them, rather than committing to a possibly-wrong first answer.&lt;/p&gt;

&lt;p&gt;None of this is worth shipping without a way to measure whether it's actually better, so v2.1.0 added a CoIR-format eval harness (ACL 2025, arXiv:2407.02883) — &lt;code&gt;trelix eval --golden &amp;lt;file&amp;gt;&lt;/code&gt; reports nDCG@10, Recall@10, and MRR, implemented as pure-Python &lt;code&gt;trelix.eval.ndcg&lt;/code&gt; with zero pandas dependency. Every query now also writes a row to a &lt;code&gt;query_telemetry&lt;/code&gt; SQLite table — latency, intent classification, result count — surfaced through &lt;code&gt;trelix telemetry&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Here's the detail that actually matters: in v2.1.0, &lt;code&gt;MultiQueryExpander&lt;/code&gt; existed as a class but nothing called it. It took until v2.3.0 (2026-07-02) for it to get wired into &lt;code&gt;_retrieve_standard&lt;/code&gt;, and even then I had to be careful about one specific line — &lt;code&gt;variants[1:]&lt;/code&gt; is used, not &lt;code&gt;variants[:]&lt;/code&gt;, so the original query never runs twice through the fusion. It's a one-character difference between "seven legs" and "seven legs, one of them redundant."&lt;/p&gt;

&lt;h2&gt;
  
  
  Teaching it to act, not just retrieve
&lt;/h2&gt;

&lt;p&gt;v2.2.0 (2026-07-01) shipped across four parallel feature branches — PRs #29 through #32, merged via release PR #33 — and it's the release where trelix stopped being purely a retrieval system.&lt;/p&gt;

&lt;p&gt;The agentic loop (&lt;code&gt;trelix/agent/&lt;/code&gt;) is CodeAct-style ReAct: instead of one retrieval-then-synthesize pass, the agent can decide it needs another lookup, run it, and fold the result back into its reasoning before answering. Alongside it, &lt;code&gt;trelix/analysis/taint.py&lt;/code&gt; and &lt;code&gt;defuse.py&lt;/code&gt; added real data-flow and taint analysis — tracing how a value flows from a source to a sink across function boundaries, which is a different kind of question than "what code is semantically similar to this query."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feb37scna97luwkhcieb6.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feb37scna97luwkhcieb6.gif" alt="A 3D animation of a chess piece moving on a digital board with neural network-style connections appearing above it" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The other two branches were retrieval-quality work: SPLADE-Code sparse retrieval (&lt;code&gt;trelix/embedder/sparse.py&lt;/code&gt;, &lt;code&gt;trelix/store/sparse_store.py&lt;/code&gt;) gives trelix a learned sparse representation to sit alongside BM25 and dense vectors, and multi-granularity indexing (&lt;code&gt;trelix/indexing/multi_granularity.py&lt;/code&gt;, MGS3-style) means the index isn't forced to choose one chunk size — function-level, class-level, and file-level granularities can all be retrieved against.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production hardening: guards built before the bugs happened
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyieb58ecklpf5tfw017y.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyieb58ecklpf5tfw017y.gif" alt="A vintage computer screen showing a 'Net Busy' error message with a 'Will call later' button" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The most interesting engineering in this window isn't a feature — it's a class of failure I preempted instead of debugging in production. &lt;code&gt;DimensionGuard&lt;/code&gt;, added at &lt;code&gt;Retriever.__init__&lt;/code&gt; in v2.3.0, checks embedding provider and dimension at startup and raises &lt;code&gt;DimensionMismatchError&lt;/code&gt; with the exact recovery command (&lt;code&gt;trelix migrate-vectors --reset&lt;/code&gt;) if they don't match what's on disk. Without it, switching from an Azure embedder (3072-dim) to a local model (384-dim) doesn't error — it silently returns wrong results, because cosine similarity between mismatched-dimension vectors still computes &lt;em&gt;a&lt;/em&gt; number, just not a meaningful one. That's the worst kind of bug: no stack trace, no crash, just quietly bad answers. v2.5.0 (2026-07-06) extended the same guard to &lt;code&gt;FileWatcher.__init__&lt;/code&gt;, so a provider mismatch fails fast at watch startup instead of at query time, days later, when nobody remembers which embedder was configured when.&lt;/p&gt;

&lt;p&gt;The MCP surface grew up in the same window. v2.3.0 added MCP Resources (&lt;code&gt;trelix://index/stats&lt;/code&gt;, &lt;code&gt;trelix://repo/{path}/manifest&lt;/code&gt;, &lt;code&gt;trelix://repo/{path}/symbols/{name}&lt;/code&gt;) and MCP Prompts (&lt;code&gt;trelix-search&lt;/code&gt;, &lt;code&gt;trelix-explain&lt;/code&gt;, &lt;code&gt;trelix-blast-radius&lt;/code&gt;) — reusable, application-addressable primitives instead of one-off tool calls. v2.5.0 went further: &lt;code&gt;trelix-mcp&lt;/code&gt; now advertises &lt;code&gt;resources.subscribe=True&lt;/code&gt;, and a thread-safe &lt;code&gt;SubscriptionRegistry&lt;/code&gt; tracks who's watching which URI, so &lt;code&gt;notify_file_changed()&lt;/code&gt; can fire &lt;code&gt;notifications/resources/updated&lt;/code&gt; the moment &lt;code&gt;watchfiles&lt;/code&gt; detects a change. That notification path had a gap of its own until v2.7.0 Phase 1 (PR #55): &lt;code&gt;FileWatcher._do_reindex&lt;/code&gt; only fired the notification on hash-identical skips, never on an actual successful re-index — the one case where a subscriber genuinely needed to know. The same release added &lt;code&gt;idx_files_rel_path&lt;/code&gt; as an index on &lt;code&gt;files.rel_path&lt;/code&gt;, eliminating a full table scan that &lt;code&gt;GraphUpdater.update_file()&lt;/code&gt; was silently paying on every single file-change event.&lt;/p&gt;

&lt;h2&gt;
  
  
  Federation, and finding your code's twin
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;DiffReviewer&lt;/code&gt; and &lt;code&gt;trelix review &amp;lt;repo&amp;gt; [--diff] [--base] [--head]&lt;/code&gt; (v2.3.0) turned trelix into something you point at a diff, not just a repo — it parses git diffs via &lt;code&gt;DiffParser.from_git()&lt;/code&gt; (subprocess with &lt;code&gt;shell=False&lt;/code&gt;, no injection surface), turns each hunk into a retrieval query, and generates review comments that are crash-safe by construction: &lt;code&gt;DiffReviewer.review()&lt;/code&gt; never raises. v2.4.0 (2026-07-04) connected that to GitHub directly — &lt;code&gt;GitHubPRClient&lt;/code&gt; plus &lt;code&gt;trelix review --pr owner/repo#N --post-comments&lt;/code&gt;, authenticating only via &lt;code&gt;GITHUB_TOKEN&lt;/code&gt;, handling all seven GitHub file-status values, and warning past a 3,000-file truncation limit. v2.7.0 Phase 3 (PR #57) closed the loop with &lt;code&gt;.github/workflows/trelix-review.yml&lt;/code&gt;, which runs that same review command on every PR and posts findings as GitHub Check annotations with file and line references — &lt;code&gt;continue-on-error: true&lt;/code&gt; on the indexing step, because CI runners without local embedding models shouldn't fail the whole workflow. The same phase shipped &lt;code&gt;workspace-vscode/&lt;/code&gt;, a VS Code extension scaffold with &lt;code&gt;trelix.search&lt;/code&gt; and &lt;code&gt;trelix.ask&lt;/code&gt; commands, talking to the existing &lt;code&gt;trelix-mcp&lt;/code&gt; package over stdio — no new backend, just a new front door.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbmo8dv72wkm8vkq2v99x.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbmo8dv72wkm8vkq2v99x.gif" alt="Bugs Bunny from Looney Tunes looking around frantically through a magnifying glass" width="600" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The other axis of growth was going multi-repo. &lt;code&gt;RepoRegistry&lt;/code&gt; (v2.3.0) manages &lt;code&gt;~/.config/trelix/repos.json&lt;/code&gt;, and &lt;code&gt;FederatedRetriever&lt;/code&gt; fans a query out across every registered repo in parallel, RRF-merges the results, and dedupes by &lt;code&gt;(file_path, symbol_id)&lt;/code&gt; — crash-safe, returning &lt;code&gt;[]&lt;/code&gt; if every repo fails rather than propagating one bad repo's exception. v2.4.0 added a SHA-256-keyed TTL cache (&lt;code&gt;cache_ttl=120.0&lt;/code&gt;) tuned for the query patterns of an actual debugging session, where you ask variations of the same question five times in ten minutes. v2.7.0 Phase 2 (PR #56) pushed federation further with &lt;code&gt;make_scip_symbol_id()&lt;/code&gt; — stable, SCIP-style cross-repo symbol IDs, sha256-truncated and pipe-separated so scoped npm packages like &lt;code&gt;@scope/pkg&lt;/code&gt; resolve unambiguously — and &lt;code&gt;DiffEmbedder&lt;/code&gt;, a CCRep-style (arXiv:2302.03924) before/after body-pair encoder for PR diff hunks. &lt;code&gt;search_similar_diffs()&lt;/code&gt; finds historically similar changes via cosine similarity, with a NaN guard and dimension-mismatch protection baked in from day one, because I'd already been burned once by silent dimension corruption.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scale work, and holding the frontier to a real bar
&lt;/h2&gt;

&lt;p&gt;v2.6.0 (2026-07-08) tackled the two things that get expensive as a codebase grows: recomputing the whole community graph on every file save, and blocking on a full index pass before you can search anything. The DF Louvain frontier heuristic (&lt;code&gt;compute_affected_frontier()&lt;/code&gt;, &lt;code&gt;detect_communities_incremental()&lt;/code&gt;, arXiv:2404.19634) reprocesses only the seed nodes, their neighbors, and their existing community members — falling back to a full recompute only when the affected frontier exceeds 50% of the graph. &lt;code&gt;TRELIX_INDEXER_STREAMING=true&lt;/code&gt; (v2.7.0 Phase 2) makes indexing itself lazy — &lt;code&gt;_iter_files()&lt;/code&gt; yields files one at a time into a bounded &lt;code&gt;Queue(maxsize=64)&lt;/code&gt;, with a &lt;code&gt;try/finally&lt;/code&gt; guarantee that the producer sentinel always gets sent even on an exception. It's off by default, and I mean that literally: zero behavior change on the path everyone is actually running.&lt;/p&gt;

&lt;p&gt;I also shipped two things I'm not willing to oversell. The XTR late-interaction reranker (NeurIPS 2023, arXiv:2304.01982) is cheaper than ColBERT/PLAID by reusing tokens you already retrieved instead of reloading every document's full token set — that's a genuinely good idea. But it's explicitly marked EXPERIMENTAL in the changelog, it emits a &lt;code&gt;UserWarning&lt;/code&gt; on first use, and it has not been benchmarked against CoIR or CoREB on code-specific retrieval. PLAID stays the production-validated default. Same discipline applies to the GroUSE-inspired synthesis harness (arXiv:2409.06595, COLING 2025) — &lt;code&gt;SynthesisEvalHarness&lt;/code&gt; scores hallucination, completeness, and faithfulness across seven failure modes, because I'd been leaning on "does GPT-4 think this answer sounds right" as an implicit quality bar, and that correlation is not a substitute for actually checking whether the citations are real.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm5rrvvufkfqneikvnadl.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm5rrvvufkfqneikvnadl.gif" alt="A chef carefully plating a dish, representing the fusion of different elements" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Test count tells the same story as the changelog: 929 unit tests at the v1.0.0 baseline, 1,467 unit plus 41 MCP tests — 1,508 total — by v2.7.0. &lt;code&gt;pip install "trelix[local]"&lt;/code&gt; still gets you a fully offline setup, and the index is still one SQLite file. Growing the surface area from three retrieval legs to seven, plus a knowledge graph, an agentic loop, and federation, didn't require growing the infrastructure footprint at all — every new leg, every graph feature, every federation layer is opt-in behind a config flag that defaults to off. That was a deliberate constraint, not an accident, and it's the one I'm least willing to relax as this keeps growing.&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>aiengineering</category>
      <category>python</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I Built trelix Because I Was Tired of Grepping My Way Through Codebases</title>
      <dc:creator>SAI RAM</dc:creator>
      <pubDate>Sun, 05 Jul 2026 11:38:03 +0000</pubDate>
      <link>https://dev.to/sai_ram_0000/i-built-trelix-because-i-was-tired-of-grepping-my-way-through-codebases-1f3b</link>
      <guid>https://dev.to/sai_ram_0000/i-built-trelix-because-i-was-tired-of-grepping-my-way-through-codebases-1f3b</guid>
      <description>&lt;p&gt;I spent my most of day's on a new team grepping through 80,000 lines of code trying to find where authentication worked.&lt;/p&gt;

&lt;p&gt;Four hours. Three teammates interrupted. Twelve dead ends across files I didn't understand. The code was fine — it was well-written, well-organized, reasonably documented. The tooling was the problem. I was using grep to understand something that wasn't a text search problem. Code has structure: call edges, import chains, type hierarchies, AST relationships. Grep ignores all of it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F38aseby7mu7x6xcg4bb3.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F38aseby7mu7x6xcg4bb3.gif" alt="A person typing quickly on a computer with a tech-focused background" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That day stuck with me. I kept running into the same pattern on different teams, different codebases, different languages. Every time I joined something new or came back to a project after six months away, the first few days were archaeology. Tracing calls manually. Reconstructing context that should have been queryable.&lt;/p&gt;

&lt;p&gt;I built trelix to fix this. It's an open-source code intelligence engine that indexes any repository with Tree-sitter, embeds every symbol, and answers natural-language questions using hybrid BM25 + vector + call-graph search. It works offline. No API key needed. Zero infrastructure.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="s2"&gt;"trelix[local]"&lt;/span&gt;
trelix index ./my-repo
trelix ask ./my-repo &lt;span class="s2"&gt;"how does the authentication middleware work?"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The Problem With Code Search
&lt;/h2&gt;

&lt;p&gt;The tools we have for understanding code — editors, grep, ctags, language servers — were designed for writing code, not for understanding it at scale. They're excellent at navigating to a known destination. They're poor at answering questions like "how does the request lifecycle work end-to-end?" or "what calls this function, and what does that caller depend on?" when you don't already know the answer.&lt;/p&gt;

&lt;p&gt;The fundamental limitation of grep is that it treats your codebase as a document corpus. It finds strings. Code isn't a document corpus — it's a graph. Functions call other functions. Modules import other modules. Classes extend other classes. When you ask "how does authentication work?", the answer isn't a file or even a few files. It's a traversal of that graph, starting from a semantic entry point and following edges to collect the relevant context.&lt;/p&gt;

&lt;p&gt;Vector search solves part of this — semantic similarity gets you closer to the right files without knowing the exact tokens. But pure vector search misses structural relationships. It doesn't know that &lt;code&gt;UserRepository.get_by_token()&lt;/code&gt; is always called by &lt;code&gt;AuthMiddleware.verify()&lt;/code&gt; which is called by every protected route handler. That's call-graph knowledge, not embedding knowledge.&lt;/p&gt;

&lt;p&gt;trelix uses both.&lt;/p&gt;

&lt;h2&gt;
  
  
  What trelix Does
&lt;/h2&gt;

&lt;p&gt;trelix indexes any repository into a single SQLite file (&lt;code&gt;.trelix/index.db&lt;/code&gt;) and then answers questions about it.&lt;/p&gt;

&lt;p&gt;The index contains: every symbol extracted via Tree-sitter (functions, classes, methods, their bodies and line spans); call edges and import edges between symbols and files; a hybrid search index combining sqlite-vec HNSW vectors with FTS5 BM25; and since v2.1.0, a Code Property Graph that unifies all of the above into a traversable NetworkX graph.&lt;/p&gt;

&lt;p&gt;A query like &lt;code&gt;trelix ask ./repo "explain how authentication works"&lt;/code&gt; goes through a 3-tier adaptive router:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tier 1 (Direct)&lt;/strong&gt; — for simple factual patterns like "what is X" or "define X", trelix skips retrieval entirely and answers from the LLM directly. No unnecessary round-trips.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tier 2 (8-intent)&lt;/strong&gt; — for most code queries, it classifies the intent into one of eight categories (symbol_lookup, feature_flow, dependency_map, blast_radius, etc.) and runs the appropriate retrieval strategy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tier 3 (Multi-step)&lt;/strong&gt; — for complex queries like "walk me through the request lifecycle end-to-end", it decomposes the question into 2-3 sub-queries, runs each independently, and merges the results.&lt;/p&gt;

&lt;p&gt;Results from all active retrieval legs are fused via Reciprocal Rank Fusion (k=60) before being assembled into the context window for LLM synthesis.&lt;/p&gt;

&lt;h2&gt;
  
  
  How It Actually Works
&lt;/h2&gt;

&lt;p&gt;The indexing pipeline runs in four phases:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 1 (Parse)&lt;/strong&gt; — Tree-sitter walks every file and extracts symbols with their source, line spans, and AST structure. Runs in parallel via ThreadPoolExecutor.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 2 (Write)&lt;/strong&gt; — Symbols and chunks are written to SQLite. Cross-file parent_id relationships are resolved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 3 (Embed)&lt;/strong&gt; — Every chunk is embedded asynchronously in batches of 4 concurrent API calls. With the local provider (sentence-transformers, no API key), this runs entirely offline.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 4 (Resolve)&lt;/strong&gt; — Cross-file call edges are resolved with a 3-priority strategy: qualified name first, then type_hint+name, then name-only fallback. This gives about 40% fewer false-positive cross-file edges compared to name-only matching.&lt;/p&gt;

&lt;p&gt;The result is a single &lt;code&gt;.trelix/index.db&lt;/code&gt; file that contains everything: vectors, BM25, call graph, import graph, symbols, file hashes for incremental updates.&lt;/p&gt;

&lt;h2&gt;
  
  
  Zero Infrastructure, Full Power
&lt;/h2&gt;

&lt;p&gt;This was a deliberate design decision and one I keep coming back to.&lt;/p&gt;

&lt;p&gt;Most code intelligence tools require running a vector database, a relational database, and often a separate API server. That's a lot of infrastructure to maintain for what is fundamentally a local developer tool. trelix's default is a single SQLite file using sqlite-vec for HNSW vector search and FTS5 for BM25. Zero external infrastructure. Works on a laptop with no internet connection.&lt;/p&gt;

&lt;p&gt;When you need to scale: LanceDB backend for 100k+ chunks (3-5× faster vector insert on ARM/Apple Silicon), Qdrant for 500k+ chunk deployments with multi-repo shared collections. But the default handles most codebases and most developers will never need to switch.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Default (sqlite) — up to ~100k chunks&lt;/span&gt;
trelix index ./my-repo

&lt;span class="c"&gt;# LanceDB — 100k+ chunks&lt;/span&gt;
&lt;span class="nv"&gt;TRELIX_STORE_BACKEND&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;lance trelix index ./my-repo

&lt;span class="c"&gt;# Qdrant — 500k+ chunks&lt;/span&gt;
&lt;span class="nv"&gt;TRELIX_STORE_BACKEND&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;qdrant trelix index ./my-repo
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw2ek3kv4zjil45o10uag.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw2ek3kv4zjil45o10uag.gif" alt="A character entering beast mode with a rage effect" width="480" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Beast Mode: Seven Retrieval Legs
&lt;/h2&gt;

&lt;p&gt;The default setup (BM25 + vector + grep + call graph) handles most questions well. But trelix has five additional retrieval legs that you can enable when you need higher recall or more sophisticated query handling:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Leg 5: File-summary semantic search&lt;/strong&gt; — RAPTOR-style (arXiv:2401.18059). At index time, trelix generates LLM summaries of every file and embeds those summaries separately. This surface is especially good for "explain this codebase" or "what files deal with payment processing?" queries — questions where the answer is at the file level, not the symbol level.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Leg 6: SPLADE-Code&lt;/strong&gt; — sparse+dense hybrid via learned sparse retrieval. SPLADE encodes queries into sparse high-dimensional token vectors, expanding vocabulary beyond exact matches in a way that complements both BM25 and dense vector search.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Leg 7: Multi-granularity&lt;/strong&gt; — indexes code at block AND statement level simultaneously. Some queries are better answered by a full function body; others are better answered by a single statement. Having both granularities in the index improves recall on precise questions.&lt;/p&gt;

&lt;p&gt;Plus query-side enhancements: &lt;strong&gt;HyDE&lt;/strong&gt; (generates a hypothetical code answer as the ANN query vector, improving recall on abstract questions), &lt;strong&gt;FLARE&lt;/strong&gt; (confidence-gated re-retrieval — when synthesis spans show uncertainty, trelix re-queries before finalizing the answer), and since v2.2.0, an &lt;strong&gt;agentic ReAct loop&lt;/strong&gt; that does multi-turn retrieve→observe→re-retrieve with self-correction.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Enable everything&lt;/span&gt;
&lt;span class="nv"&gt;TRELIX_RETRIEVAL_AGENTIC&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nv"&gt;TRELIX_GRAPH_SEARCH_ENABLED&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nv"&gt;TRELIX_RETRIEVAL_FILE_SUMMARY_LEG&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nv"&gt;TRELIX_RETRIEVAL_HYDE_FALLBACK&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nv"&gt;TRELIX_RETRIEVAL_FLARE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nv"&gt;TRELIX_RETRIEVAL_SPARSE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="nv"&gt;TRELIX_CHUNKER_MULTI_GRANULARITY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
trelix ask ./my-repo &lt;span class="s2"&gt;"explain the full request lifecycle"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The Features I'm Most Proud Of
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;GitHub PR review.&lt;/strong&gt; This is v2.4.0 and it's become one of the most-used features. &lt;code&gt;trelix review --pr owner/repo#42&lt;/code&gt; fetches the PR diff from GitHub, retrieves codebase context for each changed hunk, runs an LLM review, and can post findings back as a single batched review comment with &lt;code&gt;--post-comments&lt;/code&gt;. The key insight is that reviewing a diff without understanding the surrounding codebase is like proofreading a sentence you've never read before.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffp24hkwflpmc5gl9sz5b.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffp24hkwflpmc5gl9sz5b.gif" alt="A Pudgy Penguin nodding in approval with an " width="480" height="480"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;trelix review &lt;span class="nt"&gt;--pr&lt;/span&gt; sairam0424/trelix#42
trelix review &lt;span class="nt"&gt;--pr&lt;/span&gt; sairam0424/trelix#42 &lt;span class="nt"&gt;--post-comments&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Federated search.&lt;/strong&gt; &lt;code&gt;trelix search-all "query"&lt;/code&gt; fans out across all registered repos in parallel via ThreadPoolExecutor and RRF-merges the results. With &lt;code&gt;trelix watch-all&lt;/code&gt;, a single &lt;code&gt;watchfiles.awatch()&lt;/code&gt; call watches all registered repos simultaneously. The TTL cache on FederatedRetriever gives about 90% hit rate for typical debugging-session query patterns.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;trelix federation add api ./services/api
trelix federation add web ./services/web
trelix search-all &lt;span class="s2"&gt;"JWT validation"&lt;/span&gt;
trelix watch-all
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;MCP integration.&lt;/strong&gt; One command and trelix is available inside Claude Code, Cursor, Windsurf, and Continue.dev:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;trelix-mcp
claude mcp add trelix &lt;span class="nt"&gt;--&lt;/span&gt; trelix-mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then inside Claude Code: &lt;em&gt;"index my repo at /path/to/repo, then find how authentication works"&lt;/em&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Surprised Me Building This
&lt;/h2&gt;

&lt;p&gt;I expected the hardest part to be the embedding and retrieval architecture. It wasn't. The hardest part was making the system opinionated enough to be useful without being so opinionated that it broke on unusual codebases.&lt;/p&gt;

&lt;p&gt;The call-graph resolver was the most representative example. My first version used name-only matching for cross-file call edges — &lt;code&gt;login()&lt;/code&gt; in file A calls &lt;code&gt;login()&lt;/code&gt; in file B. This produced a dense, noisy graph with maybe 40% false-positive edges. The fix was a 3-priority resolution strategy: try qualified name first (most precise, lowest recall), then type hint + name (moderate precision), then name-only as fallback. That reduced false positives significantly while maintaining recall on codebases that don't have full type annotations.&lt;/p&gt;

&lt;p&gt;The other thing that surprised me was how much value came from the structural metadata rather than the semantic embeddings. The call graph, import graph, and type hierarchy are what make trelix's answers qualitatively different from a vector search over code files. Semantic similarity gets you to the right neighborhood. Graph traversal gets you to the right answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm Still Uncertain About
&lt;/h2&gt;

&lt;p&gt;The 3-tier query router works well for the queries I've tested it on. I'm less confident about it on very large codebases (millions of lines) where the graph becomes expensive to traverse. The current implementation caps BFS depth at 2, which is usually right but occasionally misses important connections. I'm still figuring out the right heuristics for adaptive depth.&lt;/p&gt;

&lt;p&gt;I'm also still calibrating the GraphRAG map-reduce threshold. The current default (activate at &amp;gt;20 results or &amp;gt;8k tokens) is conservative. For some query types it activates too eagerly; for others, not eagerly enough. This is the main retrieval parameter I'm watching in practice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try It
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Offline — no API key&lt;/span&gt;
pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="s2"&gt;"trelix[local]"&lt;/span&gt;
trelix index ./your-repo
trelix ask ./your-repo &lt;span class="s2"&gt;"how does your main feature work?"&lt;/span&gt;

&lt;span class="c"&gt;# With LLM synthesis&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;trelix
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;sk-...
trelix ask ./your-repo &lt;span class="s2"&gt;"explain the request lifecycle end-to-end"&lt;/span&gt;

&lt;span class="c"&gt;# MCP in Claude Code&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;trelix-mcp
claude mcp add trelix &lt;span class="nt"&gt;--&lt;/span&gt; trelix-mcp

&lt;span class="c"&gt;# Review a PR&lt;/span&gt;
trelix review &lt;span class="nt"&gt;--pr&lt;/span&gt; owner/repo#42 &lt;span class="nt"&gt;--post-comments&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Everything is MIT licensed, on PyPI, and at &lt;a href="https://github.com/sairam0424/trelix" rel="noopener noreferrer"&gt;github.com/sairam0424/trelix&lt;/a&gt;. The full documentation is in the repo README including the beast-mode activation block if you want all seven retrieval legs at once.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq72nv8bao0jzl1kbxyv0.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq72nv8bao0jzl1kbxyv0.gif" alt="An animation of a developer typing quickly with a glowing keyboard" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What's the longest you've spent trying to understand a piece of code you didn't write? I've had 4-hour archaeology sessions on codebases with good documentation. I'd like to know how much of that time you think was the code being genuinely complex versus the tooling failing you.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs6ytm8vx4fz3wv4zk32c.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs6ytm8vx4fz3wv4zk32c.gif" alt="A nostalgic Geocities-style folder icon labeled Digital Archaeology" width="480" height="270"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>python</category>
      <category>opensource</category>
      <category>ai</category>
      <category>codesearch</category>
    </item>
  </channel>
</rss>
