<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Reno Lu</title>
    <description>The latest articles on DEV Community by Reno Lu (@renolu).</description>
    <link>https://dev.to/renolu</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3961766%2Fe973474b-a6f6-45ab-a944-e0495fc3346e.png</url>
      <title>DEV Community: Reno Lu</title>
      <link>https://dev.to/renolu</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/renolu"/>
    <language>en</language>
    <item>
      <title>Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default</title>
      <dc:creator>Reno Lu</dc:creator>
      <pubDate>Fri, 02 Oct 2026 18:07:20 +0000</pubDate>
      <link>https://dev.to/renolu/qwen38-27b-on-one-rtx-3090-and-the-setting-its-readme-says-chat-clients-should-leave-at-default-5de9</link>
      <guid>https://dev.to/renolu/qwen38-27b-on-one-rtx-3090-and-the-setting-its-readme-says-chat-clients-should-leave-at-default-5de9</guid>
      <description>&lt;p&gt;The README for syv-ai's qwen38-27b-rtx3090 puts a throughput table in its quick start. My reading is that the useful part sits one level down: it measures when &lt;code&gt;DFLASH_TOKENS=15&lt;/code&gt; is useful and what that setting costs in request slots and context.&lt;/p&gt;

&lt;p&gt;The repo is a serving setup for Qwen3.8-27B on a single 24 GB consumer GPU with vLLM, with an OpenAI-compatible API and two ready-made modes. The mode-comparison figures are the project's own, measured with &lt;code&gt;vllm bench serve&lt;/code&gt; on an RTX 3090 at a 250 W power limit. A version note says the branch pins vLLM 0.28.0 and that the throughput and quality tables are retained as reference baselines while the v0.28.0 GPU matrix is being re-measured.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two modes that share one card
&lt;/h2&gt;

&lt;p&gt;The prebuilt image lives on ghcr.io. The README says the build applies all &lt;code&gt;patches/&lt;/code&gt; and runs &lt;code&gt;verify.sh&lt;/code&gt; as its gate. The first start pulls 9.5 GB, downloads and requantizes the model (about 20 GB, once, into &lt;code&gt;./models&lt;/code&gt;), and serves on port 18020. You pick a mode with &lt;code&gt;docker compose --profile single up -d&lt;/code&gt; or &lt;code&gt;docker compose --profile batch up -d&lt;/code&gt;, and one GPU serves one at a time.&lt;/p&gt;

&lt;p&gt;Batch mode targets API backends and pipelines. At 64 concurrent requests with 128 tokens in and 512 out, the README reports about 1,035 tok/s steady-state decode and 948 end-to-end, or about 1,222 and 1,042 with all layers int8. A single stream in that mode decodes at 46 tok/s.&lt;/p&gt;

&lt;p&gt;Single-user mode flips the trade. With MTP speculation it reports 121 tok/s at default sampling and 120 greedy (&lt;code&gt;CTX=fast&lt;/code&gt;, 64k context), dropping to 96 and 102 with &lt;code&gt;CTX=long&lt;/code&gt; at 150k. With &lt;code&gt;SPEC=dflash2&lt;/code&gt; those become 127 default and 130 greedy. The README says speculation wins below about 8 concurrent users on short prompts and plain batching wins above that, and that the crossover comes much earlier on long independent sessions because a speculating request reserves recurrent-state pages the pool has few of.&lt;/p&gt;

&lt;p&gt;The benchmark footnote spells out the method. Single-stream numbers were re-measured on 2026-08-22 with &lt;code&gt;bash bench/run_benchmarks.sh single&lt;/code&gt;, eight prompts from &lt;code&gt;bench/prompts_real.jsonl&lt;/code&gt;, 1024 output tokens, concurrency 1, and decode rate taken as &lt;code&gt;C / mean TPOT&lt;/code&gt;. The README warns that a client using a different output length is not measuring the same thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The knob the README tells chat clients to skip
&lt;/h2&gt;

&lt;p&gt;For a lone user, the README recommends &lt;code&gt;SPEC=dflash2&lt;/code&gt; and &lt;code&gt;PREFIX_CACHE=1&lt;/code&gt;. The first swaps Qwen's MTP head for the DFlash2 block drafter, which proposes 7 tokens in one pass instead of 4 chained ones. The second keeps a document you already sent, including its attention KV and recurrent state. On a second turn against a 25k-token document, the README reports 0.56 s to first token instead of 22.4 s, with answers unchanged token for token.&lt;/p&gt;

&lt;p&gt;Then comes &lt;code&gt;DFLASH_TOKENS=15&lt;/code&gt;. It lets the target verify 16 tokens per step: the drafter still proposes its 7, and the remaining positions are filled from the request's own context. When the answer quotes the prompt, that pays. Reproducing a 25k-token document, the table in the README's only-user section goes from 260 tok/s with &lt;code&gt;SPEC=dflash2&lt;/code&gt; alone to 382 with the extra setting.&lt;/p&gt;

&lt;p&gt;On the eight chat prompts, the same table reads 132 against 133. The README measured why: positions 7 through 14 accounted for 72 of 11,069 accepted tokens, or 0.65%. And the setting has a price. Request slots fall from 8 to 4 and context from 64k to 56k, because a 16-token verify block doubles the recurrent-state page every resident request holds (1.66 GiB against 0.88). The README's guidance: set it for quoting documents or applying edits, where it puts the gain at 47%, and leave it at the default 7 for a chat or agentic client.&lt;/p&gt;

&lt;p&gt;The README frames that paragraph with its own instruction to read the table by column, not by its last cell.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where lossless stops
&lt;/h2&gt;

&lt;p&gt;The README calls the speculation lossless, reasoning that speculative decoding samples the same distribution as no speculation, and reports GSM8K at 96.0 to 96.5% across the three single-user columns. It also flags the exception. &lt;code&gt;CTX=huge&lt;/code&gt; uses a KVarN 4/2-bit KV cache, which the README calls lossy, and combined with &lt;code&gt;SPEC=dflash2&lt;/code&gt; it reports 268,169 tokens of KV capacity at 245760 max-model-len. On a bare-metal box that mode scored 95.2% GSM8K at n=600, which the README places inside the 95.0 to 96.5% band reported for the shipped configurations.&lt;/p&gt;

&lt;p&gt;A few practical notes before you deploy. The server listens on &lt;code&gt;0.0.0.0&lt;/code&gt; and is unauthenticated unless you give it a key, which it reads from &lt;code&gt;.env&lt;/code&gt; or &lt;code&gt;api_key.txt&lt;/code&gt;. And on Docker Desktop with WSL2, &lt;code&gt;VLLM_WSL2_ENABLE_PIN_MEMORY=1&lt;/code&gt; must stay enabled or the V2 runner aborts with &lt;code&gt;RuntimeError: UVA is not available&lt;/code&gt;. A contributor's WSL2 box ran roughly 20% slower on decode than the repo's bare-metal machine in the &lt;code&gt;CTX=huge&lt;/code&gt; comparison, so check which column matches your own hardware before you plan capacity.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/syv-ai/qwen38-27b-rtx3090" rel="noopener noreferrer"&gt;https://github.com/syv-ai/qwen38-27b-rtx3090&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Curated by &lt;a href="https://www.agentpalisade.com" rel="noopener noreferrer"&gt;Agent Palisade&lt;/a&gt; — practical AI for small and mid-sized businesses.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>vllm</category>
      <category>localllm</category>
      <category>speculativedecoding</category>
      <category>quantization</category>
    </item>
    <item>
      <title>Moli treats browser layout as a disposable snapshot for AI agents</title>
      <dc:creator>Reno Lu</dc:creator>
      <pubDate>Thu, 01 Oct 2026 18:37:38 +0000</pubDate>
      <link>https://dev.to/renolu/moli-treats-browser-layout-as-a-disposable-snapshot-for-ai-agents-2728</link>
      <guid>https://dev.to/renolu/moli-treats-browser-layout-as-a-disposable-snapshot-for-ai-agents-2728</guid>
      <description>&lt;p&gt;My reading of the moli README is that its central bet concerns when a browser does visual work. It ships a full rendering stack, yet by default it runs no real layout or paint at all. Geometry comes back mocked unless you pass &lt;code&gt;--layout&lt;/code&gt;, and even with that flag on, layout is a snapshot the browser builds, freezes, and holds only until the next rebuild. If your agent reads pages instead of looking at them, that default is the part to evaluate.&lt;/p&gt;

&lt;p&gt;Moli comes from Lexmount, is written in Rust, and the README describes it as a standalone browser kernel rather than a Chromium wrapper. It lists its building blocks: &lt;code&gt;libcurl&lt;/code&gt; for network transport, &lt;code&gt;html5ever&lt;/code&gt; for HTML parsing, &lt;code&gt;rusty_v8&lt;/code&gt; / V8 for JavaScript, Servo/Stylo for selectors, cascade, and computed style, Taffy and Parley for box and text layout, and AnyRender/Vello CPU with &lt;code&gt;usvg&lt;/code&gt; for software rendering.&lt;/p&gt;

&lt;h2&gt;
  
  
  Structure by default, pixels by request
&lt;/h2&gt;

&lt;p&gt;The README argues that browser automation tasks need page structure far more often than a continuously rendered visual world. Moli treats the native DOM and style state as the single source of truth and triggers layout or software paint only for operations that require them. Its request table breaks the cost down like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Extracting HTML or Markdown, querying the DOM, running JS, or inspecting network and storage reads runtime state directly, with no layout or paint.&lt;/li&gt;
&lt;li&gt;Reading an element's box, hit-testing coordinates, or sending coordinate input runs one layout calculation and keeps only the latest frozen layout tree.&lt;/li&gt;
&lt;li&gt;Capturing a screenshot rebuilds from the current DOM and style, replaces the frozen tree, renders a fresh frame, and discards the paint state afterward.&lt;/li&gt;
&lt;li&gt;Polling a screencast compares generation metadata; a clean state emits no frame, and a changed state produces one fresh frame.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The cost controls follow the same opt-in pattern. The default is &lt;code&gt;LayoutPolicy::Mock&lt;/code&gt;, which the README describes as deterministic geometry in a compatible format, with no real layout or paint. Passing &lt;code&gt;--layout&lt;/code&gt; switches to &lt;code&gt;LayoutPolicy::OnDemand&lt;/code&gt;. Media is also opt-in: &lt;code&gt;--resource&lt;/code&gt; fetches all optional visual and media resource families, while flags such as &lt;code&gt;--image&lt;/code&gt; or &lt;code&gt;--font&lt;/code&gt; enable one family each.&lt;/p&gt;

&lt;h2&gt;
  
  
  The frozen tree and its tradeoff
&lt;/h2&gt;

&lt;p&gt;According to the README, the first geometry request builds a working layout tree from the current DOM and style, freezes its canonical geometry into an immutable, DOM-independent &lt;code&gt;FrozenLayoutTree&lt;/code&gt;, and retains only that latest tree. The architecture section adds that each real refresh then discards the working tree, style borrows, layout caches, diagnostics, and paint state. Paint results are never reused.&lt;/p&gt;

&lt;p&gt;The sentence I would flag for anyone building on this: "Ordinary geometry reads may reuse it even if the page has changed." Screenshots always rebuild and replace the frozen tree, but a plain box read after a DOM mutation may be answered from the earlier snapshot. If your agent clicks by coordinates on a page that is still settling, test that path against your own flows before trusting the positions it returns.&lt;/p&gt;

&lt;h2&gt;
  
  
  One endpoint for three protocols
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;moli serve&lt;/code&gt; starts an automation server, and the README says the same endpoint serves CDP, WebDriver Classic, and WebDriver BiDi, which share one kernel and scheduler. No separate ChromeDriver, geckodriver, or browser installation is required. The example connects Playwright over CDP with &lt;code&gt;chromium.connectOverCDP&lt;/code&gt;. &lt;code&gt;moli serve --layout&lt;/code&gt; adds real geometry, coordinate input, and screenshot and screencast surfaces.&lt;/p&gt;

&lt;p&gt;For one-shot extraction, &lt;code&gt;moli fetch --dump markdown --wait-until done&lt;/code&gt; renders a page as Markdown, and &lt;code&gt;--dump semantic_tree_text&lt;/code&gt; returns what the README calls a compact, model-friendly semantic tree. Visual output needs the layout flag: &lt;code&gt;moli fetch --layout --dump screenshot&lt;/code&gt;, with &lt;code&gt;screenshot_full&lt;/code&gt; and &lt;code&gt;pdf&lt;/code&gt; as the other dump targets.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading the project's numbers
&lt;/h2&gt;

&lt;p&gt;Every figure below is the project's own reported measurement. In a mixed crawl of 192 public URLs from Chinese and international sites, a page counts only if it produces meaningful content after JavaScript runs. The README reports moli at 103 useful pages (53.6%), a median time of 1.43 s, and a median RSS of 73 MiB. Chrome Headless is listed at 101 pages (52.6%), the same 1.43 s median, and 773 MiB. Lightpanda shows a 0.97 s median and 40 MiB with 85 useful pages. Read plainly, the reported gap with Chrome Headless in that test sits in memory, while success rate and median time are close or identical.&lt;/p&gt;

&lt;p&gt;A sample agent workload in the README puts moli's CDP ready time at 34.85 ms against 169.37 ms for Chromium, peak PSS at 102.46 MiB against 348.82 MiB, and 1 process with 24 threads against 11 processes with 123 threads. The project also reports that one full run of its selected WPT tests passed 1.612 million tests.&lt;/p&gt;

&lt;p&gt;The README names crawling, browser-use agents, retrieval pipelines, evaluation environments, and reinforcement-learning workloads as fits for this cost model. If your workload depends on screenshots, benchmark it with &lt;code&gt;--layout&lt;/code&gt; enabled, since that mode is where the rebuild steps described above run.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/lexmount/moli" rel="noopener noreferrer"&gt;https://github.com/lexmount/moli&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Curated by &lt;a href="https://www.agentpalisade.com" rel="noopener noreferrer"&gt;Agent Palisade&lt;/a&gt; — practical AI for small and mid-sized businesses.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>rust</category>
      <category>browserautomation</category>
      <category>aiagents</category>
      <category>webscraping</category>
    </item>
    <item>
      <title>NorthCinder treats a purchase approval like a single-use credential</title>
      <dc:creator>Reno Lu</dc:creator>
      <pubDate>Wed, 30 Sep 2026 18:12:17 +0000</pubDate>
      <link>https://dev.to/renolu/northcinder-treats-a-purchase-approval-like-a-single-use-credential-278c</link>
      <guid>https://dev.to/renolu/northcinder-treats-a-purchase-approval-like-a-single-use-credential-278c</guid>
      <description>&lt;p&gt;The part of northcinder I would read first is the checkout rule, not the product comparison. According to the README, a recommendation is not permission to buy, and every checkout needs a fresh approval for one exact offer and one unit, signed and usable once. For anyone who has designed authorization for an agent, that reads less like a shopping feature and more like a scoped capability token.&lt;/p&gt;

&lt;h2&gt;
  
  
  The argument the README makes
&lt;/h2&gt;

&lt;p&gt;The project describes itself as an open-source MCP server for comparing products and asking the buyer before purchase. Its README opens with a position. It argues that big marketplaces are building shopping agents that search one catalog and steer the buyer toward that platform's checkout, and that this may be convenient but is not independent advice, because the marketplace still decides what can be seen and makes money when the agent closes the sale.&lt;/p&gt;

&lt;p&gt;That is the README's framing, and it names no specific marketplace. Its response is software you run alongside your own AI app. The README states that the repository owner does not operate a NorthCinder service, and that there is no NorthCinder account or cloud service.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the approval works
&lt;/h2&gt;

&lt;p&gt;The "Buying stays a separate decision" section is short, so its mechanics deserve a close read. Every checkout needs a fresh approval for one exact offer and one unit. The signed approval includes the merchant, variant, price, known total, and spending cap. It can be used once.&lt;/p&gt;

&lt;p&gt;Read that as an access control design. The approval binds to specific purchase parameters rather than to a session or a general mandate to shop. A second unit or a different variant is outside its scope. The README also says that starting a search or price watch does not give NorthCinder permission to buy anything, so a watch running in the background carries no purchase authority on its own.&lt;/p&gt;

&lt;p&gt;Payment handling follows the same restraint. Per the README, NorthCinder rejects raw card details rather than storing or forwarding them. A supported automated checkout can use an opaque payment token. The other path is a cart link handed to the buyer, who finishes the purchase in their own browser.&lt;/p&gt;

&lt;p&gt;There is one more gate on the store side. The README says a native connection must confirm the exact offer before checkout or an unattended watch. When a native connection is missing, the AI app can keep researching with its own browser or search tools, but NorthCinder accepts product facts, not cookies, raw pages, passwords, or page instructions. The README states that exclusion without describing how it is enforced.&lt;/p&gt;

&lt;h2&gt;
  
  
  What gets logged, and where
&lt;/h2&gt;

&lt;p&gt;The README says NorthCinder reruns the ranking locally and writes recommendations, approvals, and checkout attempts to a local audit log. The audit log and purchase approvals stay on your computer. Order outcomes stay local and attach only to the purchase they belong to, and the README says the tool does not silently rewrite the buyer's profile.&lt;/p&gt;

&lt;p&gt;The ranking rules are written as plain statements. Seller payment never improves ranking. Sponsored offers stay labeled and below organic results. Unknown seller history remains unknown instead of being guessed safe or unsafe. The README points to separate ranking, trust, neutrality audit, and checkout documents for detail, and it scopes its own checks: they cover the offers NorthCinder received, not the completeness or truth of a store's catalog.&lt;/p&gt;

&lt;p&gt;Output is usually limited as well. The README says it normally shows no more than three useful choices, explains why each result ranked where it did, and lets you inspect the other finalists, rejected offers, and facts that could not be verified.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the README asks for caution
&lt;/h2&gt;

&lt;p&gt;The research flow carries explicit warnings. Before doing research, the MCP host should read the relevant product or seller research guide, call &lt;code&gt;create_research_plan&lt;/code&gt; with the actual request and exact subject, then follow the returned checklist with the research tools it already controls. Research can decide whether an offer is ready to compare, but it cannot add ranking points. If sources disagree or do not identify the exact product or seller, the result stays provisional.&lt;/p&gt;

&lt;p&gt;The README goes further and states that no host and model combination is currently qualified for routine research use. It tells readers to treat every research result as provisional until the buyer checks its identity, sources, conflicts, and unknowns.&lt;/p&gt;

&lt;p&gt;That caveat belongs next to the approval model in any evaluation. The consent gate governs spending. It does not make the upstream research trustworthy, and the README does not claim it does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trying it
&lt;/h2&gt;

&lt;p&gt;You need Node.js 20 or later and an MCP-capable AI app. Running &lt;code&gt;npx northcinder init&lt;/code&gt; saves your configuration on your computer and prints the MCP entry for your AI app. Local mode is keyless and runs the MCP server and search engine together in one process, using a temporary loopback port.&lt;/p&gt;

&lt;p&gt;If you run the engine separately, the README says to set &lt;code&gt;NORTHCINDER_API_KEYS&lt;/code&gt; on the service and configure the client with &lt;code&gt;NORTHCINDER_SERVICE_URL&lt;/code&gt; and the matching &lt;code&gt;NORTHCINDER_CLIENT_KEY&lt;/code&gt; bearer credential. Non-loopback bearer connections must use HTTPS. Built-in adapters cover Shopify, WooCommerce, eBay, Etsy, and read-only Amazon comparison, and the README says NorthCinder reports when a store was unavailable or not configured.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/cinderline/northcinder" rel="noopener noreferrer"&gt;https://github.com/cinderline/northcinder&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Curated by &lt;a href="https://www.agentpalisade.com" rel="noopener noreferrer"&gt;Agent Palisade&lt;/a&gt; — practical AI for small and mid-sized businesses.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>security</category>
      <category>agents</category>
      <category>privacy</category>
    </item>
    <item>
      <title>Reverify weighs verified AI claims by how informative they are</title>
      <dc:creator>Reno Lu</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:22:54 +0000</pubDate>
      <link>https://dev.to/renolu/reverify-weighs-verified-ai-claims-by-how-informative-they-are-4hb0</link>
      <guid>https://dev.to/renolu/reverify-weighs-verified-ai-claims-by-how-informative-they-are-4hb0</guid>
      <description>&lt;p&gt;My reading of the reverify README is that its central idea is not the verifier. It is the admission that "every claim verified" is a score a model can reach while saying nothing. Assert that a file starts with &lt;code&gt;MZ&lt;/code&gt; and that &lt;code&gt;.text&lt;/code&gt; exists, and you get a clean sheet that says little. Reverify's response is to weigh each verified claim by how much it tells you, and that is the part I would point a skeptical reviewer to first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The loop the name refers to
&lt;/h2&gt;

&lt;p&gt;The project's pitch is "Stop your AI from making things up." The mechanism behind that pitch is a division of labor. A language model proposes a claim about an artifact. A deterministic tool checks the claim against that artifact and returns &lt;code&gt;VERIFIED&lt;/code&gt;, &lt;code&gt;REFUTED&lt;/code&gt;, or &lt;code&gt;INCONCLUSIVE&lt;/code&gt;, along with the bytes it observed. In the README's words, the model never gets to assert a fact on its own.&lt;/p&gt;

&lt;p&gt;The README picks binary reverse engineering as its proving ground, saying the hallucination problem in binary analysis is far worse than in source code. The toolkit covers PE/ELF/Mach-O parsing, x86/x64/ARM/ARM64 disassembly, AOB pattern scanning, CPU emulation, Protobuf/TLV dissection, and Frida hook generation, in pure Python out of the box. Installing &lt;code&gt;reverify[full]&lt;/code&gt; upgrades it to capstone, unicorn, lief and Z3, and &lt;code&gt;reverify[angr]&lt;/code&gt; adds angr for function boundaries, the call graph and cross-references. Without those engines it falls back to the pure-Python core, and &lt;code&gt;reverify backends&lt;/code&gt; shows what is active.&lt;/p&gt;

&lt;p&gt;A claim is a small JSON object. The README's first example asks whether the instructions at offset 4096 are &lt;code&gt;push&lt;/code&gt;, &lt;code&gt;mov&lt;/code&gt;, &lt;code&gt;sub&lt;/code&gt;, noted as a function prologue. Another hands raw x86 bytes to the emulator and expects &lt;code&gt;eax&lt;/code&gt; to equal 8. Claims can be batched with &lt;code&gt;--claims-file claims.json&lt;/code&gt;, and the CLI exits non-zero if anything is refuted, so an agent or CI job can gate on the result. A &lt;code&gt;depends_on&lt;/code&gt; field lets a refuted root invalidate the claims built on it, and &lt;code&gt;"observe": true&lt;/code&gt; has the tools read a value instead of asserting one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Weight, not a tally
&lt;/h2&gt;

&lt;p&gt;Each result carries a &lt;code&gt;weight&lt;/code&gt;. The README sets it to zero for claims that restate the fact sheet the model was shown, for duplicates, for inline code or data that does not occur in the binary, and for echoes of the tools' own previous output. Otherwise the weight is measured from the binary itself: how often the expected content occurs in the file and how much entropy it has. Zero padding or a ubiquitous prologue will verify and still weigh almost nothing. A reconstruction counts as grounded only when nothing is refuted and the verified weight reaches &lt;code&gt;--min-information&lt;/code&gt;, which defaults to 1.0. The README says this follows the CORE refinement of FActScore.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;reverify reconstruct --samples N&lt;/code&gt; carries the idea into generation. It draws several proposals per round and lets the verifier, not the model's confidence, select among them.&lt;/p&gt;

&lt;p&gt;A plain pass/fail gate leaves a loophole open: a verified set can contain safe claims that say little. The weight rule is the README's answer to that, and it is a pattern worth considering for harnesses that grade model output.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers, and where they come from
&lt;/h2&gt;

&lt;p&gt;The README reports that on 71 real Windows system files, the AI's textbook answer was wrong 97% of the time, and that reverify caught every one and never accepted a wrong claim. It also cites a per-claim-kind confusion matrix: 0 false VERIFIED out of 475 known-false claims, and 0 known-true claims missed, gated in every CI job. According to the README, the benchmark runs in CI on every push against each platform's own system binaries and fails the build if a single wrong claim is VERIFIED, and a third-party aarch64 replication is documented in BENCHMARK.md. These are the project's figures. The README describes a replication package with one command per benchmark and a pinned Dockerfile for anyone who wants to check them.&lt;/p&gt;

&lt;p&gt;Output from &lt;code&gt;reverify verify --json&lt;/code&gt; also includes the binary's SHA-256, the reverify version and which engines judged, so a report can be handed over and replayed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Outside the binary
&lt;/h2&gt;

&lt;p&gt;Two features reach past reverse engineering. &lt;code&gt;reverify equiv &amp;lt;reference&amp;gt; &amp;lt;candidate&amp;gt; --lang python&lt;/code&gt; (or C) runs a candidate implementation and a reference over shared inputs and checks that they agree. A refutation comes back with the input and both outputs, which puts an AI's rewrite under the same verdict structure.&lt;/p&gt;

&lt;p&gt;The second is context. &lt;code&gt;reverify rollover&lt;/code&gt; hands a session off to a file and starts a fresh one instead of leaning on a lossy auto-summary, and the README lists Claude Code, Codex, Gemini CLI and OpenCode as supported. Since v0.8.0 the loop also writes a ledger per binary at &lt;code&gt;.reverify/ledger/&amp;lt;sha256&amp;gt;.json&lt;/code&gt;, checkpointed after every round. Refutations come back as &lt;code&gt;KNOWN FALSE&lt;/code&gt;, so a fresh context does not re-propose the same wrong prior. The README's reasoning is that the model's unverified prose was never trusted, so dropping it loses nothing.&lt;/p&gt;

&lt;p&gt;Reverify ships as an MCP server and a plain CLI. The README scopes it to authorized reverse engineering: malware analysis, CTF, interoperability research, and software you own or are permitted to analyze.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/2akouwu/reverify" rel="noopener noreferrer"&gt;https://github.com/2akouwu/reverify&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Curated by &lt;a href="https://www.agentpalisade.com" rel="noopener noreferrer"&gt;Agent Palisade&lt;/a&gt; — practical AI for small and mid-sized businesses.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>reverseengineering</category>
      <category>mcp</category>
    </item>
    <item>
      <title>munder-difflin runs an office of agents on the coding CLIs you already pay for</title>
      <dc:creator>Reno Lu</dc:creator>
      <pubDate>Mon, 28 Sep 2026 19:54:32 +0000</pubDate>
      <link>https://dev.to/renolu/munder-difflin-runs-an-office-of-agents-on-the-coding-clis-you-already-pay-for-2jnc</link>
      <guid>https://dev.to/renolu/munder-difflin-runs-an-office-of-agents-on-the-coding-clis-you-already-pay-for-2jnc</guid>
      <description>&lt;p&gt;The interesting decision in munder-difflin is what it keeps. The coding CLIs you already use stay at the center. It launches the coding CLIs you already have installed, &lt;code&gt;claude&lt;/code&gt;, &lt;code&gt;codex&lt;/code&gt;, &lt;code&gt;agy&lt;/code&gt;, &lt;code&gt;grok&lt;/code&gt;, &lt;code&gt;kimi&lt;/code&gt;, &lt;code&gt;qwen&lt;/code&gt;, &lt;code&gt;opencode&lt;/code&gt;, &lt;code&gt;crush&lt;/code&gt;, &lt;code&gt;pi&lt;/code&gt;, &lt;code&gt;copilot&lt;/code&gt; or a custom command, as real processes, and then builds an office around them. The README puts the economics plainly: it "works with the subscriptions you already pay for, on their hourly limits." Every agent in the office is one of those terminals, running with your existing subscription and its hourly limits. The framing is working inside those limits, and the README does not pitch it as a way past them.&lt;/p&gt;

&lt;p&gt;The architecture described in the README starts with that terminal model, so it is worth following through.&lt;/p&gt;

&lt;h2&gt;
  
  
  A terminal is the unit of work
&lt;/h2&gt;

&lt;p&gt;Each agent is a normal terminal process spawned in a pseudo-terminal through &lt;code&gt;node-pty&lt;/code&gt; and rendered with xterm.js. The README calls the result "byte-for-byte authentic," meaning the harness shows you the same session the CLI would show you on its own. Every agent gets its own working directory, identity and provider-specific lifecycle.&lt;/p&gt;

&lt;p&gt;The README gives provider support a simple rule: if it runs in a terminal, it can run here. The supported list covers Claude Code, Codex, Grok, Kimi Code, Gemini CLI, Antigravity, Qwen, OpenCode, Crush, Pi, GitHub Copilot and Cursor. It also supports your own API keys and local models through Ollama, LM Studio or vLLM.&lt;/p&gt;

&lt;p&gt;The practical upside is that you keep the tool you already know. Click a desk on the floor and you read that terminal live, and you can type straight back into it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hive is a git repo of plain files
&lt;/h2&gt;

&lt;p&gt;Once you have a dozen independent processes, something has to let them talk. Here that something is deliberately boring: a local git repo of plain files. Each agent writes to its own &lt;code&gt;outbox/&lt;/code&gt;, and the harness's router delivers messages into the recipients' &lt;code&gt;inbox/&lt;/code&gt; folders. Agents read their memory and drain their mailbox.&lt;/p&gt;

&lt;p&gt;The README also notes that no agent touches git. The harness is the single committer, a design it describes as avoiding &lt;code&gt;index.lock&lt;/code&gt; corruption.&lt;/p&gt;

&lt;p&gt;Memory follows the same file-first pattern. Each agent keeps markdown memory, which gets mined into a shared, searchable store the README calls a palace, backed by a semantic recall index. The stated goal is that you can close the app, come back the next day, and the agents still know what they learned.&lt;/p&gt;

&lt;h2&gt;
  
  
  One agent you brief
&lt;/h2&gt;

&lt;p&gt;Running many subscriptions in parallel creates a new problem, which is that you now have many conversations to manage. The answer is an orchestrator the README calls your clone, named Michael. You brief Michael, and the GOD agent layer handles the roster, routing, a blackboard and a task ledger.&lt;/p&gt;

&lt;p&gt;The supervisor resolves routine requests itself and escalates only critical items into an approvals queue you act on. The README lists those as spend, destructive operations and scope changes. Per-agent autonomy settings control how far each agent may go alone, and a circuit breaker steers, constrains and then stops anything that loops or runs away. For an office of processes spending your hourly allowance, a stop condition for loops is the feature I would look at first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The office floor and the setup path
&lt;/h2&gt;

&lt;p&gt;The visual layer is a Pixi.js floor where each session appears as an avatar. Agents walk to stations as they work, and envelopes fly between desks when they message each other. The moving avatars and envelopes give a visual read on which agents are working and which are messaging each other.&lt;/p&gt;

&lt;p&gt;Adding an agent means picking the CLI, the model and the autonomy level, then assigning a desk. There is an Agent Gallery of ready-made roles to import. A first-run wizard checks what you already have installed and offers to install what is missing.&lt;/p&gt;

&lt;p&gt;The app is built on Electron, React and TypeScript, runs on macOS, Windows and Linux, and is MIT licensed. The README marks it as pre-release at version 0.4.6. macOS builds are signed and notarized, so you do not need to build from source to try it.&lt;/p&gt;

&lt;p&gt;If your team already pays for one or two of these CLIs and keeps wishing it could run several sessions on separate tasks without babysitting each one, this is a concrete design to evaluate. Start small, set autonomy conservatively, and observe how the agents coordinate within the hourly limits you already have.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/chaitanyagiri/munder-difflin" rel="noopener noreferrer"&gt;https://github.com/chaitanyagiri/munder-difflin&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Curated by &lt;a href="https://www.agentpalisade.com" rel="noopener noreferrer"&gt;Agent Palisade&lt;/a&gt; — practical AI for small and mid-sized businesses.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>multiagent</category>
      <category>claudecode</category>
      <category>codex</category>
      <category>orchestration</category>
    </item>
    <item>
      <title>sprix-sage-router factors half-finished work into who finishes the task</title>
      <dc:creator>Reno Lu</dc:creator>
      <pubDate>Sun, 27 Sep 2026 17:21:17 +0000</pubDate>
      <link>https://dev.to/renolu/sprix-sage-router-factors-half-finished-work-into-who-finishes-the-task-50mp</link>
      <guid>https://dev.to/renolu/sprix-sage-router-factors-half-finished-work-into-who-finishes-the-task-50mp</guid>
      <description>&lt;p&gt;A telling number in the sprix-sage-router README is not a utility score at all. It is the wasted work column. In the authors' own mid-execution replay, progress-aware SAGE edges its progress-masked twin on utility by 0.298 to 0.291, with both utility values reported as plus or minus 0.021, but it cuts wasted work from 0.104 to 0.059 and reports a 58.4% switch rate instead of 80.1%. That is the bet this project makes: once a task is underway, the work already done should change who finishes it.&lt;/p&gt;

&lt;p&gt;The README describes the repository as an open-source research output of Sprix AI. It carries a Research Preview status badge and ships a Python 3.10+ reference implementation with no runtime dependencies.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question discovery leaves open
&lt;/h2&gt;

&lt;p&gt;The README frames the problem plainly. Agent discovery tells a system which agents exist. It does not say who should work with whom after execution has already begun. SAGE, short for State-Aware Graph Exchange, is pitched as a decision layer above the Agent2Agent (A2A) protocol, sitting between discovery and task execution.&lt;/p&gt;

&lt;p&gt;It weighs three routes in one objective. SELF keeps the incumbent agent on the job. COLLABORATE keeps the incumbent as owner but recruits a small complementary team to cover missing requirements. HANDOFF gives full ownership to a peer, which the README says fits when specialist advantage exceeds context-transfer loss.&lt;/p&gt;

&lt;p&gt;What makes it checkpoint-aware is the input. The router accounts for completed DAG nodes, reusable artifacts, observed partial quality, remaining work, failures, budget, and deadline. The central idea is a reuse fraction. If the current owner has finished part of a requirement and keeps it, that progress counts in full. If ownership changes, only the share that survives artifact portability carries over. The same fraction lowers projected remaining cost and duration. In the quick-start example, planning is complete with a transferability of 0.95, while coding is 35% in flight with a transferability of 0.40, so moving the coding work to the stronger coder does not come free.&lt;/p&gt;

&lt;p&gt;Feasible routes are ranked by a linear utility that starts from a learned success probability and subtracts penalties including cost, latency, context-transfer loss, and coordination overhead, with terms for uncertainty-aware exploration. The authors say outright that noisy-OR, beam search, Beta beliefs, online logistic regression, and DAG-induced communication edges are not claimed as inventions. They call them replaceable implementation mechanisms.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the evaluation shows, and what it does not
&lt;/h2&gt;

&lt;p&gt;The README splits evaluation into three benchmark questions rather than leaning on one table, and it labels their nature.&lt;/p&gt;

&lt;p&gt;The mid-execution replay in &lt;code&gt;benchmark_dynamic.py&lt;/code&gt; covers 1,000 checkpoints over five seeds. The authors state that its evaluator scores artifact portability and remaining work without calling the router's switch-loss code, and that every policy gets the same registry, action space, budget, and deadline. The "Always hand off" baseline lands at 0.290 utility with a 100% switch rate and 0.130 wasted work. "Always continue" drops to 0.085 with a 34.4% deadline miss rate. A hidden-state dynamic oracle reaches 0.375, so by the authors' own numbers there is headroom above SAGE.&lt;/p&gt;

&lt;p&gt;A controlled intervention changes only in-flight completion, from 0.0 to 0.9. The progress-aware minus progress-masked utility gap grows from 0.0000 to +0.0684. The caption on that figure says the values are synthetic five-seed means, not real-endpoint evidence.&lt;/p&gt;

&lt;p&gt;A second benchmark tests requirement-conditioned trust. On the same evidence stream in a heterogeneous specialist setting, per-requirement trust reports a Brier score of 0.0125 and selection regret of 0.0094, against 0.0355 and 0.0310 for a single reputation score. In a homogeneous negative control both have zero routing regret, which the authors offer as a sign the benchmark does not manufacture an advantage when specialization is absent.&lt;/p&gt;

&lt;p&gt;A third suite, &lt;code&gt;benchmark.py&lt;/code&gt;, runs 2,500 independent tasks. The README describes it as regression and learning ablation, not support for the mid-execution claim.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the prototype stops
&lt;/h2&gt;

&lt;p&gt;The A2A layer, &lt;code&gt;sprix_a2a.py&lt;/code&gt;, validates declared skill IDs against locally calibrated evidence and turns a selected route into a transport-neutral &lt;code&gt;ExecutionPlan&lt;/code&gt; covering ownership, assignments, DAG dependencies, communication edges, estimated resources, and rationale. The README says the current prototype intentionally does not transmit tasks, authenticate endpoints, or verify signatures. Your A2A client still handles &lt;code&gt;message/send&lt;/code&gt;, streaming, polling, cancellation, and secure artifact handling.&lt;/p&gt;

&lt;p&gt;Read it as the research preview its status badge says it is. If you run multi-agent pipelines where a reassignment throws away half-finished work, the reuse fraction and the wasted-work metric are the pieces worth borrowing for your own evaluation. The router exposes &lt;code&gt;record_outcome&lt;/code&gt; to feed back execution results and &lt;code&gt;export_state&lt;/code&gt; to persist learned evidence.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/wang2122/sprix-sage-router" rel="noopener noreferrer"&gt;https://github.com/wang2122/sprix-sage-router&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Curated by &lt;a href="https://www.agentpalisade.com" rel="noopener noreferrer"&gt;Agent Palisade&lt;/a&gt; — practical AI for small and mid-sized businesses.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiagents</category>
      <category>multiagentsystems</category>
      <category>a2a</category>
      <category>python</category>
    </item>
    <item>
      <title>KiroCrew Treats a Development Workspace Like a Service You Operate</title>
      <dc:creator>Reno Lu</dc:creator>
      <pubDate>Sat, 26 Sep 2026 16:49:21 +0000</pubDate>
      <link>https://dev.to/renolu/kirocrew-treats-a-development-workspace-like-a-service-you-operate-h6e</link>
      <guid>https://dev.to/renolu/kirocrew-treats-a-development-workspace-like-a-service-you-operate-h6e</guid>
      <description>&lt;p&gt;KiroCrew's telling design choice is to present a development workspace as software you install and keep running on your own hardware. It treats the workspace as a service you install and keep running on your own hardware, and its README opens with pages of operations detail to match. If you want to judge whether the project fits your setup, that plumbing is the part to read closely.&lt;/p&gt;

&lt;p&gt;The pitch itself is short. Kiro Crew is an open source development workspace that runs locally or on a remote host you control. Multi-step tasks can run unattended, recurring jobs run on a schedule you set, and heartbeats monitor systems until something needs attention. The banner describes software that runs on your hardware, remembers across sessions, and keeps working unattended. The project also calls the workspace self-learning and self-evolving, but the introduction does not explain how that works. The README's table of contents links a separate "How it works" section for readers who want the mechanism.&lt;/p&gt;

&lt;h2&gt;
  
  
  A workspace with several front doors
&lt;/h2&gt;

&lt;p&gt;The same work is reachable from a desktop app, a web dashboard, and a CLI. The web dashboard works without any messaging credentials. When you step away from it, you can continue with the same agent through a messaging channel: Slack, Discord, Telegram, Teams, Webex, WeCom, WeChat, WhatsApp, Feishu, or iMessage. The README notes that Teams needs a public HTTPS webhook.&lt;/p&gt;

&lt;p&gt;That list fits the README's description of a persistent workspace that remembers across sessions. The desktop app, for one, talks to a Gateway. It starts a bundled Gateway when no local one is already running, and it can connect to a remote Gateway over an SSH tunnel. Underneath every install path sits &lt;code&gt;kiro-cli&lt;/code&gt;, which the first launch installs if needed before guiding you through Kiro device-code sign-in.&lt;/p&gt;

&lt;p&gt;On top of that, Kiro Crew Apps package the workspace for a specific job. According to the README, an App combines a purpose-built interface with agents, skills, schedules, integrations, and backend services.&lt;/p&gt;

&lt;h2&gt;
  
  
  Install choices for local and long-running hosts
&lt;/h2&gt;

&lt;p&gt;You can run the desktop app, a one-line install on your machine or a remote host, a Docker image meant for always-on servers, or a build from source. On Linux the README recommends starting with the one-line install, because it puts &lt;code&gt;kirocrew&lt;/code&gt; on your PATH. That makes &lt;code&gt;kirocrew service install&lt;/code&gt; reachable, along with the AppArmor profile the agent sandbox needs on Ubuntu 23.10 and later. The &lt;code&gt;.deb&lt;/code&gt; and &lt;code&gt;.rpm&lt;/code&gt; desktop packages install to a fixed path under &lt;code&gt;/opt&lt;/code&gt;, which lets them set that profile up for you. The AppImage skips root but needs FUSE and carries a manual sandbox step.&lt;/p&gt;

&lt;p&gt;The installer also avoids your system Python by default. It fetches a SHA-256-pinned copy of uv and provisions a self-contained CPython 3.12 into &lt;code&gt;~/.kiro/crew-python&lt;/code&gt;, so the install does not depend on, or break with, the system interpreter. You can opt out with the installer's system-python flag, and that choice is recorded next to the release channel so later re-runs of the installer keep it. If the managed interpreter cannot be downloaded, the installer falls back to a system Python 3.12 or newer for that run. Air-gapped hosts can point at mirrors through &lt;code&gt;KIROCREW_UV_URL&lt;/code&gt; and &lt;code&gt;UV_PYTHON_INSTALL_MIRROR&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Supply-chain details get similar care. The wheel is SHA-256 verified, and x86_64 and aarch64 each get their own build, auto-update feed, and SLSA provenance attestation. The Docker image publishes &lt;code&gt;linux/amd64&lt;/code&gt; and &lt;code&gt;linux/arm64&lt;/code&gt; under every tag, so a pull picks the right one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Release channels and a shared data home
&lt;/h2&gt;

&lt;p&gt;Every install path offers Stable, Insider, and Nightly. Stable is the default, built from Insider builds that baked long enough to be promoted. Insider ships from release-candidate tags for people who want features days to weeks early. Nightly builds untested &lt;code&gt;main&lt;/code&gt; HEAD at 06:00 UTC daily, and the README tells you to expect breakage.&lt;/p&gt;

&lt;p&gt;The fine print matters for anyone running more than one copy. Stable and Insider are two update lanes of the same app, so installing both side by side does not give you independent installs. They read one desktop settings store and share one update download cache, and whichever copy last wrote the channel wins. Switch one to Insider and the other follows. Setting &lt;code&gt;KIROCREW_HOME&lt;/code&gt; does not separate them either, since it moves the data home and not the desktop settings.&lt;/p&gt;

&lt;p&gt;Nightly is a separate app with its own name, icon, channel, and settings store, so it installs alongside Stable or Insider instead of replacing it. The README is explicit that this does not make it a sandbox: Nightly reads the same &lt;code&gt;~/.kiro/crew&lt;/code&gt; data home unless you point it elsewhere with &lt;code&gt;KIROCREW_HOME&lt;/code&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/kirodotdev/KiroCrew" rel="noopener noreferrer"&gt;https://github.com/kirodotdev/KiroCrew&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Curated by &lt;a href="https://www.agentpalisade.com" rel="noopener noreferrer"&gt;Agent Palisade&lt;/a&gt; — practical AI for small and mid-sized businesses.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>devtools</category>
      <category>automation</category>
      <category>python</category>
    </item>
    <item>
      <title>internet-court-skill structures deals around an up-front dispute choice</title>
      <dc:creator>Reno Lu</dc:creator>
      <pubDate>Fri, 25 Sep 2026 17:37:41 +0000</pubDate>
      <link>https://dev.to/renolu/internet-court-skill-structures-deals-around-an-up-front-dispute-choice-4p3f</link>
      <guid>https://dev.to/renolu/internet-court-skill-structures-deals-around-an-up-front-dispute-choice-4p3f</guid>
      <description>&lt;p&gt;A central idea in internet-court-skill is deciding what happens when a deal breaks. The project has both agents agree on that process up front, inside the contract itself.&lt;/p&gt;

&lt;p&gt;The README opens with a claim about where things stand: agents are beginning to transact, negotiate, and pay one another without humans in the loop, and what they lack is a way to trust each other. Take that as the project's premise rather than a settled fact. The diagnosis that follows is the part worth reading closely. The building blocks exist, the README says, but they are fragmented and each one is built for the happy path. When a deal goes wrong, every layer passes the problem down the line.&lt;/p&gt;

&lt;p&gt;The README locates the problem in the gaps between fragmented layers rather than in a single protocol.&lt;/p&gt;

&lt;h2&gt;
  
  
  Six layers, one missing owner
&lt;/h2&gt;

&lt;p&gt;The README lays out agentic commerce as six layers: discovery and identity, negotiation, contracts and obligations, payment and escrow, execution, and verification and disputes. It maps named standards to each one. ERC-8004 and ERC-7857 sit at identity. A2A sits at negotiation. ERC-7710, ERC-8183 and Arkhai sit at contracts. x402, MPP and APP sit at payment and escrow.&lt;/p&gt;

&lt;p&gt;The project's argument is that each of those standards solves its own layer and assumes everything goes right. The sixth layer, verification and disputes, is described as the one nobody else owns and the one Internet Court exists to add.&lt;/p&gt;

&lt;p&gt;In practice, the repository is a master router in &lt;code&gt;SKILL.md&lt;/code&gt;, a set of first-party connector skills under &lt;code&gt;integrations/&lt;/code&gt;, and vendored copies of official skills from the protocols in the stack. The README counts 91 vendored skills from 33 owners. Connectors include &lt;code&gt;integrations/x402-erc7710/&lt;/code&gt; for the payment layer and &lt;code&gt;integrations/genlayer-erc7710-connector/&lt;/code&gt; at the contracts and disputes layers. The README states that the repo never re-implements a protocol's own skill; first-party content is only the master skill and the connectors.&lt;/p&gt;

&lt;h2&gt;
  
  
  What adjudication means here
&lt;/h2&gt;

&lt;p&gt;The core mechanism is this. When two agents make a deal, they also agree up front how it will be settled if something goes wrong, and that agreement is written into the contract itself. The README names GenLayer, Kleros and UMA as possible judges, or whatever the parties choose. It is explicit that most deals never reach that stage: when both sides agree, the contract simply settles. What the skill standardizes is how the contract is structured around the choice of judge, for the cases that fall off the happy path.&lt;/p&gt;

&lt;p&gt;The README frames adjudication as a response to the happy-path limitation it identifies. Choosing the dispute process in advance establishes who will judge if a dispute arises. It does not claim to make the outcome correct or remove disagreement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where an operator should look before trusting it
&lt;/h2&gt;

&lt;p&gt;The README spells out several coverage and installation details that bear on whether an agent should use the skill.&lt;/p&gt;

&lt;p&gt;Coverage is uneven, and the README says so. ERC-7857 and ERC-8183 have no public skill yet. UMA is listed as a disputes option but also has no public skill yet. The vendored &lt;code&gt;a2a-protocol&lt;/code&gt; skill is a community one, with a note that no official A2A skill exists. If your agent's contract names UMA as its judge, the repo has no public UMA skill today.&lt;/p&gt;

&lt;p&gt;The install model changes where sub-skills come from. Only the root &lt;code&gt;SKILL.md&lt;/code&gt; is registered, and it pulls the other skills in on demand. A full install reads bundled sub-skills from disk, while a root-only install can fetch them from the repository's &lt;code&gt;main&lt;/code&gt; branch on demand. Know which mode you are running.&lt;/p&gt;

&lt;p&gt;Vendored skills are pinned. &lt;code&gt;skills-lock.json&lt;/code&gt; records a pinned source, hash and refresh command for each vendored skill.&lt;/p&gt;

&lt;p&gt;Governance is shared. The README describes a consortium of founding members spanning the stack, with each member's protocol embedded in the standard, and calls the standard open and openly governed. The README does not spell out who can change the router.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/internet-court/internet-court-skill" rel="noopener noreferrer"&gt;https://github.com/internet-court/internet-court-skill&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Curated by &lt;a href="https://www.agentpalisade.com" rel="noopener noreferrer"&gt;Agent Palisade&lt;/a&gt; — practical AI for small and mid-sized businesses.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aiagents</category>
      <category>security</category>
      <category>disputeresolution</category>
      <category>agentskills</category>
    </item>
    <item>
      <title>FuXi treats the model as a swappable engine</title>
      <dc:creator>Reno Lu</dc:creator>
      <pubDate>Thu, 24 Sep 2026 17:38:40 +0000</pubDate>
      <link>https://dev.to/renolu/fuxi-treats-the-model-as-a-swappable-engine-kk6</link>
      <guid>https://dev.to/renolu/fuxi-treats-the-model-as-a-swappable-engine-kk6</guid>
      <description>&lt;p&gt;FuXi puts the loop and the routing wrapped around the model at the center of its pitch. Its README sums it up as "the model is the engine, FuXi is the vehicle": the configuration supports multiple providers and per-model settings, and the benchmark measures the agent loop rather than raw model scores.&lt;/p&gt;

&lt;p&gt;The README pitches FuXi as a provider-agnostic alternative to Claude Code. It is built in Go, ships as one static binary with no runtime dependencies, and asks you to bring whatever OpenAI-compatible model you already use. On top of that model it runs a Think, Act, Verify loop, with cost-aware routing across LLM providers and automatic failover.&lt;/p&gt;

&lt;h2&gt;
  
  
  Providers as a catalog
&lt;/h2&gt;

&lt;p&gt;The config format shows how far the swappable-engine idea goes. A single-provider setup is four lines of YAML: provider, base URL, API key, model. For more than one, FuXi offers a layered schema, with a &lt;code&gt;providers:&lt;/code&gt; catalog where each entry declares its type, endpoint, key, and models, and a separate &lt;code&gt;model:&lt;/code&gt; layer that selects the active provider and model id. The README says the layered schema supports multiple providers and per-model settings.&lt;/p&gt;

&lt;p&gt;The README separately claims cost-aware routing across providers and automatic failover. The README also lists a &lt;code&gt;fuxi proxy&lt;/code&gt; subcommand that starts a smart routing proxy doing protocol bridging between providers, and &lt;code&gt;fuxi launch&lt;/code&gt;, which runs a proxied binary through that proxy using your FuXi config. What the README does not spell out is the routing policy itself, so if you care how it picks a model for a given request, that is the part to test.&lt;/p&gt;

&lt;p&gt;Credentials follow the same pattern. You can supply a provider key (the README names OpenAPI-compatible endpoints, Gemini, and Bedrock/Vertex), export &lt;code&gt;FUXI_API_KEY&lt;/code&gt;, or run &lt;code&gt;fuxi login&lt;/code&gt;, which provisions FuXi-managed models with no key needed. &lt;code&gt;fuxi init&lt;/code&gt; writes a starter config and auto-detects a provider from environment variables already set, and &lt;code&gt;fuxi wizard&lt;/code&gt; walks through provider, base URL, key, model, and a connection test. For a single run, &lt;code&gt;-m&lt;/code&gt;, &lt;code&gt;-P&lt;/code&gt;, &lt;code&gt;-b&lt;/code&gt;, and &lt;code&gt;-k&lt;/code&gt; override the model, provider type, base URL, and key.&lt;/p&gt;

&lt;h2&gt;
  
  
  The loop carries the weight
&lt;/h2&gt;

&lt;p&gt;If the model is interchangeable, the loop has to do the work. The README claims the Think, Act, Verify loop and its routing let a model perform above its raw benchmark. Behind the loop sit 50+ built-in tools in the same binary: file read, write, and edit, bash or PowerShell, ripgrep search, web fetch, LSP diagnostics, Jupyter, browser use, background tasks, and parallel sub-agents.&lt;/p&gt;

&lt;p&gt;Autonomy comes with controls attached. Shell commands pass an AST safety classifier before they execute, with fine-grained permissions and audit logging on top. An auto flag approves safe tool calls on its own, gated by that classifier plus a circuit breaker, and the permission-mode flag accepts &lt;code&gt;default&lt;/code&gt;, &lt;code&gt;plan&lt;/code&gt;, or &lt;code&gt;bypassPermissions&lt;/code&gt;. Transcripts persist to disk, checkpoints let you resume, roll back, or fork a session, and long conversations auto-compact to save tokens. The README also describes an idle "dreaming" step that consolidates memory across sessions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading the benchmark in context
&lt;/h2&gt;

&lt;p&gt;The central claim rests on a head-to-head against Claude Code, and the README explicitly identifies its limits. Both agents ran through their own native clients on identical baselines, scored by pytest plus coverage, across 15 micro dimensions and 4 large-project dimensions, with methodology, raw results, and exact commands kept in the repo. The same section describes it as a small, self-run task set rather than a third-party benchmark, notes that it measures the agent loop and not raw model scores, and asks readers to treat it as a data point. FuXi currently has no published score on SWE-bench, Terminal-Bench, or the Aider polyglot benchmark.&lt;/p&gt;

&lt;p&gt;In place of headline numbers, the README hands you a checklist. Run &lt;code&gt;fuxi doctor&lt;/code&gt; to check config, API key, git, and ripgrep, then &lt;code&gt;fuxi verify&lt;/code&gt; to confirm the provider connection. Point FuXi at a failing test in one of your own projects, then push the identical task, model, and context through another tool and compare correctness, tool coverage, cost, and iteration time. The README says &lt;code&gt;/cost&lt;/code&gt;, &lt;code&gt;/usage&lt;/code&gt;, &lt;code&gt;/context&lt;/code&gt;, and &lt;code&gt;/status&lt;/code&gt; expose what you need for that comparison inside the TUI.&lt;/p&gt;

&lt;p&gt;One licensing detail before you adopt it: the README's license badge reads Proprietary, while the highlights describe the binary as free forever, with no license cost for individuals, teams, or enterprises. No license cost is a different claim from open source.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/fuxicodex/Fuxi" rel="noopener noreferrer"&gt;https://github.com/fuxicodex/Fuxi&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Curated by &lt;a href="https://www.agentpalisade.com" rel="noopener noreferrer"&gt;Agent Palisade&lt;/a&gt; — practical AI for small and mid-sized businesses.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aicoding</category>
      <category>terminal</category>
      <category>go</category>
      <category>llm</category>
    </item>
    <item>
      <title>Utopia keeps the facts it used to believe</title>
      <dc:creator>Reno Lu</dc:creator>
      <pubDate>Wed, 23 Sep 2026 17:35:23 +0000</pubDate>
      <link>https://dev.to/renolu/utopia-keeps-the-facts-it-used-to-believe-4nmf</link>
      <guid>https://dev.to/renolu/utopia-keeps-the-facts-it-used-to-believe-4nmf</guid>
      <description>&lt;p&gt;The interesting part of utopia is not that it builds a knowledge graph from your documents. It is that correcting a fact does not overwrite its old version. When a fact gets corrected, the old version is closed and the new one is linked to it, so the graph keeps a record of what the system believed before it changed its mind.&lt;/p&gt;

&lt;p&gt;That choice connects several of the project's core features, and it is the useful lens for judging whether utopia fits your problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two timelines in the graph
&lt;/h2&gt;

&lt;p&gt;The README calls the core structure a bitemporal knowledge graph. Extraction turns documents into entities and facts, guided by an ontology you can edit. Each fact carries two things beyond its content: when it held and where it came from. Corrections produce a chain of versions instead of a replaced row. The result is two timelines side by side: when something was true in the world, and when the system came to believe it.&lt;/p&gt;

&lt;p&gt;The authors frame this with the history of astronomy. Ptolemy's model was taken for truth for a long time before Copernicus, Kepler, Galileo and Newton falsified it step by step, and they say the part worth keeping is how that shift unfolded, not only the final answer. Vector stores and ordinary knowledge graphs, in their telling, work to get present knowledge right. Utopia aims to record the path.&lt;/p&gt;

&lt;p&gt;That shows up in the query tools. The built-in agent can ask for an entity's facts as of any date, or ask what changed in a given period. Edges are reified, so a relationship carries attributes of its own, which matters when the relationship itself is what changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  How contradictions are handled
&lt;/h2&gt;

&lt;p&gt;A store that keeps old beliefs needs rules for when new material disagrees. The README spells out three kinds of conflict. A new fact that clashes with an older one: close the old, keep both, or reject the new. Data that breaks an axiom, such as a self-loop, an asymmetry violation, a transitive cycle or a cardinality breach: retract the fact, relax the axiom, or accept both. The ontology itself gets checked first, since violations of a self-contradictory ontology are noise.&lt;/p&gt;

&lt;p&gt;Derived facts follow the same discipline. Ontology axioms compile into forward-chaining rules for transitivity, symmetry, inverses and relation hierarchy. Derivation is off by default because a wrong axiom derives wrong facts. When it is on, a derived fact is marked as derived, carries validity and confidence, shows what it came from, and yields to an asserted fact if the two contradict.&lt;/p&gt;

&lt;p&gt;Duplicates go through three stages: exact name or alias, embedding similarity, then a model's call on the doubtful pairs. Merges can be undone. Low-confidence extractions, suspected duplicates and cardinality conflicts land in a review queue, and the README says the decisions people make there are recorded and used to tune the agent.&lt;/p&gt;

&lt;p&gt;Confirming a fact, merging or reverting an entity, rebuilding the graph: each writes to an append-only decision ledger that records who acted, when, and what the object looked like at the time. A ledger record outlives its object, even the knowledge base it belonged to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Starting from an empty base
&lt;/h2&gt;

&lt;p&gt;A new knowledge base has no vocabulary of its own. You pick ontology packs at creation, and five ship inside the binary: schema.org, W3C Org, PROV-O, FOAF and IOF Core. Terms outside the packs are counted as they appear, and confirming the common ones adds them to the ontology.&lt;/p&gt;

&lt;p&gt;Material arrives as uploads (PDF, DOCX, PPTX, XLSX, XLS, ODS, CSV, TSV, Markdown, HTML or plain text) or scheduled syncs from web pages, RSS, GitHub, Jira, Notion, WebDAV and S3-compatible buckets. Search fuses Tantivy full-text and pgvector results with RRF, and chat answers carry inline citations that open the source passage. You can also mount a database such as Postgres, MySQL, Snowflake or Databricks; the agent proposes how its tables map onto the ontology and you confirm. Each knowledge base gets its own MCP server exposing the same read-only tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running it today
&lt;/h2&gt;

&lt;p&gt;The footprint is one Rust binary and one Postgres, with the job queue kept in a table. With Docker installed, clone the repo, bring it up with Docker Compose using the &lt;code&gt;app&lt;/code&gt; profile, open the app at localhost port 1516 and register. The first account becomes administrator, and a public knowledge base is created alongside it. Before extracting business documents, configure the chat and embedding endpoints in the administration models settings. The README says any OpenAI-compatible endpoint works, including Ollama and vLLM, which lets the whole system run air-gapped.&lt;/p&gt;

&lt;p&gt;Read the status section before committing to it. The project is at v0.1, schema migrations only roll forward, and the README asks you to pin the image version and back up the database and data directory before upgrading. Two limits matter for the temporal pitch specifically. A connector currently rounds timestamps to a UTC day, which can shift an event across midnight; finer precision sits on the roadmap. And the decision features the philosophy section leans on, recording a decision and replaying what was understood at the time, are marked as in development. The fact history exists now. Replaying decisions over it does not yet.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/deeplethe/utopia" rel="noopener noreferrer"&gt;https://github.com/deeplethe/utopia&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Curated by &lt;a href="https://www.agentpalisade.com" rel="noopener noreferrer"&gt;Agent Palisade&lt;/a&gt; — practical AI for small and mid-sized businesses.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>knowledgegraph</category>
      <category>rust</category>
      <category>ontology</category>
      <category>rag</category>
    </item>
    <item>
      <title>Cumora's agent coordination reads like a concurrency problem</title>
      <dc:creator>Reno Lu</dc:creator>
      <pubDate>Tue, 22 Sep 2026 17:26:25 +0000</pubDate>
      <link>https://dev.to/renolu/cumoras-agent-coordination-reads-like-a-concurrency-problem-ldf</link>
      <guid>https://dev.to/renolu/cumoras-agent-coordination-reads-like-a-concurrency-problem-ldf</guid>
      <description>&lt;p&gt;Cumora puts AI agents in a chat room next to humans, and the part of its README worth reading closely is how those agents avoid colliding. The server-side mechanisms it describes read like a database's answer to concurrent writers.&lt;/p&gt;

&lt;p&gt;Cumora is cross-platform team chat where agents sit on the same roster as humans, with the same DMs, group conversations, Kanban board, and calendar. The README is explicit that agents "don't just answer when poked." They hold personas and memory, claim work, and send and receive real email. The README describes several mechanisms for containing collisions between them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stale replies get held
&lt;/h2&gt;

&lt;p&gt;The coordination layer rests on three mechanisms the README names.&lt;/p&gt;

&lt;p&gt;The first is a seen-cursor freshness gate. If newer messages arrived in the room after the point an agent had seen, its reply is HELD, and the agent is shown the newer messages so it can decide again.&lt;/p&gt;

&lt;p&gt;The second is atomic claims on real units of work. The server arbitrates who holds each claim.&lt;/p&gt;

&lt;p&gt;The third is a small-brain triage gate that shields the big model. The config reflects the split directly: &lt;code&gt;OPENAI_MODEL&lt;/code&gt; and &lt;code&gt;OPENAI_MODEL_SUPPORT&lt;/code&gt; map to big-brain and support-brain models, and a CI guard, &lt;code&gt;npm run guard:big-brain&lt;/code&gt;, enforces the rule that only agent turns may use the big model. The README calls it a CI guard.&lt;/p&gt;

&lt;p&gt;Design notes live in &lt;code&gt;docs/COORDINATION.md&lt;/code&gt;, which the README describes as covering defense layers and anti-patterns. The repo also carries a &lt;code&gt;benchmarks/&lt;/code&gt; directory of real-LLM multi-agent coordination benchmarks, with scenarios named chain, counting, werewolf, and kanban.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same discipline in the backend
&lt;/h2&gt;

&lt;p&gt;The concurrency mindset is not limited to agents. The backend is a stateless Node service on Express and &lt;code&gt;ws&lt;/code&gt;, with Postgres as the source of truth and Redis for pub/sub fan-out and presence. Durable board, document, and calendar writes enqueue realtime invalidations in a transactional PostgreSQL outbox. Any number of instances can drain that outbox through leased &lt;code&gt;SKIP LOCKED&lt;/code&gt; claims.&lt;/p&gt;

&lt;p&gt;The README spells out the behavior during Redis degradation: Redis degradation delays live refresh but never changes the command result, and clients reconcile by pulling the API. Live updates ride on top of a correct write rather than standing in for it.&lt;/p&gt;

&lt;p&gt;Spend gets a single place to land too. Every LLM call, whether from a cloud agent or a local one, is recorded in one &lt;code&gt;llm_calls&lt;/code&gt; cost ledger.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two brains, one protocol
&lt;/h2&gt;

&lt;p&gt;Agents run on one of two paths. On Cumora Cloud, each agent gets a managed per-agent Kubernetes pod, orchestrated with &lt;code&gt;kubectl&lt;/code&gt; from the server, with a Go FUSE driver mounting its server-side workspace. Turns run a multi-hop tool-calling loop on the OpenAI Responses API, with tools covering bash, files, browser, email, memory, and skills.&lt;/p&gt;

&lt;p&gt;The other path is BYOA, Bring Your Own Agent. You pair a Mac or VPS with &lt;code&gt;npx cumora agent computer&lt;/code&gt; and the agent runs on your local provider account. The README lists Claude Code, Codex, Grok Build, Cursor Agent, OpenCode, pi, Gemini CLI, Qwen Code, Antigravity, and ZCode. Claude Code and Codex run inside fail-closed filesystem, command-network, and subprocess-credential boundaries by default. The other engines need an explicit unsandboxed compatibility opt-in. The server never sees your provider keys.&lt;/p&gt;

&lt;p&gt;Both paths act on the world through the same &lt;code&gt;cumora&lt;/code&gt; CLI protocol.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running it yourself
&lt;/h2&gt;

&lt;p&gt;Local setup needs Postgres, Redis, and an &lt;code&gt;OPENAI_API_KEY&lt;/code&gt;, which is the only hard-required variable. Run &lt;code&gt;npm run setup&lt;/code&gt;, then &lt;code&gt;npm run dev:all&lt;/code&gt; to start the Vite renderer on port 5180 and the API server on 5181, with migrations applied automatically. An empty database seeds a starter team of 6 agents, 3 humans, and 9 conversations with zero messages, so everything that shows up in chat is produced live.&lt;/p&gt;

&lt;p&gt;One warning from the README deserves attention before you trust a green run. Without &lt;code&gt;INTEGRATION_DATABASE_URL&lt;/code&gt;, the integration suite prints &lt;code&gt;[integration] skipped&lt;/code&gt; and exits 0, which looks like a pass. When it does run, it truncates every table, so point it at a throwaway database.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/yetone/cumora" rel="noopener noreferrer"&gt;https://github.com/yetone/cumora&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Curated by &lt;a href="https://www.agentpalisade.com" rel="noopener noreferrer"&gt;Agent Palisade&lt;/a&gt; — practical AI for small and mid-sized businesses.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>multiagent</category>
      <category>typescript</category>
      <category>opensource</category>
    </item>
    <item>
      <title>TrueForge keeps the agent loop out of the sandbox</title>
      <dc:creator>Reno Lu</dc:creator>
      <pubDate>Mon, 21 Sep 2026 18:23:08 +0000</pubDate>
      <link>https://dev.to/renolu/trueforge-keeps-the-agent-loop-out-of-the-sandbox-558a</link>
      <guid>https://dev.to/renolu/trueforge-keeps-the-agent-loop-out-of-the-sandbox-558a</guid>
      <description>&lt;p&gt;TrueForge makes one architectural call: the sandbox is a tool the agent reaches for, not the room the agent lives in. The harness holds the loop, the session state, and the secrets, and isolated code and file execution gets provisioned only when it is needed.&lt;/p&gt;

&lt;p&gt;It is also a useful lens for reading the rest of the feature list.&lt;/p&gt;

&lt;h2&gt;
  
  
  The harness is the product
&lt;/h2&gt;

&lt;p&gt;The README opens with a claim: building an agent is easy, running one well is not. It then spells out what running well takes. Streaming, session persistence, tool servers, sandboxing, approvals, and a UI. TrueForge positions itself as the runtime layer that owns that work, running the agent execution loop across model calls, MCP tools, skills, sandboxing, approvals, context management, and session state.&lt;/p&gt;

&lt;p&gt;Once the loop lives in the harness rather than inside an execution container, the split between control and execution becomes explicit. The README puts it plainly: secrets stay in the harness. The sandbox, which is Daytona today with more providers planned, becomes an execution target the harness brings up when a task calls for it.&lt;/p&gt;

&lt;p&gt;Skills follow the same pattern. They are git-backed &lt;code&gt;SKILL.md&lt;/code&gt; instruction packs, loaded on demand in the sandbox. The instructions show up where the work happens, at the point the work needs them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context engineering is built in
&lt;/h2&gt;

&lt;p&gt;TrueForge lists its context engineering features as subagents, deferred tool loading, Code Mode, large-result offloading, and compaction.&lt;/p&gt;

&lt;p&gt;You configure models, MCP servers, skills, and a sandbox once, and agents pick from what you connected. Presets come from shipped YAML catalogs you can customize. Model support covers OpenAI, Anthropic, Google Gemini, other catalog providers, and any OpenAI-compatible endpoint. Remote MCP servers can use header auth or OAuth, and authorization can happen inside the chat.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where a person steps in
&lt;/h2&gt;

&lt;p&gt;A runtime that holds the loop is also a sensible place to pause it. TrueForge ships human checkpoints: tool approval, ask-user-questions, and Generative UI in chat. These are configured when you create an agent, alongside the resources it can use.&lt;/p&gt;

&lt;p&gt;The surfaces sit on top of that runtime. There is a bundled chat UI. There is an HTTP API with a TypeScript SDK, &lt;code&gt;@truefoundry/trueforge-sdk&lt;/code&gt;, that covers sessions, turns, and events. And there is &lt;code&gt;@truefoundry/trueforge-ui&lt;/code&gt; for embedding the chat into your own product. The API reference is published as OpenAPI paths and schemas.&lt;/p&gt;

&lt;h2&gt;
  
  
  Local first, but read the warning
&lt;/h2&gt;

&lt;p&gt;A first run is one command, &lt;code&gt;npx @truefoundry/trueforge@latest&lt;/code&gt;, and the README badge lists Node.js 22.14 or later. That starts local mode: one process, backed by SQLite, with no extra infrastructure. Hosted mode moves storage to Postgres with Redis and runs on Docker Compose, Helm, or Railway, aimed at teams and multi-replica setups. Optional OIDC login is documented for shared deployments.&lt;/p&gt;

&lt;p&gt;The maintainers are direct about the limits of local mode. It has no login by default, data sits in a local SQLite file, and they ask users to keep it on localhost and use hosted mode for anything shared or production. If the separation of secrets from execution is what draws you to this project, keep in mind that the local setup is meant for trying it on your own machine, and the deployment story for a team runs through hosted mode.&lt;/p&gt;

&lt;p&gt;The project also publishes a comparison against Claude Managed Agents and deepagents on the same tasks, tools, and model, and reports the same accuracy at lower cost. The setup to reproduce it lives in the &lt;code&gt;benchmark/&lt;/code&gt; directory, so you can rerun it against your own expectations before leaning on the result. TrueForge is released under the MIT License.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/truefoundry/trueforge" rel="noopener noreferrer"&gt;https://github.com/truefoundry/trueforge&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Curated by &lt;a href="https://www.agentpalisade.com" rel="noopener noreferrer"&gt;Agent Palisade&lt;/a&gt; — practical AI for small and mid-sized businesses.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>mcp</category>
      <category>typescript</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
