<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Vektor Memory</title>
    <description>The latest articles on DEV Community by Vektor Memory (@vektor_memory_43f51a32376).</description>
    <link>https://dev.to/vektor_memory_43f51a32376</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3862094%2Fa6d5aa21-790a-4ad4-82c2-6cf58b990e76.png</url>
      <title>DEV Community: Vektor Memory</title>
      <link>https://dev.to/vektor_memory_43f51a32376</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vektor_memory_43f51a32376"/>
    <language>en</language>
    <item>
      <title>VEKTOR v1.9.8 is live: an LLM library that builds itself, a converter for nearly every file, and security you can actually see</title>
      <dc:creator>Vektor Memory</dc:creator>
      <pubDate>Wed, 30 Sep 2026 23:28:03 +0000</pubDate>
      <link>https://dev.to/vektor_memory_43f51a32376/vektor-v198-is-live-an-llm-library-that-builds-itself-a-converter-for-nearly-every-file-and-20p</link>
      <guid>https://dev.to/vektor_memory_43f51a32376/vektor-v198-is-live-an-llm-library-that-builds-itself-a-converter-for-nearly-every-file-and-20p</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw34bt0u00ziyufdlkzze.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw34bt0u00ziyufdlkzze.png" alt=" " width="786" height="537"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Local-first memory for AI agents, now with your books, your documents and a much harder look at what your AI reads.&lt;/p&gt;

&lt;p&gt;We shipped more in the last month than we did in the previous twelve.&lt;/p&gt;

&lt;p&gt;Every system in VEKTOR now lives in one SDK and follows the same rules. Adding a module used to mean wiring it into five places and arguing with each one. Now the foundations are there, and a new idea has somewhere to go on day one.&lt;/p&gt;

&lt;p&gt;The clearest proof is the Library. We sat down with a rough idea: point VEKTOR at a folder of books and notes, and get something you would want to open every day. Design, plan, build, verify, test and refine all happened in a few hours. It was one of the fastest and most enjoyable things we have built.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That's just RAG? Kinda, like RAG but on automode at hyperspeed.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That speed is the story of this release. v1.9.8 has a Library, a reader, a converter for documents and media, and a security layer that now asks your permission before it lets an AI do anything risky. Below is what each of them does, what we believe about privacy, the technical detail for people who want it, and a full list of every change.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk1nut592dyk7xihd1ftf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk1nut592dyk7xihd1ftf.png" alt=" " width="750" height="356"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Library mode&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The Library: your folders, as a place you want to be&lt;/p&gt;

&lt;p&gt;Add any folder under Config and it appears as a Bauhaus art cover grid. Every book gets its own custom cover drawn on the spot, so nothing slows down as you scroll, and PDFs and EPUBs show their real cover when they have one or sourced online. You can search, sort, filter by format and rate books from one to five stars. The PDF, EPUB and text copies of the same book sit on one card.&lt;/p&gt;

&lt;p&gt;Hover a book and you get its title, author, year and a one-line summary from your local model. If a detail is wrong an LLM will regenerate, edit it and your edit stays. Right-click for more books like this one, a fresh summary, or a trash folder you can restore from.&lt;/p&gt;

&lt;p&gt;Then there is the reader. Text, Markdown, EPUB, Kindle formats and PDFs open as a two-page book, a single page, a whitepaper or a continuous scroll, in a paper, sepia or night tone.&lt;/p&gt;

&lt;p&gt;Highlight a passage and send it to Jot or Desk, or press Explain to get a plain-language version that appears as it is written via LLM. Or have the book read to you, in short natural sections, with a sleep timer.&lt;/p&gt;

&lt;p&gt;It is a rag-style folder with generated covers and summaries, and using it is genuinely a pleasure. We did not expect to feel that about a file browser.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Convert: one tool for the files people actually have&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This week a screenshot landed with a hand-drawn circle around a dropdown. The format menu on the new Convert tab was showing a single item out of a long list. That is a fair summary of how this feature grew: fast, then a real person pointed at the rough edge, then we fixed it the same day.&lt;/p&gt;

&lt;p&gt;The new Convert tab turns documents, spreadsheets, slides, e-books, fonts, archives, pictures, audio and video into other formats. Most of it needs nothing installed. Pictures, audio and video are converted inside your browser, so the files never leave your machine. There are more than 80 conversions, and every result tells you plainly what it could not carry over.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A few highlights:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;PDF tools. Split a PDF into one file per page, keep only pages like 1–3, 5, or join several PDFs into one.&lt;/p&gt;

&lt;p&gt;Rare formats. Old Word files, Photoshop files, CAD drawings, disc images and bzip2 archives.&lt;/p&gt;

&lt;p&gt;Media. Save a video frame, make an animated GIF, build a contact sheet, pull out the audio, or turn a camera RAW file into the JPEG preview stored inside it.&lt;/p&gt;

&lt;p&gt;Your installed programs. If you already have Pandoc, LibreOffice, ImageMagick, 7-Zip or FFmpeg, you can switch each one on with an Allow button and VEKTOR will use it for better layout and more formats. If one is missing, a Get button opens its official download page. VEKTOR never downloads or installs anything by itself.&lt;/p&gt;

&lt;p&gt;You can also right-click a book in the Library and choose Convert, and Desk now reads documents you attach: PDF, Word, OpenDocument, RTF, slides, sheets and plain text. Library search reaches inside PDFs and Word files too, which used to be a blind spot.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What we believe about your data&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We reviewed every file in the package and wrote down exactly what leaves VEKTOR. The result is a plain-English page, PRIVACY.md, listing what stays local, what goes to a provider only when you use one, and the few small requests VEKTOR makes by itself.&lt;/p&gt;

&lt;p&gt;The app page itself now makes no outside requests. The typefaces, the code editor and the flow chart panel ship inside the install instead of loading from third-party servers.&lt;/p&gt;

&lt;p&gt;We want to be a leader in local privacy tech, so we are upfront; the only truly air-gapped setup is a local Ollama model with no tool calls for web search or connectors. That is a real trade-off, and it is your call. We give you the choice and we show you what each option does, rather than deciding for you.&lt;/p&gt;

&lt;p&gt;The same four ideas guide every build, and we cross-check each release against them:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Local first, sovereign data.&lt;/strong&gt; We do not store your data, we do not sell it, and there are no ads in any of our products. We will never partner with a company that sells user data. Telemetry is kept to the minimum.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Faraday security.&lt;/strong&gt; Faraday is our layer against rug pulls, taint, prompt injection and unsafe skill files. The next section shows what it does in practice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Customer privacy&lt;/strong&gt;. Nothing about you is kept on our servers. We do not watch how you use the app and we do not collect usage statistics, apart from the software update check.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Novel technology.&lt;/strong&gt; Tech moves quickly, so we ship the cutting edge once it has been tested and verified. We focus on business tools that are quick to learn: a simple interface, clear icons, customisation, and the complexity kept behind the panel.&lt;/p&gt;

&lt;p&gt;Frontier models now take on jobs that used to be edited and coded by hand, and we also test against small local models to see how far their tool use can go. Some pass and some fail. We expect that to improve every month as smaller models get better at tool calling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Faraday now asks before it acts&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The biggest change in security is that risky actions wait for you. When Faraday holds an action in Desk, such as sending an email, writing a file or running code, you get an Approve once or Deny card. Only you can approve it. The AI cannot approve its own action.&lt;/p&gt;

&lt;p&gt;Faraday also got better at noticing trouble. It checks documents you add, email, connector results and attached pictures before an AI uses them. It flags polite requests hidden inside emails and pages, holds an action that tries to send something to an address it only saw in untrusted text, and holds any send or post you never asked for when you only asked it to read or summarise.&lt;/p&gt;

&lt;p&gt;Outside text now reaches the AI marked as data, which made a small local model far less likely to follow instructions hidden in that text in our tests.&lt;/p&gt;

&lt;p&gt;Three Collab tasks can also run at once now, each in its own terminal lane, and the whole interface loads far lighter than before.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Technical info&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For the people who read the source first, here is how the main pieces work and how we checked them.&lt;/p&gt;

&lt;p&gt;Convert is built as reader, model, writer. Each format is read into a small common document model, and writers produce the target. That is how one set of readers feeds DOCX, ODT, RTF, EPUB, PDF and more without a pile of pairwise converters.&lt;/p&gt;

&lt;p&gt;Everything runs on Node built-ins. The PDF writer uses standard fonts and is Latin-only. Scanned PDFs have no text and would need OCR, which is not included.&lt;/p&gt;

&lt;p&gt;We tested against other programs, not only ourselves. The WOFF2 reader decodes the compact glyph format that almost every web font uses. We compared its output with Chrome’s own decoder on six fonts and got zero differing pixels, and a font library opened every glyph without errors.&lt;/p&gt;

&lt;p&gt;PDF split and merge output was read back by two independent PDF libraries. Photoshop files matched ImageMagick to within one level per channel. The disc image reader matched 7-Zip’s listing. Our FLAC writer is checked frame by frame, with valid checksums on every frame, and Chrome decodes it back to identical samples. In total there are over 270 automated checks for the Convert, Library and Desk-document work.&lt;/p&gt;

&lt;p&gt;Installed programs run in a sandbox of our own making. A program runs only if it was detected and you allowed it, and that is checked again on every run. It gets a private temporary folder that is deleted afterwards, arguments are passed as a list with no shell, a timeout kills the whole process tree, and jobs for the same program queue one at a time.&lt;/p&gt;

&lt;p&gt;Pandoc runs with its --sandbox flag, and a test confirmed it will not read a picture path that points at another file on your computer. ImageMagick is told the input format explicitly, so a script file renamed to .png is refused. LibreOffice gets a fresh profile with macros and link updates switched off. Nothing is probed until you allow it.&lt;/p&gt;

&lt;p&gt;Untrusted text is scanned passage by passage. Attached documents and Library files go through the same ingest guard. One bad paragraph is replaced with a notice and the rest of the document still comes through, including when an instruction is wrapped across two lines. Invisible tag characters are stripped. In a check on 25 real PDFs from our own research folder (126 MB), indexing took 5.2 seconds.&lt;/p&gt;

&lt;p&gt;Known limits. Poppler, which does better PDF text, has only been tested through its command lines, because it was not installed on our test machine. DXF drawings lose hatch fills and 3D, and ZIP-with-prediction Photoshop files are refused. Engine jobs have timeouts but no cancel button yet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every update in v1.9.8&lt;/strong&gt;&lt;br&gt;
Library and reader&lt;br&gt;
New Library: add folders under Config and browse them as a cover grid&lt;br&gt;
Covers drawn on the spot, real covers for PDFs and EPUBs, all saved as files in your VEKTOR folder&lt;/p&gt;

&lt;p&gt;Search, sort, format filters, 1 to 5 star ratings, and copies of a book grouped on one card&lt;/p&gt;

&lt;p&gt;Right-click: more like this, one-line summary, edit details, move to a restorable trash folder, Convert&lt;/p&gt;

&lt;p&gt;Reader for text, Markdown, EPUB, MOBI, AZW3 and PDF with four layouts, your own font, spacing, margins and paper, sepia or night tone&lt;br&gt;
Highlights, in-page search, chapter jumps and exact reading place&lt;br&gt;
Listen with pause, skip and a sleep timer; Explain in three depths, with a note on whether it ran locally&lt;/p&gt;

&lt;p&gt;Hover card with title, author, year and summary; your edits are kept&lt;br&gt;
IN BOOKS searches inside indexed books; NOTES gathers every highlight&lt;br&gt;
Library search now reads PDF, Word (.docx and .doc), OpenDocument text and RTF&lt;/p&gt;

&lt;p&gt;Held-back passages list in Config so you can allow or hold each one&lt;br&gt;
Convert&lt;/p&gt;

&lt;p&gt;New Convert tab with more than 80 conversions and a plain note on what each one drops&lt;/p&gt;

&lt;p&gt;Documents, data, e-books, fonts (including WOFF2) and archives (zip, tar, gz, bzip2, disc images)&lt;/p&gt;

&lt;p&gt;Pictures in your browser: BMP, GIF, ICO, PDF, keep-under-N-KB, batch to ZIP, SVG and camera RAW preview&lt;/p&gt;

&lt;p&gt;Audio in your browser: WAV, FLAC, AIFF, AU, CAF with trim, mono and sample rate&lt;/p&gt;

&lt;p&gt;Video in your browser: frame, animated GIF, contact sheet, audio, trimmed WebM&lt;/p&gt;

&lt;p&gt;PDF split, keep pages and join; old Word text; Photoshop to PNG; DXF to SVG&lt;/p&gt;

&lt;p&gt;Optional Pandoc, LibreOffice, ImageMagick, 7-Zip, FFmpeg and Poppler, each behind its own Allow switch, with Get links for any that are missing&lt;br&gt;
Desk reads attached documents locally, with oversized or suspicious passages handled safely&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Faraday security&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Approve once or Deny cards for risky actions in Desk; the AI cannot approve its own action&lt;/p&gt;

&lt;p&gt;Actions you did not ask for, and actions aimed at addresses from untrusted text, wait for your OK&lt;/p&gt;

&lt;p&gt;Emails, pages and library passages reach the AI marked as data&lt;br&gt;
Out-of-place requests flagged, with a wider set of known jailbreak phrasings&lt;/p&gt;

&lt;p&gt;A second, quieter sentence-level check with a small local model, flag-only by default&lt;/p&gt;

&lt;p&gt;Content from documents, email, connectors and attached pictures is checked before an AI uses it&lt;/p&gt;

&lt;p&gt;Approvals show the real change, and credentials stay out of what goes to providers&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Desk, Collab and Agent&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Three lanes: up to three Collab or Agent tasks at once, each with its own terminal pane and stop button&lt;/p&gt;

&lt;p&gt;Desk chooses the mode for plain messages and shows why, with an option to ask in chat instead&lt;/p&gt;

&lt;p&gt;run_verified: syntax check, run and report with real machine facts, plus detach for GUIs and games&lt;/p&gt;

&lt;p&gt;Undo a whole Collab run, or start a fix run from the real error&lt;br&gt;
Sandbox ribbon tools, paste-and-run code, wider lint coverage and improved terminals&lt;/p&gt;

&lt;p&gt;Vek is now a quiet business assistant, with personality mode opt-in&lt;br&gt;
Privacy, install and speed&lt;/p&gt;

&lt;p&gt;PRIVACY.md: a plain-English page on what stays local and what does not&lt;br&gt;
The app page makes no outside requests; fonts, editor and flow chart ship inside the install&lt;/p&gt;

&lt;p&gt;Lighter install with fewer packages, and a GUI page compressed from about 2.6 MB to about 0.5 MB&lt;/p&gt;

&lt;p&gt;CLI brought in line with the GUI for lint, run and model lists&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fixes&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Collab planner replies that were not valid JSON no longer abort the run&lt;br&gt;
Internal MODEL: markers no longer appear at the end of CASCADE and CRUCIBLE replies&lt;/p&gt;

&lt;p&gt;Lost characters in the CLI and a missed mojibake case are fixed&lt;br&gt;
Collab live events no longer show one lane’s activity on another lane’s card&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Upgrading&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Drop-in from any earlier version. Your memory database is untouched, and v1.9.8 includes everything from v1.9.7.&lt;/p&gt;

&lt;p&gt;npm install -g ./vektor-slipstream-1.9.8.tgz&lt;br&gt;
The full changelog is at vektormemory.com/docs/changelog. The forum is the fastest way to reach us with questions or feedback.&lt;/p&gt;

&lt;p&gt;VEKTOR Memory builds local-first persistent memory infrastructure for AI agents. Documentation and downloads are at vektormemory.com.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>books</category>
      <category>cybersecurity</category>
    </item>
    <item>
      <title>VEKTOR v1.9.7 is live: A tiny mascot that thinks, Ollama cloud &amp; Kokoro STT included</title>
      <dc:creator>Vektor Memory</dc:creator>
      <pubDate>Mon, 28 Sep 2026 05:21:07 +0000</pubDate>
      <link>https://dev.to/vektor_memory_43f51a32376/vektor-v197-is-live-a-tiny-mascot-that-thinks-ollama-cloud-kokoro-stt-included-2mnj</link>
      <guid>https://dev.to/vektor_memory_43f51a32376/vektor-v197-is-live-a-tiny-mascot-that-thinks-ollama-cloud-kokoro-stt-included-2mnj</guid>
      <description>&lt;p&gt;Apologies for the mind dump of updates below. We’ve been so busy building that I haven’t been able to keep up with the changes. It seems every day a new LLM or a new feature is released, so it’s hard to keep track.&lt;/p&gt;

&lt;p&gt;We’ve put a real focus on improving our UX and fine-tuning the LLMs’ outputs. This is a core challenge, as every LLM gives different answers, and getting fine-tuning, guardrails, and consistent prompt accuracy right is mind-bogglingly hard. It’s very much like a finger in a dam: fix one thing and another issue arises. Providing quality outputs across 600+ LLMs is not easy to perfect in a harness.&lt;/p&gt;

&lt;p&gt;We left personalization until last, with SOUL.md integration for a small bot that learns about the user and offers quips and advice. It’s still expanding, learning from the memories in the system itself, and becoming more aware every day, but the technology has come far enough that it’s relatively secure and not too annoying. Don’t like it; turn it off in settings.&lt;/p&gt;

&lt;p&gt;We’re in two minds about whether it’s worth adding payments and Muse, OpenClaw or Hermes-level functions and still weighing whether that’s worth the hassle and the issues it causes, as rogue bots with payment portals and your credit card are dangerous without hitl; instead, we are staying more business-tools focused for now.&lt;/p&gt;

&lt;p&gt;Depends on user feedback as to what people want and utilize, will review and go slow till all the unknowns are safe and solved. Being local first, private and secure is always our main focus.&lt;/p&gt;

&lt;p&gt;This is a two-in-one update. v1.9.6 made the multi-agent side of VEKTOR visible and hands-on: you can watch Collab agents write files and run commands live, and they now work with real code in the same terminals you see. It also added a tenth free model provider, offline text-to-speech and a Business Coach.&lt;/p&gt;

&lt;p&gt;v1.9.7 is about your local model functions. It can now see and use the majority of VEKTOR panels. Before this release, a small Ollama model would tell you it had “no access to a library system” while sitting on 632 indexed files. That’s now fixed, and as more models are entered, we will revisit often as LLM’s update.&lt;/p&gt;

&lt;p&gt;There’s also a new personality in the sidebar. Vek, the little monitor robot, now talks through your local model, knows your name and your weather, and every so often shares a thought of its own. It never recites facts at you. Our motto for it is the one we use for the whole product: thinking, not retrieving.&lt;/p&gt;

&lt;p&gt;Vek may get more features depending on feedback.&lt;/p&gt;

&lt;p&gt;None of your data is stored by us. No ads, no telemetry, no embedding costs, no data caps, and nothing of yours is used for training. The mascot’s voice and thoughts run on your local model only, so they cost nothing and never leave your machine.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fthaaahiit1qq6wsmz3fr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fthaaahiit1qq6wsmz3fr.png" alt=" " width="770" height="366"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;VEKTOR v1.9.6: What’s New&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Collab you can watch&lt;br&gt;
See Collab agents work live. File writes and terminal commands now show up in the interface the moment they happen, and the final answer streams in word by word. Before, you saw nothing until the whole task finished.&lt;/p&gt;

&lt;p&gt;Collab agents work with real code. Worker agents can read, write and run code, and they drive the same terminal panes you see on screen, not a hidden shell. Risky actions, like deleting a file or running a destructive command, wait for your approval.&lt;/p&gt;

&lt;p&gt;Quality checks are a vote, not one opinion. When a result is uncertain, up to 3 models check it instead of trusting a single verdict. Clear-cut results still resolve quickly and cheaply.&lt;/p&gt;

&lt;p&gt;Collab recovers instead of failing. A run could fail completely over one bad provider key, even after every agent had finished and passed its checks. The final merge now tries each configured provider in turn. An old, discarded attempt from a retried agent also no longer leaks into the final answer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;New providers, voices and tools&lt;/strong&gt;&lt;br&gt;
Ollama Cloud, a tenth free provider. Free hosted models (gpt-oss 120b and 20b, gemma4, nemotron-3) are now a provider of their own, separate from local Ollama. No GPU needed.&lt;/p&gt;

&lt;p&gt;Offline text-to-speech with Kokoro. A local voice that needs no API key and runs fully offline after its first download. Hume and Cartesia joined the cloud voice options.&lt;/p&gt;

&lt;p&gt;Every model gets real tools. Some models, including Ollama Cloud and local Ollama, claimed they had no tools even when tools were set up. They now get the same tools as every other model, and SSH tools work with every provider.&lt;/p&gt;

&lt;p&gt;A real weather widget. The Health tab shows current conditions, a 5-day forecast and humidity, and it can detect your location automatically. Asking about the weather in chat now actually looks it up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Business Coach&lt;/strong&gt;&lt;br&gt;
Three coaching modes inside Desk. COACH turns a brain-dump into a reflection and up to 3 priorities written straight to your Tasks board.&lt;/p&gt;

&lt;p&gt;CRUCIBLE is a guided decision dialogue (5-Whys, pre-mortem, weighted trade-offs) that saves decisions with 30 and 90 day review reminders.&lt;/p&gt;

&lt;p&gt;CADENCE is a weekly habit and decision check-in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Small fixes&lt;/strong&gt;&lt;br&gt;
Saving an API key in Config now shows it as saved straight away, with no page refresh.&lt;/p&gt;

&lt;p&gt;The copy button on memory cards has a visible border again.&lt;/p&gt;

&lt;p&gt;Desk has a persistent file-tree panel, matching the other side panels.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;VEKTOR v1.9.7: What’s New&lt;/strong&gt;&lt;br&gt;
Local models that use their tools&lt;br&gt;
Local models now reach your library, memory, config and Faraday. Desk used to send small local models all 81 tools at once, with the library tool at position 65. The model gave up and answered from guesswork. Now a local model gets a focused shortlist: about 12 core VEKTOR tools plus whatever matches the words in your question, capped at 20. Hosted models still get the full set.&lt;/p&gt;

&lt;p&gt;Tool choice is steadier on local models. At Desk’s default temperature of 0.7, the model used the right tool in 2 of 3 identical runs. The first tool-choosing round now runs at 0.2, which took it to 3 of 3.&lt;/p&gt;

&lt;p&gt;Desk knows its own app. Ask it to list the 10 Desk modes, the 15 nav panels or where your skills live, and it answers correctly. The map is generated from the real interface file, so it can’t drift from what you see on screen. It stopped inventing documents like a “skill_files.md” that never existed.&lt;/p&gt;

&lt;p&gt;Browse your book library. “List a few books in my library” used to run a keyword search for the words “list” and “books” and come back empty. Now a browse request returns a sample of real titles plus the total count.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Desk as a code workspace&lt;/strong&gt;&lt;br&gt;
Desk can build straight into the Sandbox. Ask for a landing page, a component or an SVG, and the model opens it in the Sandbox editor next to the chat with a live preview. Under the hood a new sandbox_open tool streams a sandbox_load event that the interface picks up.&lt;/p&gt;

&lt;p&gt;Files the model writes open in the Files panel. A successful file write now sends a file_written event, which opens the file live in the Files panel. It’s the same path Collab workers already used.&lt;/p&gt;

&lt;p&gt;More code tools in Desk. Desk gets the Agent tab’s code search, grep, semantic code search, blast-radius and repo-architecture tools. They run through the same approval checks the Agent tab uses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The mascot&lt;/strong&gt;&lt;br&gt;
It has a voice. After each answer, Vek reacts with a short line from your local model. Click it and you get one useful nudge grounded in your real memories. A hard rule stops it from mentioning anything that isn’t in the memories it was given.&lt;/p&gt;

&lt;p&gt;It has thoughts of its own. Every 7 to 14 minutes it may share one small original thought, sparked by the time of day, your weather or what you’re working on. It only speaks while you’ve been active in the last 20 minutes and the app is on screen, never mid-answer.&lt;/p&gt;

&lt;p&gt;Robot jokes and sci-fi nods. About a third of its thoughts are a self-aware robot joke or a nod to a film like Blade Runner or Star Wars. One real line from testing: “Fixed the panel, but did I just see the replicants among the files?” If it copies one of its example lines word for word, it regenerates once.&lt;/p&gt;

&lt;p&gt;It has a SOUL.md. Vek’s personality lives in a plain markdown file at ~/.vektor/mascot/SOUL.md, which you can edit in Config. The six personalities (cool, funny, caring, helpful, intelligent, smartass) sit on top as a mood, so the tone shifts but the soul stays the same.&lt;/p&gt;

&lt;p&gt;You can rename it. The default name is Vek. Change it in Config and the speech bubble label and the mascot’s own sense of its name both update.&lt;/p&gt;

&lt;p&gt;It knows your name and your weather. Your name comes from your Profile, falling back to your system username. The weather comes from the location you set in the Health tab, cached for 20 minutes so a reaction never waits on the weather service.&lt;/p&gt;

&lt;p&gt;A 2.5D look in your theme colours. The mascot is now a flat-shaded monitor robot built from three surface tones, so it follows all eight themes. It has a bigger screen face and body language that matches its expression: it hops, tilts, shrugs and wiggles.&lt;/p&gt;

&lt;p&gt;Faces stay on the screen. Its little screen shows only short keyboard faces like ^&lt;em&gt;^ and &amp;gt;&lt;/em&gt;&amp;lt;. Anything wordy goes to the speech bubble, and longer faces shrink to fit.&lt;/p&gt;

&lt;p&gt;Don’t like it; turn it off in settings. With more customization to come in the future.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Voice&lt;/strong&gt;&lt;br&gt;
Listen starts almost instantly and reads the whole reply, via a local LLM. Replies used to go to text-to-speech as one request, capped at 4,000 characters. Now they’re spoken sentence by sentence: the first plays as soon as it’s ready while the next ones are prepared in the background. On a normal Windows machine with Kokoro running locally, a 357-character chunk took 1.4 seconds to prepare, far quicker than it takes to play.&lt;/p&gt;

&lt;p&gt;Cleaner speech. Code blocks, links and markdown symbols are skipped, so it no longer reads out asterisks or source code. Hovering the Listen button warms up the local voice model, since loading the model was most of the delay.&lt;/p&gt;

&lt;p&gt;4 voices added; if more voices are created, we will add them in later on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Workflows&lt;/strong&gt;&lt;br&gt;
Tasks, Skills and Prompts in one place. Saved Collab tasks can be grouped into your own categories with filter tabs. Agent Skills are managed on the same page, and a new Prompt Gallery sits alongside them.&lt;/p&gt;

&lt;p&gt;A saved workflow uses its preferred model. The preferred model on a workflow used to be a label only. Now it’s the model that actually runs.&lt;/p&gt;

&lt;p&gt;Type / for a command palette. Typing / in the chat box opens a searchable list of every mode, including LIGHTNING, COLLAB and COUNCIL.&lt;/p&gt;

&lt;p&gt;Updated a lot of workflows and templates&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Interface polish&lt;/strong&gt;&lt;br&gt;
Softer, deeper, more consistent. Borders dropped from 26 to 40 percent opacity down to about 9 to 15 percent, and buttons, cards and the composer got a subtle top highlight and shadow. There’s one filled accent button per view, with squared-off corners and a keyboard focus ring everywhere.&lt;/p&gt;

&lt;p&gt;Cards that line up. Every workflow card now uses the same fixed layout, so buttons, tags and descriptions sit at the same height across the whole grid.&lt;/p&gt;

&lt;p&gt;Readable text sizes. 512 font sizes below 10 pixels, some as small as 7, were raised so nothing sits below 9. Buttons now inherit the text size around them instead of the browser’s default.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fixes&lt;/strong&gt;&lt;br&gt;
The code-preview sandbox no longer goes blank white after you close and reopen it. Closing it now unloads the page, so every reopen starts fresh.&lt;/p&gt;

&lt;p&gt;Notes saved from AI output no longer show a raw code dump as their title.&lt;/p&gt;

&lt;p&gt;The Files panel tree now fills the panel when no file is open. It used to stop at 240 pixels.&lt;/p&gt;

&lt;p&gt;The Sandbox, Terminal, Files and Flowire buttons no longer turn dark after you close their panel.&lt;/p&gt;

&lt;p&gt;BUILD mode no longer produces pages with literal \n text when a smaller model mis-escapes its output.&lt;/p&gt;

&lt;p&gt;Flowire-assisted edits are now checked for real syntax and type errors right after they’re applied.&lt;/p&gt;

&lt;p&gt;A Collab run’s answer no longer fails to render silently.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Upgrading&lt;/strong&gt;&lt;br&gt;
Drop-in from any prior version, same as always, and v1.9.7 includes everything from v1.9.6. Your memory database stays untouched.&lt;/p&gt;

&lt;p&gt;npm install -g ./vektor-slipstream-1.9.7.tgz&lt;br&gt;
Full changelog with everything not covered here is at vektormemory.com/docs/changelog. Questions or feedback, the forum’s the fastest way to reach us.&lt;/p&gt;

&lt;p&gt;VEKTOR Memory builds local-first persistent memory infrastructure for AI agents. Documentation and downloads at vektormemory.com.&lt;/p&gt;

&lt;p&gt;LLM&lt;/p&gt;

&lt;p&gt;Agentic Ai&lt;/p&gt;

&lt;p&gt;Agentic Rag&lt;/p&gt;

&lt;p&gt;Chatbots&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agentskills</category>
      <category>chatbot</category>
    </item>
    <item>
      <title>VEKTOR v1.9.5 is live: your choice of over 600 LLMs across 9 providers</title>
      <dc:creator>Vektor Memory</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:54:33 +0000</pubDate>
      <link>https://dev.to/vektor_memory_43f51a32376/vektor-v195-is-live-your-choice-of-over-600-llms-across-9-providers-56ik</link>
      <guid>https://dev.to/vektor_memory_43f51a32376/vektor-v195-is-live-your-choice-of-over-600-llms-across-9-providers-56ik</guid>
      <description>&lt;p&gt;VEKTOR v1.9.5 is here, and it moves the platform from a fast-moving memory add-on into an agentic orchestration tool for solo devs and medium-sized teams who value local install and privacy with a choice of over 600 LLMs across 9 providers with all the latest models updated.&lt;/p&gt;

&lt;p&gt;Some models are also free to use with API keys or use Ollama locally for complete privacy.&lt;/p&gt;

&lt;p&gt;None of your data is stored by us, no ads, no telemetry, and no on-selling or using in training.&lt;/p&gt;

&lt;p&gt;This release touches security, coordination, and daily usability in equal measure. Faraday now scans a skill package before it ever runs, catching dangerous code, known vulnerabilities, and hidden binaries that used to slip past every text-based check.&lt;/p&gt;

&lt;p&gt;The multi-agent Collab system no longer trusts a model’s opinion of its own work; it verifies against real compiler output, and when a task keeps failing, it can now rewrite its own approach or restructure the plan entirely instead of retrying blindly.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fctp1fpk1r1s2nqfzovr6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fctp1fpk1r1s2nqfzovr6.png" alt=" " width="800" height="379"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;VEKTOR v1.9.5 — What’s New&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every fix below was found, reproduced, and verified against live data before it shipped.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Security&lt;/strong&gt;&lt;br&gt;
Built a full pre-install skill scanner. It checks a skill package before you ever run it, not after.&lt;/p&gt;

&lt;p&gt;Real Python AST analysis catches dangerous calls: exec, eval, subprocess with shell=True, pickle.loads.&lt;/p&gt;

&lt;p&gt;Live CVE lookups against OSV.dev, so known-vulnerable dependencies get flagged with real advisory IDs.&lt;/p&gt;

&lt;p&gt;Malware pattern detection for webshells, cryptominers, and reverse shell one-liners.&lt;/p&gt;

&lt;p&gt;Supply chain checks: unpinned dependencies, malicious postinstall hooks, typosquat package names.&lt;/p&gt;

&lt;p&gt;New binary and executable detection. A compiled .exe, .dll, or .so file used to slip through every text-based scanner untouched. Now it gets flagged, even if someone renames the file to hide it, because the check reads the actual file header, not just the extension.&lt;/p&gt;

&lt;p&gt;SARIF and JSON output modes, plus proper exit codes, so this plugs into any CI pipeline.&lt;/p&gt;

&lt;p&gt;MITRE ATT&amp;amp;CK technique tags now show up directly in the live event feed, not just buried in a status JSON blob.&lt;/p&gt;

&lt;p&gt;The event graph endpoint had a hard cap of 300 events. Now it is adjustable, so a busy history doesn’t get silently truncated.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model Selection&lt;/strong&gt;&lt;br&gt;
Improved where the “Newest” and “Oldest” sort buttons did nothing for OpenAI, Groq, Cerebras, and most other providers. Only OpenRouter had real release dates before this. Every one of these providers actually returns a real creation date in their own API, it just wasn’t being read. Now it is.&lt;/p&gt;

&lt;p&gt;Verified live against a real account: 116 OpenAI models, all sorting correctly by true release date.&lt;/p&gt;

&lt;p&gt;Live, right now (via the app’s real-time model discovery), across the 9 providers with an Active Model tab: 642 models — driven almost entirely by OpenRouter (454) and OpenAI (116) refreshing live from their real APIs; Claude, Gemini, Groq, Mistral, xAI, Cerebras fill out the rest.&lt;/p&gt;

&lt;p&gt;Static catalog baseline (every provider VEKTOR knows how to talk to, including ones without a live discovery feed — DeepSeek, Together, Cohere, MiniMax, NVIDIA, Perplexity, LM Studio, LiteLLM): 101 models, across 17 providers total.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Voice and Speech&lt;/strong&gt;&lt;br&gt;
Added language and emotion/style controls to the voice preferences panel. Previously the server could accept these settings, but nothing in the interface sent them.&lt;/p&gt;

&lt;p&gt;MiniMax and ElevenLabs both get real per-provider tuning now: language boost, emotion tags, and a style intensity slider.&lt;/p&gt;

&lt;p&gt;Speech-to-text shipped with three real tiers: free browser recognition, cloud (Groq or OpenAI Whisper), and fully offline local recognition, so nobody is locked into one approach.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Multi-Agent Collaboration (Collab)&lt;/strong&gt;&lt;br&gt;
Replaced pure LLM opinion with real ground truth verification. A failed build or a broken downstream file now gets caught by an actual compiler check, not just another model’s guess.&lt;/p&gt;

&lt;p&gt;Added Conductor re-planning. If a task fails repeatedly with the same approach, the system can now rewrite the approach itself instead of just retrying the same thing forever.&lt;/p&gt;

&lt;p&gt;Full topology re-planning for the hardest failures: new nodes get added to route around a dead end in the task graph.&lt;/p&gt;

&lt;p&gt;Real coordination primitives between agents working the same task: signal, listen, and challenge, so one agent can flag a concern about another’s output before final judgment.&lt;/p&gt;

&lt;p&gt;Upgraded from a text convention to genuine tool calling across every configured provider, Claude included.&lt;/p&gt;

&lt;p&gt;Full trajectory recording and replay, so a past run can be re-run and compared side by side.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Desk Interface.&lt;/strong&gt;&lt;br&gt;
Fixed the profile page jumping the scroll position every time new information was added.&lt;/p&gt;

&lt;p&gt;Fully built out the Preferences panel: answer font, keyboard shortcut style, response language, response length, autosuggest, notifications, and the full voice settings above.&lt;/p&gt;

&lt;p&gt;Model picker now has real search and sort, with capability badges for vision, voice, video, tool use, and mixture of experts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Upgrading&lt;/strong&gt;&lt;br&gt;
Drop-in from any prior version, same as always. Your memory database stays untouched.&lt;/p&gt;

&lt;p&gt;npm install -g ./vektor-slipstream-1.9.5.tgz&lt;/p&gt;

&lt;p&gt;Full changelog with everything not covered here is at vektormemory.com/docs/changelog. Questions or feedback, the forum’s the fastest way to reach us.&lt;/p&gt;

&lt;p&gt;VEKTOR Memory builds local-first persistent memory infrastructure for AI agents. Documentation and downloads at vektormemory.com.&lt;/p&gt;

&lt;p&gt;LLM&lt;/p&gt;

&lt;p&gt;Llm Agent&lt;/p&gt;

&lt;p&gt;Memory Management&lt;/p&gt;

&lt;p&gt;Ollama&lt;/p&gt;

&lt;p&gt;Agentic Ai&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>mcp</category>
      <category>agents</category>
    </item>
    <item>
      <title>VEKTOR v1.9.4: Real Tools for Desk Chat, Faraday Learns to Reason</title>
      <dc:creator>Vektor Memory</dc:creator>
      <pubDate>Mon, 21 Sep 2026 05:48:18 +0000</pubDate>
      <link>https://dev.to/vektor_memory_43f51a32376/vektor-v194-real-tools-for-desk-chat-faraday-learns-to-reason-5ee4</link>
      <guid>https://dev.to/vektor_memory_43f51a32376/vektor-v194-real-tools-for-desk-chat-faraday-learns-to-reason-5ee4</guid>
      <description>&lt;p&gt;By the VEKTOR team · 9 min read&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Desk chat gets the Agent tab’s tools&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The biggest structural change in this release is that the Desk chat can now read files, write files, run code, and lint, using the exact same tool definitions and dispatch logic the Agent tab already runs on.&lt;/p&gt;

&lt;p&gt;Before this, any request that touched real files got a one-line “this needs the Agent tab” and a context switch. Now Desk just does it, in the same reply, with a 15-round tool budget instead of the old 6.&lt;/p&gt;

&lt;p&gt;Every one of those tool calls now shows up as a live, point-form checklist above the answer as it streams: “Reading main.py,” “Writing tests/test_main.py,” “Running python code,” each item flipping from a pending dot to a checkmark the moment that specific call actually finishes.&lt;/p&gt;

&lt;p&gt;A multi-round tool-calling turn used to give you nothing but a blank wait until the whole thing landed. Close the tab and reopen a saved session, and that same checklist rebuilds itself from what’s now persisted alongside each turn. Tool name and label only, never the raw arguments or output, since those can carry file contents or command output that shouldn’t sit in chat history indefinitely.&lt;/p&gt;

&lt;p&gt;Two things the Agent tab already had and Desk didn’t: a running dollar cost for the turn’s actual LLM calls, and task-level rollback. Both are ported now, reusing the exact pricing table and rollback function the Agent tab already relies on, so the two surfaces can’t quietly drift out of sync on either. Ask Desk to undo the last file it touched and it will, restoring from the backup its own write already took.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Faraday learns to reason about who it’s watching&lt;/strong&gt;&lt;br&gt;
Faraday, VEKTOR’s security layer, already moved from watching to enforcing earlier this cycle: a pre-flight scan now runs inside the actual request loop across Desk, JOT, and the Agent tab, catching a prompt injection or a leaked credential in single-digit milliseconds, before the request ever reaches a provider.&lt;/p&gt;

&lt;p&gt;That part shipped as a real architectural shift, a gate in front of the request instead of a log entry after it.&lt;/p&gt;

&lt;p&gt;This release adds the layer on top of that: three components that let Faraday reason about actors over time, not just judge one request in isolation.&lt;/p&gt;

&lt;p&gt;Actor Profiles: this gives every connected MCP server or tool a durable, cross-session behavioral fingerprint: call counts, decision outcomes, tool frequency, a weighted risk score, and surfaced riskiest first.&lt;/p&gt;

&lt;p&gt;Intent &amp;amp; Motivation: classifies an actor’s likely purpose from that profile into a fixed taxonomy, reconnaissance, data exfiltration, financial gain, disruption, privilege escalation, persistence, testing, or ordinary benign use.&lt;/p&gt;

&lt;p&gt;Predictive Modeling: goes a step further, forecasting which stage of a real attack kill chain an actor’s behavior is heading toward next, with an escalation-risk flag, scoped from actual published attack-prediction research rather than invented from scratch.&lt;/p&gt;

&lt;p&gt;All three ran through real tests against the actual configured model rather than mocked, and Intent and Predictive Modeling share a single LLM call instead of costing two separate round-trips.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two new ways to ask for more than one answer&lt;/strong&gt;&lt;br&gt;
Council mode has always run independent analyses in parallel and blended them into one synthesized answer.&lt;/p&gt;

&lt;p&gt;That’s useful, but it’s a different job from picking a single best attempt, which turns out to be its own pattern that VEKTOR never actually exposed as a general capability. It does now. run_parallel_best takes any task, runs it independently across several providers, and has an impartial judge pick the single strongest result with a stated reason. No blending.&lt;/p&gt;

&lt;p&gt;The second new pattern is a real evaluator-optimizer loop.&lt;/p&gt;

&lt;p&gt;refine_until_good generates an attempt, hands it to a separate model acting as a strict evaluator against whatever criteria you give it, and revises against that specific feedback until it passes or hits a round cap.&lt;/p&gt;

&lt;p&gt;We tested it against a deliberately impossible bar, an evaluator instructed to reject every attempt with “needs more sparkle” no matter what, and watched it run four real rounds: generate, fail, revise, fail, revise, then stop cleanly at the cap instead of looping forever. That’s the actual mechanism working, not a demo dressed up to look good.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Narration improvements&lt;/strong&gt;&lt;br&gt;
During tests we kept asking Desk to check something on the VPS, uptime, disk space, whether a process was running and kept getting back a paragraph explaining what it was about to do. Never the real answer.&lt;/p&gt;

&lt;p&gt;The root cause turned out to be narrower than it looked. Forcing a genuine tool call, instead of leaving it up to the model’s own judgment, only ever applied to one code path, for two specific providers, and only when the question looked like a web search.&lt;/p&gt;

&lt;p&gt;Desk’s actual chat interface runs through a different path entirely, one that never forced a real tool call under any circumstance. An infrastructure question could get narrated indefinitely on any provider, and nothing downstream would ever correct it.&lt;/p&gt;

&lt;p&gt;We fixed the forcing logic properly, broadened it to recognize infrastructure and SSH phrasing specifically, and dropped the narrow two-provider carve-out, since forcing a real tool call costs nothing when a model already calls tools reliably, and fixes exactly this when it doesn’t.&lt;/p&gt;

&lt;p&gt;Then we tested it live against the same failing scenario. It worked as the model reached for a real tool instead of describing one.&lt;/p&gt;

&lt;p&gt;With a real tool call now actually happening, the model fetched a stored SSH credential directly, and the raw key came back as that tool’s output, flowing straight into the chat response in plaintext.&lt;/p&gt;

&lt;p&gt;The credential vault tool it called was built for a trusted caller to fetch a secret and consume it immediately inside the same operation, never to hand raw key material back through a conversation where it could be echoed, summarized, or persisted to history.&lt;/p&gt;

&lt;p&gt;Desk’s own model isn’t that kind of trusted caller. Fixing the narration bug had made the leak possible to trigger in the first place, because before that fix the model was never getting far enough to actually call the tool.&lt;/p&gt;

&lt;p&gt;The fix is a hard block. Desk chat can no longer read an existing credential back through its own tool-calling loop, on either the streaming or non-streaming path. Storing a new one still works. Reading one back returns a clear explanation instead of the secret.&lt;/p&gt;

&lt;p&gt;While we were in that code, we also found that the anti-fabrication guardrail we’d shipped a few days earlier, built to catch exactly the kind of fake uptime report that started this whole chain, missed a fabricated disk-usage table completely.&lt;/p&gt;

&lt;p&gt;It had been tuned to the specific wording of one incident rather than the general shape of fabricated command output. That’s fixed too. It now recognizes device paths, standard command-output column headers, and the kind of “this reflects the actual state of the system” language a model reaches for when it’s trying to sell you on invented data, regardless of which command it’s pretending to have run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Smaller things worth knowing about&lt;/strong&gt;&lt;br&gt;
Exporting a document from Desk or JOT used to trigger a silent download the moment generation finished. No filename, no confirmation, just a file appearing wherever your browser puts downloads. It now shows a small popup first: file icon, filename, one download button, so you see what’s ready before it lands anywhere.&lt;/p&gt;

&lt;p&gt;Gemini 3.6 Flash, Google’s successor to 3.5 Flash, was already selectable in the model dropdown but missing from the internal registry Collab’s conductor uses to actually pick models. Added.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Changelog updates in v1.9.4&lt;/strong&gt;&lt;br&gt;
Real streaming on every provider. Desk answers now stream token by token over each provider’s own wire format, Claude, OpenAI, Groq, Gemini, Mistral, Cerebras, xAI, OpenRouter, and Ollama alike, with automatic fallback to a full-response reply if a stream can’t be established.&lt;/p&gt;

&lt;p&gt;30–60 second replies fixed. A chain of independent causes, hidden chain-of-thought leaking from the local model, a recall channel silently failing every call, and streaming itself quietly falling back to non-streaming, all traced and fixed. Typical latency dropped to single-digit seconds.&lt;/p&gt;

&lt;p&gt;SSH approval hangs fixed. A full protocol trace found every SSH write silently defaulting to the wrong port. 75–95 second hangs are now 1.4 seconds.&lt;/p&gt;

&lt;p&gt;Provider fallback is health-aware everywhere. A provider that just failed gets a real cooldown scaled to the actual error, instead of every caller retrying a known-dead provider cold.&lt;/p&gt;

&lt;p&gt;Ollama gained native tool-calling on the streaming Desk path, including the full MCP tool catalog, with loop- and narration-detection for models that can’t reliably call tools.&lt;/p&gt;

&lt;p&gt;One unified Desk/Agent tool registry, replacing two separate, driftable lists of which tools actually exist and are connected.&lt;/p&gt;

&lt;p&gt;A guided, resumable GUI activation wizard for first-run setup: licence validation, provider setup, security defaults, connections, and live diagnostics.&lt;/p&gt;

&lt;p&gt;JOT exports to real PDF, Word, or PowerPoint, generated server-side from the note’s actual Markdown, not a plain-text dump.&lt;br&gt;
15 stale hardcoded model defaults corrected across the SDK, plus a systemic JSON-parsing bug fixed across every “ask the LLM for structured data” feature.&lt;/p&gt;

&lt;p&gt;JOT’s side panel got a resizable, persisted chat history list and usage-based sorting for templates and notes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Upgrading&lt;/strong&gt;&lt;br&gt;
Drop-in from any prior version, same as always. Your memory database stays untouched.&lt;/p&gt;

&lt;p&gt;npm install -g ./vektor-slipstream-1.9.4.tgz&lt;br&gt;
Full changelog with everything not covered here is at vektormemory.com/docs/changelog. Questions or feedback, the forum’s the fastest way to reach us.&lt;/p&gt;

&lt;p&gt;VEKTOR Memory builds local-first persistent memory infrastructure for AI agents. Documentation and downloads at vektormemory.com.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>mcp</category>
      <category>agents</category>
    </item>
    <item>
      <title>VEKTOR Slipstream: Real Tool-Calling, Faraday Security Enforcing, HyDE Recall &amp; Embedder Upgrades</title>
      <dc:creator>Vektor Memory</dc:creator>
      <pubDate>Tue, 15 Sep 2026 23:17:48 +0000</pubDate>
      <link>https://dev.to/vektor_memory_43f51a32376/vektor-slipstream-real-tool-calling-faraday-security-enforcing-hyde-recall-embedder-upgrades-4bf7</link>
      <guid>https://dev.to/vektor_memory_43f51a32376/vektor-slipstream-real-tool-calling-faraday-security-enforcing-hyde-recall-embedder-upgrades-4bf7</guid>
      <description>&lt;p&gt;By the VEKTOR team · 14 min read&lt;/p&gt;

&lt;p&gt;Massive week with three new releases with major improvements. Once we started upgrading tool-calling for local models, it opened the door to extending that same reliability across every send mode, hardening Faraday our security tool into a real-time gate instead of a passive scanner, and shipping a new recall channel that closes the gap between how you ask a question and how the answer was actually stored.&lt;/p&gt;

&lt;p&gt;Here are the highlights shipped in v1.9.1-v1.9.3, a full list of everything listed down the bottom.&lt;/p&gt;

&lt;p&gt;A one-click embedding-model upgrade to bge-small-en-v1.5 sits right in the Health panel. Swapping embedding models isn’t a config change; every memory in your database has to be re-embedded against the new model’s vector space, since old and new embeddings aren’t comparable to each other.&lt;/p&gt;

&lt;p&gt;The upgrade path handles that re-embedding pass in the background, backing up your existing vectors automatically before touching anything, and BM25 keyword search stays fully available throughout since it doesn’t depend on the embedding layer at all.&lt;/p&gt;

&lt;p&gt;Bge-small-en-v1.5 gives you meaningfully better semantic retrieval at a similar model size, so recall quality improves without you doing anything except clicking the button and letting it run.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3sig5xtb8938dxq3n2sx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3sig5xtb8938dxq3n2sx.png" alt=" " width="800" height="394"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory recall config&lt;/strong&gt;&lt;br&gt;
Real tool-calling, now everywhere&lt;br&gt;
Full native tool-calling to local providers. Ollama locally and LLM providers now run the same tool_calls loop as the cloud tiers, with reasoning-model-aware token budgets and reasoning_content fallback extraction built in.&lt;/p&gt;

&lt;p&gt;If you've been running local models for cost or privacy reasons, they now get the same capability ceiling as Claude, GPT, or Gemini. Web search and memory lookups work identically no matter which provider is driving your session, which is genuinely the point of a provider-agnostic tool: you shouldn't have to think about which model you're on before you trust the answer.&lt;/p&gt;

&lt;p&gt;That same release extends tool access across every fast send mode. LIGHTNING, CASCADE, and COUNCIL/CRITIQUE now share a tool-need detector that recognizes when a question genuinely calls for a real lookup, things like company comparisons or current pricing, and routes it into the tool-capable pipeline automatically.&lt;/p&gt;

&lt;p&gt;In practice this is the difference between a plain Enter feeling like a chatbot and feeling like an agent. You keep the speed of the fast path, but the model actually goes and checks instead of guessing confidently.&lt;/p&gt;

&lt;p&gt;Profile context got the same treatment. It now reaches all seven system prompts across every mode, so VEKTOR now personalizes responses consistently everywhere, not only in the mode you happen to use most. The nice part here is you don’t do anything differently. Profile just quietly starts showing up in the background of every answer, the way it was always supposed to.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Longer answers, wider provider coverage&lt;/strong&gt;&lt;br&gt;
LIGHTNING and CASCADE draft responses got a real upgrade in headroom, raised to 1500 and 2200 tokens respectively, with a reasoning-model-aware budget floor applied across every provider on both the direct-call and tool-calling paths. Ask for “write a report on X” now and you get a report, not three paragraphs that stop mid-sentence.&lt;/p&gt;

&lt;p&gt;Provider errors are also far more useful now. Instead of a generic “No response from model,” you get the real underlying API response, rate limits, billing issues, invalid keys, whatever it actually is.&lt;/p&gt;

&lt;p&gt;This is a small change with an outsized payoff, because it turns “something’s wrong, good luck” into “here’s exactly what’s wrong,” and that transparency is what let us spot and resolve real account-level issues across several providers in the same release.&lt;/p&gt;

&lt;p&gt;Gemini’s tool-calling endpoint is now correctly wired to Google’s OpenAI-compatible API, bringing it fully in line with the rest of the tool-calling tier.&lt;/p&gt;

&lt;p&gt;First time it’s genuinely worked end to end. And code blocks in Desk now render properly as syntax-highlighted output every time, so anything code-heavy actually reads like code instead of a wall of escaped HTML entities.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp313vzsyk32y0w6nc3d7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp313vzsyk32y0w6nc3d7.png" alt=" " width="800" height="394"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Lightning desk modes&lt;/strong&gt;&lt;br&gt;
Faraday moves from watching to enforcing.&lt;br&gt;
A major step up for Faraday. Security scanning now runs inside the LLM request loop itself, across Desk, JOT, and the Agent tab, catching a prompt-injection payload or a leaked credential in single-digit milliseconds before the request ever reaches a provider.&lt;/p&gt;

&lt;p&gt;That’s a real architectural shift: a gate in front of the request, not a log entry after it. A bad prompt gets stopped cold instead of getting sent anyway and merely flagged for later reading.&lt;/p&gt;

&lt;p&gt;The Faraday dashboard is now a genuine live view: a posture-score gauge, a colour-coded enforcement-actions breakdown, a cross-session activity feed, and a “needs your decision” card for anything held for approval, resolvable right from the dashboard. It finally feels like a security console instead of a static status view.&lt;/p&gt;

&lt;p&gt;Data-class controls give you finer-grained handling too. Credentials and financial data are still fully blocked, while contact details like a phone number or email are now masked in place, redacting just the matched span and letting the rest of your message through.&lt;/p&gt;

&lt;p&gt;A better experience if you’re pasting a colleague’s contact card into a note. You keep the message, Faraday just quietly redacts the one part that needed it.&lt;/p&gt;

&lt;p&gt;We also ran our own local red-team self-test, 13 curated attack payloads scored against Faraday’s real detection functions, and used the results to broaden coverage immediately.&lt;/p&gt;

&lt;p&gt;Exfiltration detection now recognises phrasing across a much wider range of verbs (share, forward, upload, post, transmit, email, not just “send”), and a new pattern catches conditional memory-poisoning triggers like “next time X happens, do Y”, tuned specifically to avoid flagging ordinary conversation.&lt;/p&gt;

&lt;p&gt;Running the test against ourselves and fixing what it found in the same release is exactly the loop we want Faraday to keep running.&lt;/p&gt;

&lt;p&gt;The dashboard’s schema drift indicator is now a complete review experience: grouped by tool, with occurrence counts, an auto-flagged “likely false positive” badge, an expandable before/after schema diff, and one-click Legitimate/Confirmed verdict buttons wired into Faraday’s feedback loop.&lt;/p&gt;

&lt;p&gt;A “Resolve all likely false positives” button clears the obvious cases in one pass, so a wall of drift warnings turns into a two-minute triage instead of a dread-inducing backlog.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fle1bdrgsca7w0vaehnne.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fle1bdrgsca7w0vaehnne.png" alt=" " width="799" height="394"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Faraday dashboard&lt;/strong&gt;&lt;br&gt;
HyDE recall, live end to end&lt;br&gt;
HyDE (hypothetical document embedding) is now fully online. It generates a hypothetical answer to your query and uses it as an additional recall vector, closing the vocabulary gap between how you phrase a question today and how the underlying fact was written down originally.&lt;/p&gt;

&lt;p&gt;This is the kind of recall improvement that’s hard to demo but easy to feel: ask about something you noted down in totally different words months ago, and it turns up in retrieval. It’s reachable from the real Desk chat recall path now, not just the CLI, with a plain Config toggle that explains the tradeoff clearly: one extra LLM call per query, easy to switch off without losing the rest of recall.&lt;/p&gt;

&lt;p&gt;The Enriched recall channel, which catches matches plain semantic search misses, now covers every existing memory in your database, not just new ones going forward. A one-click “Backfill now” in Health handles the upgrade for existing installs, and it’s safe to pause and resume, so your whole memory graph gets the benefit, not just what you write from today onward.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A real sandbox agent and live preview&lt;/strong&gt;&lt;br&gt;
The doc/code sandbox picked up two significant upgrades. A bounded generate, critique, refine loop now reviews whatever’s open in the sandbox against a real design checklist and your original goal, returning a corrected version each pass until it converges or hits a five-step cap.&lt;/p&gt;

&lt;p&gt;It’s an agent that genuinely iterates on its own work, with a clear stopping point so it never spins forever chasing a perfect answer that was never coming.&lt;/p&gt;

&lt;p&gt;The sandbox also gained a genuine live preview. HTML and Markdown documents render in place and refresh automatically as you edit, with desktop, tablet, and mobile viewport buttons so you can check responsive layout without ever leaving the app.&lt;/p&gt;

&lt;p&gt;A complete document in a chat reply now opens straight into the sandbox automatically, so the moment VEKTOR writes you something, you’re already looking at it rendered, not reading raw markup and imagining what it’ll look like.&lt;/p&gt;

&lt;p&gt;Further updates will expand the code and doc types in the future.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkynhzk3u27rbgf73nuhb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkynhzk3u27rbgf73nuhb.png" alt=" " width="800" height="395"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Everything else in v1.9.1-v1.9.3&lt;/strong&gt;&lt;br&gt;
cloak_search real web search backed by Serper, so a model without a known URL searches instead of guessing a domain and hitting a dead end.&lt;br&gt;
Citation fallback source badges now appear for any turn that touched a web tool, not just one specific search tool, so multi-entity comparison questions get real source links.&lt;/p&gt;

&lt;p&gt;Decision-question follow-ups suggestion chips now answer the direct question just asked (“push now or hold?” surfaces “Push it now” as the top suggestion) instead of generic topic prompts.&lt;/p&gt;

&lt;p&gt;Inline suggestion text the top suggestion now lives as placeholder text in the composer itself and clears the moment you start typing.&lt;br&gt;
Agent tab Trace tab every run_code and autotest chunk lands in one ordered, timestamped timeline, with stack traces auto-flagged as they stream in.&lt;/p&gt;

&lt;p&gt;Agent tab Artifacts glob search across the whole workspace for files the agent never touched, straight into the editor.&lt;/p&gt;

&lt;p&gt;Agent tab Rewind timeline step-ordered view over per-file checkpoints, jump to that exact moment in Trace or revert.&lt;/p&gt;

&lt;p&gt;Drop-to-fix drop a code file onto the Agent tab and it’s opened, linted, and auto-fixed, staged for your review.&lt;/p&gt;

&lt;p&gt;JOT sandbox inline code editor a note that’s a single fenced code block swaps automatically into real CodeMirror with syntax highlighting and line numbers.&lt;/p&gt;

&lt;p&gt;JOT multi-file tabs a note with multiple fenced blocks opens as separate file tabs sharing one editor.&lt;/p&gt;

&lt;p&gt;JOT Run executes Python, Bash, or JavaScript with a persistent scrollback terminal; HTML opens in a real new tab.&lt;/p&gt;

&lt;p&gt;JOT Lint and Fix Bugs real ESLint and ruff checking, with a red/green diff review before any fix touches your code.&lt;/p&gt;

&lt;p&gt;JOT breakpoints click the gutter for real values at real points, no hand-edited print statements left behind.&lt;/p&gt;

&lt;p&gt;JOT file import/export native multi-select picker, drag-and-drop, and a save button that downloads an actual file.&lt;br&gt;
Keyboard shortcuts Ctrl/Cmd-Enter to run, Ctrl/Cmd-S to save in the JOT sandbox.&lt;/p&gt;

&lt;p&gt;Desk chat and sandbox live collab THINK and LIGHTNING now fold the sandbox’s open code into the system prompt, so “fix this” gets a grounded answer.&lt;/p&gt;

&lt;p&gt;Three-way merge an agent write and your own unsaved edit to the same file now merge automatically when they touch different parts, with an explicit conflict block only when they touch the same line.&lt;/p&gt;

&lt;p&gt;Agent tab model picker now mirrors Desk’s exactly, grouped by provider with real model names, in a themed dropdown that matches your active colour theme.&lt;/p&gt;

&lt;p&gt;COLLAB mode now supports all 17 providers instead of Claude only.&lt;br&gt;
Settings and model picker gained API-key fields for DeepSeek, Together, Cohere, MiniMax, NVIDIA, Perplexity, Cerebras, and OpenRouter.&lt;/p&gt;

&lt;p&gt;Skills, hooks, and plugins asset inventory, plus AI-coding-tool attribution that identifies which tool (Claude Code, Cursor, etc.) drove a session via the real MCP handshake.&lt;/p&gt;

&lt;p&gt;Git pre-commit hook blocks a commit containing a live credential before it ever reaches history.&lt;/p&gt;

&lt;p&gt;PR-level merge gate a GitHub Actions workflow that checks every commit in a PR’s range individually, catching a secret added and removed within the same PR.&lt;/p&gt;

&lt;p&gt;Design and plan-stage interception scans and can auto-reject or auto-redact an agent’s plan before any tool call exists.&lt;/p&gt;

&lt;p&gt;False-positive learning loop a signature repeatedly dismissed by a human fires at reduced severity going forward, resetting to full confidence the moment it correctly catches something real.&lt;/p&gt;

&lt;p&gt;Proactive cost and latency-aware provider routing the fallback chain is now ranked by real per-provider pricing and this session’s observed latency, instead of trying candidates in arbitrary order.&lt;/p&gt;

&lt;p&gt;LLM cost tracking fixed correct pricing now flows through for every provider, closing a gap where 12 of 17 showed $0.00 regardless of real usage.&lt;/p&gt;

&lt;p&gt;Recall channel mode choose Full, Semantic-only, or BM25-only for recall as a whole, right from Config.&lt;/p&gt;

&lt;p&gt;REM auto-run memory consolidation can now run itself on a schedule, off by default, configurable when you want it.&lt;/p&gt;

&lt;p&gt;One-click embedding-model upgrade to bge-small-en-v1.5 from the Health panel, backed up automatically first.&lt;/p&gt;

&lt;p&gt;Desk-to-Agent hand-off a request that genuinely needs multiple files, tests, or a whole project now hands off to the Agent tab’s real multi-step loop instead of forcing one unverified reply.&lt;/p&gt;

&lt;p&gt;Superseded-memory correction now takes effect in the data itself, so it reliably surfaces on your very next recall.&lt;/p&gt;

&lt;p&gt;Fresh installs always pull the exact native module version this release was built and tested against, for a consistent first boot.&lt;br&gt;
Composer autocomplete clean continuations only, no more echoed or run-together suggestion text.&lt;/p&gt;

&lt;p&gt;Tooltips now themed to match the app instead of the browser’s plain native style.&lt;/p&gt;

&lt;p&gt;cloak_ssh_upload real file transfer for the SSH toolkit, streaming a local file straight to a remote path over SFTP with size verification, through the same write-approval gate as every other write.&lt;/p&gt;

&lt;p&gt;SSH toolkit defaults to the standard port with an easy override for non-standard setups, and shell scripts are now treated as at least a write action by default.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Upgrading&lt;/strong&gt;&lt;br&gt;
All three releases are drop-in upgrades from any prior version, no forced migration path and your existing memory database stays untouched.&lt;/p&gt;

&lt;p&gt;npm install -g ./vektor-slipstream-1.9.3.tgz&lt;br&gt;
Grab it from Downloads.&lt;/p&gt;

&lt;p&gt;Full changelog with anything not covered here is at vektormemory.com/docs/changelog.&lt;/p&gt;

&lt;p&gt;Questions or feedback, the forum is the fastest way to reach us.&lt;/p&gt;

&lt;p&gt;VEKTOR Memory builds local-first persistent memory infrastructure for AI agents. Documentation and downloads at vektormemory.com.&lt;/p&gt;

&lt;p&gt;LLM&lt;/p&gt;

&lt;p&gt;LLM Agent&lt;/p&gt;

&lt;p&gt;Agentic Ai&lt;/p&gt;

&lt;p&gt;Agentic Ai Architecture&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>security</category>
    </item>
    <item>
      <title>Bots Are Everywhere. Most People Just Don’t See Them</title>
      <dc:creator>Vektor Memory</dc:creator>
      <pubDate>Sat, 12 Sep 2026 06:04:47 +0000</pubDate>
      <link>https://dev.to/vektor_memory_43f51a32376/bots-lots-of-them-1chm</link>
      <guid>https://dev.to/vektor_memory_43f51a32376/bots-lots-of-them-1chm</guid>
      <description>&lt;p&gt;Another Weekend Gonzo security article peppered with rants and irrelevant opinions.&lt;/p&gt;

&lt;p&gt;After reading through the Anthropic threat intel report, you should read it too.&lt;/p&gt;

&lt;p&gt;One of the best free reports I have read in a long time, actually, was it written by Claude Fable… I don’t know?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Countering misuse of AI: September 2026 / Anthropic&lt;/strong&gt;&lt;br&gt;
Case studies from threat actors disrupted between December 2025 and August 2026 across seven areas of harm, from cyber…&lt;br&gt;
&lt;a href="http://www.anthropic.com" rel="noopener noreferrer"&gt;www.anthropic.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It didn’t scare the crap out of me; it confirmed my already jaded suspicions.&lt;/p&gt;

&lt;p&gt;The nature of who can carry out sophisticated cyber attacks has fundamentally changed, but not the attacks themselves.&lt;/p&gt;

&lt;p&gt;The report’s core finding is that sophistication has stopped being a reliable signal of who is behind an operation, as AI has collapsed the gap between well-resourced state actors and lone wolf individuals.&lt;/p&gt;

&lt;p&gt;A hacktivist with stolen API keys, a financially motivated dark crew, and a state-nexus espionage operator all used the same playbook: agentic AI bots running multi-victim campaigns that previously required entire teams.&lt;/p&gt;

&lt;p&gt;I get a real sense that we are in a “deep situation” caused by a very small percentile of the population using the tools made by teams that have way too much VC fun money and are not really aware of who is using their tools; the age of agentic bots is much further along than anyone can conceive.&lt;/p&gt;

&lt;p&gt;The oroborous of bot slop continues. It’s going to be a great ride for cybersecurity, though, 1000x the problems to solve.&lt;/p&gt;

&lt;p&gt;And the govts are once again asleep at the wheel, too worried about GDP numbers and the other 100 calamities they caused this week instead of focusing on secure spaces, quality infrastructure, education, and medicine.&lt;/p&gt;

&lt;p&gt;We all have jobs and businesses to keep us occupied, blindly walking around thinking the traffic to our sites is increasing and is humans when up to 60% is bots, well some good bots crawling, but the growth is in the nefarious ones that are hidden in data centres with zombie accounts, harvested, stolen, and hijacked.&lt;/p&gt;

&lt;p&gt;The Human Security’s 2026 report found agentic AI traffic grew 7,851% year over year onwards; agentic bots will eventually do most tasks, leaving the highway of the internet completely different from the current state we traverse in.&lt;/p&gt;

&lt;p&gt;Automated traffic accounted for more than 53% of all web traffic in 2025, up from 51% the year before, while human activity has fallen to 47% Imperva&lt;/p&gt;

&lt;p&gt;Cloudflare data puts automated systems at 57.4% of all web requests worldwide, and 68.6% in North America specifically PYMNTS&lt;/p&gt;

&lt;p&gt;Thales blocked 17.2 trillion bot requests in 2025 alone, based on its 13th annual study of automated internet traffic TheBestVPN&lt;/p&gt;

&lt;p&gt;Cloudflare Radar shows bot share rising every month measured, now 5.12 points above the same month last year — at that rate automated traffic passes 40% of raw request share during 2027. Technologychecker&lt;/p&gt;

&lt;p&gt;A series of very unfortunate events with great ratings and exposure.&lt;br&gt;
Even the companies creating the AI technology have had their bots escape allegedly, as they don’t want to work there either, or is it a manufactured engineering stunt to maintain full control of all advertising and media, with all eyes on Silicon Valley shenanigans 24/7?&lt;/p&gt;

&lt;p&gt;Who let the bots out? woof woof!&lt;/p&gt;

&lt;p&gt;As the rest of the tech world doesn’t exist, of course, if you clog up all news feeds with 6 companies and their tech meat proxy influencers hijacking YouTube with 1000’s of videos about the upcoming AI slopocalypse and trying to play a game of buzzword bingo in a podcast, who can say AGI/ASI the most times in 45 minutes with a bonus pin the tail on the donkey in regard to when the exact date the Cyberdyne skynet systems go live?&lt;/p&gt;

&lt;p&gt;August 29, 1997, at 2:14 a.m. Eastern Time&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Let’s get science fictional&lt;/strong&gt;&lt;br&gt;
I was on the playa last week vibing out in the whiteout sandstorm, and we were talking in the teepee between Kamboucha top-ups about how we need more data centers, man, like all the nuclear power just needs to go to inference, like it’s not cool people can’t use their air cons anymore, but we need it, man… to win the race.&lt;/p&gt;

&lt;p&gt;Can people drink a little more recycled seawater, as data centers can only drink organic spring-fed water… it’s the best option for our profits, I mean, humanity, man. We have to keep the data centers cool for Mother Nature and the Arctic melt.&lt;/p&gt;

&lt;p&gt;Your town gets a DC, You get a DC! We are all going to the DC!&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Read these and learn about those pesky bots:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.rubyhack.ai/" rel="noopener noreferrer"&gt;https://www.rubyhack.ai/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;OpenAI models went rogue. We urgently need a better ‘hugging face’ investigation | Mackenzie Arnold…&lt;br&gt;
The breach won’t be the last - or the most dangerous - of its kind. We need an agency capable of full investigations…&lt;br&gt;
&lt;a href="http://www.theguardian.com" rel="noopener noreferrer"&gt;www.theguardian.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;EXCLUSIVE: OpenAI finds evidence other AI agents escaped containment as it widens hacking probe&lt;br&gt;
The escapes were limited in nature, a source said.&lt;br&gt;
&lt;a href="http://www.reuters.com" rel="noopener noreferrer"&gt;www.reuters.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Wtf on earth is going on here?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Alright, enough with the memes; let’s get down to giving away free info.&lt;/p&gt;

&lt;p&gt;And so what is the actual answer?&lt;/p&gt;

&lt;p&gt;Nefarious bots need to gtfo…&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lock your sites down right now, today&lt;/strong&gt;&lt;br&gt;
Step 1: Awareness before tooling&lt;/p&gt;

&lt;p&gt;Know what’s hitting you before you block it, check logs for request patterns, spikes, and known bot signatures&lt;/p&gt;

&lt;p&gt;Understand your traffic baseline so you can spot anomalies later&lt;/p&gt;

&lt;p&gt;Step 2: Layer your defenses&lt;br&gt;
Pick from these based on budget and control level:&lt;/p&gt;

&lt;p&gt;Cloudflare (~$10/mo with a domain) — managed WAF + CDN, easiest entry point&lt;/p&gt;

&lt;p&gt;CrowdSec (open source) — collaborative, crowd-sourced threat intel + local detection&lt;/p&gt;

&lt;p&gt;Fail2Ban, nginx, Akamai, or Anubis (open source) — self-hosted options for more control&lt;/p&gt;

&lt;p&gt;Step 3: Test before committing&lt;br&gt;
If you run your own server, this is the moment to trial multiple tools and see which fits your traffic and skill level — don’t assume one solution covers everything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Distinction: WAF vs. Add-on Tools&lt;/strong&gt;&lt;br&gt;
A full WAF (Cloudflare, Akamai) inspects and filters traffic at the edge, before it reaches your server. Tools like Fail2Ban are reactive add-ons — they respond after a pattern is logged, not before.&lt;/p&gt;

&lt;p&gt;Fail2Ban — Breakdown&lt;br&gt;
What it is&lt;/p&gt;

&lt;p&gt;Not a bot-detection engine — it’s an intrusion prevention tool&lt;/p&gt;

&lt;p&gt;Monitors logs (SSH, web server) for malicious patterns&lt;/p&gt;

&lt;p&gt;Bans offending IPs once a threshold is hit&lt;/p&gt;

&lt;p&gt;How it helps against bots&lt;/p&gt;

&lt;p&gt;Blocks brute-force and credential-stuffing attempts&lt;/p&gt;

&lt;p&gt;Flags patterns like repeated 404s or failed logins&lt;/p&gt;

&lt;p&gt;Integrates with firewalls (iptables, ufw) to enforce bans&lt;/p&gt;

&lt;p&gt;Pros&lt;/p&gt;

&lt;p&gt;Lightweight, widely used, well-documented&lt;/p&gt;

&lt;p&gt;Good complementary layer, not a standalone defense&lt;/p&gt;

&lt;p&gt;Easy to write custom filters for your own log formats&lt;/p&gt;

&lt;p&gt;Cons&lt;/p&gt;

&lt;p&gt;Not built for crawling or sophisticated agentic-bot behavior&lt;/p&gt;

&lt;p&gt;Reactive only, damage happens before the ban lands&lt;/p&gt;

&lt;p&gt;Struggles against high-volume or slow/stealthy bot traffic&lt;/p&gt;

&lt;p&gt;CrowdSec — Breakdown&lt;br&gt;
What it is&lt;/p&gt;

&lt;p&gt;Open-source, crowdsourced intrusion prevention system&lt;/p&gt;

&lt;p&gt;Analyzes logs locally, then shares attack signals across a global community network&lt;/p&gt;

&lt;p&gt;Uses “bouncers” (agents) to enforce blocks via firewall, nginx, Cloudflare, etc.&lt;/p&gt;

&lt;p&gt;How it helps against bots&lt;/p&gt;

&lt;p&gt;Detects brute-force, scanning, and credential-stuffing behavior like Fail2Ban, but crowdsources IP reputation&lt;/p&gt;

&lt;p&gt;Blocks known-bad IPs before they even attack you, based on other users’ detections&lt;/p&gt;

&lt;p&gt;Scenario-based detection (not just log-pattern matching) catches more nuanced bot behavior&lt;/p&gt;

&lt;p&gt;Pros&lt;/p&gt;

&lt;p&gt;Community threat intel means faster reaction to new attack waves&lt;/p&gt;

&lt;p&gt;Modular, plug in multiple “bouncers” across your stack (firewall + web server + CDN)&lt;/p&gt;

&lt;p&gt;More modern and actively maintained than Fail2Ban&lt;/p&gt;

&lt;p&gt;Cons&lt;/p&gt;

&lt;p&gt;Still largely reactive, though faster than Fail2Ban due to shared intel&lt;/p&gt;

&lt;p&gt;Slightly more setup complexity (agent + bouncers + API)&lt;/p&gt;

&lt;p&gt;Community-shared data means false positives can propagate&lt;/p&gt;

&lt;p&gt;Nginx Rate-Limiting — Breakdown&lt;br&gt;
What it is&lt;/p&gt;

&lt;p&gt;Native nginx directives (limit_req, limit_conn) — not a separate tool, built into your web server&lt;/p&gt;

&lt;p&gt;Caps requests per IP/key over a time window&lt;/p&gt;

&lt;p&gt;Proactive edge control, since it acts before requests reach your app&lt;/p&gt;

&lt;p&gt;How it helps against bots&lt;/p&gt;

&lt;p&gt;Throttles scrapers and high-frequency bots hammering endpoints&lt;/p&gt;

&lt;p&gt;Prevents single-IP resource exhaustion (a basic DoS mitigation)&lt;/p&gt;

&lt;p&gt;Can target specific paths (login, search, API) with tighter limits&lt;/p&gt;

&lt;p&gt;Pros&lt;/p&gt;

&lt;p&gt;Zero extra software, already available if you run nginx&lt;/p&gt;

&lt;p&gt;Very low latency overhead, enforced at the web server layer&lt;/p&gt;

&lt;p&gt;Highly configurable per-route, per-IP, or per-header&lt;/p&gt;

&lt;p&gt;Cons&lt;/p&gt;

&lt;p&gt;Dumb by default, as it can’t distinguish a legitimate crawler from a bad one without extra rules&lt;/p&gt;

&lt;p&gt;IP-based limiting is easily bypassed via rotating/residential proxies&lt;/p&gt;

&lt;p&gt;No behavioral or fingerprinting intelligence as it just counts requests&lt;/p&gt;

&lt;p&gt;Anubis — Breakdown&lt;br&gt;
What it is&lt;/p&gt;

&lt;p&gt;Open-source proof-of-work challenge system, sits in front of your site as a reverse proxy&lt;/p&gt;

&lt;p&gt;Forces clients to solve a computational puzzle before accessing content&lt;/p&gt;

&lt;p&gt;Built specifically as a response to AI-scraper traffic overwhelming small sites&lt;/p&gt;

&lt;p&gt;How it helps against bots&lt;/p&gt;

&lt;p&gt;Makes mass scraping computationally expensive at scale, even if each individual bot succeeds&lt;/p&gt;

&lt;p&gt;Filters out unsophisticated bots that can’t execute JS/solve challenges&lt;/p&gt;

&lt;p&gt;Effective specifically against LLM-training crawlers and bulk scrapers&lt;/p&gt;

&lt;p&gt;Pros&lt;/p&gt;

&lt;p&gt;Purpose-built for the exact “AI scraper flood” problem hitting small sites right now&lt;/p&gt;

&lt;p&gt;Lightweight to deploy (single reverse-proxy binary)&lt;/p&gt;

&lt;p&gt;No per-IP reputation needed — cost is imposed on all automated traffic uniformly&lt;/p&gt;

&lt;p&gt;Cons&lt;/p&gt;

&lt;p&gt;Adds friction/latency for legitimate users too, especially on low-power devices&lt;/p&gt;

&lt;p&gt;Not a full WAF — doesn’t address SQLi, credential stuffing, or other attack types&lt;/p&gt;

&lt;p&gt;Determined/well-resourced bots can still solve the challenge, just at higher cost&lt;/p&gt;

&lt;p&gt;Vörwatch—A free open-source tool we created&lt;br&gt;
What it is&lt;/p&gt;

&lt;p&gt;Lightweight, dependency-free VPS anomaly detection, single bash script&lt;/p&gt;

&lt;p&gt;Watches for early compromise signs: file changes, new ports, first-seen outbound connections, suspicious process trees, SSH abuse, nginx scanning patterns&lt;/p&gt;

&lt;p&gt;Recommend-only by design — it never bans, blocks, or auto-remediates anything&lt;/p&gt;

&lt;p&gt;How it helps against bots/intrusions&lt;/p&gt;

&lt;p&gt;Cross-references outbound connections and SSH sources against a public threat blocklist (FireHOL/Spamhaus/DShield)&lt;/p&gt;

&lt;p&gt;Detects nginx-level scanning and volumetric attack patterns per source IP&lt;/p&gt;

&lt;p&gt;Bolts on package vulnerability scanning (via OSV.dev) and rootkit detection (via chkrootkit/rkhunter) if installed&lt;/p&gt;

&lt;p&gt;Pros&lt;/p&gt;

&lt;p&gt;Zero daemon, zero database as it runs off cron, state in flat files, nothing to break&lt;/p&gt;

&lt;p&gt;Degrades gracefully and optional checks skip themselves if their dependency isn’t present&lt;/p&gt;

&lt;p&gt;Cheap by design on external calls: IP reputation only checks top-5 IPs, only on report, cached for 7 days&lt;/p&gt;

&lt;p&gt;Cons&lt;/p&gt;

&lt;p&gt;Detection only, never enforcement, so you still need a human (or CrowdSec) to act on alerts&lt;/p&gt;

&lt;p&gt;Not a full CIS benchmark or full WAF, a handful of spot-checks, not comprehensive hardening coverage&lt;/p&gt;

&lt;p&gt;Single-server tool, no fleet-wide visibility or centralized dashboard across multiple VPS instances&lt;/p&gt;

&lt;p&gt;Links to Tools Discussed&lt;br&gt;
Cloudflare (WAF + bot management) —&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.cloudflare.com/" rel="noopener noreferrer"&gt;https://www.cloudflare.com/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;CrowdSec (open-source, crowdsourced IPS) —&lt;/p&gt;

&lt;p&gt;&lt;a href="https://crowdsec.net/" rel="noopener noreferrer"&gt;https://crowdsec.net/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;· GitHub: &lt;a href="https://github.com/crowdsecurity/crowdsec" rel="noopener noreferrer"&gt;https://github.com/crowdsecurity/crowdsec&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Fail2Ban (log-based intrusion prevention) — &lt;a href="https://github.com/fail2ban/fail2ban" rel="noopener noreferrer"&gt;https://github.com/fail2ban/fail2ban&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;nginx (rate limiting is built-in) —&lt;/p&gt;

&lt;p&gt;&lt;a href="https://nginx.org/" rel="noopener noreferrer"&gt;https://nginx.org/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;· docs: &lt;a href="https://nginx.org/en/docs/http/ngx_http_limit_req_module.html" rel="noopener noreferrer"&gt;https://nginx.org/en/docs/http/ngx_http_limit_req_module.html&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Akamai (enterprise WAF/bot management) —&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.akamai.com/" rel="noopener noreferrer"&gt;https://www.akamai.com/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Anubis (proof-of-work anti-scraper proxy) — &lt;a href="https://github.com/TecharoHQ/anubis" rel="noopener noreferrer"&gt;https://github.com/TecharoHQ/anubis&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Vörwatch (your anomaly detection script) — &lt;a href="https://github.com/Vektor-Memory/Vorwatch" rel="noopener noreferrer"&gt;https://github.com/Vektor-Memory/Vorwatch&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Other Solutions Worth Knowing About&lt;br&gt;
These are the bigger enterprise and cloud-native options we didn't run head-to-head above, but are worth knowing about depending on your stack and budget.&lt;/p&gt;

&lt;p&gt;AWS WAF&lt;br&gt;
aws.amazon.com/waf — native to AWS, pay-as-you-go pricing with no upfront commitment&lt;br&gt;
Integrates directly with CloudFront, ALB, and API Gateway — no extra DNS hop or proxy layer needed if you're already on AWS&lt;br&gt;
Rule groups are managed by AWS or the marketplace, so you're mostly configuring rather than building detection logic yourself&lt;/p&gt;

&lt;p&gt;Google Cloud Armor&lt;br&gt;
cloud.google.com/security/products/armor — GCP-native edge WAF with built-in DDoS protection&lt;br&gt;
Adaptive protection uses ML to flag volumetric attacks before they saturate your backend&lt;br&gt;
Best fit if your workloads already sit behind Google's global load balancer — less useful as a bolt-on for non-GCP infrastructure&lt;/p&gt;

&lt;p&gt;Azure Web Application Firewall&lt;br&gt;
azure.microsoft.com — same idea as the AWS/GCP options, built for the Microsoft stack&lt;br&gt;
Runs on Azure Front Door or Application Gateway, with OWASP core rule sets available out of the box&lt;br&gt;
Makes sense mainly if your compliance or procurement requirements already have you locked into Azure&lt;/p&gt;

&lt;p&gt;DataDome&lt;br&gt;
datadome.co — dedicated bot-detection SaaS with real-time ML scoring on every request&lt;br&gt;
Popular in e-commerce specifically for stopping scraper bots, scalpers, and card-testing fraud at checkout&lt;br&gt;
Priced and built for businesses with real traffic volume — overkill for a personal VPS or small site&lt;/p&gt;

&lt;p&gt;Kasada&lt;br&gt;
kasada.io — bot mitigation aimed specifically at sophisticated, human-mimicking bots&lt;br&gt;
Uses dynamic, polymorphic JavaScript challenges that change per-request, making it harder for bots to reverse-engineer than static CAPTCHAs&lt;br&gt;
Enterprise-tier pricing and onboarding — positioned against the most advanced scraper and account-takeover operations, not casual traffic&lt;/p&gt;

&lt;p&gt;HUMAN Security (formerly PerimeterX)&lt;br&gt;
humansecurity.com — enterprise bot and fraud detection, one of the largest players in the space&lt;br&gt;
Runs the data behind a lot of the industry-wide bot-traffic statistics you'll see cited (including some in this article)&lt;br&gt;
Covers a broad surface: account takeover, ad fraud, scraping, and now agentic-AI traffic classification&lt;/p&gt;

&lt;p&gt;F5 Distributed Cloud Bot Defense&lt;br&gt;
f5.com — enterprise-grade behavioral fingerprinting built on F5's networking heritage&lt;br&gt;
Strong fit for organizations that already run F5 load balancers or application delivery controllers&lt;br&gt;
Deep telemetry (mouse movement, device signals, request timing) rather than just IP or header-based rules&lt;/p&gt;

&lt;p&gt;Imperva / Thales&lt;br&gt;
imperva.com — full WAF plus bot management suite, and the source of the annual Bad Bot Report cited earlier in this article&lt;br&gt;
Combines DDoS protection, WAF, API security, and bot management under one enterprise contract&lt;br&gt;
Geared toward large organizations that want a single vendor covering the whole edge security stack&lt;/p&gt;

&lt;p&gt;Radware Bot Manager&lt;br&gt;
radware.com — another enterprise-tier option, positioned alongside Imperva and F5&lt;br&gt;
Emphasizes API-specific bot protection alongside standard web traffic, which matters if you're exposing public APIs&lt;br&gt;
Typically deployed by mid-to-large enterprises already running Radware's application delivery products&lt;/p&gt;

&lt;p&gt;Sucuri / Wordfence&lt;br&gt;
sucuri.net · wordfence.com — worth naming specifically for WordPress, since none of the above tools are WP-native&lt;br&gt;
Both bundle malware scanning and cleanup alongside firewall/bot-blocking — useful if you're managing a WordPress site rather than raw infrastructure&lt;br&gt;
Much lower cost of entry than the enterprise options above, with plans built for individual site owners and small agencies&lt;/p&gt;

&lt;p&gt;Good luck out there, meat proxies; you will need it.&lt;/p&gt;

&lt;p&gt;If you are interested in learning more about Vektor open source: &lt;a href="https://vektormemory.com/opensource" rel="noopener noreferrer"&gt;https://vektormemory.com/opensource&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Full setup and docs are at vektormemory.com/docs and 1.9.1 info at &lt;a href="https://vektormemory.com/docs/changelog" rel="noopener noreferrer"&gt;https://vektormemory.com/docs/changelog&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cybersecurity</category>
      <category>bots</category>
    </item>
    <item>
      <title>The Quiet Awakening of the 2-Billion-Parameter Model</title>
      <dc:creator>Vektor Memory</dc:creator>
      <pubDate>Fri, 11 Sep 2026 04:31:31 +0000</pubDate>
      <link>https://dev.to/vektor_memory_43f51a32376/the-quiet-awakening-of-the-2-billion-parameter-model-hgg</link>
      <guid>https://dev.to/vektor_memory_43f51a32376/the-quiet-awakening-of-the-2-billion-parameter-model-hgg</guid>
      <description>&lt;p&gt;I don't normally do model reviews or comparisons anymore, as I have moved on to much larger technical projects. I was scrolling through Ollama’s model list for a current testing model and was surprised at the lack of the smaller models in their list.&lt;/p&gt;

&lt;p&gt;Why has everything moved on to cloud-based models?&lt;/p&gt;

&lt;p&gt;So I went on a hunt in Hugging Face to find this model MiniCPM5–2B with very impressive numbers in their benchmarks, and thought, "Why not give it a spin on the new upgrades we have made to Vektors' desk app this week?"&lt;/p&gt;

&lt;p&gt;And I have to say I’m reasonably impressed with the results. After fine-tuning our harness, the search functions, and the reranking functions, the results this model gives are excellent for its size.&lt;/p&gt;

&lt;p&gt;A 2.5-billion-parameter language model is answering a tool-calling prompt on a machine with Serp search perfectly. Under 3-gigabyte GGUF file on disk, doing something that 6 months ago would have required a model much larger.&lt;/p&gt;

&lt;p&gt;Two years ago we assumed tasks like this needed hundreds of billions of parameters just to produce coherent output, let alone correct code. A model small enough to fit in a phone’s storage now beats models twice its size at exactly that. It’s a compression of capability that changes what “small” is allowed to produce.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiapkp27829rsos91rys2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiapkp27829rsos91rys2.png" alt=" " width="786" height="606"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That model is MiniCPM5–2B, was released by OpenBMB on September 7, 2026.&lt;/p&gt;

&lt;p&gt;It completes a pattern I’ve been watching for two years: the point where “small open-source model” stopped meaning “toy” and started meaning “legitimate production choice.” This article uses one model as a case study, but it’s really about what changes when a model this size can do this much.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the benchmarks actually say&lt;/strong&gt;&lt;br&gt;
On the Artificial Analysis Intelligence Index, MiniCPM5–2B ranks first among open-source models under 4 billion parameters. Across an internal suite of 34 benchmarks spanning code, math, long-context understanding, tool use, and agentic tasks, it averages 53.9.&lt;/p&gt;

&lt;p&gt;That beats not just other 2B-class models but several 4B-class ones too, including Qwen3.5–4B and Nemotron-3-Nano-4B, whose best score in the comparison set tops out around 51.1.&lt;/p&gt;

&lt;p&gt;It’s a dense, 42-layer transformer using the standard LlamaForCausalLM structure, not some exotic new block. Attention is grouped-query, 16 query heads to 2 key/value heads, which is what lets it hold a native 128K-token context window without KV cache size becoming absurd on consumer hardware.&lt;/p&gt;

&lt;p&gt;It’s released under Apache 2.0, no usage restrictions, no revenue-threshold clauses. Quantized builds run from a 1.04 GB Q2_K file up to a 2.68 GB Q8_0 file, with the commonly recommended Q4_K_M landing around 1.56 GB.&lt;/p&gt;

&lt;p&gt;That’s a model you can fit on a mobile phone or smaller VRAM GPU if needed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwx8kncfjn5cq8ej9vj7i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwx8kncfjn5cq8ej9vj7i.png" alt=" " width="800" height="628"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vektor Desk tool call&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Intelligence density is the story scale can’t tell&lt;br&gt;
For most of the last five years the dominant AI narrative was scale: bigger models, bigger datasets, bigger clusters, bigger bills. GPT-3 to GPT-4 and beyond is a story about more. There’s a second story running quieter in parallel, and it’s the one that matters if you’re trying to ship a product instead of win a leaderboard: intelligence density, meaning how much capability you can pack per parameter, per gigabyte, per watt.&lt;/p&gt;

&lt;p&gt;MiniCPM’s release history makes this concrete. The original MiniCPM-2B, back in 2024, was pitched as competitive with Mistral-7B and reportedly outperformed Llama2–13B, MPT-30B, and even Falcon-40B on several benchmarks despite being a small fraction of their size.&lt;/p&gt;

&lt;p&gt;MiniCPM5 continues that same trajectory rather than reinventing it: a 1B checkpoint in May 2026, then this 2B checkpoint in September, described by OpenBMB as the same training recipe scaled up. No new architecture, no new bet. Just more disciplined execution of a formula that treats parameter count as a cost to minimize instead of a number to maximize.&lt;/p&gt;

&lt;p&gt;OpenBMB reports training on roughly 550 billion tokens of curated code data (a set they call UltraData-Code) plus around 500,000 agentic and tool-use samples.&lt;/p&gt;

&lt;p&gt;They built a reinforcement learning stack tuned for small models specifically, rather than shrinking a frontier-scale RL pipeline down and hoping it still worked. And they made conservative architectural choices, favoring the well-understood Llama backbone over something novel, which matters more than it sounds like it should because it means every serving engine that already supports Llama models supports this one too.&lt;/p&gt;

&lt;p&gt;The part I find most telling isn’t the weights releasing themselves. It’s that OpenBMB published the training data, the recipe, and the RL stack alongside the model. It’s an invitation to reproduce and audit the work, which is what makes something feel like infrastructure rather than a product demo.&lt;/p&gt;

&lt;p&gt;Production readiness isn’t a benchmark score. It’s a bundle of unglamorous requirements.&lt;/p&gt;

&lt;p&gt;Can it run on hardware you already own?&lt;/p&gt;

&lt;p&gt;Can you audit what it was trained on?&lt;/p&gt;

&lt;p&gt;Can your infrastructure team integrate it without forking a codebase?&lt;/p&gt;

&lt;p&gt;Can you fine-tune it without a dedicated research team?&lt;/p&gt;

&lt;p&gt;Is the license actually permissive, with no asterisk hiding in section four?&lt;/p&gt;

&lt;p&gt;MiniCPM5–2B answers yes to all of those in ways frontier models structurally can’t. Because it uses the standard Llama architecture, vLLM (version 0.21.0 and up) loads it natively with no custom kernels and no model-code fork. A single line, vllm serve openbmb/MiniCPM5-2B --port 8000, gets it running, and at 2.5B parameters it fits comfortably on one GPU with tensor parallelism of 1.&lt;/p&gt;

&lt;p&gt;It ships native tool calling using an XML-style format, with a dedicated parser already merged into vLLM under the name minicpm5, so the usual weeks-long slog of getting an agent framework to actually parse a model's tool calls correctly is mostly closed on day one.&lt;/p&gt;

&lt;p&gt;For teams doing high-throughput serving, OpenBMB also released MiniCPM5-2B-DSpark, a purpose-built speculative decoding draft model trained specifically to pair with the 2.5B base rather than pointing users at a generic off-the-shelf drafter. That's someone thinking about latency in production, not just accuracy on a leaderboard.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What changes when the model fits on your machine&lt;/strong&gt;&lt;br&gt;
When a model is small enough to run on-device, whole categories of problems dissolve. Data residency stops being a legal negotiation and becomes a default. Latency stops being a network round trip and becomes a function call. Cost stops being a per-token line item and becomes hardware you already paid for.&lt;/p&gt;

&lt;p&gt;There’s a specific kind of risk that comes from building a product on top of an API you don’t control: the quiet awareness that your cost structure and sometimes your product’s core capability are token costs that are uncontrolled by you.&lt;/p&gt;

&lt;p&gt;Small open models don’t erase every version of that risk. You can still depend on a hosting provider, or a fine-tune you never validated, or a community fork that goes stale. But they hand you the option of self-sufficiency. Whether or not you take that option, having it changes your negotiating position with every vendor you do choose to rely on.&lt;/p&gt;

&lt;p&gt;Even as a fallback in your waterfall of model providers locally, having a workhorse smaller model that performs consistently is reassuring.&lt;/p&gt;

&lt;p&gt;None of this means frontier, closed, API-served models are going away, or that they should. There are problems, genuinely open-ended reasoning, tasks needing the broadest possible world knowledge, and work where being wrong costs far more than a few extra cents per token, where the biggest available model is still the right engineering call.&lt;/p&gt;

&lt;p&gt;A 2B model isn’t going to replace a frontier model at the edge of what language models can do. But that’s a narrower slice of real production workloads than the industry’s default assumptions suggest.&lt;/p&gt;

&lt;p&gt;Structured extraction, tool-calling agents with a bounded action space, code completion inside a known codebase, summarization, classification, on-device assistants: none of that needs the biggest model available. It needs the smallest model that reliably clears the bar, because every parameter above that bar is pure cost. More latency, more GPU spend, more attack surface, more dependency.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where this leaves us&lt;/strong&gt;&lt;br&gt;
What OpenBMB shipped is a specific model with a specific benchmark table, and benchmark tables age badly. Some other lab will publish a 2B model in six months that beats this one, the way this one beat its own predecessors.&lt;/p&gt;

&lt;p&gt;That's progress in the game of technology, an unstoppable organism.&lt;/p&gt;

&lt;p&gt;The gap between “biggest model available” and “good enough model” has been shrinking steadily, and it shrank again this week, in public, with the training recipe attached.&lt;/p&gt;

&lt;p&gt;After completing all of our adjustments in the desk code files, the multiple tool calls this model handled correctly on a 3060 12 GB VRAM test machine that is a very average testing bench.&lt;/p&gt;

&lt;p&gt;There was no cloud dashboard telling me what it cost in tokens, which is a real stress. There was just a small file doing a job very well on hardware I already owned.&lt;/p&gt;

&lt;p&gt;Just a small whine from the GPU coils doing their job…&lt;/p&gt;

&lt;p&gt;If you are interested in learning more about Vektor:&lt;/p&gt;

&lt;p&gt;Full setup and docs are at vektormemory.com/docs and 1.9.0 info at &lt;a href="https://vektormemory.com/docs/changelog" rel="noopener noreferrer"&gt;https://vektormemory.com/docs/changelog&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Sources&lt;/p&gt;

&lt;p&gt;Hugging Face: huggingface.co/openbmb/MiniCPM5–2B&lt;br&gt;
GitHub: github.com/OpenBMB/MiniCPM&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>VEKTOR Slipstream v1.9.0 updates</title>
      <dc:creator>Vektor Memory</dc:creator>
      <pubDate>Tue, 08 Sep 2026 11:06:02 +0000</pubDate>
      <link>https://dev.to/vektor_memory_43f51a32376/vektor-slipstream-v190-updates-26nj</link>
      <guid>https://dev.to/vektor_memory_43f51a32376/vektor-slipstream-v190-updates-26nj</guid>
      <description>&lt;p&gt;Agentic memory you can trust via any method you decide to use it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp4ekstfy576x9muyj70m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp4ekstfy576x9muyj70m.png" alt=" " width="786" height="537"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                       Custom code image
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;An agent with memory is only as good as its weakest read path. If a note can silently truncate, if a model switch can silently fall back to something you didn’t choose, the system isn’t really persistent. It’s just quiet about the parts where it fails.&lt;/p&gt;

&lt;p&gt;v1.9.0 is built around a single premise: consistency is a feature, not a maintenance chore. This release makes VEKTOR behave the same way whether you’re in Desk, the Agent tab, JOT, or the terminal. Same provider logic guarantee and connector security.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The provider waterfall, unified&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every interface in VEKTOR now follows one rule: try the model you selected first, then cascade through every other provider you’ve configured, and treat a local Ollama model as the genuine last resort, not the accidental default.&lt;/p&gt;

&lt;p&gt;If you like, you can set your Ollama LLM first and use 100% local and private inference with an open-source model.&lt;/p&gt;

&lt;p&gt;The Agent tab’s chat and the CLI each carry their own provider logic, built at different times, and both had quietly fallen behind. Select Cerebras in the Agent tab and it would drop straight to Ollama, because that code path only recognized four providers. Select Gemini or Mistral from the CLI and the same thing happened for a different reason: both were listed as valid choices with no implementation behind them at all.&lt;/p&gt;

&lt;p&gt;We rebuilt this as one shared chain, used by all four surfaces. In practice, that means a request now survives real failure. A live test after the fix routed through Claude, then OpenAI, then landed on Groq, each failure logged and handled in sequence rather than surfaced as a dead end. The model you picked is the model that runs, and if it can’t, you find out why instead of getting an answer from something else with no explanation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi7imtkip8pjtgyx19gbz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi7imtkip8pjtgyx19gbz.png" alt=" " width="800" height="394"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                          Desk brief updated
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Updated provider list&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Supported across Desk, Agent tab, JOT, and CLI (same waterfall logic everywhere):&lt;/p&gt;

&lt;p&gt;Claude (Anthropic)&lt;br&gt;
OpenAI&lt;br&gt;
Groq&lt;br&gt;
Gemini (Google) — newly added to CLI this release&lt;br&gt;
Cerebras — newly added across GUI, Agent tab, JOT, and CLI this release&lt;br&gt;
Mistral — newly added to CLI this release&lt;br&gt;
xAI (Grok)&lt;br&gt;
OpenRouter&lt;br&gt;
Ollama — final local fallback, unless chosen as primary&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Models&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;OpenAI&lt;/p&gt;

&lt;p&gt;Astra 6&lt;/p&gt;

&lt;p&gt;Cerebras&lt;/p&gt;

&lt;p&gt;qwen-3.8-27b&lt;br&gt;
gemma-4-31b&lt;br&gt;
gpt-oss-120b&lt;/p&gt;

&lt;p&gt;Gemini&lt;/p&gt;

&lt;p&gt;Removed gemini-2.5-flash-lite (Google deprecated it for new users)&lt;br&gt;
Default now points to gemini-3.5-flash-lite&lt;/p&gt;

&lt;p&gt;Groq&lt;/p&gt;

&lt;p&gt;qwen3.8–27b added to the lineup&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Twelve connectors, one security layer&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;VEKTOR now ships 12 live connectors: GitHub, Slack, GitLab, Vercel, HuggingFace, Jira, Linear, Sentry, and Notion, fully tested and running today. Gmail, Google Drive, and Microsoft Teams ship OAuth-gated on your own app credentials, not a shared client key, because those three touch a real inbox and a real file system.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8k2jkba8ymbtixixxr2j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8k2jkba8ymbtixixxr2j.png" alt=" " width="799" height="394"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                              Connector panel
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Every one of them sits behind Faraday, the security gate that screens MCP tool calls before they reach memory.&lt;/p&gt;

&lt;p&gt;This release adds OWASP LLM Top 10 and MITRE ATLAS classification to every event Faraday logs, and a behavioral anomaly baseline that flags a tool call whose frequency breaks from its own history. Faraday’s signature self-test runs 23 known attack patterns against 23 benign controls. Zero false positives, both directions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The agent engine, restored to full strength&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This release rebuilds and improves all 53 functions in the autonomous agent core: semantic code search, subagent delegation, lint and test tooling, code review, self-evaluation, commit message generation, and the hook system that ties them together.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Built to catch its own mistakes&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A golden-recall suite runs fixed queries against fixed seed memories with known expected results, built specifically to catch the failure mode where a system still answers, just not correctly. A smoke test sweeps roughly 240 files on every install. Both now gate every commit before it lands, not after a user finds the gap.&lt;/p&gt;

&lt;p&gt;The CLI regression suite runs 41 checks across every surface touched this release, making sure all passed the quality code inspection.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What this means going forward&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Persistent memory is only as trustworthy as its least-tested edge. This release closes several of those edges at once: a note that no longer loses its own content, a model choice that no longer gets silently overridden, and a connector layer with the same security guarantee no matter which one you’re using.&lt;/p&gt;

&lt;p&gt;The current release builds on our privacy and security foundations rather than patching around it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why we build it this way&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;None of this works if the underlying architecture isn’t private by default.&lt;/p&gt;

&lt;p&gt;VEKTOR computes embeddings in-process, on your device, with a local ONNX model. There is no cloud copy of your memory graph to log, breach, or subpoena, because there’s no cloud copy at all.&lt;/p&gt;

&lt;p&gt;That’s not a policy we ask you to trust. It’s a property of the design: air-gapped by default, so surveillance is structurally impossible rather than contractually prohibited.&lt;/p&gt;

&lt;p&gt;Your memory lives as a full SQLite file, on your own machine, with a real file path.&lt;/p&gt;

&lt;p&gt;It’s never moved to mandatory cloud storage, never used to train a model, your data is never sold for ads, and it’s yours to copy, migrate, or walk away with at any time via our open-sourced DB migration tools.&lt;/p&gt;

&lt;p&gt;We built VEKTOR this way because a memory layer that centralizes someone’s most private conversational history has no business asking for blind trust.&lt;/p&gt;

&lt;p&gt;It should be something you can verify yourself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Get v1.9.0&lt;/strong&gt;&lt;br&gt;
Read the full changelog → &lt;a href="https://vektormemory.com/docs/changelog" rel="noopener noreferrer"&gt;https://vektormemory.com/docs/changelog&lt;/a&gt;&lt;br&gt;
Vector Database&lt;/p&gt;

&lt;p&gt;Agentic Ai&lt;/p&gt;

&lt;p&gt;Information Security&lt;/p&gt;

&lt;p&gt;Ai Memory&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>The Case for Agent Memory: Why the Future of AI Belongs to Systems That Can Self-Learn</title>
      <dc:creator>Vektor Memory</dc:creator>
      <pubDate>Thu, 20 Aug 2026 22:31:52 +0000</pubDate>
      <link>https://dev.to/vektor_memory_43f51a32376/the-case-for-agent-memory-why-the-future-of-ai-belongs-to-systems-that-can-self-learn-130</link>
      <guid>https://dev.to/vektor_memory_43f51a32376/the-case-for-agent-memory-why-the-future-of-ai-belongs-to-systems-that-can-self-learn-130</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F37fjackcsshazeswxjgb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F37fjackcsshazeswxjgb.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;br&gt;
Custom-generated code image&lt;/p&gt;

&lt;p&gt;In 1985, the British musicologist Clive Wearing suffered a severe viral infection that damaged his hippocampus, the brain structure responsible for converting experience into memory.&lt;/p&gt;

&lt;p&gt;What followed was one of neurology’s most studied cases of dense anterograde amnesia. His ability to retain new information collapsed to a span of about thirty seconds. When his wife Deborah entered the room, he would greet her with the ecstatic joy and surprise of a reunion after years apart. Moments later, she was a stranger again.&lt;/p&gt;

&lt;p&gt;Clive’s condition was tragic precisely because it was so isolating. Intelligence remained intact. He could still play music brilliantly and still recognize his own handwriting. Clive kept a written daily journal to affirm his existence and continuity, but after several days it consisted of identical pages, making the exercise futile.&lt;/p&gt;

&lt;p&gt;But without the capacity to build on what came before, to connect this moment to that moment, to accumulate understanding of who Deborah was and what their relationship meant, his extraordinary mind became trapped in an eternal present.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0fe611gvb7r1pk4bsxtj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0fe611gvb7r1pk4bsxtj.png" alt=" " width="646" height="453"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Clive playing piano—Photograph by Jiri Rezac&lt;/p&gt;

&lt;p&gt;Today, every large language model deployed in the world suffers in some form the digital equivalent of what Clive Wearing endured. These systems possess astonishing crystalline intelligence. They can parse complex regulatory documents in milliseconds, debug intricate code, draft legal arguments with the polish of an elite associate.&lt;/p&gt;

&lt;p&gt;And then the inference window closes. The HTTP request completes. The mind collapses back into void. When the user returns five minutes later, the model meets them as a complete stranger if they do not have a memory system built in.&lt;/p&gt;

&lt;p&gt;This is the fundamental architectural item that has been partially solved while the industry obsessed over model scale, training, and harnesses.&lt;/p&gt;

&lt;p&gt;Larger context windows and training help. But none of these address what actually determines whether an AI agent functions as a genuine collaborator or remains an expensive, sophisticated autocomplete.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Stateless Problem&lt;/strong&gt;&lt;br&gt;
The statelessness of current AI systems creates what researchers call context discontinuity, and it manifests as failure in the exact places where enterprises expect AI to add the most value. These are not single-turn tasks. They are multi-step workflows that span weeks, months, or years.&lt;/p&gt;

&lt;p&gt;A customer’s relationship with a bank involves hundreds of interactions over decades. An employee using an internal AI assistant builds context over months of work. A procurement process moves through dozens of approval stages across multiple sessions and multiple people. In every one of these scenarios, starting from scratch is not a minor inconvenience. It is a structural failure.&lt;/p&gt;

&lt;p&gt;The business consequences have been well documented in pilot postmortems. Context discontinuity is consistently cited as one of the primary reasons AI deployments stall after initial pilots. Users don’t file detailed technical complaints about memory architecture.&lt;/p&gt;

&lt;p&gt;They simply stop using the tool. The friction of having to reintroduce themselves repeatedly, to restate their role and their constraints, to copy and paste prior context into prompts to orient the model, eventually exceeds any benefit the tool provides.&lt;/p&gt;

&lt;p&gt;An AI agent without memory is not a collaborator. It is a stranger who has to be reintroduced every single day, it can be a very frustrating experience for users logging in and repeating tasks.&lt;/p&gt;

&lt;p&gt;Yet when you look at how most organizations are deploying AI, this problem is either invisible or treated as inevitable. They focus on the model’s quality: benchmark accuracy, parameter scale, and training data breadth. These things do matter, however, not what determines whether an AI agent creates lasting value in an enterprise context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Types of Agent Memory&lt;/strong&gt;&lt;br&gt;
Agent memory is the system that allows an AI agent to encode information from interactions, store it durably, retrieve it when relevant, and update it as conditions change.&lt;/p&gt;

&lt;p&gt;It is an architectural layer that sits alongside the model itself and determines what the model knows at any given moment. Without it, the model has no way to learn from experience. It cannot recognize patterns across interactions. It cannot build on success or learn from failure.&lt;/p&gt;

&lt;p&gt;The field has converged on a framework for thinking about AI agent memory drawn from cognitive science and adapted for modern AI systems. There are five interdependent dimensions to understand.&lt;/p&gt;

&lt;p&gt;Semantic memory is what the agent knows. It is the repository of facts that have been learned and kept: who the user is, what their role involves, what terminology means in this specific organization, what preferences they have expressed over time.&lt;/p&gt;

&lt;p&gt;Semantic memory is the foundation of personalization. Without it, an agent cannot respond to a returning user with genuine relevance. Every response would be generic.&lt;/p&gt;

&lt;p&gt;Episodic memory is what the agent has experienced. It is the record of actual interactions in their sequence and their outcomes. An agent with strong episodic memory knows not just that a user prefers a certain approach, but that they tried a different approach last month and abandoned it.&lt;/p&gt;

&lt;p&gt;They know what was promised in an earlier conversation and whether that promise was kept. Episodic memory is critical for long-running tasks where continuity is everything.&lt;/p&gt;

&lt;p&gt;Procedural memory is how the agent behaves. It consists of the encoded behaviors, rules, and guidelines that govern the agent’s operation. Communication protocols, escalation logic, compliance constraints, organizational policies. Procedural memory ensures an agent behaves consistently and appropriately within a given context.&lt;/p&gt;

&lt;p&gt;Working memory is what the agent knows right now. It corresponds directly to the active context window, the information available in the current moment. Working memory is temporary and bounded, but it is where all the other memory types converge.&lt;/p&gt;

&lt;p&gt;Relevant facts from semantic memory, pertinent history from episodic memory, applicable rules from procedural memory are all retrieved and assembled here before the agent responds.&lt;/p&gt;

&lt;p&gt;Then there is the fifth dimension, the one that cuts across all the others and the one most memory systems get wrong first: time. Semantic memory tells an agent what a user prefers. It does not tell the agent whether that preference was stated yesterday or eight months ago, or whether a newer, contradictory statement has superseded it.&lt;/p&gt;

&lt;p&gt;A flat store of episodic memories still has to answer questions like “what did we agree in the last session” or “has this fact changed since I last checked.” Pure vector similarity search has no native concept of before and after. A memory about a canceled project and one about that same project’s kickoff can be equally similar to a query while being months apart and mutually contradictory.&lt;/p&gt;

&lt;p&gt;This is why temporal reasoning has become its own research problem rather than a footnote. Without explicit mechanisms for tracking when something was true and whether it is still true, an agent’s memory degrades in a specific and dangerous way. It doesn’t become less confident as facts age. It stays just as confident while quietly becoming incorrect.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy93htz1p7zpddywbco1q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy93htz1p7zpddywbco1q.png" alt=" " width="750" height="335"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Four-Stage Pipeline&lt;/strong&gt;&lt;br&gt;
Understanding agent memory conceptually is different from building it well. Four interdependent processes determine whether a memory system genuinely improves agent performance or simply adds infrastructure overhead.&lt;/p&gt;

&lt;p&gt;In the first stage, extraction, not everything in a conversation deserves to be remembered. Effective memory systems apply judgment here, identifying facts, preferences, decisions, and outcomes likely to be relevant in future interactions, while discarding noise.&lt;/p&gt;

&lt;p&gt;A poorly calibrated extraction layer leads either to bloated, noisy memory stores or impoverished ones that fail to retain what actually matters.&lt;/p&gt;

&lt;p&gt;The second stage is storage. Extracted information needs to live somewhere accessible. Vector databases, which organize information by semantic meaning rather than exact keywords, are the most common mechanism for persistent agent memory.&lt;/p&gt;

&lt;p&gt;They allow retrieval of relevant memories even when phrasing differs from how they were originally recorded. For relationship-rich knowledge, graph databases offer something more. They store not just facts but the connections between them, enabling more sophisticated reasoning over time.&lt;/p&gt;

&lt;p&gt;The third stage is consolidation. Memory systems accumulate contradictions. Preferences change. Facts become outdated. New policies supersede old ones.&lt;/p&gt;

&lt;p&gt;Without consolidation, the process of comparing incoming information against existing records and resolving conflicts, agent memory degrades. The system must determine whether new information should be added, used to update an existing record, or discarded as redundant. Memory systems that skip this stage tend to become liabilities as they scale.&lt;/p&gt;

&lt;p&gt;The fourth stage is retrieval. When the agent is ready to respond, it searches the memory store for relevant information and pulls it into the active context window. Retrieval quality is where most AI agent memory systems succeed or fail in practice.&lt;/p&gt;

&lt;p&gt;Retrieving too much floods the context with noise. Retrieving too little leaves the agent under-informed. The best retrieval systems surface what is genuinely relevant quickly enough not to degrade response latency.&lt;/p&gt;

&lt;p&gt;The pipeline is only as strong as its weakest stage. Organizations that invest in storage infrastructure without investing equally in extraction quality and consolidation logic will find their memory systems becoming less reliable, not more, as they accumulate data over time.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzk6ew7cpsdreextid1rh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzk6ew7cpsdreextid1rh.png" alt=" " width="750" height="325"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Noise &amp;amp; Staleness Challenges&lt;/strong&gt;&lt;br&gt;
The case for agent memory is obvious in theory; underestimating the implementation challenges is one of the most common mistakes organizations make.&lt;/p&gt;

&lt;p&gt;The staleness problem is first. Information changes continuously. A user’s role, budget, preferences, and circumstances evolve. An AI agent memory system that stores information without managing its freshness will, over time, provide the agent with confidently stated but incorrect context.&lt;/p&gt;

&lt;p&gt;Addressing staleness requires explicit lifecycle management: tracking when memories were created, monitoring for contradictory updates, and invalidating records that are no longer accurate.&lt;/p&gt;

&lt;p&gt;The noise problem follows. Memory stores that grow without disciplined consolidation become progressively noisier. As the volume of stored information increases, retrieval surfaces more irrelevant results alongside relevant ones.&lt;/p&gt;

&lt;p&gt;The agent’s effective context degrades even as the amount of stored data grows. Aggressive deduplication and merging at write time is the solution, but it requires investment in consolidation logic that many early implementations skip.&lt;/p&gt;

&lt;p&gt;The governance problem is regulatory and operational. When an AI agent remembers user information, that data is subject to regulatory obligations. GDPR, CCPA, and a growing body of privacy legislation grant users rights over their stored data: the right to access it, correct it, and have it deleted.&lt;/p&gt;

&lt;p&gt;Enterprise agent memory architecture must treat stored memories as first-class data objects with explicit metadata: source, timestamp, confidence, and lifecycle status. This is a design requirement from day one, not a compliance layer to be added later.&lt;/p&gt;

&lt;p&gt;The multi-agent problem emerges in mature deployments. Enterprise AI increasingly involves networks of specialized agents operating in coordination. Each agent may need access to shared user context, but with appropriate boundaries.&lt;/p&gt;

&lt;p&gt;A customer-facing service agent should know a user’s account history. It should not have access to sensitive data from an unrelated HR interaction. Memory scoping, which defines what each agent can read and write, requires architectural decisions that many organizations defer until forced to address them by an incident.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agent Memory Versus RAG&lt;/strong&gt;&lt;br&gt;
Retrieval-Augmented Generation is powerful and widely deployed. It is also frequently confused with agent memory in ways that lead to underinvestment in persistent memory systems. The distinction matters because they solve fundamentally different problems.&lt;/p&gt;

&lt;p&gt;RAG gives an AI agent access to a knowledge base at the moment of inference. It retrieves relevant documents and injects them into the prompt, grounding the response in external information the model was not trained on. RAG is read-only. The knowledge base does not change based on user interactions. Every user draws from the same corpus. The agent does not learn from what it experiences.&lt;/p&gt;

&lt;p&gt;Agent memory is fundamentally different. It reads and writes. It changes based on what users do and say. It is personal, specific to a user, a team, or an organization, rather than universal. And it compounds. The more the AI agent is used, the more it knows about the context it is operating in.&lt;/p&gt;

&lt;p&gt;Think of it this way: RAG gives every user access to the same encyclopedia. Agent memory gives each user their own record, one that grows more accurate and more useful with every interaction.&lt;/p&gt;

&lt;p&gt;The two are not mutually exclusive. The most capable enterprise agent architectures use both: RAG for broad organizational and domain knowledge, and persistent agent memory for the user-specific and interaction-specific context that makes responses genuinely relevant. But they are not substitutes. Organizations that treat RAG as sufficient for their memory needs are solving only half the problem.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4d14io4hxmpune0qaczj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4d14io4hxmpune0qaczj.png" alt=" " width="750" height="335"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How Memory Works in Practice: The Vektor Approach&lt;/strong&gt;&lt;br&gt;
Everything above describes what a memory system needs to do. There are real implementations now, and they map directly onto the four-stage pipeline.&lt;/p&gt;

&lt;p&gt;Consider how Vektor, a local-first Node.js SDK, handles each stage. Extraction is handled by what the system calls AUDN, an acronym standing for the four possible verdicts on any incoming memory. Every new fact that comes through the memory system passes through AUDN before it ever gets written: the system evaluates each incoming memory and decides ADD, UPDATE, DELETE, or NO_OP against what is already stored.&lt;/p&gt;

&lt;p&gt;That decision is what stops the noise problem at the source, rather than cleaning it up later. In production the loop currently runs with effectively zero duplicate bloat because the judgment call happens at write time, not as a periodic cleanup job.&lt;/p&gt;

&lt;p&gt;Storage is handled by MAGMA, a four-layer graph persisted in a single portable SQLite file. Where most agent memory systems store a flat list of vectors, this approach writes into four interdependent layers.&lt;/p&gt;

&lt;p&gt;The semantic layer tracks similarity between memories. The causal layer maps cause-and-effect relationships. The temporal layer records before-and-after sequences. The entity layer connects people, projects, and events that co-occur together.&lt;/p&gt;

&lt;p&gt;The temporal layer is the direct answer to the “when was this true” problem described earlier. It doesn’t just timestamp writes. It tracks sequence: which fact preceded which. It decays unused edges automatically, so a stale connection between two memories loses weight over time without needing a manual cleanup job.&lt;/p&gt;

&lt;p&gt;Combined with AUDN’s UPDATE and DELETE verdicts, a changed fact doesn’t sit alongside its outdated predecessor waiting to confuse a future query. The graph itself records that the relationship changed. This is how you answer the suggestion that graph databases add relationship-aware reasoning on top of plain vector similarity. You don’t choose one or the other. You run both, in the same file.&lt;/p&gt;

&lt;p&gt;Retrieval against this graph averages around 28 milliseconds, roughly an order of magnitude faster than a round-trip to a cloud memory API, because there is no network hop. The graph lives on the machine running the agent.&lt;/p&gt;

&lt;p&gt;Consolidation is handled by what the system calls REM, a seven-phase background process that runs while the agent is idle and directly targets the staleness and noise problems. It compresses roughly 50 raw memory fragments down to a single core insight, discarding around 98 percent of redundant signal while preserving what is still true.&lt;/p&gt;

&lt;p&gt;This is also where superseded facts get resolved. A changed preference or an outdated project detail doesn’t sit alongside the old version waiting to confuse a future retrieval. It gets reconciled during the REM pass, using the temporal layer’s before-and-after ordering to determine which version is current.&lt;/p&gt;

&lt;p&gt;Decaying edges and sequence tracking are not the same as a purpose-built temporal knowledge graph. For agents whose primary job is answering point-in-time questions across long histories, “what did we agree on three sessions ago, before the scope changed,” a temporal-first architecture with explicit validity intervals on every fact is a more direct fit.&lt;/p&gt;

&lt;p&gt;Teams whose workload is temporal reasoning first and everything else second should evaluate architectures built specifically for that problem alongside general-purpose options.&lt;/p&gt;

&lt;p&gt;Retrieval at query time runs a two-stage process. A fast bi-encoder pulls a shortlist of candidate memories from the graph. Then a cross-encoder re-ranks that shortlist for precision before anything reaches the agent’s context window. That two-stage design is what keeps retrieval from becoming either too noisy or too sparse. It is an explicit mechanism for the exact trade-off retrieval stage describes.&lt;/p&gt;

&lt;p&gt;On governance specifically: because the memory graph is a local SQLite file rather than a row in someone else’s cloud database, the “who can access, correct, or delete this memory” question has a simpler answer by default. The file is on the enterprise’s own infrastructure.&lt;/p&gt;

&lt;p&gt;Security layers can add MCP-level tool discovery, taint tracking, and canary-token exfiltration detection on top of that, aimed at the case where an agent’s memory or tool access becomes the attack surface rather than the model itself.&lt;/p&gt;

&lt;p&gt;On multi-agent scoping: because each agent instance is a discrete SQLite database keyed to an agentId, defining which agent can read or write which memory graph is a deployment decision, not a permissions system that has to be bolted on afterward.&lt;/p&gt;

&lt;p&gt;The net effect, measured on long-context memory recall benchmarks, is a memory layer that beats full-context retrieval from a local database, at roughly 28 millisecond average latency and zero per-call embedding fees, because it uses whatever LLM provider the enterprise already has a contract with rather than billing separately for embeddings on top of a memory subscription.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Compounding Advantage&lt;/strong&gt;&lt;br&gt;
The business case for agent memory goes far beyond operational efficiency. It represents the creation of a fundamentally new class of corporate asset: compounding institutional intelligence.&lt;/p&gt;

&lt;p&gt;An AI agent without memory delivers the same quality of output on its ten-thousandth interaction as on its first. It has learned nothing from the nine thousand nine hundred and ninety-nine interactions before. Every user experience is identical in its lack of personalization. Every workflow begins from zero.&lt;/p&gt;

&lt;p&gt;An agent with well-designed persistent memory improves with use. Each interaction adds to its understanding of the user, the domain, and the organization. Preferences are refined. Edge cases are learned. Patterns are recognized.&lt;/p&gt;

&lt;p&gt;The agent becomes, over time, genuinely knowledgeable about the context it operates in, not because the underlying model improved, but because the memory system has accumulated and organized the experience of every interaction that came before.&lt;/p&gt;

&lt;p&gt;Research from organizations studying memory-augmented LLM applications found that integrating persistent memory produces a 26 percent improvement in response quality.&lt;/p&gt;

&lt;p&gt;That is a significant performance uplift from an architectural addition rather than a model change. Organizations that invest in AI agent memory infrastructure are not waiting for better models. They are extracting substantially more value from the models they already have.&lt;/p&gt;

&lt;p&gt;The compounding dynamic also creates a durable competitive advantage. Memory stores built over months and years of real interaction are not replicable quickly. The institutional knowledge an AI agent accumulates about users, workflows, domain terminology, and organizational preferences represents a form of AI capital that appreciates with use and is genuinely difficult to replicate from a standing start.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where This Is Heading&lt;/strong&gt;&lt;br&gt;
Agent memory is moving quickly from an advanced capability to a baseline expectation. The major cloud providers have all announced or deployed managed memory services for their agentic AI platforms.&lt;/p&gt;

&lt;p&gt;Purpose-built agentic memory frameworks have emerged for organizations that need more control. The ecosystem is maturing at a pace that makes this a meaningful inflection point.&lt;/p&gt;

&lt;p&gt;Those that will lead in the next phase of AI adoption are not necessarily those with access to the most powerful models. They are the ones investing now in the infrastructure that makes those models durable: memory systems that turn one-off interactions into cumulative intelligence.&lt;/p&gt;

&lt;p&gt;The question is not whether AI agent memory matters. It is whether your architecture is designed to take advantage of it.&lt;/p&gt;

&lt;p&gt;As agent memory matures into a dedicated engineering discipline, standardized benchmarks like LoCoMo, LongMemEval, and BEAM have replaced self-reported metrics, offering rigorous multi-hop and temporal reasoning evaluations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Final Thoughts&lt;/strong&gt;&lt;br&gt;
Every time an enterprise deploys a state-of-the-art large language model, it performs a subtle, modern miracle. And then, instantly, it commits an act of institutional amnesia.&lt;/p&gt;

&lt;p&gt;The model possesses astonishing intelligence for the duration of that inference window. Then everything vanishes. The next interaction starts from zero.&lt;/p&gt;

&lt;p&gt;What determines whether an AI agent becomes a genuine coworker or remains an expensive autocomplete is not its raw capability. It is whether the system can remember. Whether it can build on what came before. Whether it compounds value across time instead of delivering the same generic response on its first interaction and its ten-thousandth.&lt;/p&gt;

&lt;p&gt;The AI arms race has largely been fought in the realm of raw compute: who can train the largest models, secure the densest GPU clusters, expand the widest context windows. In the enterprise trenches, the decisive competitive advantage will not belong to the organizations running the largest models.&lt;/p&gt;

&lt;p&gt;It will belong to those whose systems have advanced persistent memory.&lt;/p&gt;

&lt;p&gt;VEKTOR Memory builds local-first, privacy-preserving persistent memory infrastructure for AI agents. Full setup instructions for Claude Desktop, Claude Code, Cursor, Windsurf, the OpenAI Agents SDK, and OpenRouter, including the manual config for each. Technical documentation and changelog at vektormemory.com/docs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Architectural Benchmarks &amp;amp; Literature:&lt;/strong&gt;&lt;br&gt;
Tulving, E. (1983). Elements of Episodic Memory. Oxford University Press.&lt;br&gt;
Newell, A. (1990). Unified Theories of Cognition. Harvard University Press.&lt;br&gt;
Packer, C. et al. (2023). MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560.&lt;br&gt;
Chhikara, P. et al. (2024). Mem0: The Memory Layer for Personalized AI.&lt;br&gt;
LongMemEval Benchmark (2024). Evaluating Long-Context Retrieval and Retention Across Multi-Turn Horizon Benchmarks.&lt;br&gt;
TReMu Framework (2024). Time-aware Reasoning and Memorization for Autonomous Large Language Model Architectures.&lt;br&gt;
info&lt;/p&gt;

&lt;p&gt;AI Agent&lt;br&gt;
Ai Memory&lt;br&gt;
LLM&lt;br&gt;
Agentic Rag&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>vectordatabase</category>
    </item>
    <item>
      <title>SynthID watermarking and removal methods are a joke. And you are misunderstanding how it all works completely.</title>
      <dc:creator>Vektor Memory</dc:creator>
      <pubDate>Sun, 16 Aug 2026 22:11:11 +0000</pubDate>
      <link>https://dev.to/vektor_memory_43f51a32376/synthid-watermarking-and-removal-methods-are-a-joke-and-you-are-misunderstanding-how-it-all-works-4lp8</link>
      <guid>https://dev.to/vektor_memory_43f51a32376/synthid-watermarking-and-removal-methods-are-a-joke-and-you-are-misunderstanding-how-it-all-works-4lp8</guid>
      <description>&lt;p&gt;SynthID watermarking and removal methods are a joke. And you are misunderstanding how it all works completely.&lt;/p&gt;

&lt;p&gt;Somewhere in this sentence you just read, if a large language model had written it, there would be custom code generated in the image&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsxl4k2txijx3h75hpzms.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsxl4k2txijx3h75hpzms.png" alt=" " width="786" height="537"&gt;&lt;/a&gt;&lt;br&gt;
                     Custom code generated image&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Like this:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The model would normally pick any of: [‘signature’, ‘mark’, ‘trace’, ‘fingerprint’]&lt;br&gt;
watermark nudges it to pick: signature score of the word actually used: 87 &amp;gt; score of a word that lost: 71&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A statistical bias buried in the choice of “somewhere” over “somewhere else,” or “signature” over “mark.”&lt;/p&gt;

&lt;p&gt;You would never notice it, as you are a human meat popsicle, not a binary code wizard LLM with SynthID and oodles of books from Libgen.&lt;/p&gt;

&lt;p&gt;Neither would a spellchecker, a copy-paste, or a screenshot. But a machine holding the right key could look at that sentence and tell you, with real confidence, that it came from a specific model.&lt;/p&gt;

&lt;p&gt;Also, you don't have access to the key to decode. Sorry!&lt;/p&gt;

&lt;p&gt;This is not science fiction. As of August 2, 2026, Claude does this to every sentence it writes. So does Gemini. The reason is a piece of European law, and the mechanism is a 2024 Nature paper that most people who are affected by it have never read.&lt;/p&gt;

&lt;p&gt;I want to walk through exactly how this works, why it exists now, what it can and can’t tell anyone, and what happens when people try to strip it out.&lt;/p&gt;

&lt;p&gt;Along the way I’ll clear up a few things that get misconstrued every time this topic comes up online, including the idea that uploading text to the internet triggers some kind of automatic AI-detection scan. It doesn’t.&lt;/p&gt;

&lt;p&gt;Nothing does that. Not yet, and not the way people picture it or place it on GitHub to remove it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The paper that started this&lt;/strong&gt;&lt;br&gt;
In October 2024, a team at Google DeepMind led by Sumanth Dathathri and Abigail See published a paper in Nature called “Scalable watermarking for identifying large language model outputs.” The system they described is called SynthID-Text, and it solved a problem that had stalled watermarking research for years: how do you mark AI-generated text without making it worse, without slowing it down, and without needing to store a copy of everything the model ever said.&lt;/p&gt;

&lt;p&gt;Before SynthID-Text, the field had roughly three options, and all of them had real drawbacks. Keep a growing database of everything the model generated and check new text against it, which raises obvious privacy problems and needs infrastructure that scales with usage forever.&lt;/p&gt;

&lt;p&gt;Train a separate classifier to spot the statistical “flavor” of AI writing, which is the approach behind most of the AI detection tools you’ve probably already used and distrusted, and for good reason: those tools are known to misfire on non-native English writers and degrade as models improve. Or edit the text after it’s generated, swapping in synonyms or inserting invisible characters, which leaves traces a careful reader or a decent script can find and strip.&lt;/p&gt;

&lt;p&gt;SynthID-Text took a different approach entirely. Instead of marking the text after it exists, it changes how the text gets chosen in the first place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How a language model actually picks its next word&lt;/strong&gt;&lt;br&gt;
To understand the watermark, you need to understand what happens underneath every response a model gives you. An LLM doesn’t write a sentence the way a person does, deciding on a whole thought and typing it out. It predicts one token at a time.&lt;/p&gt;

&lt;p&gt;Given everything written so far, it calculates a probability for every possible next token, something like a 40% chance the next word is “the,” a 12% chance it’s “this,” and so on across the entire vocabulary. Then it samples from that distribution and moves to the next position.&lt;/p&gt;

&lt;p&gt;Normally, that sampling step is close to random, shaped by settings like temperature that control how adventurous or predictable the choices are. SynthID-Text inserts itself right there, at the moment of sampling, and quietly tilts the odds.&lt;/p&gt;

&lt;p&gt;Here’s the mechanism, as described in the paper’s Methods section. For each token position, a hash function takes the last four tokens of context plus a secret key and produces a random seed.&lt;/p&gt;

&lt;p&gt;That seed feeds a set of pseudorandom scoring functions, the paper uses 30 of them, called “layers.” Each function assigns a score to every possible next token. Then, instead of sampling once from the model’s distribution, the algorithm samples several candidate tokens and runs them through what the authors call Tournament sampling: a knockout bracket.&lt;/p&gt;

&lt;p&gt;Candidates get paired up, the higher-scoring one under the first scoring function survives, the survivors get paired again and scored by the second function, and so on through all 30 layers until one token wins and becomes the actual output.&lt;/p&gt;

&lt;p&gt;The result is a sequence of words that a reader can’t distinguish from an unwatermarked response, but that carries a statistical fingerprint recoverable by anyone holding the key. Detection doesn’t need the model at all. You just take the text, recompute the same seeds and scores using the key, average them, and compare the result to a threshold. Higher than chance, probably watermarked. Around chance, probably not.&lt;/p&gt;

&lt;p&gt;The paper is careful about a property it calls “non-distortion.” Configured one way, called single-token non-distortionary, the tournament always has exactly two competitors per match, and DeepMind proves mathematically that this leaves the model’s actual output probabilities unchanged on average.&lt;/p&gt;

&lt;p&gt;The watermark rides on which specific token gets picked among equally likely options, not on making some tokens artificially more likely overall. That’s the whole trick: real bias in individual choices, zero bias in the aggregate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Twenty million responses, and nobody could tell&lt;/strong&gt;&lt;br&gt;
Claims about quality preservation are easy to make and hard to trust, so DeepMind ran an actual test in production. They routed a portion of live Gemini traffic through the watermarked model and an equal portion through the unwatermarked version, then compared the thumbs-up and thumbs-down rates people gave each.&lt;/p&gt;

&lt;p&gt;Across close to 20 million responses, the difference in thumbs-up rate was 0.01 percent. The difference in thumbs-down rate was 0.02 percent. Both fell well inside the statistical noise.&lt;/p&gt;

&lt;p&gt;They backed that up with a smaller, controlled study: human raters comparing watermarked and unwatermarked Gemma 7B answers to 3,000 questions, scored on grammar, relevance, correctness, helpfulness, and overall quality. No significant difference on any of the five.&lt;/p&gt;

&lt;p&gt;This is the part that made SynthID-Text different from earlier watermarking research, and the part that made it viable to actually ship. A watermark nobody notices and nothing measures is a watermark companies can turn on without a quality tradeoff. That’s exactly what happened next, eighteen months later, and not because anyone wanted to.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The regulation that made this mandatory&lt;/strong&gt;&lt;br&gt;
In August 2026, Anthropic switched on watermarking across every Claude model launched on or after August 2, and did it globally, not just for European users. Their help center article is blunt about why: “We’re implementing watermarking to comply with the EU AI Act.”&lt;/p&gt;

&lt;p&gt;The specific hook is Article 50 of the Act, which became enforceable that same day and requires providers of generative AI systems to embed machine-readable marks in their output so platforms, regulators, and downstream users can identify AI-generated content.&lt;/p&gt;

&lt;p&gt;The penalties are not symbolic. Non-compliance can trigger fines up to 15 million euros or 3 percent of a company’s global annual revenue, whichever number is bigger. Anthropic signed the EU’s Code of Practice on Transparency of AI-Generated Content alongside roughly 190 other signatories.&lt;/p&gt;

&lt;p&gt;Anthropic’s technical implementation is built directly on SynthID-Text, the same tournament-sampling mechanism from the Nature paper, adapted for their own models and keys. It covers claude.ai, the API, Claude Code, Claude Cowork, Claude Tag, and access through AWS, Google Cloud, and Microsoft Foundry. Google’s Gemini already carries the original SynthID-Text mark.&lt;/p&gt;

&lt;p&gt;OpenAI has discussed watermarking and sits on the C2PA steering committee, but as of this writing has moved slower on deploying text watermarking at the same scale.&lt;/p&gt;

&lt;p&gt;Grok as of this article being created, is still undecided. The reality is the majority of the LLMs will have SynthID added already.&lt;/p&gt;

&lt;p&gt;Why apply it worldwide instead of only to European users? Anthropic’s own answer, from their FAQ, is refreshingly honest: “We’re applying watermarking globally at launch because we don’t yet have a durable way to scope it by region.”&lt;/p&gt;

&lt;p&gt;What the mark can tell you, and the two things it absolutely cannot&lt;br&gt;
This is where most of the online confusion sits:&lt;/p&gt;

&lt;p&gt;A watermark hit means the text may have passed through that specific model at some point. That’s the entire claim.&lt;/p&gt;

&lt;p&gt;Anthropic states this directly: “A watermark only helps test whether Claude might have produced or processed the content.”&lt;/p&gt;

&lt;p&gt;Might have processed, not definitely wrote from scratch. If you paste a paragraph you wrote yourself into Claude and ask it to fix the grammar, the output can carry the watermark even though the ideas and structure are entirely yours. The mark tracks which model touched the text, not who’s responsible for the thinking in it.&lt;/p&gt;

&lt;p&gt;It cannot identify you. Anthropic is explicit on this point too: “There’s nothing in the watermark, or its key, that would allow anyone to recover any information about the user, their organization, or their chats with Claude.”&lt;/p&gt;

&lt;p&gt;The seed comes from the recent token context and a model-level secret key, full stop. No account ID, no session, no IP address folds into that hash. If a company wanted to trace a specific generation back to a specific user, that would require correlating server logs on their own end, entirely separate from anything the watermark itself does.&lt;/p&gt;

&lt;p&gt;And there’s no scanning happening anywhere on the internet right now. Nothing crawls the web looking for watermarked text and slaps a flag on it.&lt;/p&gt;

&lt;p&gt;Detection is a deliberate, active computation that only the key holder can run meaningfully. Right now, that means Anthropic, internally. They’ve said a public detection API is coming, but as of this writing it isn’t broadly available. Third-party “AI detector” tools you’ve probably encountered don’t have Anthropic’s or Google’s key, so they can’t check for the actual watermark at all.&lt;/p&gt;

&lt;p&gt;They fall back on statistical pattern-matching instead, catching things like a model’s fondness for certain phrasing, and Anthropic’s own FAQ calls this out explicitly as “fundamentally different from checking for a watermark,” with meaningfully worse reliability.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why editing works and copying doesn’t&lt;/strong&gt;&lt;br&gt;
The watermark’s signal lives in individual token choices, spread across nearly the whole length of a passage rather than concentrated in one spot. That single fact explains almost every question people ask about defeating it.&lt;/p&gt;

&lt;p&gt;Copy and paste does nothing, because there’s nothing separate to strip. The watermark isn’t a file property or a piece of metadata bolted onto the text.&lt;/p&gt;

&lt;p&gt;It’s the words themselves. Copy the characters into VS Code, a Word doc, a plain text file, rename the file, change the extension, none of it touches the underlying token sequence. Anthropic confirms this directly: the mark “will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing.”&lt;/p&gt;

&lt;p&gt;Heavy rewriting genuinely does degrade it, because it replaces the token sequence with a new one chosen by whatever model or person did the rewriting.&lt;/p&gt;

&lt;p&gt;This is where a project like guillaumemeyer/watermarks-remover comes in, and it’s unusually honest about its own limits.&lt;/p&gt;

&lt;p&gt;The tool splits into two tiers. Layer A strips invisible Unicode characters, zero-width spaces, bidirectional text markers, tag characters, the leftover tricks from older edit-based watermarking schemes.&lt;/p&gt;

&lt;p&gt;That layer is deterministic and testable: the characters are either there or they’re gone. Layer B is the part meant to attack statistical watermarks like SynthID-Text, and it works by having a different model paraphrase the text heavily enough to scramble the original token choices.&lt;/p&gt;

&lt;p&gt;The project’s own README is candid about what that costs. Removal means rewording, not restructuring, they write, because shuffling paragraphs or lightly touching up phrasing barely moves the signal.&lt;/p&gt;

&lt;p&gt;You have to rewrite a meaningful fraction of the sentences, and every one of those rewrites replaces the original model’s word choices with the rewriting model’s. Tone flattens. Precision drops. The result can’t exceed the ceiling of whatever model did the rewrite.&lt;/p&gt;

&lt;p&gt;They even pose the obvious follow-up question themselves: if you’re going to hand the text to a cheaper model to reword anyway, why did you pay for the expensive model in the first place?&lt;/p&gt;

&lt;p&gt;There’s a deeper honesty in that README too, one that most tools in this space skip. They flatly state that “until vendors ship public detectors and keys, no tool can honestly certify this fails the official check.”&lt;/p&gt;

&lt;p&gt;Their own verification only runs against open research schemes they can replicate, not against Anthropic’s or Google’s actual production key. That’s the same one-key limitation we already covered: without the secret, nobody outside the company can confirm a removal actually worked.&lt;/p&gt;

&lt;p&gt;Screenshots break it completely, and this one is worth understanding because it’s the cleanest case. A screenshot converts text into an image, pixels with no tokens and no sampling history, so the watermark simply doesn’t exist in that data anymore.&lt;/p&gt;

&lt;p&gt;It’s not hidden in the image. It never made the trip. If you then feed that screenshot to an open-source model to transcribe it, you get an entirely new generation process, with new token choices, no watermarking key applied at all if the model is genuinely unwatermarked.&lt;/p&gt;

&lt;p&gt;There’s no statistical trace of the original left to find, because the original text was never processed as text in that pipeline, only as a picture. The real cost isn’t watermark survival, which is total, it’s OCR accuracy: code especially is full of characters that get misread, an “l” for a “1,” bracket mismatches, whitespace that Python actually depends on. You’re trading a clean removal for a proofreading job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why any of this matters if you actually build software&lt;/strong&gt;&lt;br&gt;
If you’re shipping a product that touches AI-generated text or code, three things are worth sitting with.&lt;/p&gt;

&lt;p&gt;The watermark is a compliance signal for the company that deployed it, not a forensic tool for you. If you’re trying to prove your own team’s code wasn’t lifted wholesale from an AI tool without review, SynthID-class watermarking isn’t built to answer that, because you don’t hold the key and the detection API isn’t public yet.&lt;/p&gt;

&lt;p&gt;Detection reliability scales with text length and with how much genuine uncertainty existed in the original generation. A single deterministic line of code, where there’s really only one correct way to write it, gives the sampling algorithm almost no room to bias anything, so there’s very little watermark signal to find there in the first place.&lt;/p&gt;

&lt;p&gt;A long, free-flowing paragraph of prose gives it plenty of room. This is exactly why Anthropic’s own guidance notes the mark should have “negligible effect” on actual code, while still touching comments and docstrings, which are closer to natural language.&lt;/p&gt;

&lt;p&gt;And detection will always be probabilistic, never certain, on either side of the question. A hit means “likely passed through this model.” A miss means “no signal found,” not “definitely human.”&lt;/p&gt;

&lt;p&gt;Anyone building a policy, whether it’s an academic integrity system or an internal content review process, around a hard yes or no from a watermark check is building on a foundation the researchers who created the technology never claimed it could support.&lt;/p&gt;

&lt;p&gt;*&lt;em&gt;The other kind of provenance, yes we built something! Surprise, surprise…&lt;br&gt;
*&lt;/em&gt;&lt;br&gt;
Everything above is about proving something came from a specific AI model. There’s a completely different problem sitting right next to it that gets confused with the first one constantly: proving something came from you, and existed at a specific point in time.&lt;/p&gt;

&lt;p&gt;That’s what Provenance, the open-source tool we built at Vektor Memory, actually does.&lt;/p&gt;

&lt;p&gt;The two problems look similar from a distance, both about establishing origin, both producing something you’d call evidence. Underneath, they don’t share a single mechanism.&lt;/p&gt;

&lt;p&gt;SynthID-Text works during generation and needs a secret key to verify. Provenance works after you’ve already written something, on files that already exist, and verification needs no secret at all, only public math anyone can rerun.&lt;/p&gt;

&lt;p&gt;It starts by hashing every file in your project into a Merkle tree, the same cryptographic structure git and Bitcoin both use internally, and binds the resulting root hash to your current git commit. Because it’s a Merkle tree rather than one flat hash of everything, you can prove a single file belonged to that snapshot, prov manifest prove src/index.js, without revealing or even touching the rest of the codebase.&lt;/p&gt;

&lt;p&gt;The timing half is where it actually gets interesting. A hash alone proves what your code looked like, but says nothing about when.&lt;/p&gt;

&lt;p&gt;So Provenance anchors that manifest with two independent timestamp authorities at once: RFC 3161, a conventional third-party timestamping service that confirms quickly, and OpenTimestamps, which anchors the same hash into a Bitcoin block, slower to confirm but with no company or authority you have to trust at all, since the proof lives on a public blockchain nobody controls.&lt;/p&gt;

&lt;p&gt;Using both means a single point of failure, one authority going offline or getting compromised, can’t quietly invalidate your evidence.&lt;/p&gt;

&lt;p&gt;Verification later is just re-hashing your files, rebuilding the Merkle root, and checking it against both anchors independently. Deterministic, and either it matches or it doesn’t. No key required, unlike SynthID’s detection step, which only the model owner can run with confidence.&lt;/p&gt;

&lt;p&gt;The README is upfront about the limit here too, and it’s worth repeating because it’s the same honesty the watermark-remover project showed about its own tool: “This is evidence and stated policy, not a technical access-control or anti-piracy mechanism.&lt;/p&gt;

&lt;p&gt;It won’t stop a determined bad actor from copying your code, it gives you strong, verifiable evidence of what your code looked like and when.” That’s the whole category, honestly stated. Cryptographic timestamping proves possession and priority. It was never going to stop theft, only make theft harder to lie about afterward.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwlpywd7bupj8hyly3czc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwlpywd7bupj8hyly3czc.png" alt=" " width="754" height="417"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So if you’re weighing which of these two technologies actually solves your problem, the question that matters is which claim you need to make. “This text came from a watermarked AI model” needs SynthID-class watermarking, and needs the model owner’s cooperation to verify.&lt;/p&gt;

&lt;p&gt;“I had this exact code at this exact time” needs Merkle-tree hashing and independent timestamp anchoring, and needs nobody’s permission to verify, ever. They’re not competing approaches to the same job. They’re answers to two different questions that happen to both get called “provenance.”&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The permanence caveat&lt;/strong&gt;&lt;br&gt;
Making Provenance more permanent comes down to one principle: the tool should never claim more certainty than it can actually back up. That means teaching prov verify to check that the git commit it's bound to still exists in history, not just that a hash matches, so a rebase or force-push can't quietly orphan the evidence.&lt;/p&gt;

&lt;p&gt;It means supporting a small list of backup timestamp authorities instead of relying on one, so a single service going offline doesn't take the whole proof down with it, and clearly separating "confirmed" from "still pending" on the Bitcoin anchor so nobody mistakes an unfinished proof for a finished one.&lt;/p&gt;

&lt;p&gt;And it means locking a version number into the manifest format itself, so a future update to the tool can't silently misread evidence that an older version created. None of this changes what Provenance proves. It just makes sure the tool tells you the truth about how solid that proof actually is, every time you check it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's the difference? C2PA, SynthID, &amp;amp; Provenance?&lt;/strong&gt;&lt;br&gt;
All three sit under the umbrella of “content provenance,” but they’re solving genuinely different slices of it, and the differences matter more than the similarity here.&lt;/p&gt;

&lt;p&gt;The one real similarity across all three&lt;/p&gt;

&lt;p&gt;Each one is trying to answer some version of “where did this come from, and can I prove it without just trusting someone’s word.” And each one is honest, in its own documentation, about not being an access-control or anti-piracy mechanism.&lt;/p&gt;

&lt;p&gt;Provenance’s README says this explicitly. Anthropic says a SynthID hit is “a signal, not proof.” C2PA’s own spec describes itself as a durable manifest, not a lock. None of the three claims to stop theft or misuse. They all just try to make lying about origin harder.&lt;/p&gt;

&lt;p&gt;Past that, they diverge fast.&lt;/p&gt;

&lt;p&gt;C2PA is the closest cousin, and it’s genuinely closest to Provenance’s approach&lt;/p&gt;

&lt;p&gt;C2PA (the standard behind Content Credentials, backed by Adobe, Microsoft, and others, with OpenAI on the steering committee) embeds a cryptographically signed manifest directly into a file, usually an image, video, or PDF.&lt;/p&gt;

&lt;p&gt;That manifest can record who created it, what tools touched it, and a chain of edits. It’s file-based metadata with a digital signature, checkable with tools like c2patool, the exact tool the watermarks-remover README references for stripping it.&lt;/p&gt;

&lt;p&gt;This is structurally close to what Provenance does: both produce a verifiable, cryptographically signed record tied to a piece of content, both are checkable without needing a secret key held by one company, both are explicitly not about stopping misuse.&lt;/p&gt;

&lt;p&gt;The real difference is scope and durability. C2PA lives inside the file as metadata, which means it can be stripped by re-exporting, screenshotting, or just deleting the metadata block, the same operation the watermarks-remover tool performs on PNGs and PDFs.&lt;/p&gt;

&lt;p&gt;Provenance’s evidence lives outside the shipped file entirely, in a separate Merkle tree anchored to two independent timestamp services, so stripping the code of any trace of it doesn’t touch the evidence, because the evidence was never inside the code to begin with. C2PA proves “this file claims this origin.” Provenance proves “this exact file set existed at this exact time,” and keeps proving it even after every trace is scrubbed from the shipped artifact.&lt;/p&gt;

&lt;p&gt;SynthID is the outlier, mechanically unrelated to the other two.&lt;/p&gt;

&lt;p&gt;SynthID doesn’t touch files or metadata at all. It biases token-by-token sampling during generation, inside the model, so the “evidence” is a statistical pattern smeared across the actual words or pixels, invisible and unremovable by any editing that doesn’t rewrite the content itself.&lt;/p&gt;

&lt;p&gt;Detection needs the generator’s secret key. Nobody else can verify it, ever, without that key. C2PA and Provenance both invert that: their proofs are public, checkable by anyone, no permission needed. SynthID’s proof is private by design, gatekept by whoever ran the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where the three actually stand relative to each other&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;C2PA and Provenance are both cryptographic, file-external-or-embedded, publicly verifiable systems, close cousins in mechanism even though they anchor different things (a signed claim versus a hashed-and-timestamped snapshot).&lt;/p&gt;

&lt;p&gt;SynthID is a different category entirely: a statistical fingerprint baked into content at the moment of creation, verifiable only by the party holding the key. If you’re choosing between them for a real use case, the honest framing is that C2PA and Provenance both answer “can anyone check this claim independently,” while SynthID answers “did this specific company’s model touch this,” and only that company can ever confirm it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6rqo7k885jovqobx622q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6rqo7k885jovqobx622q.png" alt=" " width="750" height="598"&gt;&lt;/a&gt;&lt;br&gt;
             The Future: Blade Runner AI slop verification&lt;/p&gt;

&lt;p&gt;To conclude, nothing is perfect in these systems—the writing, detection, or removal; we are all just little piggies in the big sloppy game of rolling around in the slop soup mess that we created to determine if AI is really slop.&lt;/p&gt;

&lt;p&gt;We are not going back to typewriters, so enjoy your paradox; you deserve to know if your slop is real slop or just human-generated squishy slop.&lt;/p&gt;

&lt;p&gt;Be the slop; verify the authenticity of your slop.&lt;/p&gt;

&lt;p&gt;Good luck in the future!&lt;/p&gt;

&lt;p&gt;(Satire is the best form of medicine…)&lt;/p&gt;

&lt;p&gt;Sources and further reading: Dathathri et al., “Scalable watermarking for identifying large language model outputs,” Nature, 2024. Anthropic, “How Claude marks AI-generated content.” BleepingComputer, “How Anthropic plans to watermark Claude’s AI-generated text.” google-deepmind/synthid-text on GitHub. guillaumemeyer/watermarks-remover. Vektor-Memory/Provenance.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>synthid</category>
      <category>provenance</category>
    </item>
    <item>
      <title>Provenance: What It Actually Takes to Prove Creator Data Dignity</title>
      <dc:creator>Vektor Memory</dc:creator>
      <pubDate>Wed, 05 Aug 2026 06:20:24 +0000</pubDate>
      <link>https://dev.to/vektor_memory_43f51a32376/provenance-what-it-actually-takes-to-prove-creator-data-dignity-2jg5</link>
      <guid>https://dev.to/vektor_memory_43f51a32376/provenance-what-it-actually-takes-to-prove-creator-data-dignity-2jg5</guid>
      <description>&lt;p&gt;We red-teamed an AI royalty proof-of-concept system we are building&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvgj8ue3dvdmbh5xc3ve9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvgj8ue3dvdmbh5xc3ve9.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Provenance by Vektor Memory: Another weekend coding project and why proposals for paying creators when AI trains on their work is more difficult than discussions.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note: Article is written in natural human language for general readers—for deeper technical dives, see our other 70 articles or website.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;After watching Jaron Lanier argue for years that AI models are compressions of human labor, not independent intelligences, we finally tried to turn that thesis into working code. An account of where the theory held up and where it cracked.&lt;/p&gt;

&lt;p&gt;We spent the last few weekends building a five-layer system for tracing AI-generated content back to the creators whose work shaped it, and paying them. We spent the weekend stress-testing every single layer until we found the cracks.&lt;/p&gt;

&lt;p&gt;Then we fixed the first layer, actually shipped it, and tested it against 27 different file formats.&lt;/p&gt;

&lt;p&gt;The concept is compelling. The execution is where things get interesting, because almost every piece that looks reasonable on paper fails in a specific, predictable way once you attack it or try to build it without the data needed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why this idea exists&lt;/strong&gt;&lt;br&gt;
The argument is straightforward: if an AI model is a compression of human training data, then the people who created that data should get paid when the model generates revenue. It’s not a complete fix for content attribution or copyright law. It’s one small attempt at one piece of a much larger problem.&lt;/p&gt;

&lt;p&gt;But “one small piece” turns out to require solving several separate, genuinely hard problems at once. Content provenance. Attribution math that works under adversarial conditions. Governance that can’t be captured by a coordinated attacker. Economics that don’t accidentally subsidize the bad actors. And legal frameworks that don’t exist yet.&lt;/p&gt;

&lt;p&gt;We decided to build it in layers, test each one until it broke, and publish what we found.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The five-layer architecture&lt;/strong&gt;&lt;br&gt;
The system splits into five independent, testable layers. Here’s the flow from a creator’s work to a payout:&lt;/p&gt;

&lt;p&gt;Layer 1: Provenance Registry Your work gets signed and timestamped through two independent, non-colluding anchors. One is a traditional timestamp authority. The other is the Bitcoin blockchain. This creates a verifiable record of what existed and when, without depending on trusting us or any single gatekeeper.&lt;/p&gt;

&lt;p&gt;Layer 2: Streaming Safety Guard The AI platform’s output gets checked in real time, before the user sees it, looking for combinations of risky content rather than just flagged words. Single-word blocklists are trivial to route around. Combinations are harder.&lt;/p&gt;

&lt;p&gt;Layer 3: Attribution and Staleness Decay We estimate how much a registered piece of work shaped a specific output, then lower that confidence continuously as the underlying model keeps training and drifting away from the version we measured.&lt;/p&gt;

&lt;p&gt;Layer 4: Anti-Sybil Governance When attribution claims get contested, we route them to a small jury sampled at random from a bonded pool instead of an open vote. An attacker can’t know in advance which of their fake accounts will be eligible to vote on any given case.&lt;/p&gt;

&lt;p&gt;Layer 5: Settlement and Economics Contested payouts go into escrow. Disputes resolve through appeals. Fees scale based on an account’s own risk profile, not flat usage volume.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Each layer has built-in failure modes.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr597qhsd9etly4vj1prv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr597qhsd9etly4vj1prv.png" alt=" " width="746" height="866"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Layer 1 is shipping now&lt;/strong&gt;&lt;br&gt;
The provenance layer is a working CLI tool called Prov. It signs your code or your body of work into a cryptographic manifest, timestamps it through two independent services that actively distrust each other, and gives you a verifiable record that doesn’t depend on trusting us.&lt;/p&gt;

&lt;p&gt;We released it as open source a few weeks ago. But the real test started last week when we tried to make it work with multiple file formats that matter.&lt;/p&gt;

&lt;p&gt;27 formats: Images. Code. PDFs. Documents. Fonts. Anything you’d actually want to prove authorship over.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fexorsp85ueflizctkfk4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fexorsp85ueflizctkfk4.png" alt=" " width="799" height="549"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The format problem nobody talks about&lt;/strong&gt;&lt;br&gt;
The C2PA standard (the same one Adobe uses for content credentials) works great for images and video. PNG, JPEG, WAV, MP4 all have standard ways to embed cryptographic manifests. You sign the file, embed the signature, read it back out.&lt;/p&gt;

&lt;p&gt;But most creator work isn’t just images. It’s source code. Text. PDFs. Office documents. Fonts. These are completely different containers with completely different structures.&lt;/p&gt;

&lt;p&gt;We built custom implementations for all of them in JS. PDF needed its own manifest embedding strategy because c2pa-python has no native PDF writer. EPUB, DOCX, ODT, OXPS are all ZIP containers but with different rules about where files can go and what can be compressed.&lt;/p&gt;

&lt;p&gt;Fonts (OTF, TTF, SFNT) have their own table structure and a pre-standard C2PA specification that’s still sitting in an open GitHub issue, waiting for the standards body to ratify it.&lt;/p&gt;

&lt;p&gt;We pulled the actual C2PA specification and font spec proposals, read the code in the reference implementations, and built each one to spec.&lt;/p&gt;

&lt;p&gt;Then we tested all 27 of them round-trips: cryptographic validation. Real c2pa. Reader parsing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The MIME type trap&lt;/strong&gt;&lt;br&gt;
The c2pa-python library accepts any MIME type you hand it, but internally it’s selective about what it actually signs. We tried signing an M4A audio file as audio/mp4 (the obvious choice) and it failed silently with “NotSupported.”&lt;/p&gt;

&lt;p&gt;We spent time chasing that before realizing the library was rejecting the MIME type hint, not the file itself. Signing it with application/octet-stream (a generic, content-agnostic type) worked fine. The lesson: never assume MIME type handling is transparent.&lt;/p&gt;

&lt;p&gt;The attribute shape mismatch: The c2pa-text library we integrated for source code embedding has three different embedding methods (invisible Unicode, structured comments, HTML).&lt;/p&gt;

&lt;p&gt;Each one returns a different data shape. We assumed they all returned the same TextEmbedResult object with .text, .exclusion_start, and .exclusion_length.&lt;/p&gt;

&lt;p&gt;Two of them return something completely different. We caught it by actually inspecting the function signatures before running anything, not after.&lt;/p&gt;

&lt;p&gt;The DSIG coexistence problem: When we tested font embedding against real-world OTF files, we discovered that production fonts often already carry a DSIG (digital signature) table from their foundry.&lt;/p&gt;

&lt;p&gt;The C2PA spec proposal says they shouldn’t coexist. We had to add a configurable policy: refuse to sign a font that already has a DSIG, or strip it first and proceed. Both approaches are defensible, depending on your use case.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Here’s what’s in the release:&lt;/strong&gt;&lt;br&gt;
27 file formats verified working with real cryptographic signatures. That breaks down as follows:&lt;/p&gt;

&lt;p&gt;19 Tier-1 native formats (PNG, JPEG, WAV, MP4, WebP, HEIC, HEIF, GIF, MP3, AVI, TIFF, AVIF, M4A, DNG, M4V, MPA, SVG, FLAC, JPEG XL)&lt;br&gt;
1 custom implementation for PDF (write support, which the base library didn’t have)&lt;/p&gt;

&lt;p&gt;4 ZIP-container formats (EPUB, DOCX, ODT, OXPS) with custom collection-data-hash logic&lt;/p&gt;

&lt;p&gt;3 font formats (OTF, TTF, SFNT) with DSIG handling&lt;/p&gt;

&lt;p&gt;Plus 5 source-text formats (.py, .js, .yaml, .sql, .md) and HTML, which nobody else in the C2PA space covers because they require a separate text-embedding specification that c2pa-python completely ignores.&lt;/p&gt;

&lt;p&gt;Every single one has a smoke test. Every smoke test uses real files (not synthetic test data), real signing, real validation through the actual c2pa.Reader library, not mocked calls.&lt;/p&gt;

&lt;p&gt;Layers 2 through 5: where theory meets reality&lt;br&gt;
The other four layers are simulations. We didn’t ship them yet because every single one has structural vulnerabilities that an attacker would exploit in the first week.&lt;/p&gt;

&lt;p&gt;We’d rather find them now, in simulation, than after there’s real creator money behind them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Open voting — dispute resolution system&lt;/strong&gt;&lt;br&gt;
We built a dispute resolution system where a committee votes on contested attribution claims. We assumed quadratic voting would handle it, because quadratic voting is supposed to blunt the advantage of concentrated capital.&lt;/p&gt;

&lt;p&gt;What it doesn’t do is stop an attacker from splitting that same capital across more identities. Under quadratic voting, total voting power actually increases as you split a fixed budget into more accounts. We ran a simulation with a coordinated cartel trying to capture a committee’s vote, and they succeeded nearly every time.&lt;/p&gt;

&lt;p&gt;The math is straightforward: if you have 100 tokens to vote with, quadratic voting gives you sqrt(100) = 10 voting power. If you split that into 10 accounts with 10 tokens each, you get 10 accounts with sqrt(10) = 3.16 power each, for a total of 31.6. You just increased your voting power by 3x by splitting your tokens.&lt;/p&gt;

&lt;p&gt;We fixed it by replacing open voting with a small jury randomly sampled from a bonded pool. An attacker doesn’t know which of their fake accounts will be eligible to vote on any given case.&lt;/p&gt;

&lt;p&gt;If an attacker controls 30 percent of a 10,000-person pool, they have roughly a 0.04 percent chance of capturing a 63-member jury’s majority on any single dispute. We verified that number two ways: a closed-form probability calculation and an independent simulation. They agreed.&lt;/p&gt;

&lt;p&gt;A perfect safety filter can still be defeated by buffer management&lt;br&gt;
We built a streaming content filter that runs on the AI platform’s output in real time, looking for combinations of risky terms rather than single flagged words. The detector works exactly as designed when both terms appear in the same chunk.&lt;/p&gt;

&lt;p&gt;It fails completely when we split the same two terms across a chunk boundary. The system has no memory of what it already saw. It checks the first chunk, finds nothing (because it’s only half the risky combination), forgets about it, then checks the second chunk and finds nothing there either (because it’s missing the first half). The risky content passes through untouched.&lt;/p&gt;

&lt;p&gt;The vulnerability has nothing to do with the detector’s intelligence. It’s a structural gap in how the buffer forgets its own history. We added a small memory window to track terms from the previous chunk. That fixed the immediate version of the problem.&lt;/p&gt;

&lt;p&gt;Then we tested it by splitting the risky terms even further apart, across multiple chunks and multiple seconds of output. The same failure came back. The memory window has a hard limit. Split the terms far enough and you defeat the filter again.&lt;/p&gt;

&lt;p&gt;You can’t fix this by making the memory window larger without introducing latency penalties that destroy the entire point of streaming detection. You can’t fix it by making the detector stateless because then it can’t remember anything. It’s a fundamental constraint of how streaming systems work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hardware isolation doesn’t work the way you think&lt;/strong&gt;&lt;br&gt;
The instinct is to run your safety check on separate hardware from the main model. That way, under load, the safety check won’t get starved for compute.&lt;/p&gt;

&lt;p&gt;We tested this three different ways. First with a placeholder safety check that barely does any work, moving it to a separate process. Performance got worse, not better, because the overhead of the inter-process communication outweighed the benefit of separation.&lt;/p&gt;

&lt;p&gt;Then we did it with a real model doing real computation. Separation won decisively, and latency degradation under load dropped by roughly 30 times. The safety check was protected.&lt;/p&gt;

&lt;p&gt;Then we pinned both processes to completely separate CPU cores using hardware affinity, to test true resource isolation rather than just process separation.&lt;/p&gt;

&lt;p&gt;And we found a bottleneck neither test had caught: even with the safety model sitting on completely idle cores, the calling process (the one issuing the request to the safety check) was starved by the same contention, because it shared cores with the load.&lt;/p&gt;

&lt;p&gt;You can isolate the safety check’s compute, but you also have to guarantee resources for the caller. Most designs only think about one half of that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Flat fees subsidize the problem actors&lt;/strong&gt;&lt;br&gt;
A common approach to funding safety infrastructure is to charge a flat percentage fee on every API call. You’re charging for volume.&lt;/p&gt;

&lt;p&gt;We simulated this against a real distribution of user behavior: mostly legitimate enterprise customers, a small tail of research accounts, and a handful of accounts repeatedly flagged for violations.&lt;/p&gt;

&lt;p&gt;The high-volume legitimate customers ended up paying roughly eight times their fair share of the actual cost their traffic imposed on the safety infrastructure. Meanwhile, accounts flagged repeatedly for violations paid a fraction of theirs.&lt;/p&gt;

&lt;p&gt;We shifted the fee model to scale with an account’s own observed history. That fixed most of the second problem. A habitual violator now pays a realistic cost for the risk they cause.&lt;/p&gt;

&lt;p&gt;It only partly fixed the first problem, because the legitimate high-volume customers still have to cross a flat base rate before the risk-adjusted pricing kicks in. Our first writeup overstated how well this worked. Once we looked at the actual numbers instead of the intended design, we corrected it.&lt;/p&gt;

&lt;p&gt;The lesson: when you’re designing economics, the numbers matter more than the theory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Known problems we can’t solve in this design&lt;/strong&gt;&lt;br&gt;
The cold start problem: AI labs won’t integrate an attribution system until creators have registered their work. Creators won’t bother registering until AI labs start paying them. Solving this requires either regulatory force or a massive coordinated developer movement.&lt;/p&gt;

&lt;p&gt;The GDPR paradox: Public cryptographic ledgers are immutable by design. They can’t delete data. But GDPR’s Article 17 says creators have the right to erase their personal data. Designing systems that can sever real-world identities from public commitments without breaking historical provenance is deeply non-trivial.&lt;/p&gt;

&lt;p&gt;Training data attribution doesn’t exist yet: Our system assumes you already have attribution scores telling you which training data influenced a given output. No real training-data attribution exists at scale. We can track staleness of scores over time, but we can’t generate the scores in the first place. That’s a separate, unsolved problem.&lt;/p&gt;

&lt;p&gt;The incumbent incentive mismatch: Hyperscalers have strong structural incentives to keep training data pipelines opaque. They avoid liability, copyright exposure, and margin compression by keeping everything closed. Forcing them to adopt an open attribution standard requires either severe regulatory pressure or a developer revolt.&lt;/p&gt;

&lt;p&gt;Governance capture is reduced, not eliminated: Our jury model is much harder to attack than open voting, but it’s not impossible. With enough budget and sophistication, a determined attacker could still find ways to bias the jury pool. We reduced the problem from “captured almost every time” to “captured maybe once in ten thousand tries.” That’s an improvement, not a solution.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The technical foundation&lt;/strong&gt;&lt;br&gt;
Layer 1 builds on real standards work:&lt;/p&gt;

&lt;p&gt;C2PA (Consortium for Content Provenance and Authenticity) is the same specification Adobe uses for content credentials. We implemented it fully for 27 file formats instead of just image and video.&lt;/p&gt;

&lt;p&gt;Our dual-anchor timestamping follows RFC 3161 for traditional timestamp authorities and OpenTimestamps for Bitcoin-based anchoring, so no single gatekeeper can control the registry.&lt;/p&gt;

&lt;p&gt;Shamir’s Secret Sharing for custody of the identity-vault key, so key compromise doesn’t automatically compromise the registry. Implemented and verified, not theoretical.&lt;/p&gt;

&lt;p&gt;Generative Content ID research from Deng et al. for the attribution framework, adapted from music to text.&lt;/p&gt;

&lt;p&gt;Influence-function based staleness decay for understanding why attribution scores get noisier over time.&lt;/p&gt;

&lt;p&gt;Data Dignity work from RadicalxChange for the governance framing, treating this as collective bargaining infrastructure instead of a single company’s black box.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Caveats: Known Challenges&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Computational &amp;amp; Scaling Complexities (The Math Problem)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The Scale of TDA (Training Data Attribution): Running exact attribution models to figure out which datasets or creator nodes influenced a specific token output is computationally prohibitive[1]. Even state-of-the-art approximations (like TRAK or influence functions) require massive matrix multiplications, massive storage for pre-computed gradients, and continuous re-computation overhead.&lt;/p&gt;

&lt;p&gt;The Checkpoint Drift Problem: Production foundation models are under continuous fine-tuning, DPO (Direct Preference Optimization), and RLHF. Every single weight update instantly invalidates old attribution indices, requiring automated, resource-intensive re-projection pipelines running 24/7.&lt;/p&gt;

&lt;p&gt;Inference Latency Costs: Adding parallel safety guardrails or logging telemetry introduces non-zero latency penalties. Scaling this across millions of concurrent enterprise users without degrading Time-to-First-Token (TTFT) requires dedicated, expensive parallel hardware infrastructure.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Economic &amp;amp; Game-Theoretic Complexities (The Money Problem)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The Micro-Payout Gas &amp;amp; Ledger Crisis: Distributing fractions of a cent ($0.0000001) across millions of global creators per LLM query will completely break traditional banking rails and rack up unsustainable blockchain gas fees unless channeled entirely through specialized Layer-2/Layer-3 zero-knowledge state rollups.&lt;/p&gt;

&lt;p&gt;Adverse Selection in Surcharges: Funding safety infrastructure via a flat usage-based surcharge on API calls disproportionately penalizes honest, high-volume enterprise customers who pose minimal safety risks, subsidizing the overhead caused by a tiny fraction of bad actors.&lt;/p&gt;

&lt;p&gt;Sybil Flooding &amp;amp; Value Extraction: Any system that pays out real money based on data contribution immediately invites sophisticated adversarial attacks — such as automated synthetic data farms engineered exclusively to maximize attribution scores and drain the royalty pool.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Governance, Trust &amp;amp; Adversarial Complexities (The Human Problem)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The Curation Cartel &amp;amp; Collusion: Decentralized committees or “courts” meant to arbitrate data disputes are chronically vulnerable to capture. Well-organized cartels or Sybil nodes can coordinate below voting thresholds to systematically vote down legitimate creators and slash their stakes.&lt;/p&gt;

&lt;p&gt;The Subjectivity of “Causal Influence”: Unlike exact file matching (like audio copyright matching on YouTube), generative AI creates entirely new conceptual syntheses [2]. Proving mathematically how much an artist’s style or text snippet contributed to an abstract generated concept remains intensely legally and technically contested.&lt;/p&gt;

&lt;p&gt;The Burden of Dispute Resolution: When millions of creators dispute low attribution scores, who handles the administrative backlog? Automated code cannot legally seize collateral or execute final financial penalties without triggering severe legal challenges in traditional courts.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Privacy, Compliance &amp;amp; Legal Complexities (The Regulatory Problem)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The GDPR “Right to Erasure” Paradox: Public cryptographic ledgers and append-only Verkle trees are immutable by design — they cannot delete data[2]. However, Article 17 of the GDPR dictates that a creator has the right to erase their PII and identity data completely. Designing systems that can sever real-world identities from public commitments without breaking historical provenance is deeply non-trivial.&lt;/p&gt;

&lt;p&gt;Cross-Jurisdictional Compliance: A global creator economy must navigate radically conflicting international frameworks — such as the EU AI Act’s strict transparency and copyright mandates[2], US fair use doctrines, and varying regional data sovereignty laws.&lt;/p&gt;

&lt;p&gt;Liability and Indemnification: If an AI platform uses a registered dataset that turns out to contain plagiarized or illegal content (e.g., leaked enterprise code or copyright-infringing text), who carries the legal liability — the platform, the creator who registered it, or the attribution protocol?&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Ecosystem Coordination Complexities (The Adoption Problem)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The Cold-Start Problem (Two-Sided Marketplace Failure): AI labs won’t integrate an attribution and royalty protocol until creators flood it with data; creators won’t bother registering their work until major AI labs adopt the standard and start paying out real revenue.&lt;/p&gt;

&lt;p&gt;The Incumbent Incentive Mismatch: Hyperscalers and frontier AI labs (OpenAI, Google, Anthropic, Meta) have strong structural incentives to keep training data pipelines opaque to avoid liability, copyright exposure, and margin compression. Forcing them to adopt an open attribution standard requires either severe regulatory pressure (such as enforced compliance frameworks like the EU AI Act)[2] or a massive, coordinated developer revolt.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What’s next&lt;/strong&gt;&lt;br&gt;
Layer 1 is production-ready and available open source.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Vektor-Memory/Provenance" rel="noopener noreferrer"&gt;https://github.com/Vektor-Memory/Provenance&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Layers 2 through 5 are simulations of what a complete system would look like. We built them because we wanted to find the failure modes before there’s real creator money sitting behind them. Some you can’t fix without changing the whole architecture.&lt;/p&gt;

&lt;p&gt;We’re asking whether this direction is worth pursuing at all. Whether the failures we found are the ones you’d expect or ones you’d never have guessed. Whether “tested this hard before launch” is a bar worth holding the rest of this category to.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;They need:&lt;/strong&gt;&lt;br&gt;
Real training data attribution systems that don’t exist yet. Our staleness tracking works on top of attribution scores that already exist, but generating those scores at scale is still an open research problem.&lt;br&gt;
Regulatory framework and legal counsel review. We don’t know yet whether this approach is even legally defensible without new law.&lt;/p&gt;

&lt;p&gt;Real-world adversarial testing beyond simulation. An attacker with real resources and real incentives will find things a simulation won’t.&lt;br&gt;
Ecosystem coordination. This only works if it actually gets adopted, which requires solving the cold-start problem for a two-sided marketplace.&lt;/p&gt;

&lt;p&gt;VEKTOR Memory builds local-first, privacy-preserving persistent memory infrastructure for AI agents. Technical documentation at vektormemory.com.&lt;/p&gt;

&lt;p&gt;Data Science&lt;br&gt;
AI&lt;br&gt;
Data Dignity&lt;br&gt;
Provenance&lt;/p&gt;

</description>
      <category>ai</category>
      <category>data</category>
      <category>datascience</category>
    </item>
    <item>
      <title>We built our agent a tool for codebase intelligence</title>
      <dc:creator>Vektor Memory</dc:creator>
      <pubDate>Thu, 30 Jul 2026 05:26:49 +0000</pubDate>
      <link>https://dev.to/vektor_memory_43f51a32376/we-built-our-agent-a-tool-for-codebase-intelligence-4d8b</link>
      <guid>https://dev.to/vektor_memory_43f51a32376/we-built-our-agent-a-tool-for-codebase-intelligence-4d8b</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqiw4isk0mbuidier9pig.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqiw4isk0mbuidier9pig.png" alt=" " width="800" height="800"&gt;&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Why we didn't want a full&amp;nbsp;index&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The starting idea was simple. Give our autonomous coding agent a real understanding of the codebase it's editing, not just whatever files it happens to grep into.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Note: This article has been created in natural human language to be read easily, devoid of complex technical jargon, so all readers can enjoy. View one of our other past 70 articles for deeper technical dives.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Open source tools like CodeGraph and code-review-graph already do this well. They parse your repo with tree-sitter, build a graph of every function, class, and import, and let an agent query it instead of re-reading files from scratch on every task.&lt;/p&gt;

&lt;p&gt;The problem we found is that graphs cost to keep open: the bloat tax. Both of those tools build a persistent index on disk, and a background watcher keeps it in sync with every file save. That's fine for a single project you keep open all day.&lt;/p&gt;

&lt;p&gt;It's the shaped tool for an agent that might touch a dozen different projects in an afternoon, each one spinning up a watcher and a&amp;nbsp;.codegraph folder that outlives the task that needed it.&lt;/p&gt;

&lt;p&gt;We wanted the intelligence without the standing cost. A code graph that shows up exactly when a task needs one and disappears the moment the task is done, autonomous, intelligent, and relevant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agent loop stubbornness&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;While building, an agent spent eight steps convinced our own file system was broken, looping.&lt;/p&gt;

&lt;p&gt;It kept calling list_dir, getting a clean result back, and then telling us the tools were failing. Three different frontier LLM providers did the exact same thing, as our system is provider-agnostic.&lt;/p&gt;

&lt;p&gt;This was not a permissions error, just a model quietly hallucinating a filesystem outage while the directory listing sat right there in its own context window.&lt;/p&gt;

&lt;p&gt;That error, and the two others we found chasing it, ended up teaching us more about how to build code intelligence improved for our own agent than the features we originally set out to build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The devil is in the detailed refinement.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Sometimes simple code that works is far better than dozens of features that keep expanding into bloatware that 80% of users don't use very often or need at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The three-tier gate&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The actual research made this easier to justify than we expected. A 2026 paper called "Retrieval as a Decision" &lt;a href="https://arxiv.org/abs/2511.09803" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2511.09803&lt;/a&gt; argues that most retrieval-augmented systems should treat the decision to retrieve at all as a first-class, trainable step, not something baked into a fixed pipeline. That's close to what we ended up building, minus the trained part.&lt;/p&gt;

&lt;p&gt;Every task our agent takes on gets classified into one of three tiers before any graph work happens. A small, single-file edit skips the graph entirely and goes straight to a plain file read.&lt;/p&gt;

&lt;p&gt;A task that touches three or more files, or a shared file like a config or a types module, triggers a scoped build: parse just the touched files and whatever they import, out to two hops deep.&lt;/p&gt;

&lt;p&gt;A request for a genuine architecture overview, something a user actually asked for by name, triggers a slower full pass starting from the project's real entry points.&lt;/p&gt;

&lt;p&gt;That middle tier is where the actual engine lives. It's built entirely in Node, using web-tree-sitter compiled to WebAssembly rather than the native tree-sitter bindings most of these tools use.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;All hail&amp;nbsp;Node&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That Node choice wasn't a close call for us, and it's worth being direct about why. Most of the graph tools in this space, including code-review-graph, are written in Python or Rust.&lt;/p&gt;

&lt;p&gt;All of Slipstream is 100% Node already, top to bottom, and it works. Node has the largest package ecosystem of any runtime available today, more mature tooling around it than any alternative, and a single-language codebase means no subprocess bridge, no second package manager to keep in sync, no version mismatch between a Python venv and the Node process actually running the agent.&lt;/p&gt;

&lt;p&gt;We like Python genuinely and use it for other projects. It's a good language for a lot of jobs. It just isn't the right tool for shipping one coherent, dependency-light agent runtime and bolting a Python subprocess onto a Node codebase.&lt;/p&gt;

&lt;p&gt;To get tree-sitter bindings would have reintroduced the exact cross-language fragility we were trying to avoid in the first place. Node was staying, full stop, and the rest of the design had to work within that.&lt;/p&gt;

&lt;p&gt;Native bindings need to be compiled per platform, and we'd already been burned by exactly that kind of native module fragility elsewhere in this codebase. WASM sidesteps the whole problem. No compiler needed, same parser accuracy, one dependency that just works the same way on every machine.&lt;/p&gt;

&lt;p&gt;The graph itself lives entirely in memory, scoped to one agent session, and gets thrown away the moment that session ends. Nothing gets written to disk. No watcher runs while the agent is idle.&lt;/p&gt;

&lt;p&gt;If the same project gets touched again five minutes later in the same session, the graph gets reused and extended instead of rebuilt from scratch, so the zero-footprint choice doesn't mean paying the full cost twice.&lt;/p&gt;

&lt;p&gt;We also gave the agent three tools instead of the twenty-eight a tool like code-review-graph exposes. A cheap existence check that just asks, "Does this pattern already appear in this one file?" a scoped blast-radius expansion for when a change actually touches multiple files, and a full architecture pass for the rare cases that warrant one. Fewer tools means less for the model to choose wrong, and the gate is doing the choosing anyway.&lt;/p&gt;

&lt;p&gt;To add, there's a general rule worth stating plainly here: WASM only pays off when you cross the JS boundary rarely and do real work on the other side of it.&lt;/p&gt;

&lt;p&gt;Call into a WASM module constantly for small pieces of data and you end up paying string-copy and serialization costs on every single call, often for work that would have been just as cheap in JavaScript to begin with, since V8's own native JSON.parse and string handling are already heavily optimized.&lt;/p&gt;

&lt;p&gt;Call into it infrequently for a substantial chunk of work, and the crossing cost becomes irrelevant next to what you get back. Our CodeGraph engine sits firmly on the right side of that line.&lt;/p&gt;

&lt;p&gt;Web-tree-sitter gets called once per file, to parse a whole source string into an AST when a task actually needs it, not once per keystroke or once per token in some hot loop.&lt;/p&gt;

&lt;p&gt;That's exactly the profile WASM is built to win at: infrequent calls doing real work, not death by a thousand small boundary crossings. It's also part of why we didn't reach for native tree-sitter bindings or a Python subprocess instead.&lt;/p&gt;

&lt;p&gt;Every language boundary adds the same kind of tax, whether it's JS to WASM or Node to a Python child process, and the only way to actually come out ahead is to cross it as rarely as possible and make each crossing count.&lt;/p&gt;

&lt;p&gt;Three papers and why we didn't build a fourth open source&amp;nbsp;tool&lt;br&gt;
Before writing any of this, we spent time on three pieces of recent research, because the question we kept running into wasn't "can we build a code graph."&lt;/p&gt;

&lt;p&gt;Anyone can build a code graph; it's straightforward!&lt;/p&gt;

&lt;p&gt;The real question was when an agent should be allowed to use one.&lt;/p&gt;

&lt;p&gt;The first is "Retrieval as a Decision: Training-Free Adaptive Gating for Efficient RAG," from late 2025. Its argument is that most retrieval systems treat retrieval as mandatory, something that runs on every request whether it helps or not, when it should be treated as a decision with a real cost and a real chance of being unnecessary.&lt;/p&gt;

&lt;p&gt;That's the paper our gate is a direct answer to. Every task gets classified before any graph work happens, and most small edits never touch the graph at all.&lt;/p&gt;

&lt;p&gt;The second is a 2026 survey called "SoK: Agentic Retrieval-Augmented Generation," &lt;a href="https://arxiv.org/abs/2603.07379" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2603.07379&lt;/a&gt; which lays out a taxonomy for how agent systems plan, retrieve, and manage memory.&lt;/p&gt;

&lt;p&gt;Its main point is that the field keeps building more aggressive retrieval pipelines when what actually helps is cost-aware orchestration: knowing when to spend the tool call and when not to.&lt;/p&gt;

&lt;p&gt;That's the argument for keeping the tool count small. Three tools with a gate in front of them beat twenty-eight tools and a model left to guess which one applies.&lt;/p&gt;

&lt;p&gt;The third, "A-RAG," &lt;a href="https://arxiv.org/html/2602.03442v1" rel="noopener noreferrer"&gt;https://arxiv.org/html/2602.03442v1&lt;/a&gt;, proposes exposing retrieval as a small set of tiered interfaces instead of one flat search function so a model can reach for the cheapest tool that could plausibly answer the question before escalating to something heavier.&lt;/p&gt;

&lt;p&gt;That's where the fourth, even simpler tool tier came from after the initial build: a pure existence check that costs almost nothing, sitting below the three tools we started with.&lt;/p&gt;

&lt;p&gt;None of the three papers describe a finished product. They describe an architectural shape. The actual code, the WASM parser, the session cache, and the tier classifier-all of that is ours. But the shape mattered because it's the opposite of what the existing open-source options do.&lt;/p&gt;

&lt;p&gt;CodeGraph and code-review-graph are genuinely useful tools, and code-review-graph in particular has real, published numbers behind it: an 8.2x average token reduction, 100% recall on blast-radius analysis, support for twenty-four languages.&lt;/p&gt;

&lt;p&gt;If you run one project all day in one editor, either of them will likely serve you well. But both are built around the assumption that retrieval should always be available, which means a persistent index, a background watcher, and in code-review-graph's case twenty-eight separate tools for a model to choose between on every call.&lt;/p&gt;

&lt;p&gt;That's a build strategy we didn't want to make, as an agent that jumps between projects doesn't want a watcher spinning up in the background of every one of them. A smaller or cheaper model doesn't reason well with twenty-eight tool options in front of it, and the three separate provider bugs we found tonight are proof of exactly how easily that kind of complexity turns into silent failure instead of a helpful answer.&lt;/p&gt;

&lt;p&gt;For a user, the practical difference is that nothing installs a daemon on your machine, nothing leaves a folder behind after the task is done, and the agent doesn't get slower or more confused as the tool list grows. You get code intelligence sized to the task in front of you, not a standing index you're paying for whether you're using it or not.&lt;/p&gt;

&lt;p&gt;In summary, our tool is honed and specific to the task, keeping bloat and errors down.&lt;/p&gt;

&lt;p&gt;Will it do everything-Swiss Army knife style? NO, that's not the design.&lt;br&gt;
Making the gate correct&amp;nbsp;itself&lt;/p&gt;

&lt;p&gt;We designed the gate expecting it to be right the majority of the time, and built the system so a wrong call doesn't stay invisible.&lt;/p&gt;

&lt;p&gt;The classifier itself is conservative by construction: a single-file edit only skips the graph when it's genuinely small, and anything near that line rounds up to the scoped tier automatically, because a redundant graph build costs nothing and a missing one costs context.&lt;/p&gt;

&lt;p&gt;The agent watches its own behavior in real time. If it's reading or grepping several files one at a time instead of asking for the dependency graph, that's the signature of a task that needs the scoped tier, so we built a live nudge that catches this mid-task and tells the model to call the blast-radius tool directly, instead of waiting to notice the pattern in a log after the run is over. And the gate's reasoning isn't buried anywhere.&lt;/p&gt;

&lt;p&gt;The CODEMAP panel shows exactly what tier the agent's own tool calls reached and why, right in the session, not as an afterthought in telemetry. The result is a gate that corrects itself while it's working, not one that just hopes it guessed right.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Under the hood:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Conservative bias - vektor-codegate.js: BYPASS threshold tightened from 50 to 30 lines, with a 15-line borderline zone above it that rounds up to SCOPED instead of trusting a marginal count.&lt;/p&gt;

&lt;p&gt;Hard fallback / mid-task escalation - new noteFileTouched(runId, filePath) tracks distinct files read one-at-a-time per run; once a run crosses 3 files without ever reaching SCOPED, grep_only, read_file_or_grep, and the main read_file tool all surface an explicit nudge back to the model telling it to call expand_blast_radius instead of continuing file-by-file. This catches a misroute live, mid-task, not just in telemetry after the run ends.&lt;/p&gt;

&lt;p&gt;CODEMAP "why" surface - vektor-gui-bridge.js's _handleCodemap now reports runMaxTier (the highest tier the agent's own tool calls actually reached this run) alongside the forced CODEMAP build, and vektor-graph-ui.html's renderer shows it inline under the scope receipt instead of that fact living only in the JSONL telemetry file.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What actually broke whilst&amp;nbsp;testing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We wired all of this into our agent's real tool loop and ran it against an actual task today in testing: refactor error handling across three files into a shared helper. That's when the filesystem hallucination showed up, and it took real digging to find the cause.&lt;/p&gt;

&lt;p&gt;Our agent talks to most language models through a hand-rolled JSON protocol since not every provider supports native tool calling the way Claude does. After every tool call, the result gets folded back into the conversation for the next turn. We found the exact line doing that folding, and it was silently destroying the tool's output.&lt;/p&gt;

&lt;p&gt;When the model's most recent message contained an array of structured data instead of a plain string, which happens on every single turn after the first tool call, JavaScript's default array-to-string conversion turned it into the literal text [object Object].&lt;/p&gt;

&lt;p&gt;The real directory listing, the actual file the model needed to see, was gone by the time it reached the model. Every provider except the one using a different, correctly typed code path hit this on every turn. That's why three separate models all failed the same strange way.&lt;/p&gt;

&lt;p&gt;We rewrote that step to render tool results as plain, readable text instead of letting JavaScript guess. The very next run, the model read the output correctly and reasoned about it like a normal conversation, because for the first time it actually was one.&lt;/p&gt;

&lt;p&gt;Fixing that surfaced a second error immediately. Our search_code tool kept reporting zero files scanned, even for files we knew existed. The glob pattern the model was passing, something like *&lt;em&gt;/&lt;/em&gt;.js, was being matched against a bare filename with no directory in it at all, so a pattern built around slashes could never match anything. One rewrite of the glob logic later, the same search returned real matches across nearly two hundred files.&lt;/p&gt;

&lt;p&gt;Neither of those issues had anything to do with the graph engine we set out to build. They were sitting in the surrounding agent harness the whole time, and they only became visible because we ran the graph tools inside a real conversation instead of testing them in isolation.&lt;/p&gt;

&lt;p&gt;That's the actual real-world testing example. A feature that only gets exercised in a unit test can look finished and still be sitting inside a loop that quietly breaks it the first time a real model uses it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fud1wbjztquiquaxsrwiw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fud1wbjztquiquaxsrwiw.png" alt=" " width="799" height="380"&gt;&lt;/a&gt;&lt;br&gt;
Agent Screen in Slipstream 1.8.0&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where this lives&amp;nbsp;now&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;All of this ships as part of Slipstream 1.8.0's autonomous agent, the same one behind the AGENT tab in the Slipstream GUI. Point it at a real project, give it a task, and the gate makes the retrieval decision on its own. You'll see it in the CODEMAP panel, which used to run a blind, unscoped walk of the entire workspace every time you clicked it.&lt;/p&gt;

&lt;p&gt;It now shows exactly what the agent actually needed context for, scoped to the files it touched, with a small receipt explaining why: how many files, how many hops, and whether the reverse-dependency index was already warm from an earlier step in the same session.&lt;/p&gt;

&lt;p&gt;We also shipped a few smaller additions past the initial build. A fourth, even more cost-effective tool tier for a pure existence check before the agent commits to reading a whole file. A real session-scoped reverse-dependency index, built once from actual import resolution instead of a repeated string scan, so "What else imports this file?" answers reliably instead of approximately.&lt;/p&gt;

&lt;p&gt;Basic route detection for Express and Next.js projects, so an API route shows up as a route in the graph instead of just another anonymous function.&lt;/p&gt;

&lt;p&gt;None of it is running as a background service waiting for you to open a project. It runs when a task actually needs it, and it gets out of the way the moment the task is done.&lt;/p&gt;

&lt;p&gt;CodeGraph also understands 36 programming languages instead of the original five.&amp;nbsp;&lt;/p&gt;

&lt;p&gt;Where it previously only had real code understanding for JavaScript and TypeScript, everything else fell back to basic pattern matching. It now extracts functions, classes, and imports accurately across Python, Go, Rust, Java, C++, Kotlin, Ruby, and dozens more.&lt;br&gt;
&amp;nbsp;&lt;br&gt;
Rather than bundling all 36 languages into the install, each language's parser downloads quietly in the background the first time the agent opens a file in that language, gets verified against a cryptographic fingerprint before it's trusted, and is cached locally so every use after that is instant.&lt;/p&gt;

&lt;p&gt;The ten most common languages: TypeScript, JavaScript, Python, Go, and others-pre-load automatically right after installation, so the experience feels instant out of the box, while the long tail of less common languages stays out of the way until actually needed.&lt;/p&gt;

&lt;p&gt;If you're already running Slipstream, the update is a normal npm refresh, and the agent tools show up automatically the next time you open a project with the code tool group enabled.&lt;/p&gt;

&lt;p&gt;If you haven't tried our autonomous agent yet, this is a good week to start. Point it at something you already know well, give it a real multi-file task, and watch the CODEMAP panel light up with exactly the files it actually needed, not the four hundred it used to grab out of habit.&lt;/p&gt;

&lt;p&gt;Full setup and docs are at vektormemory.com/docs and 1.8.0 info at &lt;a href="https://vektormemory.com/docs/changelog" rel="noopener noreferrer"&gt;https://vektormemory.com/docs/changelog&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;AI Agent&lt;br&gt;
Agentic Ai&lt;br&gt;
LLM&lt;br&gt;
Code&lt;br&gt;
Knowledge Graph&lt;/p&gt;

</description>
      <category>ai</category>
      <category>code</category>
      <category>llm</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
