<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Jonathan Berg</title>
    <description>The latest articles on DEV Community by Jonathan Berg (@jonathanmberg).</description>
    <link>https://dev.to/jonathanmberg</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4106872%2Fb718acb7-7bc5-443a-9397-2e883bc2524d.jpg</url>
      <title>DEV Community: Jonathan Berg</title>
      <link>https://dev.to/jonathanmberg</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/jonathanmberg"/>
    <language>en</language>
    <item>
      <title>The real test for agent memory is switching agents</title>
      <dc:creator>Jonathan Berg</dc:creator>
      <pubDate>Tue, 29 Sep 2026 13:26:40 +0000</pubDate>
      <link>https://dev.to/jonathanmberg/the-real-test-for-agent-memory-is-switching-agents-1ll5</link>
      <guid>https://dev.to/jonathanmberg/the-real-test-for-agent-memory-is-switching-agents-1ll5</guid>
      <description>&lt;p&gt;I built Sirro because I kept running into the same problem while building apps for clients. I would get a section working in one project, start another, and then explain it to a coding agent all over again. The second version was close, but it was never quite the thing I had already made.&lt;/p&gt;

&lt;p&gt;The problem gets more obvious when you switch tools. You save your work in one agent's project files, then open a different agent and find yourself back at the beginning. The first tool may remember the conversation perfectly. That doesn't help much if the next tool can't use the finished work.&lt;/p&gt;

&lt;p&gt;I've been thinking about how to test whether a shared library actually fixes this. It is easy to make a demo where an agent retrieves a snippet from the same tool that saved it. The harder test is to change the agent, change the project, and see whether you still get the thing you meant to keep.&lt;/p&gt;

&lt;h2&gt;
  
  
  A test I would trust
&lt;/h2&gt;

&lt;p&gt;Start with a working piece of code in a real project. Something specific enough that rebuilding it from a prompt would lose details. Save it through one coding agent, then close that project. Open a fresh project in a different tool and ask for the same piece by name.&lt;/p&gt;

&lt;p&gt;At that point, I want to know what actually came back. Did the new agent retrieve the saved code, or did it generate something that merely looks similar? Can I see which version it used? Will it work in the new project without quietly dropping the behavior that made the original worth saving?&lt;/p&gt;

&lt;p&gt;I also want to try a third tool that wasn't part of the original save. If the library only works when I stay in the tool where I made the asset, then it is another local memory with a nicer interface. Portability has to survive the handoff.&lt;/p&gt;

&lt;p&gt;This is a test plan, not a claim that every coding agent already passes it. Sirro's beta has had real authentication problems, and getting someone into the product is part of the test. A perfect retrieval flow behind a sign-in that fails for a new user isn't useful. We spent part of this week fixing that before asking anyone to judge the library itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Save the result, then test the transfer
&lt;/h2&gt;

&lt;p&gt;A lot of agent-memory discussion starts with what an agent can remember about you. That matters, but for coding work I keep coming back to the artifact. The reason I save a section is that I want that section again, including the small decisions I don't want to explain twice.&lt;/p&gt;

&lt;p&gt;A good memory system should make it possible to tell the difference between retrieval and imitation. If an agent says it used something from my library, I should be able to inspect the original asset and the version it pulled. Otherwise a plausible output can hide the fact that the work was rebuilt from scratch.&lt;/p&gt;

&lt;p&gt;I don't think a standard format or a shared folder settles this on its own. Those can make an asset available. The useful question is whether I can move between the tools I actually use, name the work I already did, and keep its important details intact.&lt;/p&gt;

&lt;p&gt;That is the test I'm interested in now: make it once, switch agents, and see if the next one can really pick it up.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>mcp</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Agent memory needs an editor</title>
      <dc:creator>Jonathan Berg</dc:creator>
      <pubDate>Tue, 22 Sep 2026 13:29:06 +0000</pubDate>
      <link>https://dev.to/jonathanmberg/agent-memory-needs-an-editor-3c3</link>
      <guid>https://dev.to/jonathanmberg/agent-memory-needs-an-editor-3c3</guid>
      <description>&lt;p&gt;I have been thinking about what happens when several agents share the same memory.&lt;/p&gt;

&lt;p&gt;The obvious version is a database they can all read and write. A coding agent saves a useful pattern. A browser agent adds research. A personal agent records a decision. The next agent opens the workspace and starts with everything the others learned.&lt;/p&gt;

&lt;p&gt;That sounds powerful until two agents disagree.&lt;/p&gt;

&lt;p&gt;One agent marks a component as the preferred version. Another replaces it after a failed run. A third saves a summary without seeing the original decision. If every write has equal authority, the shared memory slowly becomes a pile of confident contradictions.&lt;/p&gt;

&lt;p&gt;The difficult part of multi-agent memory is not storage. It is deciding what gets to become true.&lt;/p&gt;

&lt;h2&gt;
  
  
  Shared memory creates a merge problem
&lt;/h2&gt;

&lt;p&gt;Developers already know this problem from Git. Many people can work on the same codebase because changes have authors, diffs and a path to merge. We do not let every branch silently rewrite main.&lt;/p&gt;

&lt;p&gt;Agent memory needs the same basic discipline.&lt;/p&gt;

&lt;p&gt;A useful memory entry should keep its source, the agent that proposed it, the work that produced it and the decision that accepted it. When something changes, the system should preserve the old version long enough to explain what happened.&lt;/p&gt;

&lt;p&gt;Without that history, an agent cannot tell the difference between a tested decision and a guess saved five minutes ago.&lt;/p&gt;

&lt;p&gt;This matters more as agents move beyond code. A coding agent may know that an integration failed for a specific technical reason. A browser agent may later see updated documentation and conclude that the old limitation is gone. Both observations can be valid at the time they were made. Replacing one sentence with the other throws away the part that makes the knowledge useful.&lt;/p&gt;

&lt;p&gt;The memory needs to carry the change, not only the latest text.&lt;/p&gt;

&lt;h2&gt;
  
  
  One agent should hold merge authority
&lt;/h2&gt;

&lt;p&gt;My current view is that a shared workspace needs a primary agent. That agent does not need to do every task. It needs to decide what becomes part of the durable memory.&lt;/p&gt;

&lt;p&gt;Other agents can propose additions and corrections. They can attach evidence, point out stale knowledge and suggest that an asset should be replaced. The primary agent can accept, reject or ask for more context.&lt;/p&gt;

&lt;p&gt;This gives the workspace a clear authority model. A temporary agent used for one task cannot silently rewrite a decision that every other agent will trust tomorrow.&lt;/p&gt;

&lt;p&gt;The human still sits above the system. They can pin a decision, change which agent has authority or inspect why something was accepted. The point is to make those boundaries visible instead of pretending all agent output deserves the same weight.&lt;/p&gt;

&lt;p&gt;Silent self-editing is especially dangerous here. If an agent can change its own instructions or the shared memory after a failed run, it may improve. It may also erase the evidence of why it failed. Proposed changes should be reviewable events, even when another agent handles the review automatically.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory should become more trusted as it is used
&lt;/h2&gt;

&lt;p&gt;A good shared memory should learn from reuse.&lt;/p&gt;

&lt;p&gt;If an asset is saved, retrieved and used successfully across several projects, that is evidence. If agents keep pulling it and then rewriting it, that is evidence too. Usage can tell the primary agent which parts of the library are stable and which need attention.&lt;/p&gt;

&lt;p&gt;This is more useful than treating every saved item as permanent. Some knowledge should expire. Some should be tied to a version, environment or client. Some should remain a proposal until it has survived real work.&lt;/p&gt;

&lt;p&gt;The library becomes valuable because it records decisions and their outcomes. More agents writing to it should increase that value, as long as they cannot quietly overwrite one another.&lt;/p&gt;

&lt;p&gt;I still think agents need a shared place for the work we leave behind. I just no longer think shared access is enough.&lt;/p&gt;

&lt;p&gt;The next step is shared memory with authorship, diffs and merge authority. Otherwise we are giving every agent a notebook and letting each one erase the previous page.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>security</category>
    </item>
    <item>
      <title>Your coding agent should start from your library</title>
      <dc:creator>Jonathan Berg</dc:creator>
      <pubDate>Tue, 15 Sep 2026 13:30:52 +0000</pubDate>
      <link>https://dev.to/jonathanmberg/your-coding-agent-should-start-from-your-library-5bcj</link>
      <guid>https://dev.to/jonathanmberg/your-coding-agent-should-start-from-your-library-5bcj</guid>
      <description>&lt;p&gt;Every new project used to start the same way for me. I would open Cursor or Claude Code, explain what I wanted, and then spend the next hour describing things I had already built somewhere else.&lt;/p&gt;

&lt;p&gt;The authentication flow was sitting in one repo. The pricing section I liked was in another. A good way to structure the dashboard had disappeared into an old conversation. None of it was really gone, but none of it was available either.&lt;/p&gt;

&lt;p&gt;That is a strange way to build software. The more work I shipped, the more useful material I left behind. My experience was growing while the agent still started every project from zero.&lt;/p&gt;

&lt;p&gt;This week I tried to reduce the whole problem to four words: create, connect, save, reuse.&lt;/p&gt;

&lt;p&gt;Create something worth keeping. Connect the tools you already code with. Save the finished work to a library. Reuse it from any project and any agent.&lt;/p&gt;

&lt;p&gt;It sounds obvious when the loop is drawn out. In practice, most coding agents stop after create. They help produce the component, fix or pattern, then leave it inside the conversation where it happened. The next conversation has no idea it exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  A conversation is a bad place to keep finished work
&lt;/h2&gt;

&lt;p&gt;Conversations are useful while something is uncertain. You can explore an approach, reject a few versions and explain what feels wrong. The problem is that the finished result stays mixed together with all of that process.&lt;/p&gt;

&lt;p&gt;Six weeks later, finding it means remembering which agent you used, which project was open and roughly what you said. Even if you find the thread, the agent still has to work out which part was the final version and why it mattered.&lt;/p&gt;

&lt;p&gt;This is where people reach for longer prompts. They write another paragraph explaining the old component, paste a few snippets and hope the new agent reconstructs the same idea. Sometimes it does. Sometimes it gives you something close enough that you only notice the differences after you have started building on top of them.&lt;/p&gt;

&lt;p&gt;I do not want a better guess. I want the thing I already approved.&lt;/p&gt;

&lt;h2&gt;
  
  
  Libraries change how the next project starts
&lt;/h2&gt;

&lt;p&gt;A saved asset is more than a code snippet. It is a decision you do not have to make again.&lt;/p&gt;

&lt;p&gt;The useful part may be a React component, a database pattern, a working integration or a short note about why one approach failed. What matters is that it has crossed the line from conversation into something named and reusable.&lt;/p&gt;

&lt;p&gt;Once that happens, the next project can start with the library you have built through actual work. You can ask Cursor to save a component today and pull it into Claude Code next month. The value is not tied to the chat, repo or agent that happened to produce it.&lt;/p&gt;

&lt;p&gt;This also changes the economics of getting better at AI-assisted coding. Normally the agent gets another chance to generate every time. A library lets your own work compound. Ten finished projects should make the eleventh easier because the good parts are already there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory should include things, not only facts
&lt;/h2&gt;

&lt;p&gt;A lot of agent memory is described as information about the user: preferences, instructions, summaries and facts. That is useful, but builders also need memory of the work itself.&lt;/p&gt;

&lt;p&gt;Knowing that I prefer a certain kind of pricing page is weaker than having the exact section I chose. Remembering that we solved an OAuth edge case is weaker than keeping the fix and the reason the obvious approach failed. A summary can guide another generation. An asset can become part of the next product.&lt;/p&gt;

&lt;p&gt;This is the difference I care about. The goal is not to make an agent sound more familiar with me. The goal is to stop leaving useful work behind.&lt;/p&gt;

&lt;p&gt;When the library grows, coding becomes less conversational and more compositional. You still use agents to explore and create. You just stop asking them to recreate every good idea from a description.&lt;/p&gt;

&lt;p&gt;The agent should start from what you have already built, not from what it can guess you meant.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>webdev</category>
      <category>mcp</category>
    </item>
    <item>
      <title>Bug fixes are memory too</title>
      <dc:creator>Jonathan Berg</dc:creator>
      <pubDate>Tue, 08 Sep 2026 22:53:02 +0000</pubDate>
      <link>https://dev.to/jonathanmberg/bug-fixes-are-memory-too-4mok</link>
      <guid>https://dev.to/jonathanmberg/bug-fixes-are-memory-too-4mok</guid>
      <description>&lt;p&gt;Watch someone debug with a coding agent for an hour and you see the same shape every time. Twenty exchanges of wrong hypotheses, logs pasted back and forth, three fixes that fix nothing, and then the actual answer, which is usually embarrassingly small. A missing await. A header the proxy strips. An environment variable that only exists locally. The whole afternoon compresses into three lines of diff.&lt;/p&gt;

&lt;p&gt;Then the conversation ends, and those three lines go into the codebase while everything that made them findable stays in the chat log. Two months later a different project hits the same class of bug, a fresh agent session opens with no idea any of this happened, and you pay for the whole search again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generation is cheap, diagnosis is not
&lt;/h2&gt;

&lt;p&gt;The reason fixes matter more than components comes down to an asymmetry. If you lose a component, a fresh agent regenerates a decent one from a short description, because generation has a wide target: many implementations are acceptable and the model knows the neighborhood. Diagnosis has a narrow target. The value was never the three lines. The value was eliminating everything else, and that elimination work does not transfer when only the diff survives.&lt;/p&gt;

&lt;p&gt;This is why re-encountering a bug you already solved feels so bad. You are not annoyed at writing the fix twice. You are annoyed at paying for the search twice: the same wrong turns, the same token spend, the same context window filling up with dead ends.&lt;/p&gt;

&lt;h2&gt;
  
  
  The anatomy of a fix worth saving
&lt;/h2&gt;

&lt;p&gt;A diff alone is almost useless to a future agent, because the future agent will never find it. What makes a fix an asset instead of a leftover is everything around it:&lt;/p&gt;

&lt;p&gt;The symptom, written the way you would search for it. Not "fixed auth bug" but the error message, the observable behavior, the words someone would actually type when it happens again. Retrieval fails when the label describes the solution instead of the problem.&lt;/p&gt;

&lt;p&gt;The root cause in one sentence. Not the story of the afternoon, just what was actually wrong.&lt;/p&gt;

&lt;p&gt;The failed approaches. This is the part nobody saves and the part that saves the most. Negative knowledge is what lets the next session skip the first five wrong hypotheses instead of re-walking them.&lt;/p&gt;

&lt;p&gt;The environment. Library versions, runtime, the specific combination that made this bug possible, because half of these fixes are only true for a particular stack state.&lt;/p&gt;

&lt;p&gt;Save that, and a fix stops being a scar and starts being a building block.&lt;/p&gt;

&lt;h2&gt;
  
  
  Chat logs are where fixes go to die
&lt;/h2&gt;

&lt;p&gt;The default storage for all of this right now is the conversation history, and conversation history fails in a specific way: it is write-only. The knowledge goes in and never comes back out, because there is no path from "this error looks familiar" to the right thread from six weeks ago. You cannot grep for a vibe.&lt;/p&gt;

&lt;p&gt;So the habit that matters is deliberate capture at the moment the fix lands, while the root cause and the wrong turns are still in front of you. Save the fix once, with its symptom and its negative knowledge, and retrieve it in every project after that. Bug fixes are memory too.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this compounds into
&lt;/h2&gt;

&lt;p&gt;I keep coming back to the same line: what persists between sessions is memory, and what your memory stack is should be your highest priority. Components were the obvious first asset, but fixes are where the compounding gets loud, because every saved diagnosis removes an entire class of future cost instead of just a future generation step.&lt;/p&gt;

&lt;p&gt;A team, or even one developer, that treats fixes this way ends up with a strange advantage: their agents get better at their specific stack over time. The obscure failure modes of their infrastructure stop being rediscovery exercises. Debugging starts to look like composition: recognize the symptom, pull the asset, move on.&lt;/p&gt;

&lt;p&gt;That is a large part of why I am building Sirro, and I wrote about the server side of it in &lt;a href="https://dev.to/jonathanmberg/what-shipping-a-hosted-mcp-server-taught-me-about-agent-memory-ll0"&gt;what shipping a hosted MCP server taught me about agent memory&lt;/a&gt;. The short version: the agents are interchangeable, and what they keep is not. Save the fix. Your future self, in a fresh session with a familiar-looking error, will thank you.&lt;/p&gt;

</description>
      <category>debugging</category>
      <category>ai</category>
      <category>mcp</category>
      <category>webdev</category>
    </item>
    <item>
      <title>What shipping a hosted MCP server taught me about agent memory</title>
      <dc:creator>Jonathan Berg</dc:creator>
      <pubDate>Wed, 02 Sep 2026 21:44:49 +0000</pubDate>
      <link>https://dev.to/jonathanmberg/what-shipping-a-hosted-mcp-server-taught-me-about-agent-memory-ll0</link>
      <guid>https://dev.to/jonathanmberg/what-shipping-a-hosted-mcp-server-taught-me-about-agent-memory-ll0</guid>
      <description>&lt;p&gt;Every coding-agent conversation ends the same way. The component you spent an hour getting exactly right, the fix you never want to derive again, the pattern that finally worked: all of it stays behind in a chat log you will never open again. Developers are producing huge amounts of valuable, hard-won work this way, and almost none of it survives the conversation that produced it. Grabbing those assets and giving them somewhere permanent to live is the biggest unlock I see in how we work with these tools.&lt;/p&gt;

&lt;p&gt;That belief is what I am building with Sirro: a hosted memory layer for AI coding agents, served over MCP on streamable HTTP. Agents like Cursor, Claude Code, and Codex save the work that earned persistence and pull it back in a different project, in a different agent, three weeks later. You stop letting the AI guess from zero every session and start composing from building blocks you already trust.&lt;/p&gt;

&lt;p&gt;Running it in production taught me four things the docs never mentioned. It started the week I put the server on a public directory and watched the listing bounce every client that was not Cursor, Claude Code, or Codex. The error users saw was "Unknown agent", which is a terrible thing to tell someone who just tried to install your product. The bug was not in my protocol handling. It was in a decision I had made early on, to allowlist OAuth clients by name, and fixing it forced me to learn how the MCP ecosystem actually works.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. OAuth is the real integration surface
&lt;/h2&gt;

&lt;p&gt;When you build an MCP server over stdio, authentication is someone else's problem. The moment you host it, OAuth 2.1 with dynamic client registration becomes the front door, and every client walks through it differently.&lt;/p&gt;

&lt;p&gt;My mistake was assuming clients identify themselves cleanly. They do not. Cursor, Claude Code, and Codex each present different client names, some use loopback redirect URIs, and hosted gateways like Smithery sit between your server and the user's actual client, so the identity you see is the gateway's, not the user's. An allowlist by client name works until the ecosystem ships a new client, which it does constantly. Every new client was a support ticket waiting to happen.&lt;/p&gt;

&lt;p&gt;The fix that actually holds: accept loopback redirect URIs and dynamic client registration as the default path, and gate only the parts you genuinely must gate. The spec already tells you how to do this. What it does not tell you is how many real clients deviate from what you assumed, so build the permissive path first and add restrictions only when you have a concrete reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Every token you return is a design decision
&lt;/h2&gt;

&lt;p&gt;Agents read your tool responses into their context window, which means your API output competes with the user's actual work for the scarcest resource in the system. This changed how I design responses.&lt;/p&gt;

&lt;p&gt;A naive list endpoint returns full bodies. That is fine for a REST API serving a frontend. It is careless for an MCP server, because dumping forty assets into a context window to answer "do I have something like this" burns the user's budget on data the agent will mostly discard. So list returns search snippets and metadata, get returns one full asset on demand, and a compose tool does assembly server-side instead of making the agent do it in-context.&lt;/p&gt;

&lt;p&gt;I also cap assets at 64KB. Anything larger is a file, not a memory, and it belongs in the repo, not in the memory layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. A context window is not memory
&lt;/h2&gt;

&lt;p&gt;The industry currently answers the forgetting problem with files: CLAUDE.md, cursor rules, AGENTS.md. These help and they half solve it. They are write-heavy, they are read on faith at session start, and nobody curates them, so they rot. Nathan Marz put it better than I could in a reply to me: persisted corrections are right, but CLAUDE.md-style files only half-solve the problem.&lt;/p&gt;

&lt;p&gt;The reason they half-solve it is that they confuse the container with the mechanism. A context window is session state. It dies when the thread ends, it gets truncated when it fills, and it carries no semantics about what deserves to survive. Memory needs explicit save and retrieve: the developer decides what earned persistence, and the agent retrieves it when it is relevant, not because it was appended to a file months ago. That is the difference between an agent that starts every project empty and one that starts with everything you already figured out.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Cost math is a design input, not an afterthought
&lt;/h2&gt;

&lt;p&gt;A memory layer only gets used if the economics work at small scale, so I designed against a concrete target: thousands of users, each storing dozens of assets, on boring infrastructure. That constraint shaped real choices. Postgres full-text search with trigram indexes is enough at this scale, so there is no vector database to operate. Search returns snippets because full bodies are both a context problem and an egress cost. Asset size caps keep storage predictable.&lt;/p&gt;

&lt;p&gt;The pattern here generalizes: pick your scale assumption early, price it, and let it veto features. The boring stack that fits the budget beats the impressive one that does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this goes
&lt;/h2&gt;

&lt;p&gt;Coding agents made generating software cheap, which moved the bottleneck from creation to memory. A conversation is where an asset gets created, but on its own it has no memory beyond the thread it lives in. What compounds is the loop after the conversation: save what earned it, then compose from it next time. Building blocks instead of fresh guesses. A developer who has spent six months working with coding agents should never open a new project with an empty toolbox, and giving all those closed conversations somewhere permanent to live is how we get there.&lt;/p&gt;

&lt;p&gt;That is what I am building with Sirro, and these are the constraints it runs under today. If you are building in the MCP ecosystem, I would genuinely like to hear what broke for you first.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>typescript</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
