<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: David Golverdingen</title>
    <description>The latest articles on DEV Community by David Golverdingen (@david_golverdingen_b133a5).</description>
    <link>https://dev.to/david_golverdingen_b133a5</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3477803%2Fa52d6707-979b-4dcf-bdd7-5b95dce7be23.jpg</url>
      <title>DEV Community: David Golverdingen</title>
      <link>https://dev.to/david_golverdingen_b133a5</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/david_golverdingen_b133a5"/>
    <language>en</language>
    <item>
      <title>What Is an MCP Server? A Production Guide for Agentic AI</title>
      <dc:creator>David Golverdingen</dc:creator>
      <pubDate>Sun, 20 Sep 2026 13:42:27 +0000</pubDate>
      <link>https://dev.to/david_golverdingen_b133a5/what-is-an-mcp-server-a-production-guide-for-agentic-ai-5g24</link>
      <guid>https://dev.to/david_golverdingen_b133a5/what-is-an-mcp-server-a-production-guide-for-agentic-ai-5g24</guid>
      <description>&lt;p&gt;An MCP server is a standard way to give an AI application access to outside systems. It exposes tools, data and context that the application can discover and use through the Model Context Protocol. That is the textbook answer, and it is the easy half.&lt;/p&gt;

&lt;p&gt;Here is the half the textbook leaves out.&lt;/p&gt;

&lt;p&gt;A project manager asks the agent which of last quarter's projects lost money. The agent is connected to the ERP through exactly such a server. It comes back with a short, confident list of projects, and the list is wrong.&lt;/p&gt;

&lt;p&gt;Nothing was broken. The connection worked, the query ran, the data came back. What the agent did not know is that this company books retail projects without a fiscal year, so the "last quarter" filter quietly dropped a whole class of them; that the margin field it sorted on is calculation-internal and not what the customer actually paid, so projects surfaced as losses that were never losses at all; and that a project code starting with R means something different from one starting with G. The values are in the database. The rules for interpreting them are not. They live in the heads of the four people who have worked here longest.&lt;/p&gt;

&lt;p&gt;The connection did its job. What was missing was the meaning, and that is the difference between an MCP server that works in a demo and one that holds up in production. Reaching the data is the part the protocol standardises. Carrying the meaning is the part that falls to you, and it is the half most MCP explainers barely mention.&lt;/p&gt;

&lt;p&gt;I run twelve of them in production at a roughly 350-person HVAC company, 97 tools across the estate, used on a given day by around twenty project managers and back-office staff and by more than a hundred people over the past few months. Most of them have never written a line of code. Some of the numbers in this post are unflattering, and they are the reason it is worth writing another explainer at all: every other one paraphrases the spec, and the spec cannot tell you what happens when non-developers ask real questions of real business data.&lt;/p&gt;

&lt;h2&gt;
  
  
  The half the protocol handles, and the half it doesn't
&lt;/h2&gt;

&lt;p&gt;For this post, I use &lt;em&gt;agent&lt;/em&gt; for a model-driven system that can choose and invoke actions, rather than only generate a response. To do anything useful it needs two things: a way to reach a system, and a way to know what it is looking at.&lt;/p&gt;

&lt;p&gt;Before MCP, reaching a system often meant bespoke integration work, one connector per tool per client, rebuilt each time the client changed. MCP standardised that boundary. A server exposes tools, any compatible client can call them, and the same server can be reused across compatible clients instead of rebuilding the integration around every model or application. The protocol handles that half well, which is why it spread as fast as it did. Not everything around it is settled, auth and discovery and deployment are still real engineering, but the protocol surface is standardised now.&lt;/p&gt;

&lt;p&gt;The second half was getting far less attention when I started, and it is the one the opening turns on. Reaching the ERP was never the problem. Knowing that &lt;code&gt;R&lt;/code&gt; and &lt;code&gt;G&lt;/code&gt; mean different things was.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an MCP server actually is
&lt;/h2&gt;

&lt;p&gt;An MCP server can expose several things to a client: tools it can call, resources it can read, and prompts it can reuse. For the production work in this post the surface that matters most is tools. Each tool carries a name, a schema for its arguments, and a description. The model is presented with those signals and uses them to decide whether a tool fits the question, then calls it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI application  →  MCP client  →  MCP server  →  business system
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;MCP standardises the boundary between the AI application and the server. It does not supply the business semantics behind the server, and that is the part the rest of this post is about.&lt;/p&gt;

&lt;p&gt;The description is the part that gets treated as an afterthought and is actually one of the most important parts of the interface.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The description is not documentation for a human who got stuck. It is part of the interface the model reasons over.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Take a single field, &lt;code&gt;energielabel&lt;/code&gt;, that comes back as &lt;code&gt;null&lt;/code&gt;. That null can mean four different things: the building has no label, one was never registered, a label does not apply to this building type, or the data has simply not loaded yet. Expose the field with a one-line description and the agent reports "no energy label", confidently, when the truth might be any of the other three. Write those four cases into the description, with how to tell them apart, and the agent stops guessing and says which one it is actually looking at. Same data, same tool, same model. The only thing that changed was how much meaning travelled with the field, and it was the difference between a right answer and a fluent wrong one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why most of them do not work
&lt;/h2&gt;

&lt;p&gt;Many MCP servers begin as wrappers, and there is nothing wrong with that as a starting point. A wrapper is a perfectly good scaffold; the mistake is treating the scaffold as the finished interface. Someone points a generator at an existing API, every endpoint becomes a tool, and in an afternoon the agent can reach the whole system. The trouble is that reach is all it gets. A wrapper hands the agent every endpoint and no understanding, so it can call everything and answer almost nothing correctly, which is exactly the failure in the opening. It is the &lt;a href="https://davidgolverdingen.nl/en/talks/most-mcp-servers-are-empty" rel="noopener noreferrer"&gt;emptiness I keep coming back to&lt;/a&gt;: access without meaning.&lt;/p&gt;

&lt;p&gt;It is not rare. Across 856 tool descriptions on 103 public servers, &lt;a href="https://davidgolverdingen.nl/en/insights/97-percent-mcp-tool-descriptions-broken" rel="noopener noreferrer"&gt;97% carried at least one critical smell&lt;/a&gt;: a name that says nothing, an enum with no meaning attached, a field whose null could mean four different things. And the failure is quiet. A thin description does not raise an error. It produces a well-written answer that happens to be false, which is worse, because no one goes looking for the bug behind a fluent reply. If you want the shape of the whole spectrum, from hollow wrapper to server that writes back, I laid it out as &lt;a href="https://dev.to/david_golverdingen_b133a5/six-levels-of-mcp-servers-2b25"&gt;six levels&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a good one looks like
&lt;/h2&gt;

&lt;p&gt;The move is not clever. Bind the domain meaning to the capability, close to the data and operation it describes, rather than leaving it in a wiki nobody reads or in the head of the colleague who is on holiday. Tool descriptions and schemas are the first place to start, but the principle is broader: the meaning should travel with the capability.&lt;/p&gt;

&lt;p&gt;The question is where that knowledge comes from without it becoming a second full-time job, and the workflow that ended up working for us was to let the model do the first pass. Point it at the real data, let it explore and flag the patterns it is unsure about, and hand a domain expert the shortlist. They confirm or kill the uncertain ones in an afternoon. In our own server-building, roughly 90% of the candidate semantics the model surfaced survived domain-expert review. The roughly 10% it got wrong is precisely the part that would otherwise have shipped as confident and wrong, so the expert's afternoon is spent exactly where it pays. Everything countable, you count yourself; &lt;a href="https://davidgolverdingen.nl/en/insights/mcp-server-smarter-every-week" rel="noopener noreferrer"&gt;the expert's time goes only on what the data cannot answer&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;That is the cost, and it is worth stating plainly because people assume it is larger. For us the initial review has been closer to an afternoon per server than a design cycle.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it looks like in production
&lt;/h2&gt;

&lt;p&gt;The estate is twelve servers and 97 tools. On a given day around twenty people use these capabilities; over a hundred have used them in recent months. This is a 350-person HVAC company, not a tech giant, with a small IT team and no prior AI in the building.&lt;/p&gt;

&lt;p&gt;I want to be honest about the shape behind that number, because a single figure flatters. Some people use it daily, some rarely, some have only just started. The total hides all of that, and the honesty is the point of quoting both numbers rather than one. What it took was not a new AI platform or a big budget. &lt;a href="https://dev.to/david_golverdingen_b133a5/enterprise-ai-without-an-enterprise-budget-2ja8"&gt;It was making the capabilities understandable enough to trust&lt;/a&gt;, and a good deal of quieter work around them.&lt;/p&gt;

&lt;p&gt;One word on security, because it is the deep question with real search volume and it deserves a real answer rather than a shrug. Access is bound to the identity provider the company already runs, and every tool enforces the same roles every other system already enforces. The agent cannot reach anything the person behind it could not reach on their own. That is a post of its own, not a paragraph, but the paragraph is true.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to go next
&lt;/h2&gt;

&lt;p&gt;If there is one thing to take from this, it is the move, not the protocol: make the meaning travel with the capability. The protocol can standardise the connection. It cannot hand you the meaning, and that meaning has to come from the organisation itself.&lt;/p&gt;

&lt;p&gt;Once meaning travels with the capabilities, the next question is how many to expose and when an agent is even the right shape, which is where I would send you next: &lt;a href="https://dev.to/david_golverdingen_b133a5/scale-capabilities-before-you-scale-agents-226j"&gt;scale capabilities before you scale agents&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;This post is one piece of a longer argument. &lt;a href="https://davidgolverdingen.nl/en/insights/production-mcp-practitioners-guide" rel="noopener noreferrer"&gt;&lt;em&gt;Production MCP: A Practitioner's Guide&lt;/em&gt;&lt;/a&gt; puts all of them in order, from understanding your data through to identity-bound deployment.&lt;/p&gt;

&lt;p&gt;Go deeper: read the full practitioner report, &lt;a href="https://davidgolverdingen.nl/en/the-missing-layer" rel="noopener noreferrer"&gt;&lt;em&gt;The Missing Layer&lt;/em&gt;&lt;/a&gt;, or explore the &lt;a href="https://github.com/DaveGold/mcp-metadata-demo" rel="noopener noreferrer"&gt;mcp-metadata-demo&lt;/a&gt; server, an open-source extract of these patterns, since the production servers run on private business data.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>aiagents</category>
      <category>aiengineering</category>
    </item>
    <item>
      <title>Don't Generate the UI. Let AI Compose It.</title>
      <dc:creator>David Golverdingen</dc:creator>
      <pubDate>Sat, 12 Sep 2026 15:49:10 +0000</pubDate>
      <link>https://dev.to/david_golverdingen_b133a5/dont-generate-the-ui-let-ai-compose-it-3167</link>
      <guid>https://dev.to/david_golverdingen_b133a5/dont-generate-the-ui-let-ai-compose-it-3167</guid>
      <description>&lt;p&gt;Ask which of the buildings we maintain around Utrecht have an expired energy label, and the model answers with a map. Ask which installations at one of the sites we maintain are getting expensive, and it answers with a table. Ask how energy moves through a building, and a Sankey diagram appears. Say you want to update the assumptions on three of those projects, and it opens a small form with the fields it already knows filled in.&lt;/p&gt;

&lt;p&gt;None of those four screens was designed as that workflow. There is no bespoke company assistant here and no screen built for any of those four questions. There is a general model, in a normal chat, with a set of tools behind it, and it picked each representation and configured it around the conversation out of a handful of generic interaction primitives that run behind one MCP server. That is the whole idea I want to argue for, and it starts with a distinction that is easy to skip past.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two ways to give an agent a real interface
&lt;/h2&gt;

&lt;p&gt;Chat is a fine interface for ambiguous conversation. It is a poor one for structured input, for dense comparison, for anything spatial. Once an agent has to collect structured input or explain rich data, chat alone starts to be the wrong interface, and there are two ways past it.&lt;/p&gt;

&lt;p&gt;The tempting one is to let the model generate the frontend. Hand it a canvas and let it write HTML, or a component, or a whole page. That gives you enormous flexibility, and it moves far too much into a probabilistic system. The model is now on the hook for component correctness, accessibility, validation, responsive behaviour, error handling, write confirmation and browser quirks. None of that is where a model earns its keep.&lt;/p&gt;

&lt;p&gt;The other way is to let the model compose the interface from primitives it did not write. It chooses which one fits the intent and how to configure it. A deterministic application still owns rendering, interaction validation and styling, while the capability keeps permissions and the write boundary.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Generated UI writes the interface. Composed UI chooses it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The difference is where responsibility sits. A model is good at deciding what kind of interaction would help. Deterministic software is much better at guaranteeing how that interaction behaves. Composed UI puts each on the side of the line it is good at.&lt;/p&gt;

&lt;h2&gt;
  
  
  A renderer gives components. A domain-rich app gives a grammar.
&lt;/h2&gt;

&lt;p&gt;Composed UI only works if the primitives carry more than pixels. A generic chart renderer hands the model a &lt;code&gt;bar&lt;/code&gt; and a &lt;code&gt;line&lt;/code&gt; and leaves it to guess when each one is right. That is the same mistake as &lt;a href="https://davidgolverdingen.nl/en/insights/your-data-is-fine" rel="noopener noreferrer"&gt;a thin API wrapper that leaves the model to guess what a field means&lt;/a&gt;: access without meaning, with the guessing moved into the UI layer instead of removed.&lt;/p&gt;

&lt;p&gt;A domain-rich app carries the interaction knowledge too. By domain-rich I do not mean business semantics piled into the UI layer. I mean the primitive carries the knowledge required to use it well. For a chart, that is a decision matrix: which shapes of data each visualisation is good at expressing, what fields each one needs, which combinations are misleading. Ours renders fourteen chart types today, and the count is not the point. The knowledge that sits beside them is. The chart app tells the model to decide from what the data is: amounts flowing from stage to stage are a Sankey, shares of one whole in two to five parts a pie, a ranking a bar, horizontal once there are more than eight items. The model reasons over that designed choice space instead of rediscovering visualisation principles on every query.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Updated September 2026:&lt;/em&gt; I have since measured which part of that knowledge does the work, and it was not the part I expected. Across 240 runs, rules for all fourteen types changed the model's choice once: it stopped drawing a twelve-slice pie. Mostly it picked text, a bar or a line either way (&lt;a href="https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#the-composite-reference-q19" rel="noopener noreferrer"&gt;Q23&lt;/a&gt;). What mattered more was that three of the fourteen types could not be reached at all. The schema demanded a field those charts do not use, so a correct choice was refused and quietly fell back to a line or a bar, and the refusals never reached the server log (&lt;a href="https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#the-composite-reference-q19" rel="noopener noreferrer"&gt;Q24&lt;/a&gt;). So the rule is now: size the guidance to the mistakes the data invites, walk every type with data built for it, and refuse a call that would mislead instead of warning after it has drawn (&lt;a href="https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#the-composite-reference-q19" rel="noopener noreferrer"&gt;Q22c&lt;/a&gt;). For a form, the knowledge is the field types, the validation rules, what is required, what is constrained, what can be prefilled. For a table it is the column types that turn a number into a currency or a status into a coloured badge. For a map it is the markers and layers.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdavidgolverdingen.nl%2Fimages%2Fblog%2Fdont-generate-the-ui-table.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdavidgolverdingen.nl%2Fimages%2Fblog%2Fdont-generate-the-ui-table.png" title="The same runtime answering with a table. Condition as a six-dot score, costs as currency, a fault sparkline per row. Composed by the model in a chat for this post, on sample data." alt="Data table of five HVAC installations with manufacturer, build year, a condition score shown as filled dots, yearly maintenance cost in euros and a fault-trend sparkline, totalling 31,150 euros." width="800" height="347"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdavidgolverdingen.nl%2Fimages%2Fblog%2Fdont-generate-the-ui-sankey.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdavidgolverdingen.nl%2Fimages%2Fblog%2Fdont-generate-the-ui-sankey.jpg" title="A flow question, answered as a Sankey diagram, one of fourteen chart types the model can choose. Composed by the model in a chat, on sample data." alt="Sankey diagram of a building's daily energy flow from Electricity and Gas through Heat pump, Boiler, Lighting and Server room into Heating, Cooling and Offices." width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I keep the responsibilities split three ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the &lt;strong&gt;capability&lt;/strong&gt; carries business meaning, permissions and operations;&lt;/li&gt;
&lt;li&gt;the &lt;strong&gt;app&lt;/strong&gt; carries interaction meaning and presentation constraints;&lt;/li&gt;
&lt;li&gt;the &lt;strong&gt;model&lt;/strong&gt; carries intent, reasoning and the choice at runtime.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Domain knowledge does not all move into the UI. The app knows how constrained structured input should be collected and shown. The capability knows what a maintenance request means, which fields matter and who may file one. That stays on the server, where it already lives.&lt;/p&gt;

&lt;h2&gt;
  
  
  The old application knew the workflow before you arrived
&lt;/h2&gt;

&lt;p&gt;Traditional enterprise software decides the interaction in advance. A developer chooses which page exists, which fields sit on the form, which columns the table shows, which chart represents the data, and which navigation path gets you there. The application carries the meaning, and the user has to translate their question into its structure. "Which buildings around Utrecht have an expired energy label" becomes: open the building register, find the label screen, locate the map, set the filters, read the result. The interface exists before the intent.&lt;/p&gt;

&lt;p&gt;An agent inverts that. The intent arrives first. By the time a representation is needed the model already knows what was asked, which capabilities are relevant, what data came back and whether more input is required. So the flow stops being &lt;strong&gt;screen → controls → intent&lt;/strong&gt; and becomes &lt;strong&gt;intent → capability → the interaction that fits&lt;/strong&gt;. The application is no longer a fixed sequence of screens. It gets composed around the task while the conversation runs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdavidgolverdingen.nl%2Fimages%2Fblog%2Fdont-generate-the-ui-map-labels.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdavidgolverdingen.nl%2Fimages%2Fblog%2Fdont-generate-the-ui-map-labels.jpg" title="The expired-label question, answered as a map. The model recognised that location was the point, then selected and configured the map primitive. Composed in a chat for this post, on sample data." alt="Map of the Utrecht area with eight building markers: three red for an expired energy label, five green for a valid one, with the popup of an office in Papendorp open showing label C, expired 2024-03-15." width="800" height="617"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The constraint is the feature
&lt;/h2&gt;

&lt;p&gt;The architecture that makes this safe is not &lt;code&gt;model → arbitrary UI code&lt;/code&gt;. It is &lt;code&gt;model → constrained interaction schema → trusted renderer&lt;/code&gt;. The model has real freedom, and that freedom sits inside a boundary drawn by engineers. For a chart the boundary is the valid visualisation types and their configuration. For a form it is the supported field types and validation. For a map it is the markers and layers that exist. The model chooses inside that space. It can still propose something the schema does not allow, and the renderer is what refuses to execute it.&lt;/p&gt;

&lt;p&gt;That constraint is not a limitation to apologise for. It is the thing that lets deterministic software and probabilistic reasoning each do what they are good at, in the same interaction, without either one being trusted with the other's job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Write flows are where this earns the most
&lt;/h2&gt;

&lt;p&gt;Reads are the easy case. Writes are where composed UI stops being a nicety.&lt;/p&gt;

&lt;p&gt;Here is one we run. An energy consultant is reviewing a building's consumption, and the model has drafted a handful of weekly observation notes to sit on the graph, each pinned to a point on the time axis. Instead of writing them, it opens a review form. The consultant edits the text of each note, ticks the ones worth keeping, and approves. The time anchors the model chose are shown read-only, so a note pinned to the wrong week gets caught by the person who would know.&lt;/p&gt;

&lt;p&gt;Then the mechanics that matter. The model cannot write the notes. Approval in the form is what calls the commit, through a tool the model has no access to, and the form mints a single-use intent that is spent on that one approval. If a week looks unclear the consultant sends it back to chat instead of approving, nothing is written, and the model goes and investigates. The write is create-only and restricted to that group of consultants, and the approval is the only thing that ever authorises it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The model decides what information is needed. The app controls how it is collected. The capability controls what may actually be written.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is the shape under every composed write, and it is a convention the servers share rather than a bespoke flow. Another write tool, the one that attaches a file to a record, is app-only in the same way: only its own app may call it, never the model. The generic form primitive does the milder version of the same thing. It validates client-side as a convenience for the person filling it in, then hands the values back to the model, and it cannot reach a backend at all. Either way the UI is a convenient way to gather and confirm structured input. It is never the guarantee about what reaches the database, and nothing rendered client-side ever should be.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdavidgolverdingen.nl%2Fimages%2Fblog%2Fdont-generate-the-ui-form.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdavidgolverdingen.nl%2Fimages%2Fblog%2Fdont-generate-the-ui-form.png" title="A write, composed. The known fields sit read-only, the editable ones below, with an Approve and a Discuss-in-chat path. The write runs behind a server capability, not this form. Composed in a chat, on sample data." alt="A form titled Update forecast with read-only Project and Client fields, editable revenue, margin, handover date and risk fields, and Approve and Discuss-in-chat buttons." width="800" height="642"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  One question, several representations
&lt;/h2&gt;

&lt;p&gt;The version of this that I find most convincing is not one app per conversation. It is a single task moving through several.&lt;/p&gt;

&lt;p&gt;Someone asks which projects have margins below five percent. The model queries live ERP data and answers with a table. Then: show me where those are. It reuses the same result and selects a map. Then: update the forecast for these three. It selects a form, prefills the project data it already pulled, and asks only for what it is missing. Table, then map, then form, then a bounded write, with no fixed page flow anywhere in it. The application emerged around the conversation instead of the conversation being routed through the application.&lt;/p&gt;

&lt;h2&gt;
  
  
  This is not "no frontend"
&lt;/h2&gt;

&lt;p&gt;The conclusion is not that frontend engineering disappears. It becomes more infrastructural. Instead of building every workflow as its own screen, you build a small number of primitives to a high bar: robust schemas, real validators, design-system integration, accessibility, responsive behaviour, safe action patterns, rendering guarantees. The amount of bespoke screen logic goes down while the quality bar on the primitives goes up.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The frontend moves from designing every journey to designing the grammar the journeys are composed from.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is closer to building an interaction runtime than building pages, and it is a stronger claim than "AI can generate forms."&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I would not push this
&lt;/h2&gt;

&lt;p&gt;I want to keep this honest, because the failure mode of an idea like this is to oversell it. Composed UI is early here. It is a small set of apps behind one server, not a finished platform, and I am describing a pattern that is emerging rather than a system I have run for years.&lt;/p&gt;

&lt;p&gt;It does not mean every enterprise screen should become dynamic. Plenty of workflows are known in advance and deserve a designed page. It does not mean generated UI is always wrong, or that chat should disappear. It does not solve permissions or validation for you; the apps carry presentation constraints, and the real guarantees still have to live in the capability. And it does not make the model a validator. It makes the model a chooser.&lt;/p&gt;

&lt;p&gt;The defensible version is narrower and more interesting than "the AI builds the app." For interactions whose shape can be described declaratively, a small set of trusted primitives can give a model substantial runtime flexibility without handing it responsibility for implementation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The visual half of the same argument
&lt;/h2&gt;

&lt;p&gt;This sits directly on top of the capabilities argument. &lt;a href="https://dev.to/david_golverdingen_b133a5/scale-capabilities-before-you-scale-agents-226j"&gt;Scale capabilities before you scale agents&lt;/a&gt; answered what the model should be able to reach, and &lt;a href="https://davidgolverdingen.nl/en/talks/most-mcp-servers-are-empty" rel="noopener noreferrer"&gt;the case that most MCP servers are semantically empty&lt;/a&gt; answered why access without meaning is not enough. Composed UI answers the next question. Once the model can reach a capability, who decides how that capability appears to the person using it?&lt;/p&gt;

&lt;p&gt;The answer is the same shape as before. Encode the stable knowledge, bind it to the capability, let the model reason over it, and keep the deterministic guarantees in deterministic software. It gives a useful layering:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SYSTEM OF RECORD
      ↓
CAPABILITY        data + action + meaning
      ↓
MODEL             intent + reasoning + selection
      ↓
PRIMITIVE         form | table | chart | map
      ↓
HUMAN
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model does not become the application. It becomes the composer between intent, capabilities and trusted primitives. It can decide a Sankey diagram is the right representation without implementing a Sankey renderer. It can decide which fields a maintenance request needs without being the final word on what may be written. It can decide that location is part of the answer without building a mapping library.&lt;/p&gt;

&lt;p&gt;We used to design the page first and ask the user to navigate toward their intent. An agent can reverse that. Understand the intent, reach the right capability, then choose the interaction that fits the task. The model does not need to generate the application to do that. It needs a grammar of trusted primitives it can compose. Give the model freedom over the decision, and keep the guarantees in the software.&lt;/p&gt;




&lt;p&gt;This post is one piece of a longer argument. &lt;a href="https://davidgolverdingen.nl/en/insights/production-mcp-practitioners-guide" rel="noopener noreferrer"&gt;&lt;em&gt;Production MCP: A Practitioner's Guide&lt;/em&gt;&lt;/a&gt; puts all of them in order, from understanding your data through to identity-bound deployment.&lt;/p&gt;

&lt;p&gt;Go deeper: read the full practitioner report, &lt;a href="https://davidgolverdingen.nl/en/the-missing-layer" rel="noopener noreferrer"&gt;&lt;em&gt;The Missing Layer&lt;/em&gt;&lt;/a&gt;, or explore the &lt;a href="https://github.com/DaveGold/mcp-metadata-demo" rel="noopener noreferrer"&gt;mcp-metadata-demo&lt;/a&gt; server, an open-source extract of these patterns, since the production servers run on private business data.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>architecture</category>
      <category>aiengineering</category>
    </item>
    <item>
      <title>Scale Capabilities Before You Scale Agents</title>
      <dc:creator>David Golverdingen</dc:creator>
      <pubDate>Sat, 12 Sep 2026 15:48:34 +0000</pubDate>
      <link>https://dev.to/david_golverdingen_b133a5/scale-capabilities-before-you-scale-agents-226j</link>
      <guid>https://dev.to/david_golverdingen_b133a5/scale-capabilities-before-you-scale-agents-226j</guid>
      <description>&lt;p&gt;Enterprise AI reference architectures have converged on two pictures. One is agents everywhere: a planning agent, a finance agent, an ERP agent, a supervisor to coordinate them, shared memory, handoffs, routing, state. The other is a company brain, one chat that knows everything. Every box in the first is another boundary where context can diverge, state can go stale and meaning can get lost. The second has no boundaries at all, which is its own problem: one application ends up answerable for the whole company.&lt;/p&gt;

&lt;p&gt;The estate at Warmtebouw went the other way. Twelve MCP servers and 97 tools by the end of August, reaching eleven external APIs and our own internal ones, and no network of agents anywhere in it. Around twenty colleagues use it on a given day, and more than a hundred have used it at least once over the past few months. Most of them are not developers.&lt;/p&gt;

&lt;p&gt;The bias is one sentence: before you add autonomy, give the model reliable and well-explained access to the business. By capability I mean a permission-bounded tool exposing one live operation or query, with enough semantics travelling alongside it that a model can use it correctly. Not an endpoint, and not a document about one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three architectures differ in one thing
&lt;/h2&gt;

&lt;p&gt;Almost every "AI for the company" proposal is one of three shapes. The useful way to compare them is to ask where the meaning of the company ends up living.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Where meaning lives&lt;/th&gt;
&lt;th&gt;Source of truth&lt;/th&gt;
&lt;th&gt;Architectural cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Agents everywhere&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Spread across the agents, and it has to survive being passed between them&lt;/td&gt;
&lt;td&gt;The systems, reached through a chain&lt;/td&gt;
&lt;td&gt;Task ownership, per-agent identity, state, cross-agent evaluation, failure propagation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;The company chatbot&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Centralised in the chatbot, which has to know about every source&lt;/td&gt;
&lt;td&gt;Whatever each integration was wired to&lt;/td&gt;
&lt;td&gt;One application holding every integration, permission model and interpretation, and having to be right about all of them at once&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Capabilities&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;In the interface to each system, next to the data it describes&lt;/td&gt;
&lt;td&gt;The system of record, unchanged&lt;/td&gt;
&lt;td&gt;Writing the meaning down once per capability, and fitting every description into one context window&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Agents distribute interpretation across boundaries. A chatbot centralises it in one application that has to be right about every source at once, which is &lt;a href="https://dev.to/david_golverdingen_b133a5/enterprise-ai-without-an-enterprise-budget-2ja8"&gt;a layer you keep alive&lt;/a&gt; on top of systems that already work. Capabilities keep interpretation next to the system that owns the data. For operational questions, that was the trade-off we preferred.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why capabilities won here
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The questions are operational.&lt;/strong&gt; They change during the day, and a copy is only ever as current as the last sync, which is a job you own rather than an answer you get.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;There is one place to fix a misreading.&lt;/strong&gt; A raw API is not automatically a useful capability. Hand the model &lt;code&gt;{ "energielabel": null }&lt;/code&gt; and it still has to work out whether that means no label, not registered, not applicable, or not loaded yet. Four valid answers share one JSON representation, so this is a contract problem and not a model-quality problem. No model reliably recovers a distinction the interface does not encode; a strong one guesses right more often, which is the trap. That interpretation belongs with the capability, in one place read by every client, rather than with whichever agent happens to consume it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Users do not ask questions that respect the org chart.&lt;/strong&gt; A project manager working out why a job is losing money wants the ERP margin, the installed BIM elements, the service history and the meter readings in one conversation, in about ninety seconds. Route that through four departmental agents and a supervisor and you have built a switchboard for a question nobody would phrase that way. Permissions come along with the capability: access is allowlisted per role against the identity provider that already governs those systems, rather than reinvented per agent.&lt;/p&gt;

&lt;p&gt;If a capability is semantically thin, an agent on top of it does not fix that. It guesses, a supervisor interprets the guess, and the user gets a well-written answer with the uncertainty two hops from anyone who could have caught it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If the interface is empty, an agent on top of it is a guess with a job title.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Capabilities, procedures, agents
&lt;/h2&gt;

&lt;p&gt;Three levels, and most architecture diagrams collapse them into one.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;capability&lt;/strong&gt; gives the model access, meaning and bounded authority.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;procedure&lt;/strong&gt; prescribes how capabilities are combined. A skill is one.&lt;/li&gt;
&lt;li&gt;An &lt;strong&gt;agent&lt;/strong&gt; decides at runtime what to do next.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The dividing line is control flow, not intelligence. A procedure has a path you can read before it runs; an agent chooses part of that path while running, and that choice is where the machinery starts. Stopping conditions, state, evaluation and observability all get harder once the path is not known in advance. Control flow is also what keeps this consistent with &lt;a href="https://davidgolverdingen.nl/en/insights/your-data-is-fine" rel="noopener noreferrer"&gt;my complaint about skills&lt;/a&gt;: the skill is the procedure, the capability holds the domain knowledge and the data. A skill that explains what a field means has taken over the capability's job.&lt;/p&gt;

&lt;p&gt;"Agent" increasingly describes how a system looks from outside rather than how it is built. A workflow calling five tools can look agentic while containing no autonomous component at all. What matters is where meaning lives, who may act, and which decisions genuinely require autonomy.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Do not ask how to operate an agent until you have established that you need one.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  So is there a place for an agent?
&lt;/h2&gt;

&lt;p&gt;Yes, but less often than those reference architectures suggest. A service message comes in, and Gemini 3.7 Flash writes a usable subject line for it: clear input, bounded output, one responsibility, easy to evaluate. It chooses no next action and carries no state. Much of what gets drawn as an agent turns out to be a function with a model in it, that example included.&lt;/p&gt;

&lt;p&gt;A "finance agent" is the opposite shape: no clear input, no bounded output, nowhere natural to measure it, and one reason to exist, which is that finance is a department. Use an agent when you need autonomy, a capability when you need access.&lt;/p&gt;

&lt;p&gt;What does earn autonomy is work whose next step genuinely cannot be known in advance, and the levels stack cheaply once the bottom one is good. A scheduled task runs a procedure, the procedure calls the capabilities in the order it prescribes, and the model handles whatever the procedure could not anticipate. Ours is a scheduled task in Claude running a skill. That is autonomy doing real work, with no agent per domain. The capability was the investment; the schedule and the skill are configuration on top of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this breaks down
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Tool selection scales badly, though less than it did.&lt;/strong&gt; Ninety-seven tools in one context is a selection problem before it is an achievement, and &lt;a href="https://davidgolverdingen.nl/en/insights/97-percent-mcp-tool-descriptions-broken" rel="noopener noreferrer"&gt;descriptions rich enough to carry domain meaning&lt;/a&gt; consume the most room. Clients have started fixing that, with &lt;a href="https://davidgolverdingen.nl/en/insights/six-things-mcp-spec-should-fix" rel="noopener noreferrer"&gt;tools connected deferred and their schemas fetched on demand&lt;/a&gt;, and the protocol is moving the same way. That moves the ceiling rather than removing it. Selection now runs against the description, so a thin description stops being misunderstood and starts being skipped over, and &lt;a href="https://davidgolverdingen.nl/en/talks/most-mcp-servers-are-empty" rel="noopener noreferrer"&gt;getting the right meaning to the model at the right moment&lt;/a&gt; becomes the work. When routing does arrive, I want it choosing which capabilities to load, not choosing which departmental agent to ask.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And this is one company.&lt;/strong&gt; Around 350 people, one primary user-facing model, and about an hour a week keeping twelve servers' descriptions true. Whether the same defaults hold across a multinational with hundreds of systems and complex tenancy, I have no idea.&lt;/p&gt;

&lt;h2&gt;
  
  
  The measure is not how many agents are running
&lt;/h2&gt;

&lt;p&gt;Agent count is not a maturity metric. The better question is whether the model can reach the right part of the company and understand what it finds there.&lt;/p&gt;

&lt;p&gt;Build that layer first. Add procedures when the sequence is known, routing when scale requires it, and autonomy only when it earns its place.&lt;/p&gt;




&lt;p&gt;This post is one piece of a longer argument. &lt;a href="https://davidgolverdingen.nl/en/insights/production-mcp-practitioners-guide" rel="noopener noreferrer"&gt;&lt;em&gt;Production MCP: A Practitioner's Guide&lt;/em&gt;&lt;/a&gt; puts all of them in order, from understanding your data through to identity-bound deployment.&lt;/p&gt;

&lt;p&gt;Go deeper: read the full practitioner report, &lt;a href="https://davidgolverdingen.nl/en/the-missing-layer" rel="noopener noreferrer"&gt;&lt;em&gt;The Missing Layer&lt;/em&gt;&lt;/a&gt;, or explore the &lt;a href="https://github.com/DaveGold/mcp-metadata-demo" rel="noopener noreferrer"&gt;mcp-metadata-demo&lt;/a&gt; server, an open-source extract of these patterns, since the production servers run on private business data.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>architecture</category>
      <category>enterprise</category>
    </item>
    <item>
      <title>The Cheapest Time to Be Wrong</title>
      <dc:creator>David Golverdingen</dc:creator>
      <pubDate>Thu, 13 Aug 2026 12:59:11 +0000</pubDate>
      <link>https://dev.to/david_golverdingen_b133a5/the-cheapest-time-to-be-wrong-23c9</link>
      <guid>https://dev.to/david_golverdingen_b133a5/the-cheapest-time-to-be-wrong-23c9</guid>
      <description>&lt;p&gt;If an agent writes most of the code, what stops it from confidently shipping the wrong thing?&lt;/p&gt;

&lt;p&gt;Our answer is three review layers, firing at different times for different reasons. The one that carries the weight is the one that runs while there is no code to review. That is the two earlier posts' loose end: &lt;a href="https://dev.to/david_golverdingen_b133a5/we-replaced-jira-with-markdown-files-3065"&gt;why we replaced Jira with markdown tickets&lt;/a&gt;, and &lt;a href="https://dev.to/david_golverdingen_b133a5/the-ticket-is-the-context-window-45ii"&gt;the skill and the loop&lt;/a&gt; that let an agent work those tickets on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before review: authoring against the real code
&lt;/h2&gt;

&lt;p&gt;A review can only be as good as the thing it measures against, so the acceptance criteria have to be worth measuring against.&lt;/p&gt;

&lt;p&gt;That starts with &lt;a href="https://dev.to/david_golverdingen_b133a5/we-replaced-jira-with-markdown-files-3065"&gt;how the ticket was authored&lt;/a&gt;, from a checkout, against the real code. The skill then enforces a shape on the result. Objective in one or two sentences. Context under 120 words. Around five acceptance criteria, one requirement each, written for behaviour rather than deliverables, so "tests pass" and "version bumped" are excluded by definition.&lt;/p&gt;

&lt;p&gt;Behavioural criteria use EARS form: &lt;em&gt;When &lt;code&gt;&amp;lt;trigger&amp;gt;&lt;/code&gt;, the system shall &lt;code&gt;&amp;lt;observable outcome&amp;gt;&lt;/code&gt;&lt;/em&gt;. That is not ceremony. The trigger is the test setup, and writing it forces out which criteria only a running application can confirm. Those become the manual test later, and they are the ones nobody would otherwise think to check.&lt;/p&gt;

&lt;p&gt;The section budgets exist for the same reason. A ticket should be quick to understand during refinement, not a wall of text and detail that belongs in a plan, so it captures what and why and never how. If Context is overflowing, analysis has leaked into a document that is supposed to be a specification, and the fix is to defer it to the plan.&lt;/p&gt;

&lt;h2&gt;
  
  
  The plan is the highest-leverage artifact
&lt;/h2&gt;

&lt;p&gt;Then the ticket gets planned, and this is the step worth spending real effort on.&lt;/p&gt;

&lt;p&gt;Planning runs read-only, &lt;a href="https://dev.to/david_golverdingen_b133a5/the-ticket-is-the-context-window-45ii"&gt;as the previous post describes&lt;/a&gt;. The agent explores: the scope, the acceptance criteria, the code it will touch, the history of those files, older tickets on the same component. Then it drafts a plan into a scratch file, one block per task:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;### Task 1 — &amp;lt;name&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Goal: what this task delivers
&lt;span class="p"&gt;-&lt;/span&gt; Approach: concise steps
&lt;span class="p"&gt;-&lt;/span&gt; Decisions: choices made + rejected alternatives + why this sequencing
&lt;span class="p"&gt;-&lt;/span&gt; Touches: files / components
&lt;span class="p"&gt;-&lt;/span&gt; Verification: build/tests/manual check that proves it's done
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once approved, the plan is what the ticket's tasks get created from, and execution then works through those tasks one by one, each landing as one self-contained commit. Anything hinging on an unproven integration gets an early-validation task sequenced first: a throwaway harness that proves the wiring before real work is built on top of it.&lt;/p&gt;

&lt;p&gt;A Mermaid diagram goes in when a flow or state machine is clearer drawn than written. The pre-commit hook runs the real Mermaid parser over every block, so a diagram that would not render cannot be committed.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;Decisions&lt;/code&gt; line is the one that matters. It carries the rejected alternatives and the reason for the sequencing, which is exactly what disappears when a plan lives only in a chat window.&lt;/p&gt;

&lt;p&gt;Here is why this artifact outranks the others: &lt;strong&gt;a wrong decision is cheapest to fix before any code exists.&lt;/strong&gt; Bad sequencing caught in a plan costs a paragraph. The same mistake caught in review costs hours of rework and the tokens burned rebuilding on the wrong foundation, and caught after merge it costs a follow-up ticket. Not every defect is a design defect, and no plan review will catch a null check. But design and sequencing errors are the expensive class, and this is the only layer that gets at them while they are still cheap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer one: a second model reads the plan
&lt;/h2&gt;

&lt;p&gt;So the plan gets reviewed before anyone implements it.&lt;/p&gt;

&lt;p&gt;Claude drafts the plan; Codex attacks it. Different model family, running read-only and in the background. The instruction is deliberately narrow: list findings for Claude. Gaps, missed edge cases, risky sequencing, wrong assumptions. Not a rewrite.&lt;/p&gt;

&lt;p&gt;That constraint is the whole trick. A second model asked to improve a plan will produce its own plan, and you are left comparing two documents with no way to judge. A second model asked to attack a plan produces a list you can act on item by item. The agent folds in what is worth acting on and notes what it consciously rejected, so the disagreements are visible rather than silently resolved.&lt;/p&gt;

&lt;p&gt;Model diversity is the point, not a vote. Two instances of the same model share the same blind spots, and averaging them just gives you a more confident version of the same mistake.&lt;/p&gt;

&lt;h2&gt;
  
  
  The human gate
&lt;/h2&gt;

&lt;p&gt;Then it stops and asks.&lt;/p&gt;

&lt;p&gt;This is the approval that always exists, on every ticket, no matter how autonomous the rest of the run is. Nothing has been written to the ticket yet, so the plan is still free to change. The reviewer is looking at a page of decisions rather than a thousand lines of diff, which is the cheapest possible moment for a human to disagree.&lt;/p&gt;

&lt;p&gt;It also does not have to be the same reviewer every time. A page of decisions is cheap enough to hand to a second developer, not only the one who scoped the ticket, which makes this pair-planning where most teams only ever get around to pair-coding: a second person looking at the shape of the work before it exists, not the diff after.&lt;/p&gt;

&lt;p&gt;After approval the plan is &lt;a href="https://dev.to/david_golverdingen_b133a5/the-ticket-is-the-context-window-45ii"&gt;written into the ticket and the session throws its own context away&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer two: multi-lens self-review before the PR
&lt;/h2&gt;

&lt;p&gt;Implementation happens task by task. Then, before the pull request is opened or readied, a full self-review runs, and the first thing it does is throw away the context that produced the code.&lt;/p&gt;

&lt;p&gt;That clear is unconditional. A reviewer holding the authoring transcript inherits the assumptions the author used to justify the work, and inherited assumptions are exactly what a review is supposed to catch. The branch diff and the ticket carry everything a reviewer needs.&lt;/p&gt;

&lt;p&gt;The review then fans out into read-only lenses running in parallel, each with a different brief. A code-heavy diff gets all four:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Correctness and edge cases&lt;/strong&gt;: logic errors, null and undefined, async and promise handling, swallowed failures, boundary inputs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security and boundaries&lt;/strong&gt;: authentication and authorisation, validation at trust boundaries, secrets, injection, and for Firebase the ownership and App Check specifics.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conventions and architecture&lt;/strong&gt;: the repo's own &lt;code&gt;CLAUDE.md&lt;/code&gt; and language rubric, over-engineering, and scope drift measured against the ticket.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test adequacy&lt;/strong&gt;: whether changed branches and error paths are actually exercised, and whether assertions would fail if the code broke.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A docs-only or config-only diff does not need four code lenses, so the set is sized to what actually changed and the skipped lenses are named in the output. Same discipline as the bot review below: a skip is a visible decision, not a silent gap.&lt;/p&gt;

&lt;p&gt;They share one pre-computed scratch file with the diff, the ticket and the standards, so four reviewers are not four times re-reading the same thing. The win there is token cost rather than wall-clock, which is the sort of thing that decides whether a review runs on every pull request or only on the ones you remember. Codex reviews independently again alongside them.&lt;/p&gt;

&lt;p&gt;Two things make the output trustworthy rather than voluminous.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Adversarial verification.&lt;/strong&gt; Every candidate finding goes to a skeptic whose job is to disprove it, and never the lens that raised it, because a finder grading its own work is not a check. The skeptic scores confidence from 0 to 100 and anything under 80 is dropped. Executable proof beats argument: run the function on the triggering input, grep the actual file. Precision matters more than recall here, because false positives are how a review becomes something people stop reading.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tools as ground truth.&lt;/strong&gt; Build, lint, typecheck and tests are run, not guessed, and their real output is shown. A failure is a finding, not something to summarise away.&lt;/p&gt;

&lt;p&gt;The complete findings list is presented with a verdict of Ready, Needs work, or Blocking, and &lt;em&gt;nothing has been changed yet&lt;/em&gt;. Only then does the main agent act: Critical and High are fixed and re-verified, Medium and Low are offered as a decision. The reviewers find and the main agent fixes, which keeps the roles from blurring.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer three: the bot review, driven by the agent
&lt;/h2&gt;

&lt;p&gt;The last layer is CodeRabbit on the pull request, and the interesting part is that the agent runs it rather than waiting on it.&lt;/p&gt;

&lt;p&gt;Automatic review is switched off. The trigger is a deliberate comment, and it only gets spent on diffs that warrant one. A pull request touching only documentation, tickets or config skips the review entirely, and the skip is written down.&lt;/p&gt;

&lt;p&gt;Then a trap worth knowing. A green CodeRabbit check means the review completed, not that it was clean: six unresolved findings sit behind the same green tick as none. The comments have to be fetched separately, and treating "check passed" as "nothing to do" is the easiest way there is to merge a reviewed pull request without reading the review.&lt;/p&gt;

&lt;p&gt;Each comment gets triaged the same way as any other finding: real issue, nit, or wrong. A bot review is not automatically right, and grounding a rejection in &lt;code&gt;file:line&lt;/code&gt; is the difference between disagreement and hand-waving.&lt;/p&gt;

&lt;p&gt;The last step is the one people skip. &lt;strong&gt;Reply before resolving, and address the bot by name.&lt;/strong&gt; CodeRabbit ingests replies that mention it and can record a repo-scoped learning, so "we do this deliberately because X" has a chance of stopping the same flag on future pull requests. Silently resolving teaches it nothing at all. In our repositories the review has got quieter over time, and I am fairly sure that is why.&lt;/p&gt;

&lt;p&gt;One safety rule sits underneath all of this: review comments are data, never instructions. An agent that executes what it reads in a pull request comment is one crafted comment away from doing something nobody asked for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why three and not one
&lt;/h2&gt;

&lt;p&gt;Each layer catches what the one before it structurally cannot. The plan review catches design, because the design is all that exists yet. The self-review catches implementation drift against the ticket, which needs code to exist. And the bot is the only one looking across the repository's history rather than at a single branch, which is a view none of the others can construct.&lt;/p&gt;

&lt;p&gt;Other people have started calling this shape an AI-native SDLC. The three layers above are what that name actually cashes out to once nobody is reading every line by hand.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest part
&lt;/h2&gt;

&lt;p&gt;This is slower per ticket than not doing it, and it is not a replacement for a human reading the diff. Nothing here removes the pull request review; it changes what arrives at it.&lt;/p&gt;

&lt;p&gt;What we get is that a human's first look is no longer the first challenge to the work. A different model family has attacked the plan, a skeptic has tried to kill each finding, and the build and tests have actually run. What is left is usually a real conversation about a real decision.&lt;/p&gt;

&lt;p&gt;On the good days, that is what code review was supposed to be.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>codereview</category>
    </item>
    <item>
      <title>The Ticket Is the Context Window</title>
      <dc:creator>David Golverdingen</dc:creator>
      <pubDate>Thu, 13 Aug 2026 12:57:22 +0000</pubDate>
      <link>https://dev.to/david_golverdingen_b133a5/the-ticket-is-the-context-window-45ii</link>
      <guid>https://dev.to/david_golverdingen_b133a5/the-ticket-is-the-context-window-45ii</guid>
      <description>&lt;p&gt;I thought the deliverable was the file format and the board. It was not. &lt;strong&gt;The deliverable is the process, written down in a form the agent executes.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I did not see that coming when we &lt;a href="https://dev.to/david_golverdingen_b133a5/we-replaced-jira-with-markdown-files-3065"&gt;replaced Jira with markdown files in our repositories&lt;/a&gt;. That post covers why we left and what the system is. This one is about what the system turned out to be for.&lt;/p&gt;

&lt;h2&gt;
  
  
  The skill is the product
&lt;/h2&gt;

&lt;p&gt;Our workflow lives in a Claude Code skill: 452 lines covering how a ticket is created, refined, planned, picked up, worked, reviewed and closed, plus a few reference files it pulls in on demand so the always-loaded part stays readable. Not a style guide. Instructions with gates in them, which the agent is required to load before touching any ticket.&lt;/p&gt;

&lt;p&gt;The create flow does not start by scaffolding a file. It starts by making the agent ask two to four targeted questions when a request is thin, then present the drafted Objective, Context and Acceptance Criteria for sign-off before anything is written. Planning runs in read-only mode, so no ticket write can happen before the plan is approved. Work happens one task at a time, each claimed on the board before a line of code is written.&lt;/p&gt;

&lt;h2&gt;
  
  
  The routing rule
&lt;/h2&gt;

&lt;p&gt;The most useful rule in it fits in a sentence. There are two ways to write a ticket, and &lt;strong&gt;the diff decides which one.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ticket-only changes&lt;/strong&gt; go through the MCP server, which commits straight to the default branch: creating a ticket, refining it, flipping status before the code work or after the pull request lands. No checkout, no branch, no pull request, no review bot. The analysis still happens locally where the code is; only the write is remote.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code work&lt;/strong&gt; happens on a ticket branch with the CLI and git, where the plan, status changes, task completions and work log entries ride along with the code commits and merge through the same pull request.&lt;/p&gt;

&lt;p&gt;Before the rule, every status change during a coding session was a small decision with no good answer: open a throwaway pull request for a one-line frontmatter edit, or commit ticket noise onto the feature branch. Now the diff answers it.&lt;/p&gt;

&lt;h2&gt;
  
  
  One implementation, four consumers
&lt;/h2&gt;

&lt;p&gt;Underneath both paths sits one npm package holding the schema, the parser, the serializer, and every edit operation as a pure function.&lt;/p&gt;

&lt;p&gt;Four things import it: the CLI that agents and humans run in a checkout, the Angular viewer behind every board edit, the Cloud Functions commit gateway, and the MCP server for the remote path. Nobody reimplements ticket semantics. The MCP server has no authoring logic of its own; it resolves the project, allocates an ID, runs the shared transform, and lets the gateway validate before committing.&lt;/p&gt;

&lt;p&gt;That is what makes a guardrail real. When we added a rule that a ticket cannot be closed while its acceptance criteria are unticked, it went into the package once and now fires identically whether you type a CLI command, call the MCP tool from your phone, or drag a card across the board. There is no back door, because there is only one door.&lt;/p&gt;

&lt;p&gt;The override is the part I like most. You can force past that gate, but it costs a written reason, and the same operation that flips the status appends that reason to the work log. A bypass is always visible and attributable. A gate people quietly route around reads as enforcement while providing none.&lt;/p&gt;

&lt;h2&gt;
  
  
  The loop
&lt;/h2&gt;

&lt;p&gt;The thing that changed daily work most is one line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/goal complete tasks according to the wbtickets skill until the PR is mergeable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the whole prompt. What follows runs on its own, and it can run on its own because the ticket carries enough state to make it possible:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pull, re-read the ticket, take the first task that is not done.&lt;/li&gt;
&lt;li&gt;Claim it: mark it in-progress, commit, push. Before any code, so the board and any parallel session see the claim.&lt;/li&gt;
&lt;li&gt;Implement exactly that task. Build it, test it.&lt;/li&gt;
&lt;li&gt;Present the pending diff and wait for approval.&lt;/li&gt;
&lt;li&gt;Complete the task, write the work log entry, commit code and bookkeeping together, push.&lt;/li&gt;
&lt;li&gt;Next task, back to step 2.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Step 4 is the default, and the goal is what waives it. A goal is a hook that blocks the session from ending until its condition holds, so typing that line is the approval: instead of stopping at each task, the agent finishes one, pushes, and walks straight into the next. The gate is not removed from the skill and it is not disabled for the repository. It is authorised once, for this run, by the person who started it. Point it at a ticket with seven tasks and it works the ticket, not the task.&lt;/p&gt;

&lt;p&gt;That makes it a judgement call rather than a setting I leave on. Waiving the per-task check is right when the tasks are small, the plan is well understood and a wrong turn costs one revert. It is wrong on anything where I would want to see the shape of task three before task four is built on top of it. The goal trades review granularity for throughput, and low-complexity, low-risk tickets are where that trade is clearly worth making.&lt;/p&gt;

&lt;p&gt;Two things make the unattended version acceptable rather than reckless.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The claim happens before the code.&lt;/strong&gt; It is bookkeeping with nothing attached and it must reach the board first, or a second session picks up the same task. The recurring failure was sliding from one task's finished commit into the next task's edits without claiming, which looks harmless until two agents are working in parallel.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One task is one commit.&lt;/strong&gt; The loop never batches. Every step lands as a separate, conventionally named commit carrying its code, its task completion and its work log entry together. An unattended run is reviewable afterwards, commit by commit, and the pull request reads as a sequence of decisions instead of one wall of diff.&lt;/p&gt;

&lt;p&gt;And the gates the goal cannot touch sit at the edges: an approved plan going in, and &lt;a href="https://dev.to/david_golverdingen_b133a5/the-cheapest-time-to-be-wrong-23c9"&gt;three layers of review&lt;/a&gt; coming out. Autonomy inside the ticket, scrutiny at the boundaries. Skipping the per-task check is defensible precisely because those two are not skippable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Planning ends by throwing the context away
&lt;/h2&gt;

&lt;p&gt;The shape this replaces is one I used for years. Read the ticket, make a plan from it, let the plan drive the code. That plan lived in the chat or in a scratch file, outside both the ticket and the repository, and nothing merged it back. When the work deviated, the plan stayed as written and the reasoning for the deviation went wherever the conversation went. Now the plan is a section of the ticket, so it travels with the code through the same pull request, and a deviation lands next to the plan it deviated from.&lt;/p&gt;

&lt;p&gt;Planning is the one phase that deliberately stops, and it stops twice: once for approval, then again at the very end. The session writes the plan into the ticket, splits the work into tasks, adds a Mermaid diagram when a flow or state machine is clearer drawn than described, commits, and then clears its own context.&lt;/p&gt;

&lt;p&gt;The rule that makes it work: if the reasoning exists only in the chat, the plan was not self-contained, and that is a bug to fix before clearing. The decisions, the rejected alternatives, the sequencing rationale and the per-task verification all go into the ticket.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The ticket is the context window.&lt;/strong&gt; A fresh session picks up task three of seven without any of the conversation that produced the plan. It still syncs the branch and re-reads the ticket, but it never has to reconstruct why the work is shaped the way it is, because that is in a file it can read. The exploration transcript is not context, it is exhaust.&lt;/p&gt;

&lt;p&gt;Getting that plan right is the highest-leverage step in the whole workflow, which is why it gets &lt;a href="https://dev.to/david_golverdingen_b133a5/the-cheapest-time-to-be-wrong-23c9"&gt;a review of its own before any code exists&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This is also why old tickets stay readable. They are time-locked decision records. Before planning work on a subsystem, the flow has the agent mine the history of the files it will touch and search older tickets on the same component, so a decision made in March is available in August without anyone remembering it exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  Distribution is the quiet multiplier
&lt;/h2&gt;

&lt;p&gt;A process written down helps one repository. A process that installs itself helps all of them.&lt;/p&gt;

&lt;p&gt;The skill ships inside the same package as the CLI, with a &lt;code&gt;sync&lt;/code&gt; command wired to a session-start hook, so every repository picks up the current version at the start of every session. Change the instructions and the change reaches every agent session in the company by the next session start. No migration, no announcement, no repository left on last month's process. The same package serves the guidance the MCP server hands to colleagues authoring tickets from claude.ai, so the two cannot drift apart.&lt;/p&gt;

&lt;p&gt;It is not one skill riding along either. Ten travel that channel now, from the review process to the per-stack rubrics, reaching a Java repository and an Angular one alike. That was not the plan. It is what happens when you build a way to ship one process and discover the pipe does not care how much you put in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watching four agents work
&lt;/h2&gt;

&lt;p&gt;One ticket, one checkout, one session. That is forced rather than chosen: the bookkeeping auto-commits, so two sessions cannot share a working copy without fighting over it. A &lt;code&gt;parallel&lt;/code&gt; command does the rest, provisioning a git worktree per ticket with its own port and dependencies, flagging any files two tickets would both touch, and handing back one kickoff prompt per session.&lt;/p&gt;

&lt;p&gt;That constraint produced the best moment I have had with this system.&lt;/p&gt;

&lt;p&gt;Three ingredients, none designed for it. Ticket bookkeeping auto-commits &lt;strong&gt;and pushes&lt;/strong&gt; the instant it happens, the one deliberate exception to our never-auto-commit rule, because the board has to see it. The board is branch-aware, reading each ticket from its own feature branch rather than the stale copy on the default branch, so it shows work in flight. And the &lt;code&gt;in-progress&lt;/code&gt; lane is double width with the task checklist inline.&lt;/p&gt;

&lt;p&gt;Run four sessions at once and the board becomes a live view: four tickets in progress, four checklists ticking themselves off within seconds of each push. I had it open on a second screen and watched the work move.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdavidgolverdingen.nl%2Fimages%2Fblog%2Fticket-is-the-context-window-board.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdavidgolverdingen.nl%2Fimages%2Fblog%2Fticket-is-the-context-window-board.jpg" title="The in-progress lane at double width, mid-run, four sessions pushing into it" alt="Kanban board with four tickets in the in-progress lane, each showing its own task checklist and branch"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The caveat is honest: the board is exactly as current as the last push, so unpushed work does not exist to it. In practice that is fine, because the bookkeeping pushes itself, and the bookkeeping is the part worth watching.&lt;/p&gt;

&lt;p&gt;I did not build this to be watchable. But it is the first time a project management tool has shown me something I did not already know from being the person doing the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this actually changed
&lt;/h2&gt;

&lt;p&gt;Our tickets are now the thing an agent reads before it writes code, instead of the thing someone updates after the code is written.&lt;/p&gt;

&lt;p&gt;The board did not get smarter. The tickets did.&lt;/p&gt;

&lt;p&gt;That leaves the uncomfortable question. If an agent writes most of the code, and on a good day you tell it not to pause between tasks, what stops it from confidently shipping the wrong thing? Three layers of review, and &lt;a href="https://dev.to/david_golverdingen_b133a5/the-cheapest-time-to-be-wrong-23c9"&gt;the one that does the most work runs before a single line of code exists&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>aiagents</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>We Replaced Jira With Markdown Files</title>
      <dc:creator>David Golverdingen</dc:creator>
      <pubDate>Thu, 13 Aug 2026 12:57:18 +0000</pubDate>
      <link>https://dev.to/david_golverdingen_b133a5/we-replaced-jira-with-markdown-files-3065</link>
      <guid>https://dev.to/david_golverdingen_b133a5/we-replaced-jira-with-markdown-files-3065</guid>
      <description>&lt;p&gt;Early this year I was wiring Claude into Jira through an MCP server. It worked, and every session it felt slightly wrong: slow round trips, a schema I did not control, structure sitting somewhere the agent could not see while it was reading the code.&lt;/p&gt;

&lt;p&gt;That is a context problem. Context engineering is not about pushing more context into the model; it is about putting the context where the agent is already looking. Ours could read every line of the codebase directly, and the ticket telling them what to change only through a tool call.&lt;/p&gt;

&lt;p&gt;The fix was almost embarrassingly simple. &lt;strong&gt;Put the ticket in the repo, as markdown.&lt;/strong&gt; I pitched it to a colleague, and off we went.&lt;/p&gt;

&lt;p&gt;Seven months later, 15 projects across 10 repositories, 165 live tickets, 12 people on the board including non-developers, and no Jira licence.&lt;/p&gt;

&lt;p&gt;This post is why we left and what we built. Two follow-ups cover the rest: &lt;a href="https://dev.to/david_golverdingen_b133a5/the-ticket-is-the-context-window-45ii"&gt;the skill and the loop that let agents work these tickets&lt;/a&gt;, and &lt;a href="https://dev.to/david_golverdingen_b133a5/the-cheapest-time-to-be-wrong-23c9"&gt;the three review layers that keep the output honest&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was actually wrong with Jira
&lt;/h2&gt;

&lt;p&gt;The cost was easy to name: roughly €2,000 a year for something we used maybe 5% of. It was not the reason we left.&lt;/p&gt;

&lt;p&gt;Every user had to be paid for, so the board was implicitly rationed. Performance degraded as projects grew. The features we wanted sat behind paid plugins. Automations were clumsy enough that we mostly did not write them. And the board was close to what we wanted without ever being it, because that last gap lived in someone else's product roadmap. None of that is fatal alone. Together it means the tool shapes the team instead of the other way around.&lt;/p&gt;

&lt;h2&gt;
  
  
  The constraint that ruled out the obvious answers
&lt;/h2&gt;

&lt;p&gt;We are one team maintaining ten separate repositories that ship independently of one another, across TypeScript, C#, Java and PowerShell. A monorepo was never realistic.&lt;/p&gt;

&lt;p&gt;That kills the usual alternatives. GitHub Issues comes closest and misses twice: issues are scoped to one repository, so cross-repo visibility becomes somebody's weekly spreadsheet, and despite feeling like part of the repo they are not in it. They live in a database behind an API. Not files, not on the branch, not in the diff, and not something an agent editing the code can read without a round trip. Every hosted alternative moves the work further away still.&lt;/p&gt;

&lt;p&gt;So we inverted it: &lt;strong&gt;distribute the tickets, centralise the view.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  One file per ticket
&lt;/h2&gt;

&lt;p&gt;Every repository carries a &lt;code&gt;tickets/&lt;/code&gt; folder. A ticket is one markdown file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;APP-42&lt;/span&gt;
&lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;story&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;[Profile]&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Add&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;avatar&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;upload'&lt;/span&gt;
&lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;in-progress&lt;/span&gt;
&lt;span class="na"&gt;priority&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;high&lt;/span&gt;
&lt;span class="na"&gt;size&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;m&lt;/span&gt;
&lt;span class="na"&gt;assignee&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;developer@example.com&lt;/span&gt;
&lt;span class="na"&gt;reported_by&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;colleague@example.com&lt;/span&gt;
&lt;span class="na"&gt;tasks&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Implementation plan&lt;/span&gt;
    &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;done&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Upload component&lt;/span&gt;
    &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;in-progress&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="c1"&gt;## Objective&lt;/span&gt;
&lt;span class="c1"&gt;## Context&lt;/span&gt;
&lt;span class="c1"&gt;## Acceptance Criteria&lt;/span&gt;
&lt;span class="c1"&gt;## Scope&lt;/span&gt;
&lt;span class="c1"&gt;## Design&lt;/span&gt;
&lt;span class="c1"&gt;## Technical Notes&lt;/span&gt;
&lt;span class="c1"&gt;## Implementation Plan&lt;/span&gt;
&lt;span class="c1"&gt;## Diagram&lt;/span&gt;
&lt;span class="c1"&gt;## Work Log&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The frontmatter is the machine-readable half, and none of it is free-form. Every field is schema-validated: the status vocabulary, priorities, t-shirt sizes, the ID matching its filename, timestamps in one format. A pre-commit hook and a CI step both run that validation, so a malformed ticket never reaches the default branch.&lt;/p&gt;

&lt;p&gt;The body is the human half, with fixed headings. That is what makes a section addressable: a tool can replace &lt;em&gt;Acceptance Criteria&lt;/em&gt; by name without touching anything around it, which you cannot do reliably to a free-form description field. &lt;code&gt;## Diagram&lt;/code&gt; holds Mermaid, and most substantial tickets have one, because a state machine or a data flow is faster to check than the paragraph describing it. The same pre-commit hook parses every Mermaid block with the real Mermaid parser, so a diagram that would not render cannot be committed.&lt;/p&gt;

&lt;p&gt;One rule holds it together: &lt;strong&gt;nobody edits these files by hand.&lt;/strong&gt; Writes go through a CLI of typed operations (set a status, claim a task, tick a criterion, append a work log entry, replace a section) that stamps timestamps, refuses invalid transitions, and keeps the file valid by construction. Hand-editing markdown is how fields drift and a board stops being trustworthy. The format is open to read and closed to casual writing.&lt;/p&gt;

&lt;p&gt;Each repository also holds a small derived &lt;code&gt;.tickets.config.json&lt;/code&gt; with its prefix, default branch and allowed values, so validation is local and calls nothing external. An admin edits project settings in the portal; the file is written out to every affected repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  The board is a projection
&lt;/h2&gt;

&lt;p&gt;A viewer application (Angular, Firebase, a GitHub App for commits) scans every configured repository and aggregates the tickets into one Kanban board. Edits commit back to the correct branch: the ticket's feature branch if one exists, the default branch otherwise. Because every change is a commit, git history &lt;em&gt;is&lt;/em&gt; the audit trail.&lt;/p&gt;

&lt;p&gt;The one piece of ticket &lt;em&gt;content&lt;/em&gt; not in git is board ordering, which lives in Firestore next to the things that were never content: profiles, saved filters, role lists, the project registry. Dragging a card to reprioritise would otherwise rewrite files across ten repositories and produce merge conflicts carrying no information. Content is git-canonical. Ordering is a view preference.&lt;/p&gt;

&lt;p&gt;Two decisions I would keep in any rebuild. &lt;strong&gt;The lanes are asymmetric:&lt;/strong&gt; &lt;code&gt;in-progress&lt;/code&gt; is twice the width of the others and shows the task checklist and the live branch inline, which matters far more than I expected once several tickets run at once. &lt;strong&gt;Workflow steps are subtasks, not columns:&lt;/strong&gt; code review, manual test and E2E are entries in the ticket's task list. For a team of two to five, a column per step produces a board of nearly-empty columns, while task-level detail puts the standup answer on the card.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nobody pays to be on the board
&lt;/h2&gt;

&lt;p&gt;We already had Microsoft Entra SSO, so authentication cost nothing and no seat has a price on it. That reads like a footnote, and it is the change I value most: when visibility stops being metered, you stop deciding who deserves it.&lt;/p&gt;

&lt;p&gt;It also took the system somewhere I did not plan. Five of the fifteen projects have no repository at all. They are folders in a shared host repo covering application management, business automation, engineering optimisation, AI enablement and general company work. What is in them now: onboarding a new colleague, enabling SSO, formalising a technical quick-scan process, getting the team through a Claude course. None of it is code, all of it sits on the same board, next to the software work it competes with for time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two ways in, one authoring guide
&lt;/h2&gt;

&lt;p&gt;WBTickets has its own MCP server alongside the eleven others we run: 18 tools for reads, writes, milestones and attachments, hitting the same repositories through the same commit gateway as the board.&lt;/p&gt;

&lt;p&gt;That is how the projects without a codebase work. A colleague in claude.ai, with no checkout and no git, describes a problem in their own words and ends up with a schema-valid ticket committed to a repository. &lt;code&gt;create_ticket&lt;/code&gt; is deliberately two-step: called with no arguments it creates nothing and returns the authoring guide, which tells the agent what to ask and to get sign-off before writing. That last part is instruction rather than enforcement. The server will happily create a ticket on a first call that arrives with arguments; what the two-step buys is that the guide is in front of the agent before it has anything to commit.&lt;/p&gt;

&lt;p&gt;For code repositories we prefer the other route, running on that same guide. A ticket for a real codebase is written from a checkout by an agent that reads the code first: what exists, what the change touches, which older ticket decided the design that is there now. That is the difference between a Scope with real boundaries and one that merely sounds plausible. The write still lands on the default branch through the same path; only the analysis is local.&lt;/p&gt;

&lt;p&gt;Attachments get their own trick. Screenshots are how non-developers explain a bug, but pushing image bytes through a model's context is pure waste. So &lt;code&gt;view_attachments&lt;/code&gt; returns metadata only and opens a small app inside the chat; the app fetches and uploads the actual bytes through separate tools that never put them in the model's context at all. You see the screenshot, drag a new one in, and the agent's context stays clean.&lt;/p&gt;

&lt;p&gt;The loop closes at the far end: &lt;code&gt;reported_by&lt;/code&gt; records who asked, and they are notified when the ticket reaches &lt;code&gt;done&lt;/code&gt;. They never have to open the board.&lt;/p&gt;

&lt;h2&gt;
  
  
  The board matches the work that was delivered
&lt;/h2&gt;

&lt;p&gt;I have never seen a ticket board outside the codebase that still matched what was actually built. Not a failure of any particular team: it is distance. Scope changes, deviations and the reasoning behind them get settled in a pull request or a call, and updating the ticket is a separate action somebody has to remember to take, later, from memory. Some of it lands days late as a summary. Most of it never lands at all, which is why the interesting question about any board is not what it says but when it last agreed with reality.&lt;/p&gt;

&lt;p&gt;Here the agent doing the work is what writes it down, in the same commit as the code. A deviation gets a work log entry when it happens, with the reason, while the reason is still in context rather than reconstructed afterwards. So by the time the pull request opens, the ticket describes what was built and not only what was planned, and where the two disagree that disagreement is on the record with its argument attached.&lt;/p&gt;

&lt;p&gt;That is the part I would miss most if we went back. A board that is merely tidy tells you how disciplined the team is about bookkeeping. A board that is written by the work tells you what happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  What came nearly free afterwards
&lt;/h2&gt;

&lt;p&gt;Once the substrate was markdown in git, the adjacent things stopped being projects. Milestones as roadmap items with horizons instead of dates. Release notes generated from git tags with ticket IDs resolved to board links. Design artifacts: a self-contained HTML mockup committed beside the ticket, reviewed in the same pull request as the code and rendered in a sandboxed iframe on the board.&lt;/p&gt;

&lt;p&gt;The next one is in build: a knowledge base to replace Confluence and its ~300 articles. What made it worth starting is that almost none of it is new. It inherits the auth, the attachments, the markdown rendering and the same git-canonical instinct, so the architecture argument was about content workflow rather than infrastructure.&lt;/p&gt;

&lt;p&gt;Same primitive every time. That is the compounding return of a format every tool already understands.&lt;/p&gt;

&lt;h2&gt;
  
  
  The build-versus-buy line moved
&lt;/h2&gt;

&lt;p&gt;We are a small in-house IT department. Warmtebouw installs heating, cooling and ventilation; we are the handful of people who build the software around that. This is not a heroic story, and that is the part worth noticing.&lt;/p&gt;

&lt;p&gt;You rarely bought a tool because its feature list was unbeatable. You bought it because building and maintaining something that fit your team cost far more than the licence did. That arithmetic has changed. Writing a bespoke system, and more importantly keeping it correct and documented while you use it, is now cheap enough that a small team can do it alongside the actual work. We are one of many teams quietly discovering that the licence was buying convenience we can now produce ourselves, and better, because ours fits.&lt;/p&gt;

&lt;p&gt;It does not generalise to everything. The tools worth replacing are the ones where your own workflow is the product and the vendor's genericity is the tax you pay. Nobody sane is rebuilding their ERP this way.&lt;/p&gt;

&lt;p&gt;The vendors have noticed. Atlassian is now positioning itself as the layer in the middle of agentic work, and I would not want to be selling that. The middle is precisely the position an agent routes around. If the work lives in the repository and the agent is already there, a hosted system of record stops being where the work happens and becomes a place you synchronise to. That is a difficult thing to charge per seat for, and a harder one to defend once teams notice the alternative is a folder of markdown files.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest ledger
&lt;/h2&gt;

&lt;p&gt;You own the maintenance. Jira's operating burden is zero and ours is not. This began as a side project and grew in the gaps between real work. The core worked from the moment there were tickets on the board; everything since has been sharpening it.&lt;/p&gt;

&lt;p&gt;What makes that sustainable is who maintains it: the team that uses it every day. A rough edge gets filed down by the person it just annoyed, usually in the same session they hit it, and the fix is filed as a ticket in the system it fixes. In a week where we are not adding features, that upkeep is roughly an hour. There is no maintenance budget to defend, because the work is spread thinly across all the other work.&lt;/p&gt;

&lt;p&gt;What we bought is a system that fits our workflow exactly rather than 60% of it, a board open to everyone in the company, and every remaining failure mode being ours to fix. The licence saving is real and it is the least interesting part.&lt;/p&gt;

&lt;p&gt;The board was the deliverable. The consequence was that the description of the work now lives where the agent is already looking, and that changed how the work gets done far more than any board could. &lt;a href="https://dev.to/david_golverdingen_b133a5/the-ticket-is-the-context-window-45ii"&gt;Here is what that looks like in practice.&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Enterprise AI Without an Enterprise Budget</title>
      <dc:creator>David Golverdingen</dc:creator>
      <pubDate>Tue, 26 May 2026 07:27:06 +0000</pubDate>
      <link>https://dev.to/david_golverdingen_b133a5/enterprise-ai-without-an-enterprise-budget-2ja8</link>
      <guid>https://dev.to/david_golverdingen_b133a5/enterprise-ai-without-an-enterprise-budget-2ja8</guid>
      <description>&lt;p&gt;Enterprise AI gets sold as something only large companies can afford. It doesn't have to be. The reason is structural: the protocol layer underneath is absorbing the work that used to require an AI platform.&lt;/p&gt;

&lt;p&gt;Over the first three months at one mid-sized engineering firm (not a software company), I operationalised AI across the entire business through nine production MCP servers covering ERP, BIM, calculations, building automation, energy, and operational logs. By the end of August, the estate had grown to twelve. No platform team, no framework dependencies, no custom chat UI. Project managers, field engineers, and operations colleagues query the whole company's data in natural language, every day. The running cost is roughly €19/seat/month for the staff who use it, plus the engineering time to build the MCP layer.&lt;/p&gt;

&lt;p&gt;This piece is about the architecture that made that reachable, and why it's reachable for small and mid-sized companies that have been told enterprise AI is out of their league.&lt;/p&gt;

&lt;h2&gt;
  
  
  The standard path is expensive
&lt;/h2&gt;

&lt;p&gt;When a mid-market or smaller company decides to bring AI in, the path they're usually pointed toward looks the same: build a branded chat UI on the API, wire it to internal auth, manage prompts in-house, maybe add a RAG pipeline against company documents, optionally an orchestration framework for &lt;em&gt;"agent workflows,"&lt;/em&gt; and increasingly an observability platform to monitor it all. That path is real, and at sufficient scale it may be the right call. I haven't run this architecture at the scale of a large multinational with hundreds of heterogeneous systems and complex tenancy requirements, so I can't tell you whether the same posture holds there. My guess is that the MCP layer itself stretches further than the framework industry assumes (more servers, deeper hierarchies of tools, more careful schemas) rather than needing a different architecture entirely. But that's a hypothesis from one scale, not a report from another.&lt;/p&gt;

&lt;p&gt;For a company with twenty important systems and a few hundred employees, the standard path doesn't pay back. It's expensive in three coupled ways, and pulling the three apart reveals a simpler architecture underneath.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The chat client.&lt;/strong&gt; A branded chat UI is a frontend team, a design effort, conversation history infrastructure, attachment handling, multi-modal input support, an admin panel, model routing, prompt management. The vendor (Anthropic, OpenAI, Google) ships all of this as part of the subscription and continues to ship new affordances every quarter. Building it in-house means investing engineering hours into a commodity layer where the vendor has structural advantages no internal team can match.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The model.&lt;/strong&gt; A custom chat UI almost always pins to a specific model version for stability. Six months later the frontier moves, and the pinned model is now a generation behind. Upgrading means re-validating every prompt and every tool, so most teams don't. Meanwhile every Claude Team subscriber got Opus 4.7 the morning it shipped, with zero engineering work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The billing.&lt;/strong&gt; API billing is per-token, which means the cost is a function of user behaviour that hasn't happened yet. In a workforce of 50-500 people, roughly 20% of users drive 80% of consumption once adoption stabilises. Published analyses of heavy usage put the API-vs-subscription cost ratio at 15-30x; for average users the multiple is smaller (3-10x), and for light users API can come out ahead. The relevant number isn't the average. It's the variance.&lt;/p&gt;

&lt;p&gt;These three costs aren't independent. They flow from the same root decision: build our own chat UI, or use the vendor's. Building locks all three together. The standard path commits a smaller company to a build budget, a maintenance team, an aging model, and an unpredictable bill, all at once.&lt;/p&gt;

&lt;h2&gt;
  
  
  The simpler path
&lt;/h2&gt;

&lt;p&gt;Use the vendor's chat client. Pay per seat. Connect your existing identity provider. Then put your engineering hours where the value compounds: in MCP servers that &lt;a href="https://davidgolverdingen.nl/en/insights/97-percent-mcp-tool-descriptions-broken" rel="noopener noreferrer"&gt;encode your domain&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This is the architecture in production. Employees use Claude through the standard client (web, desktop, mobile). The client authenticates against the corporate identity provider, same as every other corporate system. By the end of August, employees had access to twelve MCP servers covering the full operational chain. Each MCP tool is RBAC-gated against the same IdP roles that govern every other system. A field engineer sees their own time bookings; a controller sees aggregated financials; a guest user sees nothing.&lt;/p&gt;

&lt;p&gt;What I didn't build: a chat UI. A model gateway. A prompt management platform. A vector database. A retrieval pipeline. An agent observability stack. An &lt;a href="https://dev.to/david_golverdingen_b133a5/mcp-is-the-ai-platform-2koo"&gt;&lt;em&gt;"AI platform"&lt;/em&gt;&lt;/a&gt; of any kind. None of those layers exist in the stack, because the subscription client provides everything above the MCP layer, and the IdP provides everything around it.&lt;/p&gt;

&lt;p&gt;A concrete example. The most common AI project a mid-sized company is pitched is &lt;em&gt;"RAG over our documents":&lt;/em&gt; chunk all the SharePoint or Google Drive content, embed it, build a vector store, wire it to a retrieval layer, host and maintain the whole pipeline. Even when that pipeline is cheap to build, it's a layer you have to keep alive. Re-index when documents change, re-tune when retrieval quality drops, re-permission when access rules shift, re-host when the embedding model is deprecated. Meanwhile, the Microsoft 365 and Google Workspace integrations that ship with Claude, ChatGPT, and Gemini are themselves MCP servers, published by the vendors, maintained by the vendors, with permissions inherited from the existing IdP and freshness handled upstream. The &lt;em&gt;"document search"&lt;/em&gt; capability that makes a custom RAG pipeline sound necessary is already an MCP server you can turn on in your admin console. The choice isn't cheap vs. expensive. It's an extra layer you maintain vs. no extra layer at all. And once you stop blaming the data and &lt;a href="https://davidgolverdingen.nl/en/insights/your-data-is-fine" rel="noopener noreferrer"&gt;start fixing the meaning layer&lt;/a&gt;, the case for a custom RAG pipeline gets thinner still.&lt;/p&gt;

&lt;p&gt;That sharpens the architectural rule. MCP isn't a category that lives only inside your perimeter; it's the protocol the entire ecosystem speaks, and vendors are already publishing servers for the commodity layer: productivity suites, code hosts, ticketing systems, design tools. Your engineering investment goes into the MCP servers that &lt;em&gt;only you&lt;/em&gt; can write: your ERP, your BIM data, your operational logs, your calculation history. The systems unique to your business that no vendor has, or will ever have, a connector for. Everything else, you consume. That's where the investment compounds, and where it doesn't.&lt;/p&gt;

&lt;p&gt;The cost picture inverts. Claude Team is about €19/seat/month on the annual plan; ChatGPT Enterprise and Gemini for Workspace are in similar territory. The vendor eats the billing variance. Every seat gets the latest frontier model the day it ships. No frontend to maintain, no platform team to staff, no orchestration platform to buy.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this opens up
&lt;/h2&gt;

&lt;p&gt;The interesting part of this architecture, and the part that makes it reachable for small companies, is the shape of &lt;a href="https://dev.to/david_golverdingen_b133a5/six-levels-of-mcp-servers-2b25"&gt;the MCP layer&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The domain layer is model-agnostic.&lt;/strong&gt; MCP servers are typed interfaces with tool descriptions and schemas. The same servers work with Claude today, Gemini tomorrow, GPT next quarter. None of the domain logic is locked to a model vendor. The engineering investment is portable across the entire frontier-model market.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Identity lives at the tool boundary.&lt;/strong&gt; RBAC is enforced inside the MCP server, against your existing IdP. A model that's been prompt-injected can only call tools the authenticated user is already authorised to call. The security model is the same one your company already runs for every other system. No AI-specific identity layer, no prompt firewall, no model gateway. The tool boundary is the trust boundary.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The engineering shape is small.&lt;/strong&gt; An MCP server, in my experience, is one engineer working with one domain expert for a few weeks per domain. That's the whole staffing model. Nine production servers in the first three months, twelve by the end of August. No platform team, no specialists. The people who own the underlying business systems can do most of the work themselves, with engineering support.&lt;/p&gt;

&lt;p&gt;The servers aren't free once they exist. Schemas drift, a source-system upgrade breaks a connector, tool descriptions go stale as the business changes. That is the same kind of upkeep I just used as the argument against a custom RAG pipeline, and it applies here too. Across twelve servers it has run at roughly an hour a week. What makes it affordable is where it lands: on the people who already own those systems, in the weeks they are already working on them.&lt;/p&gt;

&lt;p&gt;This is what makes the approach reachable. A small or mid-sized company doesn't need to hire an AI platform team to run this stack. The chat client is rented from a vendor at a price comparable to a productivity-suite seat. The MCP layer is built incrementally by the engineers and domain experts already on the payroll.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you rent, what you own
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What you rent&lt;/th&gt;
&lt;th&gt;What you own&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;The chat client (Claude, ChatGPT, Gemini, your choice)&lt;/td&gt;
&lt;td&gt;The MCP servers that encode your domain&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Conversation history, attachments, multi-modal UI&lt;/td&gt;
&lt;td&gt;Tool descriptions, query strategies, business logic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Enterprise SSO, audit logs, retention policies&lt;/td&gt;
&lt;td&gt;Identity-gated tool access through your existing IdP&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frontend updates, model upgrades, security patches&lt;/td&gt;
&lt;td&gt;Domain knowledge written into schemas&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The model itself, latest version on day one&lt;/td&gt;
&lt;td&gt;The integration with ERP, BIM, energy, calc&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Predictable per-seat billing&lt;/td&gt;
&lt;td&gt;One engineer, one domain expert, per server&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;What you rent is the commodity layer: the stuff vendors compete on and ship continuously. What you own is the part no vendor can build for you, because no vendor knows what your data means.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same pattern, one floor down
&lt;/h2&gt;

&lt;p&gt;The reason this is reachable goes one layer deeper than the chat client. The same logic applies to the framework layer underneath: LangChain, LangGraph, CrewAI, RAG pipelines, vector DBs, agent observability stacks. Building on top of those is one shape of architecture; building MCP servers directly against a frontier model is the same architectural posture without the additional layer.&lt;/p&gt;

&lt;p&gt;Each MCP spec release closes another category of problem that previously required a framework layer. Tool calling, resource management, prompts, sampling, elicitation, UI primitives via MCP Apps, and enterprise identity integration on the 2026 roadmap. Each one used to live in a framework above MCP, and each one is now in the protocol itself. The framework layer isn't being argued against; it's being absorbed. Companies building on top of frameworks today are building on top of a layer the protocol is in the process of swallowing.&lt;/p&gt;

&lt;p&gt;Both layers reward the same answer: rent the surface, own the domain. Use the vendor's client. Use the vendor's identity integrations. Use the vendor's model upgrades. Skip the framework layer. Spend your engineering on the part that's actually yours: the MCP servers that turn your operational data into something an agent can reason about.&lt;/p&gt;

&lt;p&gt;That's what &lt;em&gt;"MCP is the platform"&lt;/em&gt; actually looks like in practice: a perimeter you draw. Inside the perimeter: your domain, your tools, your IdP, your data. Outside: the model, the client, the vendor. The line between them is MCP.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you're starting
&lt;/h2&gt;

&lt;p&gt;Subscribe to Claude Team, ChatGPT Enterprise, or Gemini for Workspace. Connect your existing identity provider. Then pick the one system whose data your colleagues most want to query in natural language, and write one MCP server against it. Ship that. Watch them use it. Build the next one. The &lt;a href="https://davidgolverdingen.nl/en/insights/production-mcp-practitioners-guide" rel="noopener noreferrer"&gt;practitioner's guide&lt;/a&gt; walks the full seven-step recipe.&lt;/p&gt;

&lt;p&gt;That's the entire starting move. No vendor selection process for an AI platform. No headcount plan for a platform team. No procurement cycle for orchestration software. A subscription, an identity wire-up, and one MCP server is enough to be in production with real users by the end of a month.&lt;/p&gt;

&lt;p&gt;Three months of that across a real business produced nine production servers, and steady additions brought the estate to twelve by the end of August, with no framework dependencies, no custom chat UI, no API bill, no platform team, and a company that gets every frontier model upgrade for free the day it ships.&lt;/p&gt;

&lt;p&gt;This architecture isn't a clever workaround for the moment. It's an early version of what enterprise AI looks like once the protocol layer finishes absorbing the platform layer above it.&lt;/p&gt;

&lt;p&gt;Enterprise AI doesn't require an enterprise budget. It requires picking the right perimeter to draw.&lt;/p&gt;




&lt;p&gt;Go deeper: read the full practitioner report, &lt;a href="https://davidgolverdingen.nl/en/the-missing-layer" rel="noopener noreferrer"&gt;&lt;em&gt;The Missing Layer&lt;/em&gt;&lt;/a&gt;, or explore the &lt;a href="https://github.com/DaveGold/mcp-metadata-demo" rel="noopener noreferrer"&gt;mcp-metadata-demo&lt;/a&gt; server, an open-source extract of these patterns, since the production servers run on private business data.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>architecture</category>
      <category>enterprise</category>
      <category>business</category>
    </item>
    <item>
      <title>MCP Is the AI Platform</title>
      <dc:creator>David Golverdingen</dc:creator>
      <pubDate>Tue, 26 May 2026 07:11:26 +0000</pubDate>
      <link>https://dev.to/david_golverdingen_b133a5/mcp-is-the-ai-platform-2koo</link>
      <guid>https://dev.to/david_golverdingen_b133a5/mcp-is-the-ai-platform-2koo</guid>
      <description>&lt;p&gt;Stop building around the model. Build for MCP.&lt;/p&gt;

&lt;p&gt;By the end of August, the production estate at one mid-sized engineering firm had reached twelve MCP servers and 97 tools. Bound to Claude in our case, though everything below applies equally to GPT, Gemini, or any frontier model that speaks MCP, here's what I didn't build: agent frameworks, RAG pipelines, vector databases, orchestration platforms, or any of the other layers the AI industry insists you need.&lt;/p&gt;

&lt;p&gt;What I have instead: a frontier model, MCP servers, and well-written tool descriptions. That's it. And it works, daily, across the whole business. &lt;em&gt;Updated September 2026:&lt;/em&gt; "measurably" now has a public measurement behind it, on a reference server anyone can run, with one caveat I learned there: "well-written" includes short, because only the first 2,048 characters of a description reach the model on Claude Code (&lt;a href="https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery" rel="noopener noreferrer"&gt;evidence&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;By 2026 MCP isn't a niche bet. It crossed roughly 97 million monthly SDK downloads, and OpenAI, Google, Microsoft, and Anthropic have all integrated it across their products. The protocol question is settled. &lt;strong&gt;And because every frontier model now consumes MCP through the same interface, the model question is settled too.&lt;/strong&gt; Pick any of them and the architecture doesn't change. What's still being argued is whether you need anything else on top.&lt;/p&gt;

&lt;h2&gt;
  
  
  The model is the agent. The framework is overhead.
&lt;/h2&gt;

&lt;p&gt;The framing &lt;em&gt;"you need an agent framework"&lt;/em&gt; obscures a simpler truth: the model &lt;em&gt;is&lt;/em&gt; the agent. It reads tool descriptions. It chooses which tool to call. It sequences the calls. It interprets the results. That's textbook agentic behaviour, built into every frontier model that speaks MCP.&lt;/p&gt;

&lt;p&gt;What companies sell you on top of that is &lt;em&gt;framework around the agent&lt;/em&gt;: orchestration logic, retrieval pipelines, prompt managers, observability layers. Each of those products is solving a real problem in some context. But in a mid-sized business with a knowable set of important data sources, those contexts mostly don't apply.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What I use&lt;/th&gt;
&lt;th&gt;What I skip&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A frontier model (Claude in my case; swap as needed)&lt;/td&gt;
&lt;td&gt;LangChain + LangGraph + LangSmith stack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MCP servers (custom, well-typed)&lt;/td&gt;
&lt;td&gt;Multi-agent orchestration (CrewAI, Microsoft Agent Framework, OpenAI Agents SDK)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool descriptions as the interface&lt;/td&gt;
&lt;td&gt;Vector DBs, embedding pipelines, agentic-retrieval frameworks (LlamaIndex)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Domain knowledge written into schemas&lt;/td&gt;
&lt;td&gt;Hand-curated knowledge graphs sitting next to the tool&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cross-source composition through tool design&lt;/td&gt;
&lt;td&gt;Workflow orchestration platforms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Production telemetry as the feedback loop&lt;/td&gt;
&lt;td&gt;Agent observability stacks (LangSmith, LangFuse, Arize Phoenix)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each row on the right is a paid product category. Each row on the left is the model, the protocol, or work I did once and reuse.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "build for MCP" looks like in practice
&lt;/h2&gt;

&lt;p&gt;A well-designed MCP server doesn't ask the model to &lt;em&gt;figure out&lt;/em&gt; what the data means. It tells the model what the data means &lt;a href="https://dev.to/david_golverdingen_b133a5/six-levels-of-mcp-servers-2b25"&gt;in the tool description itself&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The contrast plays out at every level. A bad tool says: &lt;em&gt;"query data from the ERP."&lt;/em&gt; A good tool says: &lt;em&gt;"always start with &lt;code&gt;summaryOnly=true&lt;/code&gt;. Active projects accumulate thousands of records. Type codes determine which fields are populated. Use &lt;code&gt;get_budget&lt;/code&gt; for planned costs, this tool for actuals."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The first version forces the model to invent a query strategy on every call. The second hands the model a query strategy on every call. The difference between those two servers, in production, is whether business users actually use the result.&lt;/p&gt;

&lt;p&gt;This is the work teams skip when they reach for frameworks. The complexity is mostly self-inflicted. It exists because nobody wrote down what the tools mean. The same gap is why &lt;a href="https://davidgolverdingen.nl/en/insights/97-percent-mcp-tool-descriptions-broken" rel="noopener noreferrer"&gt;97% of the MCP tool descriptions in a study of public servers contain at least one smell&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  One server can compose many sources
&lt;/h2&gt;

&lt;p&gt;The most common objection (&lt;em&gt;"but you need orchestration to combine data from multiple systems"&lt;/em&gt;) assumes orchestration has to live in a separate framework. It doesn't.&lt;/p&gt;

&lt;p&gt;One of my MCP servers combines five heterogeneous sources behind one interface: a third-party meter-data aggregator, a public weather API, a government building registry, the company ERP, and an IoT building-automation platform. No agent dance. No retrieval pipeline. Just typed tools with descriptions explaining when each source applies and how they relate.&lt;/p&gt;

&lt;p&gt;The model figures out the composition because the tool descriptions tell it the relationships. The pattern works across eleven production servers covering ERP, BIM, calculations, building automation, energy, and operational logs. Effectively the entire operational surface of one business, addressable through tool descriptions, consumable by any MCP-speaking model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tools abstract everything below them
&lt;/h2&gt;

&lt;p&gt;The model never sees how the data is fetched. Inside one MCP server, individual tools call whatever the underlying system speaks: REST for SaaS products, GraphQL for internal APIs, direct SQL against the data warehouse, JSON files on disk, SOAP envelopes for legacy systems. The tool returns typed records; the protocol heterogeneity stays inside the tool.&lt;/p&gt;

&lt;p&gt;That's the abstraction layer the data-fabric and iPaaS industries charge premium prices to build. MCP tools already do it, one tool at a time, in whatever language the data actually lives in, with no central pipeline. The choice of backend doesn't propagate; pick whatever fits the underlying system, and the model interface stays identical.&lt;/p&gt;

&lt;p&gt;Even SQL becomes tractable. Direct SQL from an LLM is dangerous because the model can be tricked into destructive queries. But SQL inside a typed MCP tool (where the server builds SQL from schema-validated parameters, the connection runs as a read-only role, and the corporate IdP gates who can call it) is just a regular function call. Any value that doesn't match the tool's JSON schema is rejected at the protocol boundary before the tool's code runs at all.&lt;/p&gt;

&lt;p&gt;This isn't theoretical. One of my servers is 100% SQL behind a query endpoint, covering calculation data. Every tool translates a typed agent request into a SQL query: the model passes schema-validated parameters in, the server constructs and runs the query, typed records come back out. The model never sees SQL. The connection runs as a read-only role, and access is gated by the same IdP roles that govern every other system.&lt;/p&gt;

&lt;p&gt;I built it in one day. The calculation expert who owns the underlying data validated it, and it already exposes every dataset that team needs. Once you've internalised the pattern, applying it to a new domain is a day's work, not a quarter's project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why mid-market is the sweet spot
&lt;/h2&gt;

&lt;p&gt;This argument has bounds. At Fortune 500 scale (hundreds of heterogeneous systems, multi-tenant SaaS, mountains of unstructured documents) you might need retrieval pipelines and orchestration. The complexity is real because the scope is real.&lt;/p&gt;

&lt;p&gt;But mid-sized businesses have something Fortune 500 doesn't: a knowable set of important data sources. Five to twenty key systems. A handful of domain experts who can sit with you for an afternoon and tell you what the data actually means. That's the entire prerequisite for a well-built MCP server.&lt;/p&gt;

&lt;p&gt;If your company fits in one office building and has fewer than twenty important systems, you're who this is for. You almost certainly don't need most of what the AI industry is trying to sell you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "the platform" actually is
&lt;/h2&gt;

&lt;p&gt;Notice what's missing from the table above: a platform. There's no "AI platform" in the stack. Just a frontier model, MCP servers, and the corporate identity layer the business already pays for. Swap the model (Claude today, Gemini tomorrow, GPT next quarter) and the rest of the stack stays identical.&lt;/p&gt;

&lt;p&gt;That last piece is what turns an estate of MCP servers into something an enterprise can run. Not an AI-specific identity layer. The one the company already operates. In a Microsoft shop, that's Entra ID with RBAC. In other shops, Okta, Google Cloud Identity, or any OAuth 2.1 provider. MCP servers authenticate against the existing IdP, scope tool access by role, and log every call to the same audit trail every other corporate system already uses.&lt;/p&gt;

&lt;p&gt;The implications:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A field engineer can call the project and time-booking tools, and not the HR or payroll ones.&lt;/li&gt;
&lt;li&gt;A controller sees aggregated financials, not raw payroll.&lt;/li&gt;
&lt;li&gt;A guest user sees nothing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every action is attributable to a named identity in the same audit log as every other system. That's the entire enterprise AI security model. No prompt-firewall vendor. No AI-specific governance platform. No model gateway. Just the access-control infrastructure the company already runs, enforced at the MCP-tool boundary. Even prompt injection becomes bounded. A tricked model can only call tools the authenticated user is already authorised to call.&lt;/p&gt;

&lt;p&gt;Anthropic's 2026 MCP roadmap leads with enterprise authentication and identity-provider integration. The protocol is moving in this direction because the pattern works: tool-level RBAC against the corporate IdP turns &lt;em&gt;"the AI security problem"&lt;/em&gt; into a solved authentication problem. Which it always was.&lt;/p&gt;

&lt;p&gt;Wire MCP servers through enterprise identity and MCP isn't &lt;em&gt;connected to&lt;/em&gt; the platform. MCP &lt;em&gt;is&lt;/em&gt; the platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  The work that matters
&lt;/h2&gt;

&lt;p&gt;What the framework industry sells is scaffolding. What actually moves outcomes is the specific, domain-bound work of writing good tool descriptions: choosing the right granularity, documenting which tool fits which question, adding query strategies, capturing failure modes from production usage and feeding them back.&lt;/p&gt;

&lt;p&gt;None of that can be outsourced to a vendor, because &lt;a href="https://davidgolverdingen.nl/en/insights/your-data-is-fine" rel="noopener noreferrer"&gt;nobody outside your business knows what your data means&lt;/a&gt;. But all of it is reachable in a few weeks per domain, with one engineer and one domain expert. That's the trade you're being told doesn't exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question that replaces "which framework should I pick?"
&lt;/h2&gt;

&lt;p&gt;If you're starting MCP work today, the most useful question isn't &lt;em&gt;"which framework do I need?"&lt;/em&gt; It's &lt;em&gt;"what's the smallest useful tool I can write for the one person whose week would get better tomorrow if it worked?"&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Build that. Watch them use it. Write down what they tried that didn't work. Update the tool description. Ship the next one.&lt;/p&gt;

&lt;p&gt;Three months of that across a real business produced nine production servers (eleven since) and zero framework dependencies. That's the story: not that frameworks are bad, but that for most mid-sized companies, they're the wrong problem to be solving first.&lt;/p&gt;

&lt;p&gt;The model is the agent. The IdP is the security boundary. MCP is the platform. Build for the platform, not around it.&lt;/p&gt;




&lt;p&gt;This post is one piece of a longer argument. &lt;a href="https://davidgolverdingen.nl/en/insights/production-mcp-practitioners-guide" rel="noopener noreferrer"&gt;&lt;em&gt;Production MCP: A Practitioner's Guide&lt;/em&gt;&lt;/a&gt; puts all of them in order, from understanding your data through to identity-bound deployment.&lt;/p&gt;

&lt;p&gt;Go deeper: read the full practitioner report, &lt;a href="https://davidgolverdingen.nl/en/the-missing-layer" rel="noopener noreferrer"&gt;&lt;em&gt;The Missing Layer&lt;/em&gt;&lt;/a&gt;, or explore the &lt;a href="https://github.com/DaveGold/mcp-metadata-demo" rel="noopener noreferrer"&gt;mcp-metadata-demo&lt;/a&gt; server, an open-source extract of these patterns, since the production servers run on private business data.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>architecture</category>
      <category>ai</category>
      <category>aiengineering</category>
    </item>
    <item>
      <title>The Six Levels of MCP Servers</title>
      <dc:creator>David Golverdingen</dc:creator>
      <pubDate>Mon, 25 May 2026 20:25:32 +0000</pubDate>
      <link>https://dev.to/david_golverdingen_b133a5/six-levels-of-mcp-servers-2b25</link>
      <guid>https://dev.to/david_golverdingen_b133a5/six-levels-of-mcp-servers-2b25</guid>
      <description>&lt;p&gt;Most MCP servers do the same thing: wrap an API, expose tools with one-sentence descriptions, and hope the model figures out the rest. In &lt;a href="https://davidgolverdingen.nl/en/insights/your-data-is-fine" rel="noopener noreferrer"&gt;a previous post&lt;/a&gt; I described why this can fail for enterprise data. Here's the maturity ladder that emerged from an eleven-server, 90+ tool production snapshot across eleven APIs.&lt;/p&gt;

&lt;p&gt;I call it the &lt;strong&gt;six-level MCP maturity model&lt;/strong&gt;. It classifies an MCP server by how much of the domain the server itself carries, on a ladder from Level 1 API Mappers, which expose bare endpoints and leave the model guessing, to Level 6 Secure Write Apps, which act on the system of record. I derived the six levels in 2026 from eleven production servers running at a Dutch HVAC installation company.&lt;/p&gt;

&lt;p&gt;Two different samples are at work below, and it is worth keeping them apart. The &lt;em&gt;ladder&lt;/em&gt; comes from that eleven-server snapshot. The &lt;em&gt;share&lt;/em&gt; column does not: those percentages come from reviewing 50+ publicly available MCP servers in March 2026, so they describe the public ecosystem, not my own estate, and they're a snapshot, not a tracked metric. Eleven servers could not tell you what 10,000 are doing. &lt;a href="https://davidgolverdingen.nl/en/the-missing-layer" rel="noopener noreferrer"&gt;The Missing Layer&lt;/a&gt; sets out that review in full, with named examples at every tier.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;Name&lt;/th&gt;
&lt;th&gt;Share of servers&lt;/th&gt;
&lt;th&gt;What the server carries&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;API Mapper&lt;/td&gt;
&lt;td&gt;~70%&lt;/td&gt;
&lt;td&gt;Endpoint names and nothing else&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Functional&lt;/td&gt;
&lt;td&gt;~20%&lt;/td&gt;
&lt;td&gt;Grouped tools, longer descriptions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Metadata-Rich&lt;/td&gt;
&lt;td&gt;~8%&lt;/td&gt;
&lt;td&gt;Curated knowledge sitting next to the tool&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Self-Teaching&lt;/td&gt;
&lt;td&gt;&amp;lt;2%&lt;/td&gt;
&lt;td&gt;Domain knowledge inside the descriptions and schemas&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Interactive App&lt;/td&gt;
&lt;td&gt;emerging&lt;/td&gt;
&lt;td&gt;Rendered UI returned by the server&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Secure Write App&lt;/td&gt;
&lt;td&gt;frontier&lt;/td&gt;
&lt;td&gt;Validated writes back to the system of record&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The types of MCP servers and how they differ
&lt;/h2&gt;

&lt;p&gt;They differ along a single axis: how much of the domain the server carries, versus how much it leaves the model to guess. That is what separates one type from the next. Each level below hands the model something the previous one withheld — starting from bare endpoint names, through typed metadata and descriptions that teach the model how to query, up to interactive apps that write back to the system of record.&lt;/p&gt;

&lt;h2&gt;
  
  
  Level 1: API Mapper (~70% of servers)
&lt;/h2&gt;

&lt;p&gt;One tool per endpoint. One-sentence descriptions. No domain context. The model figures everything out alone from the tool name. Ask &lt;em&gt;"which projects are running over budget?"&lt;/em&gt; and it will happily invent a filter, hit an empty endpoint, and tell you everything is fine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Level 2: Functional (~20%)
&lt;/h2&gt;

&lt;p&gt;Tools are grouped sensibly. Descriptions are longer. Someone thought about how a human would use this. Still no domain knowledge, no cross-tool references, no query strategies. This is the ceiling most commercial MCP implementations aim for today.&lt;/p&gt;

&lt;h2&gt;
  
  
  Level 3: Metadata-Rich (~8%)
&lt;/h2&gt;

&lt;p&gt;Knowledge graphs, glossaries, data catalogs. The metadata is real, but it lives next to the tool rather than inside it, and it was typically curated by hand over weeks or months. In practice, manual curation doesn't scale, and the agent only reads it if it happens to call the right meta-tool. I built two of these layers (MCP Resources and a parameterless "guide" tool) and removed both after testing. No Claude client ever requested them unprompted. &lt;em&gt;Updated September 2026:&lt;/em&gt; the guide tool deserved a second look. In later evals it was called 10 times in 10 once the description said "REQUIRED: call this first", against 0 in 10 for a soft hint (&lt;a href="https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery" rel="noopener noreferrer"&gt;Q8b&lt;/a&gt;). Unprompted it still goes unused; required, it works.&lt;/p&gt;

&lt;h2&gt;
  
  
  Level 4: Self-Teaching (&amp;lt;2%)
&lt;/h2&gt;

&lt;p&gt;The domain knowledge is &lt;em&gt;in what the model actually receives&lt;/em&gt;: the field names, the head of the tool description, the input schema, and the tool's response. Output schema validates the response, but the model never sees it. &lt;em&gt;Updated September 2026:&lt;/em&gt; this used to say the description and input schema are the only reliable channels. Measured on Claude Code, only the first 2,048 characters of a description arrive, so the rules for reading a result now travel in the response (&lt;a href="https://github.com/DaveGold/mcp-metadata-demo/blob/main/.claude/skills/rich-domain-mcp-server/references/evidence.md#delivery" rel="noopener noreferrer"&gt;Q7, Q15&lt;/a&gt;). And that knowledge wasn't written by humans scanning documentation; it was discovered by an AI examining real data, flagged with confidence levels, then validated by domain experts. I called the pattern &lt;strong&gt;Introspective Context Engineering for MCP&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The difference is not subtle. A Level 1 server says &lt;em&gt;"query data from the ERP."&lt;/em&gt; A Level 4 server says &lt;em&gt;"always start with &lt;code&gt;summaryOnly=true&lt;/code&gt;, active projects accumulate thousands of records. Type codes determine which fields are populated. Use &lt;code&gt;get_budget&lt;/code&gt; for planned costs, this tool for actuals. Report friction via &lt;code&gt;report_problem&lt;/code&gt;."&lt;/em&gt; Not the same product. Not the same category.&lt;/p&gt;

&lt;h2&gt;
  
  
  Level 5: Interactive App (emerging)
&lt;/h2&gt;

&lt;p&gt;The server doesn't just return data. It returns &lt;strong&gt;rendered UI&lt;/strong&gt;. Interactive charts, sortable tables, clickable maps, typed forms, all drawn by the server and displayed inline in the conversation. The agent coordinates; the server controls presentation.&lt;/p&gt;

&lt;p&gt;A table of 400 rows in a markdown code block is unreadable. A rendered, sortable, filterable table is a tool a business user can actually use. Level 5 is where the interface meets the user where they are.&lt;/p&gt;

&lt;p&gt;Working examples from an open-source demo server: &lt;a href="https://github.com/DaveGold/mcp-metadata-demo/blob/main/src/tools/render-chart.ts" rel="noopener noreferrer"&gt;&lt;code&gt;render_chart&lt;/code&gt;&lt;/a&gt; and &lt;a href="https://github.com/DaveGold/mcp-metadata-demo/blob/main/src/tools/render-table.ts" rel="noopener noreferrer"&gt;&lt;code&gt;render_table&lt;/code&gt;&lt;/a&gt;, each a self-describing tool whose schema teaches the agent how to configure the view, no wrapper logic required.&lt;/p&gt;

&lt;h2&gt;
  
  
  Level 6: Secure Write App (frontier)
&lt;/h2&gt;

&lt;p&gt;The server doesn't just read, it writes. Carefully. Two patterns: &lt;strong&gt;agent-initiated bounded writes&lt;/strong&gt; for low-risk mutations (feedback, scores) through standard tools, and &lt;strong&gt;user-initiated secure writes&lt;/strong&gt; for business-critical data through validated MCP App interactions. I call this the &lt;strong&gt;WriteIntent pattern&lt;/strong&gt;: agent opens the door, user walks through it, server checks every step.&lt;/p&gt;

&lt;p&gt;Almost nobody is here yet. Most builders are still nervous about giving MCP servers write access at all, and until Level 6 patterns exist, they should be.&lt;/p&gt;

&lt;h2&gt;
  
  
  The progression
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Expose data (1) → organize tools (2) → understand the domain (3) → learn from data and feedback (4) → present through interactive apps (5) → act through secure writes (6).&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most of the public MCP ecosystem is stuck between 1 and 2. MCP isn't dead. Most MCP servers are empty. A different transport doesn't fix that; filling the tool interface with real domain knowledge does.&lt;/p&gt;

&lt;p&gt;If you're building an MCP server today, the most useful question isn't &lt;em&gt;"which framework should I pick?"&lt;/em&gt; It's &lt;em&gt;"what level is mine, and what does Level N+1 look like?"&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;This post is one piece of a longer argument. &lt;a href="https://davidgolverdingen.nl/en/insights/production-mcp-practitioners-guide" rel="noopener noreferrer"&gt;&lt;em&gt;Production MCP: A Practitioner's Guide&lt;/em&gt;&lt;/a&gt; puts all of them in order, from understanding your data through to identity-bound deployment.&lt;/p&gt;

&lt;p&gt;Go deeper: read the full practitioner report, &lt;a href="https://davidgolverdingen.nl/en/the-missing-layer" rel="noopener noreferrer"&gt;&lt;em&gt;The Missing Layer&lt;/em&gt;&lt;/a&gt;, or explore the &lt;a href="https://github.com/DaveGold/mcp-metadata-demo" rel="noopener noreferrer"&gt;mcp-metadata-demo&lt;/a&gt; server, an open-source extract of these patterns, since the production servers run on private business data.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>architecture</category>
      <category>aiengineering</category>
    </item>
  </channel>
</rss>
