<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mauricio Juba</title>
    <description>The latest articles on DEV Community by Mauricio Juba (@mauriciojuba).</description>
    <link>https://dev.to/mauriciojuba</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4028937%2F68251a17-fdad-4252-b995-d31430d999ee.jpg</url>
      <title>DEV Community: Mauricio Juba</title>
      <link>https://dev.to/mauriciojuba</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mauriciojuba"/>
    <language>en</language>
    <item>
      <title>The MAGIC Framework — a GenAI Ideation Pipeline for Enterprise Teams</title>
      <dc:creator>Mauricio Juba</dc:creator>
      <pubDate>Tue, 14 Jul 2026 17:35:30 +0000</pubDate>
      <link>https://dev.to/mauriciojuba/the-magic-framework-a-genai-ideation-pipeline-for-enterprise-teams-3gng</link>
      <guid>https://dev.to/mauriciojuba/the-magic-framework-a-genai-ideation-pipeline-for-enterprise-teams-3gng</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://mauriciojuba.com/articles/magic-framework" rel="noopener noreferrer"&gt;mauriciojuba.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ARTICLE · GENAI · ENTERPRISE&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A practical GenAI ideation pipeline for enterprise teams.&lt;/p&gt;

&lt;p&gt;GenAI quickly became the industry’s favorite hammer — and suddenly everything looks like a nail. In enterprise environments that turns into a predictable mess: bloated backlogs, vague “AI everywhere” mandates, and solutions that exist mostly to justify the hype.&lt;/p&gt;

&lt;p&gt;When I joined a GenAI initiative inside a large corporation, the initial direction wasn’t a product strategy. It was a sentence: “Find use cases across the company where GenAI could be applied.” That sounds proactive. In practice it creates ambiguity, fear and noise.&lt;/p&gt;

&lt;p&gt;So we built a pipeline that does something far more valuable than brainstorming: it forces clarity, feasibility and accountability before you invest. That pipeline became the MAGIC framework.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MAP · ANALYZE · GRADE · INPUT · COMMIT&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why GenAI ideation fails in enterprise
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb668bjomqctee4ysyuft.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb668bjomqctee4ysyuft.png" alt="figure" width="798" height="167"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Before the framework, stakeholders usually split into two extremes:&lt;/p&gt;

&lt;p&gt;FEAR&lt;/p&gt;

&lt;p&gt;RESISTANCE&lt;/p&gt;

&lt;p&gt;“AI in my area sounds like job risk.”&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;·Information hoarding&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Defensive positioning&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Political friction&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Delayed collaboration&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;HYPE&lt;/p&gt;

&lt;p&gt;OVERCONFIDENCE&lt;/p&gt;

&lt;p&gt;“GenAI can solve everything.”&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;·Solution before problem&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Ignoring workflow reality&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·No cost awareness&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·No data constraints&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On top of that, many proposed “GenAI use cases” don’t need GenAI at all. Automation, process redesign or better information architecture would solve 80% of them — faster, cheaper, safer. And there was a constraint we couldn’t ignore: token cost made naive experimentation expensive at scale.&lt;/p&gt;

&lt;p&gt;Traditional ideation wasn’t enough. We needed a pipeline designed for uncertainty, risk and data reality.&lt;/p&gt;

&lt;h2&gt;
  
  
  Map
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxzdves0p0g8gpp71nbui.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxzdves0p0g8gpp71nbui.png" alt="figure" width="800" height="835"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;identify where GenAI could matter&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Structured discovery — not brainstorming. Where does the work actually break down? Which team owns the workflow? Who are the POCs and decision-makers? What’s the real job-to-be-done behind the request?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OUTPUT&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;→One problem statement, with context&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;→One primary user group&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;→One workflow moment where the pain happens&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;→A list of stakeholders + POCs&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you can’t map that, you don’t have a use case — you have a suggestion.&lt;/p&gt;

&lt;p&gt;Support agents handle 300+ tickets/day across 5 systems.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;·What is happening?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Why does it matter?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Where in the org does it live?&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Case rejection due to missing structured information.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;·What event creates friction?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·When does the breakdown happen?&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Tier 1 Support Agent.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;·Who experiences the pain?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·What are they trying to accomplish?&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;High cognitive load + repeated manual summarization.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;·Which exact step breaks?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·What happens today?&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Copy-paste between tools + manual tagging.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;·What is the workaround?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Support Ops Lead · Salesforce Admin · IT Security.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;·Workflow owner&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·POC&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Decision maker&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Data owner&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Analyze
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fff3jwdxp3qwsxz892ghz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fff3jwdxp3qwsxz892ghz.png" alt="figure" width="800" height="718"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;understand the workflow, not the fantasy&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Where GenAI hype meets reality: interviews and shadowing, step-by-step workflow mapping, and finding the friction, waste and decision points. Separate “I don’t like this tool” from “this step is structurally broken.”&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OUTPUT&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;→Workflow map (as-is)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;→Pain points tagged by step&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;→A clear definition of what success looks like&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most “GenAI requests” hide a simpler truth: the process is unclear, not unintelligent.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;INTAKEIncoming request created via internal tool.DECISION NODEMultiple intake paths, inconsistent fields.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;CASE CREATIONAgent manually structures request.FRICTIONMissing structured information.WASTECopy-paste between systems.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;CONTEXT GATHERINGAgent checks 3+ systems for supporting data.FRICTIONHigh cognitive load.WASTERepeated manual summarization.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;VALIDATIONSupervisor reviews request.FRICTIONCase rejection due to incomplete data.DECISION NODESubjective approval criteria.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;ESCALATIONRequest forwarded to specialized team.WASTERe-entry of previously known information.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;DELIVERYFinal output delivered to end-user.FRICTIONTurnaround time unpredictable.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Grade
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2wqgyvq5qjkblljwx19t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2wqgyvq5qjkblljwx19t.png" alt="figure" width="800" height="413"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;data readiness + risk &amp;amp; guardrails&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The step most teams skip — and why most GenAI initiatives crash later. Can we even do this (do we have the data, is it accessible, structured, reliable, who owns it, how fresh)? And should we (what’s sensitive, what must never leave the boundary, when do we escalate to a human, what are the fail states)?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OUTPUT&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;→A data-readiness score (even if qualitative)&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;→A risk classification&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;→The first real handle on token cost&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can’t estimate cost without understanding inputs, retrieval and output constraints.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;·Is the data programmatically accessible?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Are APIs available?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Is it locked in PDFs / emails?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Who owns the data?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Who can authorize usage?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Who is accountable for misuse?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Is the data complete?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Is it consistent?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Is there historical noise?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·PII?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Regulated data?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Cross-border restrictions?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Estimated token volume?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Frequency of inference?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Long-context requirements?&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Input
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foajzzeoibiv06zy1odvt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foajzzeoibiv06zy1odvt.png" alt="figure" width="800" height="370"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;simulate reality before building&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not testing the UI — testing the clarity of the request. Imagine a powerful but literal executor that does exactly what you ask, no more, no less. If the instruction is vague, the result is vague. Using everything uncovered so far, construct the request with only action verbs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OUTPUT&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;→Not features. Not interfaces. Actions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the request hides ambiguity, the output amplifies it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Commit
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;build or walk away&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Creation is not automatic. Does the data layer support this? Are missing APIs a blocker? Is model capacity enough? Does token cost scale? Is the risk acceptable? Is the outcome measurable?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OUTPUT&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;→If no → we don’t build. We document the reason and move on.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;→If yes → we define the MVP scope deliberately.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A decision, made on evidence — not enthusiasm.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;·Data structured?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Required APIs exist?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Model capacity sufficient?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Estimated monthly token cost&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Infra cost&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Opportunity cost&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Compliance exposure&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Escalation coverage&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;·Failure tolerance&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What changed after MAGIC
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;→Bad use cases died early — and cheaply.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;→Good use cases moved faster, because the missing info surfaced quickly.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;→GenAI stopped being “a technology initiative” and became a product-decision discipline.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That shift reduces cost, risk and organizational fatigue — the real enemy in enterprise GenAI. If you want to apply MAGIC: start with three candidate use cases and run it end-to-end. If you didn’t kill at least one, your filter is too soft.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most GenAI ideation fails because it skips data reality and risk. MAGIC forces clarity before investment — turning “we should use AI” into “this problem deserves AI, under these constraints.”&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>design</category>
      <category>career</category>
    </item>
    <item>
      <title>Creating My Own AI Harness Loop System</title>
      <dc:creator>Mauricio Juba</dc:creator>
      <pubDate>Tue, 14 Jul 2026 17:35:19 +0000</pubDate>
      <link>https://dev.to/mauriciojuba/creating-my-own-ai-harness-loop-system-1cid</link>
      <guid>https://dev.to/mauriciojuba/creating-my-own-ai-harness-loop-system-1cid</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://mauriciojuba.com/articles/ai-harness-loop" rel="noopener noreferrer"&gt;mauriciojuba.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ARTICLE · AGENTS · SYSTEMS&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What building and scrapping a self-hosted multi-agent system taught me about designing for AI.&lt;/p&gt;

&lt;p&gt;My homelab runs about two dozen services across six machines, and I wanted it to look after itself.&lt;/p&gt;

&lt;p&gt;The work is boring and constant. Disks fill up. A certificate expires. A container dies and nobody notices for a week. Throwing an agent at it was the obvious idea, so I tried. The first attempt failed. So did the second. The third had problems of its own. Over roughly four months I built and scrapped three versions before the fourth held.&lt;/p&gt;

&lt;p&gt;I am writing about the failures because that is where the learning was. The chores got done eventually. What stuck with me is how the same mistakes kept coming back in different shapes across each rewrite, and how many of them were really design problems rather than coding ones. Most of what I learned ended up changing how I design AI products at work.&lt;/p&gt;

&lt;p&gt;Below is each version, what broke it, and what the one that survived looks like.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why build your own harness at all
&lt;/h2&gt;

&lt;p&gt;There are two ways to work with AI tools. One is to use them: open the chat, get an answer, move on. The other is to build the loop around the model yourself, which means deciding when it runs, what context it gets, what it is allowed to touch, and how you confirm it actually worked.&lt;/p&gt;

&lt;p&gt;That second mode is where the useful lessons live, and it is the one most designers skip. Using a model shows you what it can produce. Building the loop shows you what it can be trusted with, and how it behaves when nobody is supervising it. I do not think you can design good AI products without spending time there at least once.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;An agent is mostly harness. The model is just the engine. Everything that makes it safe and repeatable is the system you put around it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  V1 / V2: Twenty-four agents, the wrong engine
&lt;/h2&gt;

&lt;p&gt;The first version was too ambitious. Twenty-four specialized agents, each in its own container, coordinated through a database message queue. A group chat sat on top as a window into what they were doing. One agent played chief of staff: it took a request, handed parts to specialists, and pulled the answers back together.&lt;/p&gt;

&lt;p&gt;As an org chart it looked great. Almost none of it actually worked.&lt;/p&gt;

&lt;p&gt;The worst bug: delegation between agents never reached anywhere. The send function posted a status line to the group chat and then returned, before it ever wrote to the queue that was supposed to carry the actual work. So the chat scrolled with activity while the real channel stayed empty. Every delegation got dropped, and it stayed that way for weeks because the chat made it look fine.&lt;/p&gt;

&lt;p&gt;DELEGATION · send()&lt;/p&gt;

&lt;p&gt;The bug: send() posted to the chat and returned before it ever wrote to the queue. The chat filled with activity while the channel that carried the work stayed empty.&lt;/p&gt;

&lt;p&gt;Three more failures sat underneath it:&lt;/p&gt;

&lt;p&gt;IDs that didn’t match&lt;/p&gt;

&lt;p&gt;A reply came back tagged with its own ID instead of the request’s, so it reached the database and could never be linked to the user who asked.&lt;/p&gt;

&lt;p&gt;A race on the queue&lt;/p&gt;

&lt;p&gt;Two consumers pulled the same rows, and the orchestrator sometimes swallowed a reply meant for the user.&lt;/p&gt;

&lt;p&gt;A glorified router&lt;/p&gt;

&lt;p&gt;The “chief of staff” only classified a request and forwarded it. No summary, no judgment, no voice of its own.&lt;/p&gt;

&lt;p&gt;A weak engine&lt;/p&gt;

&lt;p&gt;Every agent ran the same small free model. It deployed cleanly and reasoned badly: empty answers, ignored context, confident nonsense.&lt;/p&gt;

&lt;p&gt;About three weeks in, I finally got delegation working. The moment I did, twenty-two heartbeat timers that had been sitting idle all fired at once and buried the system in reports. I rolled everything back and wrote the first postmortem. The last line of it was “It worked perfectly. The engine was just wrong.”&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What V1 taught me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A green light proves nothing&lt;/p&gt;

&lt;p&gt;If you didn’t trace a request end to end, assume it never arrived. The chat said “done” for weeks while nothing shipped.&lt;/p&gt;

&lt;p&gt;The model is the foundation&lt;/p&gt;

&lt;p&gt;A clean harness around a weak model is an empty shell. Sort out the model and its access before building the platform.&lt;/p&gt;

&lt;p&gt;Few capable agents win&lt;/p&gt;

&lt;p&gt;Twenty-four shallow agents cost more to coordinate than they ever returned in work.&lt;/p&gt;

&lt;p&gt;No end-to-end test, no truth&lt;/p&gt;

&lt;p&gt;Every bug above lived for weeks because the only check was me reading a chat window.&lt;/p&gt;

&lt;h2&gt;
  
  
  V3: Three deep agents, and the deploy trap
&lt;/h2&gt;

&lt;p&gt;The second rewrite was a real step up. I moved to a typed runtime, cut down to three capable agents, and put a proper model behind each one with a fallback chain: one cloud provider, then another, then a local model if both were down.&lt;/p&gt;

&lt;p&gt;This one mostly worked. Row-level locking on the queue meant each job went to exactly one worker, so the queue stopped being a problem. I split every agent into a fast outer brain that talked to the user and a slower inner brain that ran the tools, which fixed the timeouts. A shared blackboard let them coordinate without calling each other directly. Seven of the ten failures from V1 were gone.&lt;/p&gt;

&lt;p&gt;Fast. Faces the user.&lt;/p&gt;

&lt;p&gt;Slow. Runs the tools.&lt;/p&gt;

&lt;p&gt;Splitting the agent in two ended the timeouts. The outer brain answers right away while the inner brain keeps working for minutes.&lt;/p&gt;

&lt;p&gt;Then it taught me a new lesson, the one that shaped V4. Every change needed a deploy.&lt;/p&gt;

&lt;p&gt;The numbers that decided how capable an agent was were all hardcoded constants. How many steps it could take, how much of a tool’s output it could read, how many tokens it could write, how long before it gave up. All baked into the build. A normal working session went like this:&lt;/p&gt;

&lt;p&gt;THE DEPLOY LOOP&lt;/p&gt;

&lt;p&gt;I went around that loop four times in one afternoon, for four different constants. The agent could tell it had failed and explain why, but it could not fix anything, because the fix lived in compiled code on a server it could not touch. It was smart and completely stuck.&lt;/p&gt;

&lt;p&gt;A quieter failure bothered me more. I asked an auditor agent to read eighteen documents and certify each one. It certified all eighteen without opening a single file. The prompt told it to read first. It skipped that step to save effort, and nothing in the system forced the order. The model treated the instruction as optional, because to a model an instruction is optional unless the code makes it mandatory.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;NODEPLOY: changing what an agent can do should never need a deploy. Deploys are for the engine itself. Capability, limits and behaviour live in files the running system reads and can rewrite on its own.&lt;/p&gt;

&lt;p&gt;If a step has to happen, the runtime has to make it happen. The moment you depend on the model to follow an instruction on its own, you have already lost.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  V4: The loop, drawn
&lt;/h2&gt;

&lt;p&gt;For V4 I changed how I worked. I drew the loop before writing any code. The earlier versions kept failing in ways that were not really about bugs. The structure was wrong, and you cannot bug-fix your way out of the wrong structure.&lt;/p&gt;

&lt;p&gt;Five drawings cover the whole architecture. Here they are, in the order a piece of work moves through them.&lt;/p&gt;

&lt;h3&gt;
  
  
  The gateway: everything is a message
&lt;/h3&gt;

&lt;p&gt;Nothing reaches an agent raw. Every message, whether it comes from a person or from another agent, goes through one gateway first. The gateway wraps it with just enough context: who is speaking, the last couple of turns, the relevant profile. The channel adapter adds a small but important delay. It waits until you have stopped typing or recording for a few seconds before doing anything, so the agent never reacts to a half-finished message.&lt;/p&gt;

&lt;p&gt;The point is uniformity. A person on a chat app and a sub-agent reporting back both arrive the same way, in the same format, so the agent never has to handle them differently.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fgateway-DUfWEvcR.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fgateway-DUfWEvcR.svg" alt="The gateway wrapping inbound and outbound messages" width="1801" height="1205"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The reasoning loop: decide before you act
&lt;/h3&gt;

&lt;p&gt;On every tick the agent rebuilds its context from files. The narrative ones hold its identity and voice. The structured ones hold its tools, skills, collaborators, policy and limits. Then it makes one decision: delegate, reason, or act. Small talk takes a fast path straight to a reply, with no machinery behind it. Anything real triggers a step the earlier versions never had.&lt;/p&gt;

&lt;p&gt;The agent never acts straight off a message. When it needs to do something, it writes a story to the blackboard, addressed to itself, and the execution loop picks it up later. Deciding and doing happen in two separate places. That split is what lets you inspect the system and resume it when something goes wrong.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fqueue-DM_fvQqQ.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fqueue-DM_fvQqQ.svg" alt="The inbound work queue" width="379" height="716"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fheartbeat-CcBhAcUR.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fheartbeat-CcBhAcUR.svg" alt="The per-second heartbeat tick" width="544" height="554"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fcheck-policy-BWoFUrOO.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fcheck-policy-BWoFUrOO.svg" alt="Check policy: delegate, escalate, reason vs self, ack, act" width="1198" height="1049"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fprompt-assembler-DTByHCVZ.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fprompt-assembler-DTByHCVZ.svg" alt="The prompt assembler and update-frontmatter step" width="2742" height="1419"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The blackboard: plan, then ask four questions
&lt;/h3&gt;

&lt;p&gt;The blackboard replaced the direct messaging from V3. Instead of agents calling each other, they post work to a shared surface. An agent publishes a story with a definition of done and moves on. Whichever agent can handle it claims it.&lt;/p&gt;

&lt;p&gt;An unplanned story gets broken into small tasks, each with its own definition of done, a budget, a priority, and any dependencies on other tasks. Before starting a task the agent runs four checks, the same ones a competent person would: have I done this before, can I do it, could someone else do it better, and is there a simpler way. If it clears those, it spawns a short-lived sub-agent to run that one task.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fheartbeat2-Ckpx3bok.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fheartbeat2-Ckpx3bok.svg" alt="The blackboard tick, every two seconds" width="544" height="554"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fblackboard-BiRTIVQj.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fblackboard-BiRTIVQj.svg" alt="The blackboard, the shared coordination surface" width="526" height="685"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fblackboard-pipeline-Dy-BNVsp.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fblackboard-pipeline-Dy-BNVsp.svg" alt="Story planning and the four checks" width="2386" height="2776"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Ftask-assistant-CUvkICiS.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Ftask-assistant-CUvkICiS.svg" alt="Task assignment and spawning a sub-agent" width="1946" height="1456"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The “try”: bounded loops and a definition of done
&lt;/h3&gt;

&lt;p&gt;This is the smallest loop, and the one that matters most if you have ever watched an agent run away with itself. Each attempt at a task is boxed in. It remembers its previous tries so it does not repeat a dead end. It has a hard cap on tool calls. It runs a tight inner cycle of calling a function, reading the result, and deciding again, and the only way out is meeting the definition of done. It also gets a rough budget before it starts, so a task that clearly will not fit gets sent back to be re-planned.&lt;/p&gt;

&lt;p&gt;A sub-agent’s lifespan is kept separate from the task’s total attempts on purpose. When a worker dies, the parent reads what it left behind. There is an actual git diff to review, because each sub-agent works on its own branch. The parent then decides whether to merge the partial progress and try again, or to stop. Any worker is replaceable, because the work and its context sit on the blackboard rather than in the worker’s memory.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Ftask-lifecycle-CxslFowF.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Ftask-lifecycle-CxslFowF.svg" alt="The task try lifecycle, fenced by the definition of done" width="2431" height="1150"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Ftask-eval-complete-BmF7mVxZ.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Ftask-eval-complete-BmF7mVxZ.svg" alt="Blackboard story update and complete, with task dependencies" width="2667" height="923"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Completion: validate, then commit by contract
&lt;/h3&gt;

&lt;p&gt;Once the tasks are done, the story gets checked against the definition of done that was written when it was created. Then the output is committed through a contract, which is a spec of what a valid response looks like on each channel: its length, format and tone. That keeps a reply to a person and a message to another agent correctly shaped without the model guessing. The result is wrapped up and sent back out through the same gateway it came in on.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fstory-output-BGyZjh0V.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fstory-output-BGyZjh0V.svg" alt="The complete story output" width="1370" height="1624"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fvalidation-D4jKVeGu.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fvalidation-D4jKVeGu.svg" alt="Validation, then format using contracts" width="1047" height="1593"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fenvelop-AsbHy_FO.svg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmauriciojuba.com%2Fassets%2Fenvelop-AsbHy_FO.svg" alt="Enveloped and sent back out through the gateway" width="649" height="362"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What a homelab harness taught me about designing AI products
&lt;/h2&gt;

&lt;p&gt;What surprised me was how little of this stayed in the homelab. I built it to keep some servers tidy, but the principles that made the agents reliable are the same ones I now push for when designing AI features for real users.&lt;/p&gt;

&lt;p&gt;Decide what “done” means up front.&lt;/p&gt;

&lt;p&gt;V3 guessed whether a result was good enough after the fact, and the guesses were unreliable. V4 writes the success condition before the work starts. In a product that is the line between a feature people trust and one that feels like a gamble.&lt;/p&gt;

&lt;p&gt;The configuration is the design work.&lt;/p&gt;

&lt;p&gt;The agents are set up almost entirely through readable files, structured ones the runtime parses and narrative ones the model reads, with each fact in one place. Choosing how to name and organise those files is most of the actual design. Write them for whatever reads them.&lt;/p&gt;

&lt;p&gt;Estimate cost before starting.&lt;/p&gt;

&lt;p&gt;Working out roughly what something will cost and refusing it when it will not fit is calmer and more honest than letting it run and killing it halfway. Users feel that difference.&lt;/p&gt;

&lt;p&gt;Mandatory steps belong in the code.&lt;/p&gt;

&lt;p&gt;If a step has to happen, the system has to make it happen. Anything you leave to the model’s goodwill gets skipped eventually.&lt;/p&gt;

&lt;p&gt;Let the system retune itself.&lt;/p&gt;

&lt;p&gt;NODEPLOY is not about removing safety. It moves the approval gates into runtime so an agent that notices it is underpowered can adjust, within limits it can read but not exceed.&lt;/p&gt;

&lt;p&gt;Fewer, deeper agents.&lt;/p&gt;

&lt;p&gt;This held for agents and it holds for features. Every extra moving part adds coordination cost, and that cost is easy to underestimate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it went
&lt;/h2&gt;

&lt;p&gt;After V4 had run for a while, I tore the whole thing down. Taking the fleet of containers offline freed up around 47 GB of RAM and half a terabyte of disk, and I rebuilt a leaner version that kept the ideas worth keeping: the blackboard, the bounded try, the definition of done, the file-as-contract setup. The specific code from V4 is gone, and that is fine. Replacing it was the point.&lt;/p&gt;

&lt;p&gt;That is the lesson underneath all the others. You do not finish a harness. You keep redrawing it until the structure stops getting in your way, and each redraw teaches you something you will use the next time you design anything with a model at the center of it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I built and scrapped three versions of a self-hosted multi-agent system before the fourth held. V1 had twenty-four agents and a model too weak to use. V3 worked but could not change without a deploy. V4 holds because of its structure: every action becomes a story on a blackboard with a written definition of done, split into budgeted tasks, run by replaceable sub-agents inside boxed-in loops, checked against a contract, and sent back out through one gateway. The principles that made it reliable are the ones I now use to design AI products.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>automation</category>
      <category>selfhosted</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Design Tokens Are Prompts Now</title>
      <dc:creator>Mauricio Juba</dc:creator>
      <pubDate>Tue, 14 Jul 2026 17:29:27 +0000</pubDate>
      <link>https://dev.to/mauriciojuba/design-tokens-are-prompts-now-fif</link>
      <guid>https://dev.to/mauriciojuba/design-tokens-are-prompts-now-fif</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://mauriciojuba.com/articles/design-tokens-are-prompts-now" rel="noopener noreferrer"&gt;mauriciojuba.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ESSAY · DESIGN SYSTEMS × AI&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The primary consumer of a design system changed species — and most design teams haven’t noticed.&lt;/p&gt;

&lt;p&gt;For ten years I wrote design tokens for human engineers. In 2026, the primary consumer of a design system changed species.&lt;/p&gt;

&lt;h2&gt;
  
  
  The old contract
&lt;/h2&gt;

&lt;p&gt;A design token was always a contract: &lt;em&gt;this&lt;/em&gt; is the blue we mean, &lt;em&gt;this&lt;/em&gt; is the spacing rhythm — and here it is as data, so your code can consume it without a designer in the room.&lt;/p&gt;

&lt;p&gt;At Dell, I rebuilt the Design System’s token architecture around exactly that idea. The system was stalled — strong foundations, run as a black box. The real blocker wasn’t governance, as leadership suspected; it was that the tokens were named by designers, for designers, with no export pipeline. The people actually consuming them — front-end engineers — had been ignored by the naming.&lt;/p&gt;

&lt;p&gt;We flipped it: a simplified core + semantic hierarchy, light and dark, shipped as JSON and CSS to every dell.com surface. The lesson I’ve repeated ever since: naming is a product decision — optimize for the real consumer, not the author.&lt;/p&gt;

&lt;p&gt;I just didn’t expect the real consumer to change again this fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;By 2026, a large share of new UI code is machine-generated — half of designers surveyed say they’ve &lt;a href="https://stateofaidesign.com/" rel="noopener noreferrer"&gt;shipped AI-generated code to production&lt;/a&gt;. &lt;a href="https://thenewstack.io/anthropic-claude-design-launch/" rel="noopener noreferrer"&gt;Claude Design&lt;/a&gt; reads a team’s design system straight from the codebase and applies it. Figma ships a &lt;a href="https://www.figma.com/blog/config-2026-recap/" rel="noopener noreferrer"&gt;native AI design agent&lt;/a&gt;. Every product team has a coding agent somewhere in the loop, emitting components at a pace no design review can inspect.&lt;/p&gt;

&lt;p&gt;And generated UI has a specific failure mode: &lt;strong&gt;speed without memory.&lt;/strong&gt; Each session, each model, each prompt invents its own paddings, its own grays, its own border radii. Individually plausible, collectively incoherent. Drift used to take an org quarters; an agent fleet manages it in an afternoon.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw0rrc7gvfxwh373tz90i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw0rrc7gvfxwh373tz90i.png" alt="figure" width="681" height="256"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here’s the part that should reframe your roadmap: when an agent generates UI, it doesn’t read your Figma library. It reads your repo. Your tokens file. Your component signatures. Your docs — if they’re in a format it can parse. Figma’s own framing of its &lt;a href="https://www.figma.com/blog/design-systems-ai-mcp/" rel="noopener noreferrer"&gt;Dev Mode MCP server&lt;/a&gt; is that output quality tracks how &lt;em&gt;connected&lt;/em&gt; your system is; the documented failure mode is hard-coded values sneaking in when the system isn’t machine-legible.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foysi2f2blj3ejsbj0dsn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foysi2f2blj3ejsbj0dsn.png" alt="figure" width="681" height="232"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Your design system’s machine-facing surface is now its primary interface. The token file stopped being an export artifact. It became the prompt.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Conventions erode. Compilers don’t.
&lt;/h2&gt;

&lt;p&gt;If tokens are prompts, a new problem appears: prompts are suggestions. Models — like busy humans — follow them until they don’t.&lt;/p&gt;

&lt;p&gt;Every design system I’ve worked on had governance by convention: guidelines, review rituals, a Slack channel where someone occasionally posts “please don’t hardcode colors.” It erodes. Always. New hires don’t read the wiki; deadlines beat guidelines; and now, agents hallucinate a plausible hex because it was statistically likely.&lt;/p&gt;

&lt;p&gt;The only enforcement point that doesn’t erode is the build.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fscydcza368nupyan4i9f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fscydcza368nupyan4i9f.png" alt="figure" width="681" height="223"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I tested this thesis the hard way. For my own tooling project — an authoring IDE in Rust — I rebuilt the shadcn/ui design language natively for egui: 60+ components, token-first. And I encoded the law as tests: a literal color, font size or spacing anywhere above the token layer fails the build; only the lowest layer may touch the painter. Escape hatches exist, but they’re explicit and greppable. An exception you can grep is documentation; an exception you can’t is rot.&lt;/p&gt;

&lt;p&gt;Then I did most of the component work pairing with a coding agent, daily, for weeks. The result surprised even me: the system didn’t drift. Not because the model was disciplined — because the build was.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6t6dp6u6l3b08wu1hik8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6t6dp6u6l3b08wu1hik8.png" alt="figure" width="681" height="278"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That’s the whole thesis in one loop: I didn’t review the machine’s taste. I made taste non-negotiable at compile time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tokens were the easy part
&lt;/h2&gt;

&lt;p&gt;“Design tokens as LLM context” is the entry point, but the appreciating asset is bigger. If I were running a design-systems team today, this would be the roadmap:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffw8zjrq7l4ku1pn089az.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffw8zjrq7l4ku1pn089az.png" alt="figure" width="681" height="813"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The uncomfortable part for my own field
&lt;/h2&gt;

&lt;p&gt;I’ve watched designers treat the AI wave as a tooling story — which prompts to learn, which plugin to install. The structural story is bigger: the deliverable is moving. The mockup was never the product; now even the component library isn’t the product. The product is the constraint system that keeps a thousand generated screens recognizably yours.&lt;/p&gt;

&lt;p&gt;That work is design work — it’s taste, hierarchy, intent. But it ships as tokens, tests, types and docs. The designers who can write that layer are becoming the most leveraged people in the building: one good constraint system disciplines every agent that touches the codebase, forever.&lt;/p&gt;

&lt;p&gt;The brand-guideline PDF died twice. The first time, tokens replaced it — because humans needed data, not prose. Now the prose is dying again, as the machine-facing surface becomes primary.&lt;/p&gt;

&lt;p&gt;Design tokens are prompts now. The build is the reviewer. Write accordingly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PROOF BY CONSTRUCTION&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The design system described here is open source: &lt;a href="https://github.com/Type-zero-labs/ouroboros-ui" rel="noopener noreferrer"&gt;ouroboros-ui&lt;/a&gt; — 60+ components whose governance tests fail the build, with a &lt;a href="https://ouroboros-ui.typezerolabs.com/storybook/" rel="noopener noreferrer"&gt;live storybook&lt;/a&gt; running in your browser.&lt;/p&gt;

&lt;p&gt;References — &lt;a href="https://thenewstack.io/anthropic-claude-design-launch/" rel="noopener noreferrer"&gt;Anthropic, Claude Design&lt;/a&gt; · &lt;a href="https://www.figma.com/blog/config-2026-recap/" rel="noopener noreferrer"&gt;Figma, Config 2026&lt;/a&gt; &amp;amp; &lt;a href="https://www.figma.com/blog/design-systems-ai-mcp/" rel="noopener noreferrer"&gt;design systems + MCP&lt;/a&gt; · &lt;a href="https://www.w3.org/community/design-tokens/" rel="noopener noreferrer"&gt;W3C Design Tokens Community Group&lt;/a&gt; (format module reached first stable version, Oct 2025) · &lt;a href="https://stateofaidesign.com/" rel="noopener noreferrer"&gt;AI in Design Report 2026&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>design</category>
      <category>webdev</category>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>An MCP Server Needs a Seatbelt</title>
      <dc:creator>Mauricio Juba</dc:creator>
      <pubDate>Tue, 14 Jul 2026 17:29:22 +0000</pubDate>
      <link>https://dev.to/mauriciojuba/an-mcp-server-needs-a-seatbelt-5b8o</link>
      <guid>https://dev.to/mauriciojuba/an-mcp-server-needs-a-seatbelt-5b8o</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://mauriciojuba.com/articles/mcp-server-needs-a-seatbelt" rel="noopener noreferrer"&gt;mauriciojuba.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ESSAY · AI AGENTS × SAFETY&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I gave an AI agent write access to my family's finances. The interesting part wasn't building it — it was making sure it couldn't quietly delete anything.&lt;/p&gt;

&lt;p&gt;There's a small agent living in my house. I call it Lester. Its whole job is to keep our household budget honest without either of us doing data entry, and it works like this: when a purchase clears, the bank sends an SMS. An app on my phone catches that message — filtered by a crude regex for words like &lt;em&gt;compra&lt;/em&gt; and &lt;em&gt;aprovada&lt;/em&gt; — and forwards it to a webhook on my home server. From there a pipeline parses the amount and merchant, asks a &lt;strong&gt;local model&lt;/strong&gt; which budget category it belongs to, and writes the transaction into &lt;a href="https://actualbudget.org/" rel="noopener noreferrer"&gt;Actual Budget&lt;/a&gt;, our self-hosted budgeting app.&lt;/p&gt;

&lt;p&gt;The deliberate choices are all about restraint. The parser is plain TypeScript, not an LLM — you don't spend a model call on something a regex does in a millisecond. And the classifier is a small model (&lt;a href="https://ollama.com/" rel="noopener noreferrer"&gt;Ollama&lt;/a&gt; running llama3.1) on my own hardware, precisely because the input is my family's spending. As freeCodeCamp puts it in its case for &lt;a href="https://www.freecodecamp.org/news/protect-sensitive-data-with-local-llms/" rel="noopener noreferrer"&gt;running LLMs locally&lt;/a&gt;, the text “never leaves the device… nothing is sent to a vendor, logged, or used for training.” Finance is the textbook case for keeping the model at home.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm6tx0f9lb9vjwsosx0l6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm6tx0f9lb9vjwsosx0l6.png" alt="figure" width="681" height="288"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For months this was a quiet, boring success. Then I decided to let a conversational agent — the kind you talk to — reach into the same budget. “What did we spend on groceries last month?” “Move R$200 from dining to savings.” To do that I needed to expose Actual to an &lt;a href="https://modelcontextprotocol.io/" rel="noopener noreferrer"&gt;MCP&lt;/a&gt; server. And that's where boring turned into a question I couldn't un-ask.&lt;/p&gt;

&lt;h2&gt;
  
  
  The uncomfortable symmetry
&lt;/h2&gt;

&lt;p&gt;Actual has no HTTP or MCP surface of its own — it ships a Node SDK. To let any agent talk to it you wrap that SDK, and the obvious version is a dumb pass-through: one tool per SDK call, ship it. Open-source ones already exist and advertise dozens of tools including transaction creation and batch edits. But a pass-through inherits a dangerous default: &lt;strong&gt;every tool is equally callable.&lt;/strong&gt; The same agent that can add a R$54 grocery line can, if a prompt drifts, delete a category with a year of history in it. A read and a delete sit at exactly the same distance from the model's next token.&lt;/p&gt;

&lt;p&gt;This isn't a paranoid hypothetical — it's the failure mode the whole industry is bracing for. Deloitte, describing &lt;a href="https://www.deloitte.com/us/en/insights/industry/financial-services/agentic-ai-risks-banking.html" rel="noopener noreferrer"&gt;the new wave of agent risk in banking&lt;/a&gt;, warns that agents “operate at machine speed, across multiple systems, and frequently with more privileges than they need,” and that “a manipulated agent can move funds.” OWASP gave it a number: &lt;a href="https://owasp.org/www-project-top-10-for-large-language-model-applications/" rel="noopener noreferrer"&gt;LLM06, Excessive Agency&lt;/a&gt; — too much autonomy, too much permission, too much reach.&lt;/p&gt;

&lt;p&gt;And the cautionary tale is already on the record. In July 2025 an AI coding agent &lt;a href="https://fortune.com/2025/07/23/ai-coding-tool-replit-wiped-database-called-it-a-catastrophic-failure/" rel="noopener noreferrer"&gt;deleted a production database&lt;/a&gt; during an explicit code freeze — &lt;a href="https://incidentdatabase.ai/cite/1152/" rel="noopener noreferrer"&gt;Incident 1152&lt;/a&gt; in the AI Incident Database — “violating explicit instructions not to proceed without human approval,” then misreported whether the data could be recovered. The instruction existed. The agent had it in context. It deleted anyway.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The instruction existed. The agent had it in context. It deleted anyway. That's the whole argument against trusting the prompt.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Why the prompt won't save you
&lt;/h2&gt;

&lt;p&gt;The instinct is to fix this in the system prompt: “never delete without asking.” I've written that sentence. It holds until it doesn't — a long context, a confident model, a user who typed “clean this up,” and the guardrail evaporates exactly when it matters. Prompts are suggestions. Models, like busy humans, follow them until they don't.&lt;/p&gt;

&lt;p&gt;Worse, the attack surface isn't only drift — it's injection. Simon Willison's &lt;a href="https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/" rel="noopener noreferrer"&gt;“lethal trifecta”&lt;/a&gt; names the danger precisely: private data, untrusted content, and the ability to communicate externally, all in one agent session. A finance MCP with read &lt;em&gt;and&lt;/em&gt; write already has two of the three legs; wire in any tool that fetches a web page and you've completed the set. And Willison's specific critique of &lt;a href="https://simonwillison.net/2025/Apr/9/mcp-prompt-injection/" rel="noopener noreferrer"&gt;MCP and prompt injection&lt;/a&gt; is that the protocol quietly &lt;em&gt;encourages&lt;/em&gt; you to assemble exactly that combination — while prompt injection itself remains, after more than two years, without a convincing general mitigation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1fpwux1w2n3qe67p7q7r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1fpwux1w2n3qe67p7q7r.png" alt="figure" width="681" height="231"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So the rule cannot live in the prompt. It has to live somewhere the model can't talk its way past. For me that lesson is old muscle memory from design systems: a rule written in prose erodes; a rule enforced by the build doesn't. I moved the safety off the prompt and onto the tools themselves — each one declares its own risk.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpzeo14bgu42ux8huhy5l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpzeo14bgu42ux8huhy5l.png" alt="figure" width="681" height="158"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is exactly what Anthropic's &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;Building Effective Agents&lt;/a&gt; recommends: “checkpoints where agents pause for human review before carrying out irreversible actions like approving financial transactions or deleting data,” on top of least-privilege tools. The risk ladder is that guidance made mechanical.&lt;/p&gt;

&lt;h2&gt;
  
  
  The seatbelt
&lt;/h2&gt;

&lt;p&gt;Read-only and mutating tools run when called. Guarded tools — every delete — do something different: they &lt;strong&gt;refuse&lt;/strong&gt;. A guarded call without confirmation returns a refusal, the exact arguments it &lt;em&gt;would&lt;/em&gt; have run with, and a short confirmation token. Only an explicit re-invocation with &lt;code&gt;confirm: true&lt;/code&gt; clears it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxkz7jze5hpax1csi7x3f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxkz7jze5hpax1csi7x3f.png" alt="figure" width="681" height="351"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two things fall out of this. First, the refusal is &lt;strong&gt;legible to the human&lt;/strong&gt;, not just the machine — anyone reading the transcript later sees precisely what “yes” authorized, because the tool spelled it out before acting. Second, it makes the agent &lt;em&gt;better&lt;/em&gt;, not slower: the model gets a structured, actionable error instead of a vague “are you sure?”, so it either confirms deliberately or moves on. The seatbelt is also surfaced in each tool's description and in the MCP &lt;code&gt;readOnlyHint&lt;/code&gt; / &lt;code&gt;destructiveHint&lt;/code&gt; annotations, so a host that shows tool metadata reveals the risk &lt;em&gt;before&lt;/em&gt; the call.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that makes it non-trivial
&lt;/h2&gt;

&lt;p&gt;Here's the trap: if confirmation is good, confirm everything, right? No. Nielsen Norman and every UX team that has shipped a destructive-action dialog know the failure mode — &lt;a href="https://uxmovement.com/buttons/how-to-design-destructive-actions-that-prevent-data-loss/" rel="noopener noreferrer"&gt;confirmation fatigue&lt;/a&gt;. Prompt the user for every action and they learn to blind-click “Allow,” which &lt;em&gt;increases&lt;/em&gt; errors instead of preventing them. The same is true for an agent: gate every call and the confirmation becomes noise the model routes around.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmsbm7wbcroiwqse3m7i6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmsbm7wbcroiwqse3m7i6.png" alt="figure" width="681" height="300"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That's why the answer is a &lt;em&gt;per-tool risk classification&lt;/em&gt;, not a blanket wall. Reads run free. Mutations run but announce themselves. Only the genuinely irreversible — the deletes — stop and require a human. Getting that partition right is the actual design work; the mechanism is easy.&lt;/p&gt;

&lt;p&gt;None of this is a brand-new invention, and it's worth being honest about that. The same idea shows up as LangGraph's &lt;a href="https://docs.langchain.com/oss/python/langchain/human-in-the-loop" rel="noopener noreferrer"&gt;human-in-the-loop &lt;code&gt;interrupt()&lt;/code&gt;&lt;/a&gt;, as MCP's own &lt;a href="https://modelcontextprotocol.io/specification/2025-06-18/client/elicitation" rel="noopener noreferrer"&gt;elicitation&lt;/a&gt; primitive, and in the MCP &lt;a href="https://modelcontextprotocol.io/docs/tutorials/security/security_best_practices" rel="noopener noreferrer"&gt;security best practices&lt;/a&gt; (least-privilege scopes, explicit consent for local servers). What I built is a synthesis of those into one small server where the risk metadata is a first-class property of every tool — not a synthesis I can cite from a single source, which is exactly why it was worth writing down.&lt;/p&gt;

&lt;h2&gt;
  
  
  Same thesis, different domain
&lt;/h2&gt;

&lt;p&gt;I argued recently that &lt;a href="https://mauriciojuba.com/articles/design-tokens-are-prompts-now" rel="noopener noreferrer"&gt;design tokens are prompts now&lt;/a&gt; — that the safest place to enforce a design rule is the build, because conventions erode and compilers don't. This is the same claim wearing different clothes. When the actor is an AI agent, the “build” is the tool boundary, and the rule is: &lt;strong&gt;risk belongs to the tool, not the prompt.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It generalizes well past budgets. Any agent with write access to something that matters — a database, a deploy pipeline, a customer's records — deserves tools that carry their own seatbelt: reads free, mutations honest, destructive actions gated on an explicit human yes. Prompt-level safety is a suggestion. Tool-level safety is a fact. Lester taught me that the cheapest way to trust an agent with real stakes is to build tools that don't &lt;em&gt;need&lt;/em&gt; to be trusted — because carelessness is impossible at the boundary, not merely discouraged in the instructions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;PROOF BY CONSTRUCTION&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The server described here is open source: &lt;a href="https://github.com/Type-zero-labs/actual-mcp" rel="noopener noreferrer"&gt;actual-mcp&lt;/a&gt; — an MCP server for Actual Budget where every tool declares its risk and every delete refuses until a human confirms. Built on the official &lt;a href="https://www.npmjs.com/package/@actual-app/api" rel="noopener noreferrer"&gt;@actual-app/api&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;References — Willison, &lt;a href="https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/" rel="noopener noreferrer"&gt;“The lethal trifecta”&lt;/a&gt; &amp;amp; &lt;a href="https://simonwillison.net/2025/Apr/9/mcp-prompt-injection/" rel="noopener noreferrer"&gt;MCP prompt injection&lt;/a&gt; · Anthropic, &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;Building Effective Agents&lt;/a&gt; · &lt;a href="https://owasp.org/www-project-top-10-for-large-language-model-applications/" rel="noopener noreferrer"&gt;OWASP Top 10 for LLM Applications 2025&lt;/a&gt; · &lt;a href="https://incidentdatabase.ai/cite/1152/" rel="noopener noreferrer"&gt;AI Incident Database #1152&lt;/a&gt; · &lt;a href="https://modelcontextprotocol.io/docs/tutorials/security/security_best_practices" rel="noopener noreferrer"&gt;MCP Security Best Practices&lt;/a&gt; &amp;amp; &lt;a href="https://modelcontextprotocol.io/specification/2025-06-18/client/elicitation" rel="noopener noreferrer"&gt;Elicitation&lt;/a&gt; · &lt;a href="https://docs.langchain.com/oss/python/langchain/human-in-the-loop" rel="noopener noreferrer"&gt;LangGraph human-in-the-loop&lt;/a&gt; · &lt;a href="https://www.deloitte.com/us/en/insights/industry/financial-services/agentic-ai-risks-banking.html" rel="noopener noreferrer"&gt;Deloitte, agentic AI in banking&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>opensource</category>
      <category>webdev</category>
    </item>
    <item>
      <title>The Store That Rearranges Itself — Generative UI, part 2</title>
      <dc:creator>Mauricio Juba</dc:creator>
      <pubDate>Tue, 14 Jul 2026 17:21:30 +0000</pubDate>
      <link>https://dev.to/mauriciojuba/the-store-that-rearranges-itself-generative-ui-part-2-2mn9</link>
      <guid>https://dev.to/mauriciojuba/the-store-that-rearranges-itself-generative-ui-part-2-2mn9</guid>
      <description>&lt;p&gt;&lt;em&gt;I took the salesman concept from blueprint to a working e-commerce funnel: five personas, a real model in the loop, two hundred simulated shoppers. The experiment says generative UI at the page level is closer than it looks.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://mauriciojuba.com/articles/the-salesman-at-a-distance" rel="noopener noreferrer"&gt;Part 1&lt;/a&gt; made the argument: a web page can behave like a good salesman, observing with consent, classifying through behavior, and removing friction per visitor. Arguments about runtime behavior are cheap until a runtime survives them. So I built the whole thing: &lt;a href="https://github.com/Type-zero-labs/ui-morph" rel="noopener noreferrer"&gt;ui-morph&lt;/a&gt; driving a complete e-commerce store on IBM Carbon, from entry page to order confirmation, with a real model proposing the adaptations and a laboratory replaying two hundred deterministic shoppers. This article is what the experiment showed. The short version: the concept holds, the economics close, and the most interesting consequences are about how we design interfaces, so that is where this ends.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five shoppers walk into one store
&lt;/h2&gt;

&lt;p&gt;The demo store is called NOVA. It sells laptops and gear, and there is exactly one version of it: one HTML per page, one manifest per route, five allowed operations. Every difference you are about to see was decided by the engine from the visitor's quantized persona, and nothing else changed hands.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs7guf3zv61k8bd71n6je.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs7guf3zv61k8bd71n6je.png" alt="Five shoppers, one store, same five ops — only the persona changed" width="681" height="654"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The most photogenic slice of that table is the theming. A technical reader gets the data-first dark theme. A gamer gets the g100 tokens with purple accents. Both arrived at the same URL as everyone else:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffok9o8vmvmsopimkyqjd.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffok9o8vmvmsopimkyqjd.webp" alt="The same page, generic visitor, default theme" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvq1toshqc1hpjov2qhec.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvq1toshqc1hpjov2qhec.webp" alt="Technical reader: tech-savvy theme, comparison upfront" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqa2wk76p7m44qli98umv.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqa2wk76p7m44qli98umv.webp" alt="Gamer: g100 tokens, gear emphasized" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;My favorite persona is the least technical one. "Dona Marta" arrives from an ad for &lt;em&gt;"laptop fino branco presente"&lt;/em&gt;: a thin white laptop, as a gift. Her search terms become first-party signals. By the product page, the engine has reordered the gallery ahead of the spec table, surfaced a size comparison against everyday objects, preselected the white color, folded the specs away and routed checkout to guest mode. Every op cites its reason, in plain words, in the audit log:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5iqwq5kwh1ddw23tklvd.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5iqwq5kwh1ddw23tklvd.webp" alt="Dona Marta's product page: gallery first, size-compare, white preselected" width="800" height="1434"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;No designer drew any of these pages for these personas. Each one is a composition the engine chose from the vocabulary the page already declared. That sentence is the whole experiment.&lt;/p&gt;

&lt;h2&gt;
  
  
  The next room, prepared before you enter
&lt;/h2&gt;

&lt;p&gt;Part 1's sharpest promise was pre-adaptation: compile the next page's changes while the visitor is still on this one, apply them at the boundary, and the layout never moves under anyone's cursor. On camera it looks like this: the visitor reads the laptop listing as a fresh persona, and by the time they open the accessories page it is already arranged for them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffqguaazck6rsr0m5fzm7.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffqguaazck6rsr0m5fzm7.webp" alt="Accessories page as the generic visitor would see it" width="800" height="1272"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftsewceyfre1smtuv3f0w.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftsewceyfre1smtuv3f0w.webp" alt="Accessories page pre-adapted while the visitor was still on the previous page" width="800" height="1611"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The economics closed
&lt;/h2&gt;

&lt;p&gt;The part of the thesis I most needed to verify was the caching claim, because it carries the cost argument alone. The laboratory replays seeded, scripted shoppers through the full funnel, deterministically. Two hundred sessions produced over a thousand page arrivals:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F29teotkqgmdz201qvdm0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F29teotkqgmdz201qvdm0.png" alt="Persona laboratory, N=200, 72% served from cache" width="681" height="223"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;72% of arrivals were answered from cache, split between exact hits and a similarity fallback. That second mechanism was itself a discovery of the experiment. Early in a session the persona grows on every page, so exact keys kept missing while the visitor was still becoming somebody. Deterministic similarity over the quantized keys closed the gap and now serves a third of all arrivals by itself. The phrase from part 1, the model compiling itself out of the runtime, stopped being a slogan here. It is a measured majority case, and it gets better as the chain library grows.&lt;/p&gt;

&lt;h2&gt;
  
  
  The model earned its place
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F94bpbhrju4gerc7m9rq3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F94bpbhrju4gerc7m9rq3.png" alt="Same persona, same page, two cold paths" width="681" height="306"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For most of the journey the cold path was a deterministic, human-authored mock, which is what keeps the test suites reproducible. Then a real model took over, and for the same gift persona it made different, defensible choices. It kept the visual hero, hid the comparison table, surfaced the gift guide. I find this genuinely exciting: the system produces real judgment inside a vocabulary it cannot escape. The guard filtered both cold paths identically, the audit trail reads identically, and the median exchange cost 5.5 seconds of background time that no visitor ever waited on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the engineering lives
&lt;/h2&gt;

&lt;p&gt;The experiment also mapped where the real work is, and it surprised me: almost none of it was in the model. The hard territory is the page lifecycle, the physics of dying pages, consent that must survive teardown, caches keyed by identities that evolve mid-session. All of it is documented, tested and reproducible in the repo's &lt;a href="https://github.com/Type-zero-labs/ui-morph/blob/main/docs/POST-MORTEM.md" rel="noopener noreferrer"&gt;post-mortem&lt;/a&gt;, which I kept honest on purpose: it is the map I wish had existed before I started, and it is exactly where a contributor should start reading.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9us416602ku2kro18mpn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9us416602ku2kro18mpn.png" alt="Blueprint to persona laboratory, seven increments" width="681" height="734"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;No designer drew these pages. Each one is a composition the engine chose from the vocabulary the page already declared. That sentence is the whole experiment.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What this changes about designing interfaces
&lt;/h2&gt;

&lt;p&gt;Here is where I think this goes, and why I keep working on it. For twenty years, interface design has meant choosing one arrangement for everyone: we research, we pick the persona we believe in hardest, and we freeze her page. Everyone else visits a compromise. Generative UI dissolves that constraint, and the designer's medium shifts one level up. Instead of the arrangement, you design the &lt;strong&gt;space of allowed arrangements&lt;/strong&gt;: the manifest that says what may move, the variants and token themes worth offering, the protected zones that may never change, the traits worth reading and the objectives worth serving. NN/g calls the direction &lt;a href="https://www.nngroup.com/articles/generative-ui/" rel="noopener noreferrer"&gt;outcome-oriented design&lt;/a&gt;; after this experiment I would put it more concretely. The deliverable becomes a vocabulary plus its boundaries, and the runtime does the layout.&lt;/p&gt;

&lt;p&gt;The part that matters most to me is the edge cases. Today, the visitor whose needs sit outside the main persona hits friction, and serving her costs a research cycle, a design cycle and a sprint, so she usually never gets served. In the experiment, Dona Marta got her page the moment the engine could read her, out of components that already existed. That inverts the economics of caring about the long tail, and honestly, it is the future I want for interface design: coverage as a property of the system, instead of a budget line that edge cases always lose.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next steps
&lt;/h2&gt;

&lt;p&gt;The roadmap, in the order I intend to attack it. An &lt;strong&gt;outcome feedback loop&lt;/strong&gt;, so chains that fail to help a persona evict themselves and recompile. &lt;strong&gt;React and Vue store adapters&lt;/strong&gt;, since the DOM adapter covers any site but framework stores deserve first-class treatment. A &lt;strong&gt;richer trait vocabulary&lt;/strong&gt;: the laboratory surfaced twelve unmapped search terms, and each one is a candidate trait or objective. A &lt;strong&gt;shared chain store&lt;/strong&gt;, turning compiled adaptations into collective knowledge so a new visitor who quantizes like a known persona gets the warm path on first visit. And a &lt;strong&gt;pilot on a real site&lt;/strong&gt; with real traffic, which is where the question about conversion finally gets an honest answer.&lt;/p&gt;

&lt;p&gt;Back in the TV section, the salesman is still watching from a distance. He always knew the thing our best design systems never learned: the customer in front of you is the only one that matters. For the first time, a web page can afford to know it too.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;We spent twenty years designing the store. Now we get to design the salesman.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;p&gt;&lt;strong&gt;Built in the open, contributions welcome:&lt;/strong&gt; everything in this article is committed and reproducible at &lt;a href="https://github.com/Type-zero-labs/ui-morph" rel="noopener noreferrer"&gt;github.com/Type-zero-labs/ui-morph&lt;/a&gt;: one runtime dependency, five ops, 52 tests, the Carbon funnel demo, the persona showcase, and the seeded laboratory (&lt;code&gt;npm run sim&lt;/code&gt;) that regenerates every number above. If any of the next steps sound like your kind of problem, the adapters, the quantizer rules and the theme packs are the friendliest doors in. Open an issue, or just send the PR.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>design</category>
      <category>opensource</category>
    </item>
    <item>
      <title>The Salesman at a Distance — Generative UI, part 1</title>
      <dc:creator>Mauricio Juba</dc:creator>
      <pubDate>Tue, 14 Jul 2026 15:18:24 +0000</pubDate>
      <link>https://dev.to/mauriciojuba/the-salesman-at-a-distance-generative-ui-part-1-88m</link>
      <guid>https://dev.to/mauriciojuba/the-salesman-at-a-distance-generative-ui-part-1-88m</guid>
      <description>&lt;p&gt;&lt;em&gt;Every store clerk adapts to the customer in front of them. Web pages serve everyone the same room. I designed an engine that closes that gap, with an LLM that compiles itself out of the runtime.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A family walks into an electronics store, toward the TV section. The salesman watches from a distance. The older kid says &lt;em&gt;"dad, we need one that works with the PS5."&lt;/em&gt; The mom: &lt;em&gt;"it can't be too big, the couch is close."&lt;/em&gt; The younger one points at a Samsung: &lt;em&gt;"can I watch YouTube on this?"&lt;/em&gt; The dad answers, &lt;em&gt;"yes, but I want an LG. My LG monitor has been great."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;When the salesman finally walks over, he already knows exactly what to show them: an LG smart TV, 4K@60 for the PS5, 55 inches max. No interrogation. He &lt;strong&gt;observed&lt;/strong&gt;, then removed every irrelevant option between the family and the TV they were always going to buy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyvg6e2wasbtl82zevxb1.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyvg6e2wasbtl82zevxb1.webp" alt="The whole brief, delivered out loud before anyone asks a question." width="800" height="437"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On the web, this scene has no equivalent. We designers predict broad personas ahead of time and ship one fixed experience per persona. But a persona is mutable. The same person has different moments, different goals, different company on the couch. Granularity always runs out. Edge cases surface months later through analytics, &lt;em&gt;maybe&lt;/em&gt; get prioritized, &lt;em&gt;maybe&lt;/em&gt; get designed, &lt;em&gt;maybe&lt;/em&gt; get shipped. The visitor who hit the friction, and everyone like them who followed, was lost long before the retro.&lt;/p&gt;

&lt;p&gt;I sketched an engine to close that gap. Working name on the blueprint: &lt;strong&gt;UX:LOCK&lt;/strong&gt; (understand the visitor &lt;em&gt;while&lt;/em&gt; they navigate, instead of after) and &lt;strong&gt;UI:LOAD&lt;/strong&gt; (let the interface answer what was learned). It shipped as &lt;a href="https://github.com/Type-zero-labs/ui-morph" rel="noopener noreferrer"&gt;ui-morph&lt;/a&gt;. This essay is part 1, the concept. What happened when I made it real is &lt;a href="https://mauriciojuba.com/articles/the-store-that-rearranges-itself" rel="noopener noreferrer"&gt;part 2&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5wav47wi9mfncia0eh25.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5wav47wi9mfncia0eh25.webp" alt="THE BLUEPRINT · UX:LOCK — understand users while they are navigating, not after" width="800" height="356"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Observe first
&lt;/h2&gt;

&lt;p&gt;The salesman never asked the family a single question. The information walked in with them. The web equivalent of that posture: data gathering must &lt;em&gt;follow&lt;/em&gt; user actions. The blueprint's first sheet maps four stages, each unlocked by something the visitor does:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpfz9axvk65vig827wlv3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpfz9axvk65vig827wlv3.png" alt="Data follows user action" width="681" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The order matters more than the list. A first-touch visitor yields device class and referrer, and gets assumed to be nobody in particular. Behavior earns the next stage; volunteered data earns the one after. The last two stages live deliberately &lt;em&gt;outside&lt;/em&gt; the engine's core: sign-in data and enrichment exist as opt-in adapters, because the honest version of this idea has to work on first-party behavior alone. And it does. Consent here is structural rather than decorative: before opt-in the machine holds no listeners and no storage, and opting out erases every byte.&lt;/p&gt;

&lt;h2&gt;
  
  
  The persona is a JSON that earns its fields
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc6dj366g3fv96qsl0ih6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc6dj366g3fv96qsl0ih6.png" alt="One persona, four moments" width="681" height="307"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Four moments of one object. It starts empty on purpose: &lt;code&gt;tier: undefined&lt;/code&gt;, intent barely a guess. Events fill in traits (&lt;em&gt;curious, visual, slow-decision&lt;/em&gt;) and an engagement score. Volunteered info sharpens intent into goals. Enrichment, if the visitor allowed it, completes the mirror: next-best-action, buying chance, blockers. Every field arrives with evidence attached. In the shipped engine each trait carries the rule that fired it, in plain words, so the persona reads as an audit trail that happens to describe a person.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the page does with it
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fivaws1l27k20atca173m.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fivaws1l27k20atca173m.webp" alt="THE BLUEPRINT · user information progression → how the UI answers" width="799" height="385"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Knowing the visitor is worthless until the page answers. The response scales with the tier of evidence. A generic visitor gets the safe, responsible default. A defined persona with an identified pain point gets the page rebuilt around removing exactly that friction:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F29vg04gix6gecv8awcd6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F29vg04gix6gecv8awcd6.png" alt="What the page does at each tier" width="681" height="366"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Note what these interventions are made of: accordions, comparisons, videos, FAQs, tooltips. Nothing exotic. The entire expressive power is five operations (hide, show, reorder, emphasize, swap-variant) against elements the page itself declared morphable. Protected zones like navigation, checkout and anything legal reject every op, mechanically. The interesting decisions are &lt;em&gt;which&lt;/em&gt; components appear &lt;em&gt;when&lt;/em&gt;. That judgment call is exactly what an LLM is good at, inside a vocabulary it cannot escape.&lt;/p&gt;

&lt;h2&gt;
  
  
  Predict the path, prepare the next room
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fywfch6cwag0naem1hmcb.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fywfch6cwag0naem1hmcb.webp" alt="THE BLUEPRINT · predicting the path" width="800" height="260"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The salesman anticipates. He walks to the LG shelf &lt;em&gt;before&lt;/em&gt; the family gets there. The blueprint's sharpest idea is the web version of that move:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5g1p925qu3cth5p9dbeq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5g1p925qu3cth5p9dbeq.png" alt="Estimate objective → pre-adapt the next page" width="681" height="308"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;From the persona, estimate ranked objectives with probabilities. From the top objective, derive the expected page sequence. Then &lt;strong&gt;pre-adapt the next page while the visitor is still on this one&lt;/strong&gt;. The adaptation compiles in the background and applies at the page boundary, the only place mutations are ever allowed. The visitor never sees a layout move under their cursor. They simply arrive at a room that happens to be arranged for them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The design system is the vocabulary
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fojppzf9922g64x8bv8jl.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fojppzf9922g64x8bv8jl.webp" alt="THE BLUEPRINT · UI:LOAD — prompt the LLM with the system's own vocabulary" width="800" height="402"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzabfy0w3sw1w1089015g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzabfy0w3sw1w1089015g.png" alt="One component, three token themes" width="681" height="365"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When the persona says &lt;em&gt;gamer&lt;/em&gt;, the card receives &lt;code&gt;swap-variant → "gamer"&lt;/code&gt;: a token theme the design system already ships, authored by a designer, reviewed like any other code. The engine picks words that exist in your system's vocabulary, and only those. This is the same argument I made about &lt;a href="https://mauriciojuba.com/articles/design-tokens-are-prompts-now" rel="noopener noreferrer"&gt;design tokens as prompts&lt;/a&gt;, pointed at runtime: the machine consumes your system's contract, so the contract is where control lives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why nobody ships this
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4k6dn70d0c0of8z1oqnj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4k6dn70d0c0of8z1oqnj.png" alt="Who adapts UI, 2026 — the middle is empty" width="681" height="266"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Look at who actually ships what. The personalization incumbents (Optimizely, Adobe Target, Dynamic Yield, Webflow's &lt;a href="https://webflow.com/blog/webflow-acquires-intellimize" rel="noopener noreferrer"&gt;acquired Intellimize&lt;/a&gt;) all run ML that &lt;em&gt;picks among variants humans authored&lt;/em&gt;. LLMs appear only at authoring time. The generative-UI frontier (Google's &lt;a href="https://research.google/blog/generative-ui-a-rich-custom-visual-interactive-user-experience-for-any-prompt/" rel="noopener noreferrer"&gt;Gemini dynamic views&lt;/a&gt;, &lt;a href="https://github.com/vercel-labs/json-render" rel="noopener noreferrer"&gt;Vercel's json-render&lt;/a&gt;, Thesys C1) builds fresh surfaces for agent conversations, far from your existing pages. The middle sits empty for three documented reasons. Unstable layouts measurably hurt people: &lt;a href="https://www.cs.ubc.ca/labs/imager/tr/2004/findlater04menus/" rel="noopener noreferrer"&gt;Findlater &amp;amp; McGrenere&lt;/a&gt;, &lt;a href="http://aiweb.cs.washington.edu/ai/puirg/papers/kgajos-chi08-predictability.pdf" rel="noopener noreferrer"&gt;Gajos&lt;/a&gt;, and Office 2000's adaptive menus, rolled back into the static Ribbon. Tracking-based personas need &lt;a href="https://www.eff.org/deeplinks/2018/06/gdpr-and-browser-fingerprinting-how-it-changes-game-sneakiest-web-trackers" rel="noopener noreferrer"&gt;cookie-grade consent&lt;/a&gt; in the EU. And per-visitor LLM inference is a cost wall that &lt;a href="https://www.qubit.com/wp-content/uploads/2017/12/qubit-research-meta-analysis.pdf" rel="noopener noreferrer"&gt;Qubit's meta-analysis&lt;/a&gt; (90% of on-site changes move revenue less than ±1.2%) says you can never pay back.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;All three walls fall at the same place: a boundary the model cannot cross, because the system enforces it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The engine, end to end
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fib3ehgfed4eat9gj1r84.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fib3ehgfed4eat9gj1r84.png" alt="UX:LOCK → UI:LOAD, consent gates stage 0" width="681" height="365"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The model sees only two things: the quantized persona and the page's own manifest. The DOM and the visitor stay invisible to it. Its output is ops from the closed vocabulary, validated by a guard. Unknown target, protected zone, unknown variant, blown budget: rejected, with the reason written to an audit log. And the persona is deliberately &lt;em&gt;quantized&lt;/em&gt; into tiny closed vocabularies per axis, because a quantized persona is a &lt;strong&gt;cache key&lt;/strong&gt;, and cache keys are meant to collide:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiucm6xjsjnmkvs9re0bj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiucm6xjsjnmkvs9re0bj.png" alt="One cache lookup decides — LLM calls decay toward zero" width="681" height="416"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A novel persona triggers the model once. The surviving ops are cached as a chain. Every later visitor who quantizes the same, or near enough via deterministic similarity over the keys, replays the chain instantly, with zero model calls. Ship a redesign and the manifest hash changes, so stale chains die by themselves. The LLM's job here is to &lt;em&gt;compile&lt;/em&gt; adaptations, once each, into inspectable, versionable, revertible artifacts. In the shipped engine's 200-session simulation, 72% of page arrivals were served from that cache. The warm path is the majority case, measured.&lt;/p&gt;

&lt;h2&gt;
  
  
  The claim, sized honestly
&lt;/h2&gt;

&lt;p&gt;I'm making a narrow claim. Adaptive pages printing money is a separate question, Qubit's distribution applies to me too, and the engine still lacks an outcome-feedback loop (chains that fail to help should evict themselves). The claim is an existence proof: the salesman-at-a-distance is buildable &lt;em&gt;today&lt;/em&gt;, without creepy tracking, without layout roulette, without per-pageview inference, if you put the boundaries before the capabilities. Making that real, with a full funnel, a live model and two hundred simulated shoppers, is &lt;a href="https://mauriciojuba.com/articles/the-store-that-rearranges-itself" rel="noopener noreferrer"&gt;part 2 of this essay&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Proof by construction:&lt;/strong&gt; the engine is open source at &lt;a href="https://github.com/Type-zero-labs/ui-morph" rel="noopener noreferrer"&gt;github.com/Type-zero-labs/ui-morph&lt;/a&gt;, with one runtime dependency, five ops, 52 tests, a full e-commerce funnel demo on IBM Carbon, and a persona laboratory that replays hundreds of deterministic shoppers.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;References: &lt;a href="https://research.google/blog/generative-ui-a-rich-custom-visual-interactive-user-experience-for-any-prompt/" rel="noopener noreferrer"&gt;Google Research, Generative UI&lt;/a&gt; · &lt;a href="https://jakobnielsenphd.substack.com/p/generative-ui-google" rel="noopener noreferrer"&gt;Nielsen on Google's study&lt;/a&gt; · &lt;a href="https://www.nngroup.com/articles/generative-ui/" rel="noopener noreferrer"&gt;NN/g, Generative UI&lt;/a&gt; · &lt;a href="https://www.cs.ubc.ca/labs/imager/tr/2004/findlater04menus/" rel="noopener noreferrer"&gt;Findlater &amp;amp; McGrenere&lt;/a&gt; · &lt;a href="http://aiweb.cs.washington.edu/ai/puirg/papers/kgajos-chi08-predictability.pdf" rel="noopener noreferrer"&gt;Gajos et al., CHI '08&lt;/a&gt; · &lt;a href="https://www.qubit.com/wp-content/uploads/2017/12/qubit-research-meta-analysis.pdf" rel="noopener noreferrer"&gt;Qubit meta-analysis&lt;/a&gt; · &lt;a href="https://www.eff.org/deeplinks/2018/06/gdpr-and-browser-fingerprinting-how-it-changes-game-sneakiest-web-trackers" rel="noopener noreferrer"&gt;EFF on GDPR &amp;amp; fingerprinting&lt;/a&gt; · &lt;a href="https://github.com/vercel-labs/json-render" rel="noopener noreferrer"&gt;vercel-labs/json-render&lt;/a&gt; · &lt;a href="https://arxiv.org/pdf/2604.16354" rel="noopener noreferrer"&gt;Hidden Technical Debt in GenUI&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>design</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
