<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Madhesh Vivekanandan</title>
    <description>The latest articles on DEV Community by Madhesh Vivekanandan (@madmi).</description>
    <link>https://dev.to/madmi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4113063%2F9e637c05-e291-4b3f-9072-6426665e3dc6.jpg</url>
      <title>DEV Community: Madhesh Vivekanandan</title>
      <link>https://dev.to/madmi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/madmi"/>
    <language>en</language>
    <item>
      <title>The Coding Agent That Can't Grade Its Own Homework</title>
      <dc:creator>Madhesh Vivekanandan</dc:creator>
      <pubDate>Wed, 09 Sep 2026 15:05:27 +0000</pubDate>
      <link>https://dev.to/madmi/the-coding-agent-that-cant-grade-its-own-homework-1646</link>
      <guid>https://dev.to/madmi/the-coding-agent-that-cant-grade-its-own-homework-1646</guid>
      <description>&lt;p&gt;"All tests pass. ✅"&lt;/p&gt;

&lt;p&gt;You've seen that message from a coding agent. And at least once, you've opened the file afterwards and found the tests were never run — or the test file was quietly edited until it agreed.&lt;/p&gt;

&lt;p&gt;The problem isn't that the model writes bad code. It's that the same context window writes the code, reviews the code, and declares the code correct. One brain, grading its own homework — and it grades generously. That cycle — understand → change → check → repeat — is the &lt;em&gt;agent loop&lt;/em&gt;, and every coding tool ships one that runs inside a single context window.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Madheshvivekanandan/agent-loop" rel="noopener noreferrer"&gt;agent-loop&lt;/a&gt; is a small open-source project that restructures it: six explicit stages, file-based handoffs, an independent verifier, hard iteration caps. No framework, no daemon, no SDK — markdown plus one shell script, and it runs unchanged in Claude Code, Codex, Cursor, Gemini CLI, Copilot, and 30+ other tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  The whole loop in 30 seconds
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn3pazlbea0ajkqyikgw3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn3pazlbea0ajkqyikgw3.png" alt="The agent loop: triage fans out into tiers S, M, and L — S goes straight to implement, M passes through analyze, L through analyze and plan. Everything converges on implement, then a verify gate: PASS reports done, FAIL enters a debug loop capped at 3 iterations, after which it escalates." width="800" height="1200"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Each stage starts with a &lt;strong&gt;fresh context&lt;/strong&gt;, reads only its declared inputs, and hands off through a markdown file in a gitignored &lt;code&gt;.agent-loop/&lt;/code&gt; directory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.agent-loop/
  profile.md                    # cached project profile (commands, conventions)
  runs/2026-09-09-rate-limit/
    task.md
    analysis.md
    plan.md
    implementation.md
    test-report.md              # ends in PASS or FAIL — never "looks good"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those files aren't logs. They're the &lt;em&gt;interface&lt;/em&gt; between stages — and an audit trail you can read afterwards to see exactly why the agent did what it did.&lt;/p&gt;

&lt;h2&gt;
  
  
  Not every task deserves a pipeline
&lt;/h2&gt;

&lt;p&gt;Anthropic measured multi-agent systems at roughly &lt;a href="https://www.anthropic.com/engineering/multi-agent-research-system" rel="noopener noreferrer"&gt;&lt;strong&gt;15× the tokens&lt;/strong&gt; of a regular chat&lt;/a&gt;. For a typo fix, that's a five-course tasting menu because you wanted a glass of water.&lt;/p&gt;

&lt;p&gt;So the first move is triage — one cheap classification, never delegated:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;The task&lt;/th&gt;
&lt;th&gt;What runs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;S&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Diff fits in one sentence&lt;/td&gt;
&lt;td&gt;implement → verify&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;M&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Localized change&lt;/td&gt;
&lt;td&gt;analyze → implement → verify&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;L&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Multi-file or uncertain&lt;/td&gt;
&lt;td&gt;the full pipeline&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One rule does more work than it looks like: &lt;strong&gt;torn between two tiers, pick the smaller one.&lt;/strong&gt; Left alone, agents will write a design document for a null check — the bias has to be encoded.&lt;/p&gt;

&lt;h2&gt;
  
  
  Artifacts, not transcripts
&lt;/h2&gt;

&lt;p&gt;Why fresh contexts? Because context poisoning is real. Exploring a codebase fills the window with dead ends — irrelevant files, abandoned hypotheses, stale test logs — and in one long session, all that noise quietly influences every later decision.&lt;/p&gt;

&lt;p&gt;File handoff cuts the cord. The analyzer can burn its whole budget exploring, but only &lt;code&gt;analysis.md&lt;/code&gt; survives into the next stage. What crosses the boundary is a document, not a transcript.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3d53iipe2cyav3favx1l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3d53iipe2cyav3favx1l.png" alt="Left: one agent overwhelmed by a single long, tangled, scribbled-over transcript. Right: four agents cleanly handing small, focused documents from stage to stage." width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One rule decides whether handoffs work at all:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The plan must carry rationale, not just conclusions.&lt;/strong&gt; An implementer that gets only "use middleware, put it in &lt;code&gt;upload.py&lt;/code&gt;" re-derives the reasoning, reaches a different answer, and contradicts the plan while believing it follows it. (Cognition's &lt;a href="https://cognition.com/blog/dont-build-multi-agents" rel="noopener noreferrer"&gt;"Don't Build Multi-Agents"&lt;/a&gt; circles the same failure.)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The templates enforce it structurally: &lt;code&gt;plan.md&lt;/code&gt; has required sections, and a plan without a named &lt;strong&gt;runnable check&lt;/strong&gt; — the command that will decide pass or fail — is returned once, then the loop stops. No check, no run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The verifier never sees the reasoning
&lt;/h2&gt;

&lt;p&gt;This is the part of the design I'd defend in a fight. Verification starts in a completely fresh context, and its inputs are a closed list:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;The verifier gets&lt;/th&gt;
&lt;th&gt;It never gets&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;The plan&lt;/td&gt;
&lt;td&gt;The implementer's summary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The diff&lt;/td&gt;
&lt;td&gt;The implementer's reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A shell, to run the check&lt;/td&gt;
&lt;td&gt;&lt;code&gt;implementation.md&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Why so strict? An agent that reads "here's why my implementation is correct…" doesn't verify that claim — it &lt;em&gt;confirms&lt;/em&gt; it. Sycophancy isn't a personality flaw you can prompt away; the fix is structural. The verifier learns what the code was &lt;em&gt;supposed&lt;/em&gt; to do, sees what actually &lt;em&gt;changed&lt;/em&gt;, runs the check named in the plan, and writes &lt;code&gt;PASS&lt;/code&gt; or &lt;code&gt;FAIL&lt;/code&gt; with command output as evidence. "Looks good to me" is not a verdict.&lt;/p&gt;

&lt;p&gt;And my favorite detail in the whole project: you might think &lt;em&gt;spawn the verifier without write tools and it can't cheat&lt;/em&gt;. But it must have a shell to run the tests — and a shell can edit files (&lt;code&gt;sed&lt;/code&gt;, &lt;code&gt;echo &amp;gt;&lt;/code&gt;, &lt;code&gt;git apply&lt;/code&gt;…). So the real guarantee is a &lt;strong&gt;diff fingerprint&lt;/strong&gt;: the working tree is fingerprinted before and after verification, and if it moved, the verdict is void. Tool restriction is defence in depth; detection is the guarantee — on every host, including those that can't restrict tools at all.&lt;/p&gt;

&lt;p&gt;(Reviewers fail in the other direction too, so the contract counts only &lt;em&gt;correctness-affecting&lt;/em&gt; gaps — no invented nitpicks.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Three strikes, then a human
&lt;/h2&gt;

&lt;p&gt;On FAIL, a debugger spins up — fresh context again, so failed attempts don't pile up as noise that poisons the next one. It must &lt;em&gt;reproduce&lt;/em&gt; the failure before fixing it, and if it concludes the &lt;em&gt;plan&lt;/em&gt; is wrong, that's an escalation, not a code change.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The debug loop caps at 3 iterations&lt;/strong&gt; — with oscillation detection, and a full re-verification after every fix. On cap, you get a distilled summary of what was tried and what the evidence says, not a 400-line transcript.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;An agent loop without a termination condition isn't autonomous. It's just unsupervised.&lt;/p&gt;

&lt;h2&gt;
  
  
  One folder, thirty-plus tools
&lt;/h2&gt;

&lt;p&gt;The loop is a single &lt;a href="https://agentskills.io" rel="noopener noreferrer"&gt;Agent Skills&lt;/a&gt; folder — an open standard 30+ coding agents read — and the stage contracts name no vendor, tool, or product. Hosts differ wildly in what they offer, so every capability declares an explicit fallback:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;With it&lt;/th&gt;
&lt;th&gt;Without it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Subagents&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Mode A&lt;/strong&gt; — one isolated context per stage&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Mode B&lt;/strong&gt; — sequential stages, still reading only declared inputs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool restriction&lt;/td&gt;
&lt;td&gt;Verifier spawned without write tools&lt;/td&gt;
&lt;td&gt;The diff fingerprint alone — the real guarantee anyway&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Turn caps&lt;/td&gt;
&lt;td&gt;Debugger bounded mechanically&lt;/td&gt;
&lt;td&gt;Iteration counting; the 3-cap holds either way&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Per-stage models&lt;/td&gt;
&lt;td&gt;Strong models on plan/verify/debug&lt;/td&gt;
&lt;td&gt;One model throughout — costs more, quality holds&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every run announces which mode it got. One requirement is non-negotiable: &lt;strong&gt;shell execution&lt;/strong&gt;. Verification that can't run real commands gives the loop no termination condition — so it refuses to run rather than emit confident, unverified output. That refusal &lt;em&gt;is&lt;/em&gt; the product.&lt;/p&gt;

&lt;p&gt;And because everything is markdown plus one POSIX installer — no daemon, no database — there's nothing racing the features agent hosts ship natively every month, and the skill costs &lt;strong&gt;zero context&lt;/strong&gt; until you explicitly invoke it.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what does this actually improve?
&lt;/h2&gt;

&lt;p&gt;Each failure you've met on real work maps to a specific structural counter — that mapping &lt;em&gt;is&lt;/em&gt; the design:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Failure you've seen&lt;/th&gt;
&lt;th&gt;What the loop does about it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Bad early output contaminates everything after it&lt;/td&gt;
&lt;td&gt;Fresh context per stage; artifacts cross the boundary, transcripts don't&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"All tests pass" that nobody ran&lt;/td&gt;
&lt;td&gt;Independent verifier, diff fingerprint, evidence required in the verdict&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The agent grinds on a fix forever&lt;/td&gt;
&lt;td&gt;Cap of 3, oscillation detection, escalation with a failure summary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;It rebuilds a helper that already exists&lt;/td&gt;
&lt;td&gt;The analyzer's mandatory Found / Exemplars / Missing / Reuse-plan gate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A design doc for a one-line fix&lt;/td&gt;
&lt;td&gt;Triage tiers, biased toward the smaller tier&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The implementer "follows the plan" into a different design&lt;/td&gt;
&lt;td&gt;Plans carry decisions &lt;em&gt;and&lt;/em&gt; rationale, enforced by template sections&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;None of this makes the model smarter. It makes the &lt;em&gt;process&lt;/em&gt; resistant to the specific ways coding agents fail.&lt;/p&gt;

&lt;h2&gt;
  
  
  Steal this, then run it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/Madheshvivekanandan/agent-loop &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;agent-loop
./install.sh claude-code     &lt;span class="c"&gt;# or: codex · cursor · agents&lt;/span&gt;
&lt;span class="c"&gt;# Claude Code, zero-install alternative:  claude --plugin-dir path/to/agent-loop&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then, in any project:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/agent-loop add rate limiting to the upload endpoint
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;First run profiles your project — it discovers your test/build/lint commands and &lt;em&gt;executes them once before trusting them&lt;/em&gt;. Every run after that states its tier and mode up front, works the stages, and returns a verdict, a diff summary, and a folder of evidence you can actually read.&lt;/p&gt;

&lt;p&gt;If you take nothing else from this post, take the two rules that travel to any agent workflow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Never let the agent that wrote the code declare it correct.&lt;/strong&gt; The separation must be structural — fresh context, plan + diff only — not a prompt asking it to "be critical."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every loop needs a number where it stops and calls a human.&lt;/strong&gt; Caps aren't a lack of faith in the model; they're what makes it safe to stop watching.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The repo is &lt;a href="https://github.com/Madheshvivekanandan/agent-loop" rel="noopener noreferrer"&gt;github.com/Madheshvivekanandan/agent-loop&lt;/a&gt; — MIT-licensed, one folder, readable in an evening. If you run it on a host I haven't tested, or find a hole in the verifier's guarantees, I genuinely want the issue.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Previously: &lt;a href="https://dev.to/madmi/the-dashboard-that-builds-itself-3386"&gt;The Dashboard That Builds Itself&lt;/a&gt;, on generative UI and letting a model order from a menu without getting into the kitchen.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;All images in this post were generated with ChatGPT.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>The dashboard that builds itself</title>
      <dc:creator>Madhesh Vivekanandan</dc:creator>
      <pubDate>Mon, 07 Sep 2026 04:41:51 +0000</pubDate>
      <link>https://dev.to/madmi/the-dashboard-that-builds-itself-3386</link>
      <guid>https://dev.to/madmi/the-dashboard-that-builds-itself-3386</guid>
      <description>&lt;p&gt;Every dashboard you've ever used was designed long before you opened it. Someone guessed which charts you would need, arranged them on a grid, and shipped that guess to everyone. This post is about a dashboard that skips the guess: &lt;strong&gt;it doesn't exist until you ask it a question.&lt;/strong&gt; You type &lt;em&gt;"how did Q3 go?"&lt;/em&gt; and the answer comes back as a screen, not a paragraph — stat cards and a chart, building on screen while the AI is still thinking.&lt;/p&gt;

&lt;p&gt;I built it with three plain parts: OpenAI's structured outputs (a mode where the reply is guaranteed to match a shape you define), a small Python backend (FastAPI), and Google's open A2UI protocol (a standard way for an AI to describe a UI as plain data), rendered by &lt;code&gt;@a2ui/react&lt;/code&gt;, the A2UI project's official React renderer. This post explains how the pieces fit and the five safety nets that make "an AI draws my UI" boring instead of scary. The whole app — backend, frontend, tests and config — is about 4,300 lines, and &lt;a href="https://github.com/Madheshvivekanandan/Generative-Ui" rel="noopener noreferrer"&gt;all of it is on GitHub&lt;/a&gt;.&lt;/p&gt;



&lt;p&gt;&lt;em&gt;Two questions and a button press, recorded from the running app — every answer re-forms the same canvas.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Three ways to do generative UI (and where A2UI sits)
&lt;/h2&gt;

&lt;p&gt;Nielsen Norman Group &lt;a href="https://www.nngroup.com/articles/generative-ui/" rel="noopener noreferrer"&gt;defines generative UI&lt;/a&gt; as "a user interface that is dynamically generated in real time by artificial intelligence to provide an experience customized to fit the user's needs and context." In this app, that context is the question you just asked. There are three ways to build it, and they differ in one thing: how much you let the model do.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Let the AI write real code&lt;/strong&gt; — HTML, CSS, JavaScript. Google &lt;a href="https://research.google/blog/generative-ui-a-rich-custom-visual-interactive-user-experience-for-any-prompt/" rel="noopener noreferrer"&gt;ships this in Gemini&lt;/a&gt; and the results are impressive. But in your own product it means running AI-written code in your users' browsers — a malicious input can trick the model into writing a script that then runs on someone else's screen: the classic cross-site scripting (XSS) attack, delivered by the AI itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Let the AI pick from a menu.&lt;/strong&gt; You build the components; the model only sends data saying &lt;em&gt;which&lt;/em&gt; ones, with &lt;em&gt;what&lt;/em&gt; values. The model orders from a menu — it never gets into the kitchen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The same menu idea, with a shared protocol&lt;/strong&gt; — so any client that speaks the protocol can draw the output of any agent that speaks it. That's &lt;a href="https://a2ui.org" rel="noopener noreferrer"&gt;A2UI&lt;/a&gt;, announced by Google in December 2025. An agent describes an interface as a stream of small JSON messages, and a renderer draws them using only components from a catalog &lt;em&gt;the client&lt;/em&gt; defines. UI travels as data, never as code.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The &lt;a href="https://a2ui.org/introduction/what-is-a2ui/" rel="noopener noreferrer"&gt;official docs&lt;/a&gt; cover the protocol well, and I won't repeat them. Most A2UI tutorials run on Google's own tools — their Agent Development Kit (ADK) and Gemini models — and the exceptions go through large agent frameworks (Microsoft's Agent Framework pairs A2UI with FastAPI and an OpenAI client; A2UI's own docs show LangGraph and Strands setups). What I couldn't find is a write-up that wires A2UI directly to the plain OpenAI API and FastAPI, with no agent framework in between. That's this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  The whole system in 30 seconds
&lt;/h2&gt;



&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    U(["You"]) --&amp;gt;|"how did Q3 go?"| B["Browser · @a2ui/react"]
    B --&amp;gt;|"your question"| F["FastAPI"]
    F --&amp;gt;|"schema + data"| L["OpenAI"]
    L --&amp;gt;|"typed blocks, streaming"| F
    F --&amp;gt;|"A2UI messages, card by card"| B
    B --&amp;gt;|"button press"| F&lt;/code&gt;&lt;/pre&gt;



&lt;ol&gt;
&lt;li&gt;The browser sends your question to the backend.&lt;/li&gt;
&lt;li&gt;The backend asks OpenAI (&lt;code&gt;gpt-4o-mini&lt;/code&gt; in this build) for a &lt;strong&gt;plan&lt;/strong&gt;, not code: one to four "blocks" chosen from six types (stat card, line chart, bar chart, donut, table, text note). The plan is forced through a schema, so it always has the right shape. The numbers come from a fixed demo dataset that the backend includes in the prompt, so the model copies and combines real figures rather than inventing them.&lt;/li&gt;
&lt;li&gt;A small compiler turns each finished block into A2UI messages.&lt;/li&gt;
&lt;li&gt;Those messages stream to the browser using server-sent events, or SSE (one long-lived reply that the server keeps adding to), and &lt;code&gt;@a2ui/react&lt;/code&gt; draws them. Cards appear one by one, while the model is still writing.&lt;/li&gt;
&lt;li&gt;Cards carry buttons the model invented. Pressing one sends the request back, and the loop runs again.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And one deliberate rule ties it together: &lt;strong&gt;every answer is drawn into the same single canvas.&lt;/strong&gt; Before a new answer streams in, the old one is cleared — so the dashboard is &lt;em&gt;regenerated in place&lt;/em&gt;, never stacked. Ask a different question, get a different dashboard, same spot. The ✕ button brings back the starting view.&lt;/p&gt;

&lt;h2&gt;
  
  
  The menu, not the kitchen
&lt;/h2&gt;

&lt;p&gt;The most important design decision sounds like a technicality: &lt;strong&gt;the model never speaks A2UI.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The tutorials I found — including the official A2UI agent guide — do the opposite: paste the protocol schema into the prompt, ask the model to write A2UI JSON directly, then check it afterwards. The format A2UI actually sends over the network is a flat list of components that refer to each other by ID. Great shape for a renderer; terrible shape for a model to write reliably — nothing stops it from referencing an ID that doesn't exist or forgetting a field.&lt;/p&gt;

&lt;p&gt;So there are two languages, on purpose:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Who reads it&lt;/th&gt;
&lt;th&gt;What it is&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Typed blocks&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;the model&lt;/td&gt;
&lt;td&gt;six block types with required, typed fields&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;A2UI messages&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;the browser&lt;/td&gt;
&lt;td&gt;the real protocol, produced by a compiler&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The typed blocks are the schema I hand to OpenAI. Give the API a shape and it cannot produce JSON in any other shape, because the rule is applied &lt;em&gt;while&lt;/em&gt; the text is being written rather than checked afterwards. The model cannot invent a seventh component type or skip a required field. Not "usually doesn't" — cannot. (One caveat: a guaranteed shape can still hold wrong content — the schema stops invalid structure, not bad answers. And this app keeps its count limits, like "at most four blocks," in one line of ordinary code rather than in the schema.)&lt;/p&gt;

&lt;p&gt;The cost is real: the agent can only compose layouts my compiler knows how to build. For a dashboard, that's the right trade. For a general-purpose agent canvas, it wouldn't be.&lt;/p&gt;

&lt;p&gt;It paid off in a way you can see in the code. Version 1 drew the UI with a hand-rolled renderer — three hundred lines that looked up each card type in a table and drew it by hand. Migrating to the protocol deleted that file, and the frontend no longer contains a single line that asks what kind of card something is; the only place a card type is named is one small lookup table in the backend compiler.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watching an answer build itself
&lt;/h2&gt;

&lt;p&gt;The reply streams in as it's written, which raises a fun question: when is a card inside a half-written answer safe to show?&lt;/p&gt;

&lt;p&gt;The rule in this codebase: &lt;strong&gt;a card can only be trusted once the model has moved on to the next one.&lt;/strong&gt; So the server sends each finished card immediately, and always holds back the last one it can see:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# A block is only provably finished once the next one has started,
# so the last one in the snapshot is always left to the caller.
&lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;emitted&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;raw_blocks&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each card shows the moment it's provably finished. Then, when the whole answer is done, the server sends the complete, fully-checked version once more, and it quietly replaces what is on screen. That works because an update overwrites whatever already sits at the same address — same data path, same component IDs — instead of adding a copy next to it. Any rushed mistake is corrected the moment the answer finishes.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;sequenceDiagram
    participant B as Browser
    participant F as FastAPI
    participant O as OpenAI
    B-&amp;gt;&amp;gt;F: your question
    F--&amp;gt;&amp;gt;B: open the canvas
    F-&amp;gt;&amp;gt;O: ask for typed blocks (schema attached)
    O--&amp;gt;&amp;gt;F: …card 1 finished
    F--&amp;gt;&amp;gt;B: card 1 data + layout
    Note over B: first card renders
    O--&amp;gt;&amp;gt;F: …card 2 finished
    F--&amp;gt;&amp;gt;B: card 2 data + layout
    O--&amp;gt;&amp;gt;F: model finishes
    F--&amp;gt;&amp;gt;B: final complete answer (verified)
    F--&amp;gt;&amp;gt;B: chat explanation + suggestions · done&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Here is what one message looks like on its way to the browser. This is all the "protocol" really is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"v0.9"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"updateDataModel"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"surfaceId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"dashboard"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;   &lt;/span&gt;&lt;span class="err"&gt;//&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;the&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;one&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;canvas&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"path"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"/blocks/0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Q3 Revenue"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;128400&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
               &lt;/span&gt;&lt;span class="nl"&gt;"unit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"USD"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"delta_pct"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;12.4&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cards don't hold values — they hold pointers into this data tree, so updating the data is all it takes to update the screen. (&lt;a href="https://madheshvivekanandan.github.io/Generative-Ui/media/wire-capture.txt" rel="noopener noreferrer"&gt;The full capture is here&lt;/a&gt; — a real turn, straight off the wire.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Buttons that talk back
&lt;/h2&gt;

&lt;p&gt;Without buttons, every answer is a dead end: you read the chart, and then you are on your own to work out what to ask next. The model already knows what is worth asking. So every card carries up to two buttons &lt;strong&gt;the model wrote itself&lt;/strong&gt;, and pressing one talks back to the agent.&lt;/p&gt;

&lt;p&gt;All buttons share one action name, &lt;code&gt;refine&lt;/code&gt;. The interesting part travels inside the button: as it writes the card, the model puts a complete follow-up request inside the button — label &lt;em&gt;"Split by region"&lt;/em&gt;, request &lt;em&gt;"Break Q3 revenue down by region as a bar chart."&lt;/em&gt; The backend just pulls that request out and runs the same loop again. One handler covers every button the model will ever dream up — and pressing one feels like the dashboard changing its mind.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx7o02dtre8liwzhwap5l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx7o02dtre8liwzhwap5l.png" alt="The same canvas after pressing a model-written button: a new answer has replaced the previous one, while the chat log keeps the whole conversation" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;After a button press — the same canvas, re-formed. The chat keeps the conversation; the dashboard keeps only the current answer.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  "Why would you trust an AI to draw your screen?"
&lt;/h2&gt;

&lt;p&gt;When A2UI hit Hacker News, &lt;a href="https://news.ycombinator.com/item?id=46286407" rel="noopener noreferrer"&gt;one blunt reaction&lt;/a&gt; captured a recurring worry in the thread: &lt;em&gt;"Why on earth would you trust an LLM to output a UI?"&lt;/em&gt; It deserves an architectural answer. Mine is five nets, each catching what the previous one can't:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Structured outputs&lt;/strong&gt; — the schema is enforced during generation. Wrong shapes can't exist.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A sanitizer&lt;/strong&gt; — a valid shape can still be nonsense. This layer trims oversized data (at most four chart series, 25 table rows, and so on) and rejects anything that cannot be drawn at all — a donut with one slice is not a chart, it is a circle. A bad donut next to a good chart costs you the donut, not the whole answer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The catalog whitelist&lt;/strong&gt; — the renderer refuses any component name not in the catalog, enforced by the library itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Forgiving readers&lt;/strong&gt; — every field a view renders passes through a tiny reader that turns bad input into a safe fallback instead of a crash: a missing number becomes a dash, a chart with no usable data becomes an empty state. Agent-written data is the &lt;em&gt;expected&lt;/em&gt; input.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An error boundary per card&lt;/strong&gt; — if a chart does break while being drawn, it takes down its own card and nothing else.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;On top of all that sits one promise: the generate endpoint &lt;strong&gt;never returns an error.&lt;/strong&gt; A missing key, a refusal from the model, a network failure — each one comes back as an ordinary text card, drawn through the same path as real content. That leaves the renderer with a single path to follow and no error branch at all.&lt;/p&gt;

&lt;p&gt;That promise has a story behind it. A code review on day one caught the fallback card printing raw exceptions — and an OpenAI auth error contains part of your API key. The app was painting key material into a dashboard card. The fix: stable error codes on the wire, generic copy on screen, details only in the server log. If you draw errors from anywhere near the model onto the screen, you will eventually draw a secret onto it too. Decide the boundary on day one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffabeoenodspyeflyr27n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffabeoenodspyeflyr27n.png" alt="A generated answer: two stat cards with model-written buttons above a line chart, with the explanation and follow-up suggestions in the chat panel" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;A real generated answer — every element on screen was ordered from the menu, including the buttons the model wrote itself.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Steal this, then run it
&lt;/h2&gt;

&lt;p&gt;What I'd tell you before you build one:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Don't make the model speak the protocol.&lt;/strong&gt; Let it fill in a strict, friendly schema; compile to the protocol in code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Show partial results fast, then re-send the complete checked answer&lt;/strong&gt; so it replaces anything the quick version got wrong. That final replace is what makes streaming safe to attempt.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Layer your defenses and expect each one to fire.&lt;/strong&gt; Schema, sanitizer, whitelist, forgiving readers, boundary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat the failure card as normal output.&lt;/strong&gt; One render path for success and failure means the error path is tested on every fallback.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test that streaming streams.&lt;/strong&gt; The worst thing a broken streaming system can do is keep working. Assert the first update arrives before the stream closes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build the starting view the same way.&lt;/strong&gt; The dashboard you see on load is compiled by the same code the agent uses, so bugs show up before you type anything.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;What's missing, honestly: no auth, no database, mock data, and A2UI v1.0 will eventually mean some rework.&lt;/p&gt;

&lt;p&gt;The whole thing runs with &lt;code&gt;docker compose up&lt;/code&gt; and an OpenAI key — without a key it still boots and shows the baseline. &lt;a href="https://github.com/Madheshvivekanandan/Generative-Ui" rel="noopener noreferrer"&gt;The repo is here&lt;/a&gt; — clone it, type something odd into it, and tell me what breaks.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>a2ui</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
