A while ago I finished a piece of work I was proud of: a design system built the way I'd argue a design system should be built. Tokens as data. Components as a versioned public API. Everything published to npm and verified in CI. Not perfect, but solid engineering to build on — and an unforgettable professional adventure.
I built it before agents started writing our UI.
Then one day, out of nowhere, I asked myself whether any of it made a difference to what a coding agent produces. Not whether it should. Whether it does.
I had a good answer ready. Of course it does — the agent gets semantic tokens instead of guessing hex values, a typed component API instead of reinventing a table, a written contract instead of scraping my source. I'd said versions of that sentence in several meetings. I believed it.
I just didn't have a single number behind it.
That bothered me more than it should have. I spend my working life asking teams for evidence, and there I was, holding a set of beliefs about the role of design systems in the age of agentic coding.
So I built the thing that would tell me.
What I actually built
Everything here is public — design-system-blueprint — and it's a real system, not a demo shaped like one.
Tokens, three tiers, enforced. Tier 1 is options: the raw palette, the scale, the primitives. Tier 2 is decisions — --ds-decisions-*, the semantic layer, the only tier an application may touch. Tier 3 is component-internal and private, and it changes without a major version. Style Dictionary builds all of it, and the published CSS entry point exposes tier 2 only. The boundary isn't a convention in a wiki; it's what you physically get when you install the package.
The Figma file is generated from the tokens, not the other way round. Values live in tokens.json; the design file's variable collections are produced from it, and a test fails when the export falls behind the source. Before that existed, keeping the two in step was somebody's memory — and an audit found the design file more than a hundred variables behind, a gap that had been widening for months because nothing checked.
A component library that consumes those tokens. Fourteen Angular components — accordion, avatar, breadcrumbs, button, checkbox, dropdown, footer, header, input, list, modal, radio, table, tag — standalone, packaged with ng-packagr, developed and documented in Storybook. Tests across every layer: token contracts, unit, interaction, accessibility, visual regression.
Contracts written for agents, not for people. This is the part I was most curious about. Type declarations that ship inside the tarball, so an agent that installs the package can read the real prop surface without touching my source. AGENTS.md at the repo root for anything working on the system. And llms.client.txt — about 3 KB, shipped inside the npm package — for anything working with it: which selectors exist, that behaviour match beats style match, that tier 2 is the only tier you may reference, and that if no semantic token fits the intent you report no-coverage rather than grabbing the nearest colour.
A token resolver over MCP. @jablonowski/dsb-tokens-mcp — a tool an agent can call to turn an intent ("background for the primary action") into the right decision token, instead of pattern-matching its way there.
Three npm packages, released by GitHub Actions, with verification on every push and pull request. Consuming an update is a dependency bump, not a copy-paste.
I could defend every one of those layers at a whiteboard. That was the problem. Defending something at a whiteboard is not the same as knowing whether it works.
The experiment, briefly
A coding agent is not a compiler. Give it the same prompt twice and you get two different applications, both plausible, both buildable — so one output tells you nothing, and you need repeats that differ in exactly one thing.
You also can't judge it by looking. Models produce competent-looking UI from nothing at all; they'll hand you something that looks deliberate whether or not you gave them a design system. The difference worth measuring is underneath the screenshot: how much CSS the application ended up owning, how many components it reimplemented, how many raw hex values are sitting where a token belonged, and what the run cost.
So: one specification, one application to build — a small Angular dashboard with a login, a dashboard and a users table. Byte-identical prompt every time, a fresh agent session in an empty directory outside any repository, and the same Figma frames served from a recording so every run sees exactly the same bytes.
Only one thing changes between runs: how much of the design system the agent is holding.
| Arm | What the agent has |
|---|---|
| A | the specification and the Figma frames, nothing else |
| A′ | plus a styleguide — palette, scale, component CSS — written out as prose |
| B | plus the npm packages installed; READMEs and type declarations readable |
| C | plus llms.client.txt, the guide written for agents |
| D | plus the token resolver, connected as an MCP server |
Each rung adds exactly one thing, which is what makes the gaps readable. A→B isolates what shipping the system as an installable package buys you. B→C isolates the agent-facing document on top of it. C→D isolates the resolver. And A′ is the uncomfortable one: a design system that has been written down but not shipped — which, let's be honest, is where a lot of organisations actually are. A Figma file, a Confluence page, and a paragraph in the onboarding doc.
Scorers were committed before the first run, and so were the falsifiers — the results that would have told me I was wrong. That matters, because the author of the design system is also the author of the evaluation, and I'm not going to pretend that's a neutral position.
What came out
Thirty scored runs: five arms, five runs each on the primary model, one each on a second model as a sanity check.
Five things I'd say out loud now. The mechanisms are the interesting part and each needs its own post, so here are the shapes.
Shipping the system is a step function. Writing it down is not. No overlap between conditions, none. Without the package, the agent hand-writes twelve or thirteen of the thirteen places where a library component had a natural home, and the application quietly acquires around ninety custom properties it now maintains. With the package: zero, and zero. The prose-styleguide arm — the one that looks most like where most teams actually are — didn't land anywhere near it.
The most useful thing the system does turned out to be accessibility, not consistency. The agent reads the palette out of Figma and gets the colours right. It then ships serious contrast failures, repeatably, on both models. Values are one thing; which value belongs on which surface is another, and a design file does not carry the second one. This is the result I did not expect and the one I'd defend hardest.
Agents did not hallucinate my component API. Not once in thirty runs — including the arms with no library at all. The industry's favourite scare story is not what I observed. An agent doesn't invent a component whose existence it doesn't suspect; it writes its own, and nothing in the diff looks like an error. That's the worse failure mode, because it's invisible in review.
It costs more, not less. I assumed a library would be cheaper per generation — the agent stops reinventing the wheel. The opposite is true. Using a library means reading it, and everything read stays in context and is paid for again on every turn. Any payoff has to come from rework, review and maintenance, which I haven't measured yet and won't claim until I do.
And the layer I was proudest of earned nothing measurable. The MCP resolver, compared against the plain text file that precedes it, added no improvement I can detect. That comparison was pre-registered as the decisive one, which is exactly why it's published with the same prominence as the results I liked.
If I compress all of it into one sentence: the design system does not make the agent smarter or more correct. It decides who ends up owning the code it writes — and it carries the decisions nobody wrote down anywhere the agent could read them.
All of that is one application, one framework, one model vendor: a case study with committed scorers, not a benchmark. Two measurements I want — drift, and cost per accepted screen — haven't run yet. Every per-run record, every score with its detail, every generated application and a protocol logging each rule I changed mid-study is in demo-blueprint, including the change that nearly reversed one of the conclusions above before I caught it.
Over to you
Let me know what you think. And if there's a measurement that would change your mind about design systems and I didn't take it, say so — some are still cheap to add.
I'm an engineering manager, eighteen years in IT. I build the things I argue for, then measure whether the argument survives. This one did — in a different shape than I expected.
Top comments (0)