DEV Community

Cover image for Tests Prove Behavior. Boundaries Prove Architecture.
qnbs
qnbs

Posted on AI-assisted

Tests Prove Behavior. Boundaries Prove Architecture.

A reader of my article on keeping AI an optional capability asked the best question of the whole comment thread: how do you keep an architectural seam from eroding as features accumulate — do you enforce it with tests, or with another kind of guardrail?

The honest answer starts with admitting the question contains two different propositions. "The code works" and "the architecture holds" are not the same claim, and they do not admit the same proof. A test suite can be fully green while the architecture quietly rots — tests sample behavior, and behavior can stay correct for months after the structure that was supposed to constrain it has been bypassed. In WorldScript Studio, the open-source writing app this series dissects, one seam is protected by both layers and another — candidly — by only one. The contrast is the lesson. Code references are from the repository at commit 8b329633 (2026-09-28), release v1.28.8.

1. What a working boundary looks like

The project is a web app that also ships as a Tauri desktop app. The architectural rule: application code must reach desktop capabilities through a desktopPlatform contract, so a future Qt or GPUI adapter slots in without touching consumers. This rule is not documented and hoped for — it is enforced by a CI step running a small checker with three design decisions worth stealing.

First, zero tolerance. Any real @tauri-apps/* import specifier in application source outside the approved locations fails the build. The script's header explicitly contrasts this with the project's suppression ratchet elsewhere: boundaries don't ratchet, because a ratcheted boundary is a schedule for losing it.

Second, parse real imports, not text. Whole-line comments are masked before scanning, so a JSDoc mention of @tauri-apps/api doesn't fail the build — only actual import, import(), and require() specifiers count. The header even documents the residual edge case (a block comment placed mid-line on real code is not masked and would flag) — the maintainers know exactly where their parser is conservative, and chose the direction that fails loud rather than the one that fails silent.

Third, the allowlist is the documentation. Approved locations plus a short list of exceptions, and every exception carries a written reason — two files are marked PERMANENT because the HTTP facet is deliberately not part of the platform contract, with a pointer to the roadmap section. An exception without a reason doesn't exist. Reviewing the boundary means reviewing a twenty-line allowlist, not re-auditing the codebase — and the gate runs as an ordinary CI step, so the review happens on every pull request for free.

2. What a strong behavioral layer looks like

The project's AI seam has no such gate. What it has instead is a behavioral layer that is genuinely strong — worth describing precisely, because "we have tests" is usually where the erosion conversation gets sloppy.

Application features reach AI only through a unified provider service; the factory that constructs provider clients is imported by nothing outside the provider layer — verifiable in the import graph in one grep. The factory's mapping table is fail-closed: providers nobody has vetted map to unsupported, and one test pins that a superficially OpenAI-compatible provider gets its own kind instead of being folded into the generic path. Policy gates (cloud-allowed, local-only) and the null-returning fallback registry each have their own suites — in total over two hundred cases across service, factory, policy, outbound-request shape, and fallback semantics.

This is real enforcement. It is also exclusively behavioral: it proves what happens when code uses the seam. It says nothing about code that routes around it.

3. The leak that proves the point

Here is where I get to be honest in the way the reader's question deserves. Nothing mechanically stops a new file from importing a vendor SDK directly. Today the inventory is small and known — six runtime files, four of them deliberate services-layer surfaces. The other two are the interesting ones. One is a feature thunk that imports Type from @google/genai — not to call a provider, but to use Gemini's schema vocabulary (Type.OBJECT, Type.STRING) when declaring the shape of a structured response; the actual generation request still goes through the shared service, through the policy gates, through everything. The other is a React hook that imports useCompletion from @ai-sdk/react — but points it at a virtual worldscript-internal://completion URL whose transport is the seam itself. Neither calls a provider directly. No bypass, no bug.

And yet the thunk is exactly the thing a boundary would flag. Vendor vocabulary has leaked into feature code, which means the next feature author who copies that thunk as a template inherits a Gemini-shaped assumption in code that is supposed to be provider-agnostic. Today the list of such files is policed by convention and code review. Conventions are how "six known files" becomes "ten, probably" over a year.

4. Building the second layer

The fix is not a new idea — it is the pattern transfer of section 1, and I want to be explicit that this is a recommendation, not scheduled work:

  1. Write the boundary as an allowlist of reality. Every current vendor-SDK import gets an entry with a reason — sanctioned layer, documented services exception, "client surface wired to the internal completion URL" for the hook, or, for the thunk, "known leak, candidate for re-export through the provider layer."
  2. Parse real specifiers only. Mask whole-line comments; scan for import/import()/require(); choose fail-loud on parser edge cases.
  3. Zero tolerance from day one. New vendor import outside the allowlist fails the build. No ratchet — ratchets are for suppressions, not walls.
  4. Wire it into the cheapest CI step. Seconds, not minutes, and before anything expensive runs.
  5. Review the allowlist, not the codebase. The allowlist diff in a PR is the architectural review.

5. A checklist for your own seams

  1. Name your seams — the places where a vendor, a platform, or a subsystem is supposed to be swappable.
  2. For each, ask which layer you have: behavioral (tests), structural (a gate), or both.
  3. Treat "tests + convention" as a documented gap, not a solved problem — and say so in the docs.
  4. When you add the gate, make the allowlist the documentation: every exception with a reason, or it doesn't merge.
  5. Keep the gate cheap and zero-tolerance; a slow or forgiving boundary gets routed around socially long before it gets routed around technically.

Tests prove what happens when code uses the seam. A boundary proves that new code cannot route around it. The green checkmark is the floor. The wall is what keeps it meaning something next year.


This article describes WorldScript Studio at commit 8b329633 (release v1.28.8); simplified excerpts are labeled. Part of the series "Engineering WorldScript Studio." Written with AI assistance; all technical claims verified against the linked source.

Top comments (0)