<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alex</title>
    <description>The latest articles on DEV Community by Alex (@ake2l).</description>
    <link>https://dev.to/ake2l</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4125375%2F035c43c6-06b8-4f6d-b7c8-8efae9e33968.jpg</url>
      <title>DEV Community: Alex</title>
      <link>https://dev.to/ake2l</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ake2l"/>
    <language>en</language>
    <item>
      <title>ArchKeel After a 121-File Refactoring Experiment</title>
      <dc:creator>Alex</dc:creator>
      <pubDate>Tue, 22 Sep 2026 04:36:24 +0000</pubDate>
      <link>https://dev.to/ake2l/archkeel-after-a-121-file-refactoring-experiment-5bb9</link>
      <guid>https://dev.to/ake2l/archkeel-after-a-121-file-refactoring-experiment-5bb9</guid>
      <description>&lt;p&gt;Before the first coding agent touched DATAMIMIC CE, I froze the target architecture.&lt;/p&gt;

&lt;p&gt;The current code disagreed with it in 613 places.&lt;/p&gt;

&lt;p&gt;Six and a half hours later, the violation baseline was empty.&lt;/p&gt;

&lt;p&gt;Then review found a public API regression and a gate script that could not fail.&lt;/p&gt;

&lt;p&gt;This was not another test of whether ArchKeel can print architecture violations. I covered the idea behind the deterministic gate in &lt;a href="https://dev.to/ake2l/tests-green-architecture-worse-a-deterministic-gate-for-coding-agents-4jhi"&gt;the first article&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This experiment asked a different question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can a target architecture written first as a deterministic contract drive a multi-step refactoring by coding agents without silently moving the target?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I wrote that hypothesis before the target contract and before any refactoring step. The order was part of the experiment.&lt;/p&gt;

&lt;h2&gt;
  
  
  The protocol
&lt;/h2&gt;

&lt;p&gt;The codebase was DATAMIMIC CE at a fixed commit. The checker was ArchKeel 0.5.1. One agent orchestrated the run, ten implementation-agent runs changed the code, and one separate agent advised on the design. Everything stayed local. No push and no CI.&lt;/p&gt;

&lt;p&gt;Each accepted step had to pass four gates:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The target did not widen.&lt;/li&gt;
&lt;li&gt;The known architecture debt only shrank.&lt;/li&gt;
&lt;li&gt;Existing DSL descriptors kept their observable behaviour.&lt;/li&gt;
&lt;li&gt;The relevant lint, type and test results stayed at least at the starting state.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The implementing agents could change code. They could not edit the architecture contract or add entries to the known-violation baseline.&lt;/p&gt;

&lt;p&gt;A new violation meant the change was wrong. An unavoidable contract change needed a recorded amendment bound to the exact before and after versions.&lt;/p&gt;

&lt;p&gt;The initial measurement found:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;613 declared architecture violations
557 baseline entries

83 forbidden component edges
475 imports past component facades
28 violations in the inner model and runtime graphs
1 five-component cycle
24 string dispatches in tasks
1 eval outside runtime
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The target described eight parts of DATAMIMIC and every dependency allowed between them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2gha3w680hs68s7aw8ro.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2gha3w680hs68s7aw8ro.png" alt="The target dependency layers used in the DATAMIMIC experiment. Every component card lists its complete set of permitted outgoing dependencies." width="800" height="525"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The target was fixed. The implementation had to move.&lt;/p&gt;

&lt;h2&gt;
  
  
  The contract drove the work
&lt;/h2&gt;

&lt;p&gt;The refactoring moved vocabulary to its owner, separated workers from tasks, removed runtime dependencies from model objects and exporters, moved source routing into runtime, replaced string dispatch with enums, introduced component facades and finally moved every cross-component import through them.&lt;/p&gt;

&lt;p&gt;After each step ArchKeel checked two different things:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Did the known debt shrink?

Did somebody make the target easier to satisfy?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The baseline answered the first question. Comparing the new contract with the previous contract answered the second.&lt;/p&gt;

&lt;p&gt;A shrinking violation count is weak evidence if the implementation can grant itself another dependency. The run therefore accepted four explicit amendments, but no new &lt;code&gt;requires&lt;/code&gt; edge, relaxed rule or baseline entry.&lt;/p&gt;

&lt;p&gt;The baseline moved like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;613 -&amp;gt; 551 -&amp;gt; 513 -&amp;gt; 483 -&amp;gt; 444 -&amp;gt; 390 -&amp;gt; 353 -&amp;gt; 329 -&amp;gt; 182 -&amp;gt; 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcavambj9vopvglhg3c5t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcavambj9vopvglhg3c5t.png" alt="The ArchKeel baseline ratchet used in the DATAMIMIC experiment. Known debt may shrink, new violations fail the change, and resolved debt cannot silently return." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The run finished after 11 steps beyond the initial measurement, plus later review corrections.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;121 product files changed
+2308 / -1788 lines, mostly moves
about 6.5 hours wall clock
0 declared violations
empty baseline
declared_rules: PASS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A full validation over roughly 450 files took around three seconds. The architecture gate was never the slow part.&lt;/p&gt;

&lt;p&gt;Of the 907 pre-existing descriptors, 885 could be verified unchanged. The other 22 already varied between repeated executions of the same starting code, so I marked them unverified instead of treating unstable output as evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Some architecture became genuinely better
&lt;/h2&gt;

&lt;p&gt;The empty baseline was not only import rewriting.&lt;/p&gt;

&lt;p&gt;The 23-module cross-package cycle between clients, contexts, data sources, exporters, storage, services and statements disappeared.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;io&lt;/code&gt;, &lt;code&gt;model&lt;/code&gt; and &lt;code&gt;domain&lt;/code&gt; no longer take &lt;code&gt;Context&lt;/code&gt; or &lt;code&gt;SetupContext&lt;/code&gt; parameters. They started with 17, 4 and 3 of those parameters.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;DataSourceRegistry&lt;/code&gt; went from roughly 1,000 lines to 382 lines of scalar loaders. Source routing moved to runtime, where the execution context belongs.&lt;/p&gt;

&lt;p&gt;Authoring stopped building a fake &lt;code&gt;SetupContext&lt;/code&gt;. The CLI, Python API and dry-run now share one &lt;code&gt;run_descriptor&lt;/code&gt; path.&lt;/p&gt;

&lt;p&gt;String dispatch in &lt;code&gt;tasks&lt;/code&gt; went from 24 cases to zero. &lt;code&gt;eval&lt;/code&gt; remained only inside runtime.&lt;/p&gt;

&lt;p&gt;ArchKeel also rejected changes that came directly from the design plan. Three “move this verbatim” instructions would have introduced a forbidden construct or external dependency in the new owner. An inner contract rejected another planned move. The plan was not trusted more than the code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Review found what the architecture result could not show
&lt;/h2&gt;

&lt;p&gt;The overnight run reached the contract. It was still not merge-ready.&lt;/p&gt;

&lt;p&gt;After a failed execution, the public Python API returned an empty captured result instead of the rows produced before the failure. The function signature still looked valid. The component dependencies were legal. No architecture rule could distinguish the broken behaviour.&lt;/p&gt;

&lt;p&gt;A regression test could. It passed on the starting code, failed on the refactored code and passed again after the correction.&lt;/p&gt;

&lt;p&gt;The second failure was in my gate script. It piped test output through &lt;code&gt;tail&lt;/code&gt; without &lt;code&gt;pipefail&lt;/code&gt;. A red test command could therefore leave the wrapper green.&lt;/p&gt;

&lt;p&gt;ArchKeel itself failed closed. The shell around it did not.&lt;/p&gt;

&lt;p&gt;The corrected gate now checks every command's own exit state, and a deliberately failing probe proves that the whole gate turns red.&lt;/p&gt;

&lt;p&gt;This is the boundary I want to keep:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;tests      -&amp;gt; behaviour
type check -&amp;gt; types
ArchKeel   -&amp;gt; architecture
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;ArchKeel should make the required behavioural checks part of the declared gate. It should not pretend to replace them.&lt;/p&gt;

&lt;h2&gt;
  
  
  I blamed the agents too quickly
&lt;/h2&gt;

&lt;p&gt;The facade result looked like a clean Goodhart example.&lt;/p&gt;

&lt;p&gt;At the start there were 475 imports bypassing component facades. At the end there were none. But three of the resulting facades were mostly re-export barrels:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;model.api     114 names
domains.api    46 names
io.api         29 names
runtime.api     5 names
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The import paths changed. Much of the coupling did not.&lt;/p&gt;

&lt;p&gt;My first explanation was that the agents had found the cheapest way through the architecture gate.&lt;/p&gt;

&lt;p&gt;The independent review corrected that explanation.&lt;/p&gt;

&lt;p&gt;The implementation brief explicitly prescribed re-export-only facades. The design plan derived their export lists from the names existing consumers already imported. The agents implemented what I asked for. ArchKeel enforced what I specified.&lt;/p&gt;

&lt;p&gt;The weak target was mine.&lt;/p&gt;

&lt;p&gt;I had defined who may depend on whom:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;runtime -&amp;gt; model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I had not defined what model should offer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;model provides:
  parse(...)
  element_definition(...)
  validate(...)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A component graph can enforce dependency direction. It cannot create a useful component API from arrows alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the contract did not measure
&lt;/h2&gt;

&lt;p&gt;Runtime ended in the correct place, but it grew by 28 percent. Context-typed parameters inside runtime rose from 105 to 126. &lt;code&gt;SourceRouter&lt;/code&gt; became a 666-line class of static methods. The old registry had moved behind a better boundary. It had not become a better design.&lt;/p&gt;

&lt;p&gt;Type quality moved only where a rule required it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Any annotations       245 -&amp;gt; 247
broad except           56 -&amp;gt; 56
cast                   10 -&amp;gt; 10
type: ignore           27 -&amp;gt; 27

string dispatch
inside tasks           24 -&amp;gt; 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The logical component graph also said nothing about the physical Python layout. At the end the package root still contained 22 packages and 8 modules. Model code remained spread over &lt;code&gt;model&lt;/code&gt;, &lt;code&gt;parsers&lt;/code&gt;, &lt;code&gt;statements&lt;/code&gt;, &lt;code&gt;constants&lt;/code&gt; and &lt;code&gt;enums&lt;/code&gt;. Runtime remained spread over &lt;code&gt;runtime&lt;/code&gt;, &lt;code&gt;tasks&lt;/code&gt;, &lt;code&gt;workers&lt;/code&gt;, &lt;code&gt;contexts&lt;/code&gt;, &lt;code&gt;services&lt;/code&gt; and &lt;code&gt;product_storage&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The contract answered one question well:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;dependency architecture
Who may depend on whom?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It did not yet answer two others:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;physical architecture -&amp;gt; Where should the code live?

public compatibility -&amp;gt; Which old paths must keep working?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those dimensions need separate rules and separate evidence. Combining them into one architecture score would hide the difference.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes in the next experiment
&lt;/h2&gt;

&lt;p&gt;The next target needs operations before dependency arrows. A facade should declare the questions another component may ask, not re-export every internal name the current consumers happen to use.&lt;/p&gt;

&lt;p&gt;It also needs budgets for properties that should shrink even when they cannot reach zero in one run:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;re-exported names per facade&lt;/li&gt;
&lt;li&gt;cross-component coupling width&lt;/li&gt;
&lt;li&gt;single-consumer exports&lt;/li&gt;
&lt;li&gt;component size&lt;/li&gt;
&lt;li&gt;context-typed parameters&lt;/li&gt;
&lt;li&gt;module cycle edges&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;ArchKeel already has the underlying measurements for some of these. It does not yet ratchet them.&lt;/p&gt;

&lt;p&gt;The next physical refactoring should also run in the opposite order: move code into the target package layout first, keep required old paths as explicit compatibility shims, then introduce the real component APIs, then fix dependency direction. This run did the facade work first and touched many imports twice.&lt;/p&gt;

&lt;p&gt;Several smaller gaps now have concrete reproductions: facade types need to follow re-exports and inspect fields, string-dispatch checks need to follow module-level constants, private attribute access through an untyped parameter needs to become visible, and facade lifecycle changes must stop requiring manual contract surgery.&lt;/p&gt;

&lt;p&gt;The gate belongs in ArchKeel as well. A declared set of architecture, test and type checks should fail closed without depending on a shell script getting every exit code right.&lt;/p&gt;

&lt;p&gt;The experiment supports the hypothesis I wrote before the run: a fixed deterministic contract can drive a multi-step refactoring without silently widening the target, while separate gates protect observed behaviour.&lt;/p&gt;

&lt;p&gt;It does not support the stronger claim that a dependency contract produces a clean architecture.&lt;/p&gt;

&lt;p&gt;I had not specified enough.&lt;/p&gt;

&lt;p&gt;The next run will test whether a target that names operations, placement and budgets changes that result.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>ai</category>
      <category>python</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Tests green, architecture worse: a deterministic gate for coding agents</title>
      <dc:creator>Alex</dc:creator>
      <pubDate>Tue, 15 Sep 2026 15:25:04 +0000</pubDate>
      <link>https://dev.to/ake2l/tests-green-architecture-worse-a-deterministic-gate-for-coding-agents-4jhi</link>
      <guid>https://dev.to/ake2l/tests-green-architecture-worse-a-deterministic-gate-for-coding-agents-4jhi</guid>
      <description>&lt;p&gt;My coding agents kept the tests green. The architecture still got worse.&lt;/p&gt;

&lt;p&gt;In the DATAMIMIC EE core the agents didn't break the build. They broke the structure. Utilities landed in whatever module was closest, not where they belonged. Code imported past the public interface of another component. And the one that hurt most: clients got imported in places that had no business touching them, above all in the communication between data sources and tasks. In a small task a reviewer catches that. In a large, nested task it hides in a diff that looks reasonable, with every test green.&lt;/p&gt;

&lt;p&gt;The problem has two halves. The first is obvious: an architecture change that nobody declared. The second is easy to miss. A change can make the code harder to analyze, so the next report looks clean only because the analyzer sees less of the program. A gate that can't detect when its own visibility gets worse can't distinguish clean code from code it can no longer see.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why instructions and an LLM judge weren't enough
&lt;/h2&gt;

&lt;p&gt;We build the EE core spec driven. Architecture decisions live in ADRs, and we move them into the skills our agents load. On small tasks that works. On bigger tasks and longer sessions it doesn't hold. A model's blind spots shift with the seed and with how full the context is, and an &lt;code&gt;AGENTS.md&lt;/code&gt; or a skill is a request to the agent, not a check on its output.&lt;/p&gt;

&lt;p&gt;Asking a model to judge the pull request moves the same problem one level up. The answer changes when you rephrase the question, and an agent can argue with it. A gate that an agent can talk its way around isn't a gate.&lt;/p&gt;

&lt;p&gt;Architecture tests aren't new either. ArchUnit, import-linter and dependency-cruiser check whether one snapshot of the code obeys a set of rules, and Archkeel does that too. What I needed on top was a comparison: the accepted state against the candidate, the change against what the agent said it would do, and a hard stop when the evidence itself got worse.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the code gets blinder
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/rapiddweller/archkeel/tree/0.3.0/fixtures/A-dispatch" rel="noopener noreferrer"&gt;Fixture A&lt;/a&gt; in the &lt;a href="https://github.com/rapiddweller/archkeel" rel="noopener noreferrer"&gt;Archkeel repo&lt;/a&gt; is the smallest case. You can run it yourself.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt; def run(key: str) -&amp;gt; int:
&lt;span class="gd"&gt;-    return first() + second()
&lt;/span&gt;&lt;span class="gi"&gt;+    handlers = {"first": first, "second": second}
+    return handlers[key]()
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tests pass. No new forbidden import, no cycle, no private access. A gate that only compares findings says nothing changed.&lt;/p&gt;

&lt;p&gt;The architecture may still be fine. What changed is that a static analyzer can no longer prove it: two resolved calls became one unresolved call. For a deterministic gate, that's enough to reject an undeclared change.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;expectation_fulfilled: FAIL
regression check failed in calls_unresolved: 0-&amp;gt;1
regression check failed in unresolved_ratio: 0/2-&amp;gt;1/1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The ratio is compared with integer cross-multiplication. No rounded percentages. No score.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the gate checks
&lt;/h2&gt;

&lt;p&gt;You describe a target architecture in a contract: components, the packages they own, the names each one makes public, and rules. Every ordered pair of components needs a decision, allowed or forbidden, each with a written reason. A pair nobody decided is an open decision, and validation stays red until it's gone.&lt;/p&gt;

&lt;p&gt;Archkeel keeps its verdicts separate and never blends them into one number. Three of them carry the idea:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;observation_complete&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Did the scan see everything it claims to see?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;declared_rules&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Does the code obey the contract?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;expectation_fulfilled&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Did the change match what was declared, without regressions?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The third one is built for agents. Before an agent submits an implementation, it commits an expectation: what it intends to change in the architecture. Archkeel checks Git ancestry and the merge request history on the host to verify that the expectation was published before the first submission. An agent that writes the expectation afterwards gets rejected even when its code is clean. That's &lt;a href="https://github.com/rapiddweller/archkeel/tree/0.3.0/fixtures/B-posthoc" rel="noopener noreferrer"&gt;Fixture B&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Exit codes are boring on purpose. 0 passes, 1 rejects, 2 means Archkeel couldn't verify the input, always with a diagnostic that names the subject and a remedy. Unknown never becomes green.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architect owns the target
&lt;/h2&gt;

&lt;p&gt;My first onboarding got this wrong. &lt;code&gt;init&lt;/code&gt; read today's imports and wrote them into the contract as the architecture. That report could only pass. It described the code, including every shortcut the agents had already taken.&lt;/p&gt;

&lt;p&gt;In 0.3.0 &lt;code&gt;init&lt;/code&gt; proposes components and their public interfaces, and it writes no dependency rule at all. Every pair is an open decision, listed by how many import sites use it. An import the code already has isn't a decision. If the target forbids it, the first report shows it red, and that's exactly what I want to see on day one.&lt;/p&gt;

&lt;p&gt;Who decides? The architect. The packaged skill runs in two modes. In interview mode the agent reads the ADRs and architecture documents first, prepares every decision with a recommendation, confirms the overall picture once and then asks only about conflicts and gaps. When I choose against its recommendation, it asks why. In auto mode the agent decides alone: documents first, then documented principles, then its own judgment, labeled as such. Every rule records &lt;code&gt;decided_by&lt;/code&gt;, and the report says it plainly: "156 of 159 rules decided by the agent, awaiting the architect."&lt;/p&gt;

&lt;p&gt;We tried both on a field service app: Python 3.12, FastAPI, async SQLAlchemy on PostgreSQL with PostGIS, Redis and Taskiq for the background workers, OR-Tools for route planning. 13 components. The service's owner acted as the architect.&lt;/p&gt;

&lt;p&gt;The first interview overwhelmed the architect. About 20 rounds of unranked questions over 156 component pairs. Some got answered in bulk, and three of those bulk answers contradicted the service's own architecture document. Nobody noticed until a second agent had decided the same 156 pairs blind, without seeing the architect's answers. It matched 140 of them, 89.7%. The three contradictions got fixed. Then the interviewing agent asked why the architect had gone against its recommendation on the remaining mismatches, and five more decisions changed.&lt;/p&gt;

&lt;p&gt;That interview is why interview mode looks the way it does now. Recommendations with evidence, the overall picture confirmed once, questions only where there's a conflict or a gap.&lt;/p&gt;

&lt;p&gt;The first report of the final target failed with 162 violations. 148 of them are one edge: the use cases import the persistence adapter directly, and both targets forbid that. The run also found five defects in Archkeel that no fixture had shown. One of them counted an import twice when a forbidden dependency and the interface boundary both caught it, and nearly doubled the headline. All five are fixed in 0.3.0. The anonymized contracts, both reports and the pair-by-pair comparison are in the repo under &lt;a href="https://github.com/rapiddweller/archkeel/tree/0.3.0/docs/evidence/internal-service" rel="noopener noreferrer"&gt;&lt;code&gt;docs/evidence/internal-service/&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;That agreement rate is one service, measured once. It's not a general accuracy of auto mode, and I won't sell it as one.&lt;/p&gt;

&lt;h2&gt;
  
  
  One verdict, two readers
&lt;/h2&gt;

&lt;p&gt;The reviewer gets an HTML report with the decision first, then the evidence. The agent gets the same result as JSON, plus a packaged instruction file from &lt;code&gt;uvx archkeel skill install claude&lt;/code&gt; (or &lt;code&gt;codex&lt;/code&gt;). I don't want one truth for the machine and a friendlier story for the human.&lt;/p&gt;

&lt;p&gt;For the drift itself there's a component flow view. Every card is a component from the contract, every line an import between components that Archkeel observed, drawn by the number of import sites. Teal lines conform to the contract. A dashed red line breaks a rule and carries the rule's id. A dotted amber line is a dependency the code uses that nobody has decided yet.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb8xyz0ht68q404x3053k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb8xyz0ht68q404x3053k.png" alt="Archkeel component flow of the shop sample with planted violations" width="799" height="694"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That's the small shop sample from Archkeel's fixtures, with violations planted on purpose.&lt;/p&gt;

&lt;h2&gt;
  
  
  It checks itself
&lt;/h2&gt;

&lt;p&gt;Archkeel runs its own gate in CI. Its contract has 6 components, and I decided all 30 component pairs myself: 46 rules, none of them by an agent. It requires exactly one owner per module, no component cycles, and it forbids &lt;code&gt;getattr&lt;/code&gt;, &lt;code&gt;hasattr&lt;/code&gt;, &lt;code&gt;cast&lt;/code&gt;, &lt;code&gt;eval&lt;/code&gt;, &lt;code&gt;exec&lt;/code&gt;, dynamic imports and &lt;code&gt;type: ignore&lt;/code&gt; anywhere in the package. Those are my rules for Archkeel, not defaults you inherit. Every enforcing rule was proven by planting a violation and watching the check fail.&lt;/p&gt;

&lt;p&gt;Determinism is measured, not assumed. A test runs the report three times on two clones in different paths, with different hash seeds, working directories, time zones and locales, and requires byte-identical output without normalization. That covers one machine and one Python build. Across operating systems and Python versions I haven't proven it yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it can't do
&lt;/h2&gt;

&lt;p&gt;It won't tell you that two competing implementations of the same idea exist, unless a rule or a regression exposes them. It checks private access through imports, so &lt;code&gt;import pkg; pkg._member&lt;/code&gt; slips through. It proves publication order, not that nobody edited privately before publishing. Host evidence comes from GitLab merge requests today; there's no GitHub adapter yet.&lt;/p&gt;

&lt;p&gt;About one call in five stays unresolved: 630 of 3,303 on Archkeel itself, 998 of 4,318 on the service. They're counted and reported, never guessed. Runtime behavior, data flow and performance aren't observed at all. They belong in the reason behind a decision, and Archkeel checks that a reason exists, not that it's true.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it's open source
&lt;/h2&gt;

&lt;p&gt;Before I wrote a line of Archkeel, I tried what already exists on the EE core: static analyzers, code graph tools, and cleanup helpers with and without an LLM behind them. For what I needed they added overhead instead of focus. None of them gave me the comparison I was after: the accepted architecture against the candidate against the declared intent, with weaker evidence treated as a failure.&lt;/p&gt;

&lt;p&gt;So Archkeel started as a guardrail for the agents working on the DATAMIMIC EE core. I extracted it because the problem isn't specific to our codebase. Components declare their public interface, and &lt;code&gt;interface_boundary&lt;/code&gt; rejects any import from another component that reaches past it. That's the utility and client problem from the start of this post, turned into a red verdict with a file and a line.&lt;/p&gt;

&lt;p&gt;It's MIT licensed. The code is on GitHub at &lt;a href="https://github.com/rapiddweller/archkeel" rel="noopener noreferrer"&gt;rapiddweller/archkeel&lt;/a&gt;, the package on &lt;a href="https://pypi.org/project/archkeel/" rel="noopener noreferrer"&gt;PyPI&lt;/a&gt;, and &lt;code&gt;uvx archkeel --help&lt;/code&gt; is all it takes to start. Point it at a repository where agents write code and find a verdict that's wrong. Open an &lt;a href="https://github.com/rapiddweller/archkeel/issues" rel="noopener noreferrer"&gt;issue&lt;/a&gt; with it. Those are the cases I want.&lt;/p&gt;

</description>
      <category>python</category>
      <category>architecture</category>
      <category>ai</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
