<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Egor Kraev</title>
    <description>The latest articles on DEV Community by Egor Kraev (@egor_kraev_0e262330fbb44b).</description>
    <link>https://dev.to/egor_kraev_0e262330fbb44b</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3562737%2F8f1d17b6-c236-43bc-8580-d3b2fc317733.png</url>
      <title>DEV Community: Egor Kraev</title>
      <link>https://dev.to/egor_kraev_0e262330fbb44b</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/egor_kraev_0e262330fbb44b"/>
    <language>en</language>
    <item>
      <title>Living architecture, enforced</title>
      <dc:creator>Egor Kraev</dc:creator>
      <pubDate>Thu, 17 Sep 2026 12:12:52 +0000</pubDate>
      <link>https://dev.to/egor_kraev_0e262330fbb44b/living-architecture-enforced-2ffh</link>
      <guid>https://dev.to/egor_kraev_0e262330fbb44b/living-architecture-enforced-2ffh</guid>
      <description>&lt;p&gt;In my &lt;a href="https://dev.to/egor_kraev_0e262330fbb44b/how-i-enforce-architecture-on-coding-agents-141a"&gt;previous post&lt;/a&gt;, I described the two components of what makes code “good” for me from a high-level structure perspective, namely modularity and making sure it realizes the high-level intent of the codebase. &lt;/p&gt;

&lt;p&gt;Here is how I make those constraints bite: the &lt;a href="https://github.com/MotleyAI/slayer/tree/main/architecture" rel="noopener noreferrer"&gt;architecture&lt;/a&gt; folder of the SLayer repo contains a &lt;a href="https://likec4.dev/tutorial/" rel="noopener noreferrer"&gt;LikeC4&lt;/a&gt; &lt;a href="https://github.com/MotleyAI/slayer/blob/main/architecture/model/slayer.c4" rel="noopener noreferrer"&gt;model&lt;/a&gt; of the major components of the codebase, and the dependencies between them. This is the binding high-level map of the system components. Here is how it is enforced:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;architecture/index.yaml&lt;/code&gt; maps the names of the LikeC4 components to Python packages in the codebase. Eventually the goal is for those to align, but for now some of the LikeC4 components are ‘virtual’, mapped to a set of several modules. &lt;/p&gt;

&lt;p&gt;Enforcement is one law, in &lt;code&gt;tools/arch_check.py&lt;/code&gt;: the LikeC4 relations must exactly match the AST-measured import edges, at every granularity the model declares — where a component spells out its child modules, the arrows are checked child-by-child too. A declared arrow is the only thing that licenses an import; silence is a ban. So the code can’t drift from the declared granularity model, as a hard deterministic constraint. Grandfathered crossings live right in the model as dashed &lt;code&gt;#legacy&lt;/code&gt; arrows, and arch_check pins their count to a baseline in &lt;code&gt;architecture/index.yaml&lt;/code&gt; that may only ever be lowered — a PR can shrink that list but never quietly grow it. &lt;/p&gt;

&lt;p&gt;To begin with, the LikeC4 model is quite coarse, and is being refined step by step as I touch the relevant areas of the codebase. The idea is not to have a fully granular, eternal grand plan, but rather to build up gradually the set of top-down architectural constraints to complement the bottom-up constraints provided by unit and integration tests, to make sure the codebase stays within the high-level design shape. &lt;/p&gt;

&lt;p&gt;This takes care of the modularity structure, but the structure itself doesn’t say what each module is for, what principles and constraints it must observe. That’s the second half of the living architecture docs, contained in the .arc42.md files. Most of these files correspond to a node in the LikeC4 model, and describe the principles that this module has to follow. The relevant part of the LikeC4 model is auto-compiled into a Mermaid diagram (with dashed edges indicating deprecated dependencies) and included in the respective .arc42.md - see for example &lt;a href="https://github.com/MotleyAI/slayer/blob/main/architecture/sql.arc42.md" rel="noopener noreferrer"&gt;the one for the sql module&lt;/a&gt;. &lt;/p&gt;

&lt;p&gt;Both the LikeC4 and the &lt;a href="https://arc42.org/" rel="noopener noreferrer"&gt;arc42&lt;/a&gt; files are treated by all the skills in the same way as tests, that is any change to them must be specifically and explicitly be approved by me.&lt;/p&gt;

&lt;p&gt;Importantly, each principle must come with a tag that denotes its status, either citing which tests enforce it, or that it’s believed to hold pending review, or that it still needs enforcing (citing relevant issue IDs). That way, the bigger vision can be represented coherently in one place, even if not all of it may be currently realized.&lt;/p&gt;

&lt;p&gt;This highlights an important and natural division of labour between the OpenSpec files and these - OpenSpec files refer to specific features and functionalities, and are descriptive rather than normative. They describe what is, or in the case of change specs, what we’re about to do. On the other hand, these architectural docs map to actual code structure and are normative, they describe what should hold, on a sometimes fairly abstract principle level. &lt;/p&gt;

&lt;p&gt;As the OpenSpec artefacts relate to specific features, some of them will cut across many architectural components. The link is formal: &lt;code&gt;architecture/index.yaml&lt;/code&gt; maps every OpenSpec spec group either to its owning node or to a &lt;code&gt;cross_cutting_specs&lt;/code&gt; entry listing the nodes it touches, and arch_check fails on any spec group that is unmapped, mapped but missing on disk, or empty.&lt;/p&gt;

&lt;p&gt;How are the .arc42.md used? During the planning stage, the agent reads the architecture docs for the relevant modules so it can make sure any changes it makes align with them - and also as it addresses review comments. During the test-writing phase, the agent can check which of the principles are relevant for the code being planned, and make sure its adherence to the relevant principles are covered by tests. &lt;/p&gt;

&lt;p&gt;And on the other hand, they give me a clear place to express and review the high-level laws that hold for a given component - that, for me, is a large part of what code quality is about. &lt;/p&gt;

&lt;p&gt;I believe this combination of modularity enforced via LikeC4, and principle-based planning and review for each module, allows me to achieve a degree of high-level coherence and architectural cleanliness superior to that that my hand-written code ever had back in the day (and will keep improving gradually, as I refine it) - just like my PR workflow gives me better corner case handling and test coverage than I ever achieved on my own. &lt;/p&gt;

&lt;p&gt;If you believe this combination still misses some important facet of what good code is about, please tell me what it is - I’d love to learn!&lt;/p&gt;

&lt;p&gt;And do let me know if you'd like me to refactor that into a reusable collection of skills + scripts.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>architecture</category>
      <category>programming</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>How I enforce architecture on coding agents</title>
      <dc:creator>Egor Kraev</dc:creator>
      <pubDate>Tue, 15 Sep 2026 14:06:14 +0000</pubDate>
      <link>https://dev.to/egor_kraev_0e262330fbb44b/how-i-enforce-architecture-on-coding-agents-141a</link>
      <guid>https://dev.to/egor_kraev_0e262330fbb44b/how-i-enforce-architecture-on-coding-agents-141a</guid>
      <description>&lt;p&gt;My &lt;a href="https://dev.to/egor_kraev_0e262330fbb44b/why-i-no-longer-read-code-much-1d59"&gt;previous post&lt;/a&gt; described my workflow for individual PR’s, and reasons why I believe it is better at delivering “locally good” PRs than I would have been writing by hand. However, that left open a big-picture gap: even if each individual PR is perfectly reasonable, how do I make sure the overall architecture doesn’t drift into spagetti gradually, by small increments?&lt;/p&gt;

&lt;p&gt;That is, how do I make sure the code remains modular and well-organized, and expresses my intent, at a high level? That is what this post is about. &lt;/p&gt;

&lt;p&gt;The short answer is simple: I spell out explicitly both the high-level structure of the code and the principles I expect it to embody, and then enforce both against the code. The first part is reasonably easy as far as it goes, and was being done as long as there is software - it’s called documentation. The second part is the hard bit, and the interesting one.&lt;/p&gt;

&lt;p&gt;What does “good code” mean in the first place? Every one who has ever worked with developers knows this is an unexpectedly personal question - one person’s polished code is another person’s mess. My definition centers on two characteristics: firstly, it must be maintainable; and secondly, it must correctly express the intent behind the application/library. Both take some unwrapping.&lt;/p&gt;

&lt;p&gt;For me, “maintainable” mostly means “modular” - composed of components with well defined and documented jobs and interfaces. As long as that’s the case, it’s reasonably easy to locate the component responsible for the bug or feature one cares about, and then one can turn the job over to coding agents who by now are quite good at doing smallish-scale optimization/problem solving. Of course, for that to work the components must map well to the overall purpose of the system - that’s the hard design part. The other part of “maintainable” is “thoroughly coverd by tests”, which again is easier in a modular design. The nice thing about this modular structure is that once in place, it’s fairly easy to enforce programmatically, even as new code is added to it. &lt;/p&gt;

&lt;p&gt;The second part of what makes code “good” for me is the subtler one: does the code express the overall intent of the library/product? To some extent this is what spec-driven development tries to do, but that places the onus on the spec writer to make a lot of fine-grained choices to align each spec with the big picture. This is also the Achilles heel of good test coverage considered standalone:  we may be testing everything, but who writes the tests, and how do we make sure each test is mandating the behavior that it actually should?&lt;/p&gt;

&lt;p&gt;The solution to this for me is something I’d call principle-based development: explicitly spelling out, at different levels of the architecture, the principles that the software (or a particular component) must embody. These are not descriptive but normative: that is, their purpose is not to describe the behavior of current code,  but to mandate what all relevant code, present and future, should fulfill. &lt;/p&gt;

&lt;p&gt;Then each individual PR can consult these principles and make sure its plan aligns with them (and on a brownfield project, also fix any related violations it comes across near to the bits it touches). &lt;/p&gt;

&lt;p&gt;This was the bit that was prohibitively tedious befor the LLM era - but now, it can just be a routine step in the planning and review stage of each PR. In fact, when the agent asks me to make a nontrivial design choice, I can just ask it to analyze which of the options aligns better with the relevant principles, and it will do so for me, with quotes and references. &lt;/p&gt;

&lt;p&gt;In my next post, I will describe how I operationalized both of these in my development workflow for the &lt;a href="https://github.com/MotleyAI/slayer" rel="noopener noreferrer"&gt;SLayer repo&lt;/a&gt;, using &lt;a href="https://arc42.org/" rel="noopener noreferrer"&gt;Arc42&lt;/a&gt;, &lt;a href="https://likec4.dev/" rel="noopener noreferrer"&gt;LikeC4&lt;/a&gt; and some custom verification logic.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Why I no longer read code (much)</title>
      <dc:creator>Egor Kraev</dc:creator>
      <pubDate>Wed, 09 Sep 2026 11:58:57 +0000</pubDate>
      <link>https://dev.to/egor_kraev_0e262330fbb44b/why-i-no-longer-read-code-much-1d59</link>
      <guid>https://dev.to/egor_kraev_0e262330fbb44b/why-i-no-longer-read-code-much-1d59</guid>
      <description>&lt;p&gt;One of the big divides of 2026 among people who are responsible for producing code (whatever their title) is whether reading (or at least skimming) every bit of actual code produced by the agent is a good idea. &lt;/p&gt;

&lt;p&gt;Some vehemently argue for that, citing the atrophying of mental muscles if one doesn’t, dependence on Anthropic, architectural erosion of the codebase, feeling responsible for the code they ship, and other sound reasons; others are happy with their automated code-producing and -reviewing setup and rejoice in the increased productivity it brings.&lt;/p&gt;

&lt;p&gt;I come down strongly on the latter side at the moment. To explain why, it’s best to first walk you through my code-generating workflow first. &lt;/p&gt;

&lt;p&gt;This consists of four stages, each materialized as skills and helper scripts, each ran in a new clean session. The first stage is planning: I brain-dump every detail of what I know about what I want to achieve, goals, implementation details, anything I can think of that relates to the task. Claude (using Fable) then interviews me about it, focussing on design decisions, corner cases, and gotchas.&lt;/p&gt;

&lt;p&gt;It shows to me the resulting plan and once I’m happy sends that to Codex for review. It then validates the points brought up by Codex (typically accepting most but not all, explaining to me why for those it refuses, and asking for my decision on those it deems borderline). &lt;/p&gt;

&lt;p&gt;The result of that stage is a full set of &lt;a href="https://github.com/Fission-AI/OpenSpec/blob/main/docs/concepts.md" rel="noopener noreferrer"&gt;OpenSpec artefacts&lt;/a&gt; for the change, reviewed by Codex and formally validated by the OpenSpec machinery. At this point I reset the session and start the test-writing phase.&lt;/p&gt;

&lt;p&gt;In the test-writing phase (ran by Fable 5 or Opus-4.8, as are the following stages, chosen depending on task complexity) Claude writes tests (which the OpenSpec delta in changes/specs/spec.md makes easy to both plan and verify) and then has the tests reviewed by Codex, again folding in all valid feedback.&lt;/p&gt;

&lt;p&gt;After that I reset the session again and start the implementation phase. The goals and design are available to the agent from the OpenSpec artefacts, and the tests are already written. At the end of this phase, the agent asks for my permission to push and PR (GitHub is not touched until this moment).&lt;/p&gt;

&lt;p&gt;In contrast to the planning phase, which is very interactive, the test-writing and implementation phases typically require no involvement from me, unless they run into a gotcha they need my judgement to resolve.&lt;/p&gt;

&lt;p&gt;Once the PR has been created, I reset the session again and run a review loop. At each iteration, I check CI results of course, plus review feedback from &lt;a href="https://openai.com/codex/" rel="noopener noreferrer"&gt;Codex&lt;/a&gt;, &lt;a href="https://www.sonarsource.com/" rel="noopener noreferrer"&gt;Sonar&lt;/a&gt;, &lt;a href="https://www.coderabbit.ai/" rel="noopener noreferrer"&gt;CodeRabbit&lt;/a&gt;, and some deterministic scripts enforcing internal conventions (such as imports at the top only and a maximum inline-comment-to-source-code-ratio). &lt;/p&gt;

&lt;p&gt;Once feedback has been collected from all the sources, it’s triaged by Claude Code. This phase is more interactive than the previous two, mostly because I need to decide whether some of the issues uncovered by the reviews deserve a follow-up issue, should be handled on the spot, or can just be dismissed. &lt;/p&gt;

&lt;p&gt;All the triaged results are addressed by Claude Code and the result pushed. This goes on until every single gate returns no issues. Then the openspec artefacts are archived in the same branch, and the result merged. &lt;/p&gt;

&lt;p&gt;I believe that as far as gotchas, corner cases, and code validity is concerned, the resulting code is far more reliable than the code I wrote by hand ever was. This doesn’t mean it needs no involvement from me - after the workflow is done, it’s still imporant to kick the tires so to speak, that is to try using the code for the purpose for which it was written - but that doesn’t require my reading the code either. &lt;/p&gt;

&lt;p&gt;I’m not claiming the result is perfect - but let’s be honest, the code I wrote by hand in the olden days wasn’t perfect either, though perhaps with different failure modes. Certainly this flow has pointed out (and then addressed) more corner cases and choices that I hadn’t thought about when formulating the task, than I can count. &lt;/p&gt;

&lt;p&gt;One important challenge that the above process doesn NOT address is architectural erosion - even if every single PR is sound, over time they might well add up to unmaintainable spaghetti code (also known as the lava flow or lava layer anti-pattern). At the moment, I deal with that by doing periodic interactive reviews and refactors as separate PRs, formulating high-level principles I expect the code to comply with, and using Fable to analyze current code against the principles and plan towards alignment - it’s perfectly able to reason at that level of abstraction if that’s what it’s asked to do.&lt;/p&gt;

&lt;p&gt;In the long term that does not seem like a sufficiently clean or reproducible approach (well, spelling out and enforcing principles does, but relying on ad hoc Fable analyses to make them stick doesn’t), so I’m now looking at firstly, a setup using a combination of &lt;a href="https://arc42.org/" rel="noopener noreferrer"&gt;arc42&lt;/a&gt; and &lt;a href="https://likec4.dev/tutorial/" rel="noopener noreferrer"&gt;LikeC4&lt;/a&gt; to automatically maintain, and enforce against the source code, a high-level representation of the architecture; and secondly, a principles-first design - I’ll report on how that works out, in a later post. &lt;/p&gt;

&lt;p&gt;Fundamentally, I find using agents to code not that different from running a team - you have to set up the right processes, and then trust them to get the right result, and keep adjusting the processes to address every new failure mode you come across. &lt;/p&gt;

&lt;p&gt;I do realize not reading code is anathema to many - so explain to me why!&lt;/p&gt;

</description>
      <category>agents</category>
      <category>coding</category>
      <category>agentskills</category>
    </item>
  </channel>
</rss>
