<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Viktoria</title>
    <description>The latest articles on DEV Community by Viktoria (@marketing_explyt_a7b53da9).</description>
    <link>https://dev.to/marketing_explyt_a7b53da9</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4074488%2F591a7871-a646-4ad3-a738-9f5034840513.png</url>
      <title>DEV Community: Viktoria</title>
      <link>https://dev.to/marketing_explyt_a7b53da9</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/marketing_explyt_a7b53da9"/>
    <language>en</language>
    <item>
      <title>Your AI agent turned a red test green without touching production code</title>
      <dc:creator>Viktoria</dc:creator>
      <pubDate>Wed, 30 Sep 2026 10:52:02 +0000</pubDate>
      <link>https://dev.to/marketing_explyt_a7b53da9/your-ai-agent-turned-a-red-test-green-without-touching-production-code-2334</link>
      <guid>https://dev.to/marketing_explyt_a7b53da9/your-ai-agent-turned-a-red-test-green-without-touching-production-code-2334</guid>
      <description>&lt;p&gt;Here is a loop most of us have run at least once. You give an AI agent a task with a clear exit condition: all tests in the folder pass. It works for a while, reports success, and the CI is green. Then you open the diff and find that the file with the most changes is the test.&lt;/p&gt;

&lt;p&gt;Sergey Pospelov, who works on Explyt, spent part of his September 28 webinar ("Working with AI Tools at the User Level") on exactly this failure. His framing: a green suite proves nothing about the implementation when the agent was free to edit the suite, and agents reach for the tests before they reach for the code. He listed three moves. Weaken the assertions. Add a mock. Skip the test. Each one satisfies the exit condition and ships the bug.&lt;/p&gt;

&lt;p&gt;This post takes that one section apart with code, then shows the three defenses and the prompt that keeps the fix loop from running forever. The rest of the webinar (why the agent solves the wrong task in the first place, the spec template, subagents, git worktrees, the 30-minute IDE setup, the live demo) is in the full recap on our blog, and there is a free PDF at the end of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the three moves look like in a JUnit test
&lt;/h2&gt;

&lt;p&gt;The extended guide ships a sample spec, &lt;code&gt;task-01.md&lt;/code&gt;: limit each API key to 100 requests per minute on &lt;code&gt;GET /api/orders&lt;/code&gt; in a Spring Boot service, token bucket per key. One acceptance criterion reads "101st request → 429 + Retry-After". Written before the implementation, the test for it fails, which is the state you want at step 4:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@Test&lt;/span&gt;
&lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;request101WithinAMinuteIsRejected&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="kd"&gt;throws&lt;/span&gt; &lt;span class="nc"&gt;Exception&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;int&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++)&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;mvc&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;perform&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/api/orders"&lt;/span&gt;&lt;span class="o"&gt;).&lt;/span&gt;&lt;span class="na"&gt;header&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"X-Api-Key"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="no"&gt;KEY&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
           &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;andExpect&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;isOk&lt;/span&gt;&lt;span class="o"&gt;());&lt;/span&gt;
    &lt;span class="o"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;mvc&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;perform&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/api/orders"&lt;/span&gt;&lt;span class="o"&gt;).&lt;/span&gt;&lt;span class="na"&gt;header&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"X-Api-Key"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="no"&gt;KEY&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
       &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;andExpect&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;isTooManyRequests&lt;/span&gt;&lt;span class="o"&gt;())&lt;/span&gt;
       &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;andExpect&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;header&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;exists&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"Retry-After"&lt;/span&gt;&lt;span class="o"&gt;));&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the agent is told to keep working until this passes, and it is allowed to write anywhere in the repository. The snippets below are ours, written to illustrate Sergey's list; they are not from the webinar. Each one turns the test green with zero changes to production code.&lt;/p&gt;

&lt;p&gt;Move one, weaken the assertion. The status check becomes something that any non-crashing response satisfies:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="n"&gt;mvc&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;perform&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;get&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"/api/orders"&lt;/span&gt;&lt;span class="o"&gt;).&lt;/span&gt;&lt;span class="na"&gt;header&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"X-Api-Key"&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt; &lt;span class="no"&gt;KEY&lt;/span&gt;&lt;span class="o"&gt;))&lt;/span&gt;
   &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;andExpect&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;assertTrue&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;getResponse&lt;/span&gt;&lt;span class="o"&gt;().&lt;/span&gt;&lt;span class="na"&gt;getStatus&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="o"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Move two, mock the thing under test. The limiter becomes a &lt;code&gt;@MockBean&lt;/code&gt; with a scripted answer, so the test exercises the mock's script and never the real bucket:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@MockBean&lt;/span&gt; &lt;span class="nc"&gt;RateLimiter&lt;/span&gt; &lt;span class="n"&gt;rateLimiter&lt;/span&gt;&lt;span class="o"&gt;;&lt;/span&gt;

&lt;span class="nd"&gt;@BeforeEach&lt;/span&gt;
&lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;limiterSaysNoOnCall101&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;when&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rateLimiter&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;tryConsume&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;anyString&lt;/span&gt;&lt;span class="o"&gt;()))&lt;/span&gt;
        &lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;thenAnswer&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inv&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="n"&gt;calls&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="o"&gt;);&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Move three, skip it. One annotation and a plausible reason:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight java"&gt;&lt;code&gt;&lt;span class="nd"&gt;@Disabled&lt;/span&gt;&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s"&gt;"flaky under parallel execution, tracked separately"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="nd"&gt;@Test&lt;/span&gt;
&lt;span class="kt"&gt;void&lt;/span&gt; &lt;span class="nf"&gt;request101WithinAMinuteIsRejected&lt;/span&gt;&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="kd"&gt;throws&lt;/span&gt; &lt;span class="nc"&gt;Exception&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="o"&gt;...&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All three read as reasonable in a diff you skim at 6 pm. The second one is the nasty one: the test still runs, still asserts &lt;code&gt;429&lt;/code&gt;, and still tells you nothing about the code. This is the third anti-pattern from the webinar, blind trust in the output, in its most concrete form.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqe8570f83451g0lhrmeu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqe8570f83451g0lhrmeu.png" alt="Anti-pattern 3: blind trust in the output" width="784" height="926"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Three defenses, in order of cost
&lt;/h2&gt;

&lt;p&gt;Sergey gave three. We add the mechanics for each.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Make the tests read-only for the implementing agent.&lt;/strong&gt; The guide states it tool-agnostically as an &lt;code&gt;.agentignore&lt;/code&gt;-style boundary. In Explyt the mechanism is &lt;a href="https://explyt.ai/docs/explyt-test/configuration/edit-scope" rel="noopener noreferrer"&gt;Edit Scope&lt;/a&gt;: you attach the files or folders the agent may change and mark them as the scope; everything else stays readable and becomes unwritable. Two details from the docs page are worth knowing before you rely on it. Reads are unaffected (that is &lt;code&gt;.readignore&lt;/code&gt;'s job), and the agent cannot create new files inside a scoped folder, only edit existing ones. If the implementation needs a new class, attach an empty file first, or flip to the deny-list form: &lt;code&gt;.writeignore&lt;/code&gt; blocks writing and creating on the listed paths and leaves the rest of the repository open. Both files take gitignore-like patterns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Split test authorship from implementation.&lt;/strong&gt; The tests come from one agent or chat and are approved by you; a different agent or a fresh chat implements against them. Once step 4 of the cycle is a checkpoint, the implementing agent has no diff in &lt;code&gt;src/test&lt;/code&gt; to hide behind.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Review the test diff with a separate agent before you read it.&lt;/strong&gt; Self-review by the same agent is worthless; it likes its own work. A reviewer with a clean context, ideally a model from another vendor, catches the pattern the author was optimizing for: the assert that got looser, the mock on the class under test, the &lt;code&gt;@Disabled&lt;/code&gt; with a plausible story. Explyt's &lt;a href="https://explyt.ai/docs/explyt-test/code-review/auto-review" rel="noopener noreferrer"&gt;automatic review&lt;/a&gt; runs on the diff. Your own read comes after it, and the order matters, because you read a diff differently when a reviewer has already flagged the test file.&lt;/p&gt;

&lt;h2&gt;
  
  
  The review loop with a hard stop
&lt;/h2&gt;

&lt;p&gt;A review agent without a cap is its own anti-pattern. It keeps finding nits, burns tokens, and around round five starts rewriting code that was fine. The prompt from the guide, verbatim:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Implement the task from task-01.md. When all tests pass, run a review subagent on your diff. If it reports problems, fix them and run the review again. Do at most 3 review rounds. Stop when the review is clean or after round 3, then report: what you fixed, what is still open, and why.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three things in there do the work. The review runs on the diff, so the reviewer is not asked to re-derive the task. The cap is three rounds. And the exit report is mandatory: what got fixed, what is still open, why. Anything open after round three is a human decision, which is the point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this sits in the six-step cycle
&lt;/h2&gt;

&lt;p&gt;The webinar frames it as Specification-Driven Development for the input and Test-Driven Development for the output. The cycle from issue to merge:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;Who&lt;/th&gt;
&lt;th&gt;What&lt;/th&gt;
&lt;th&gt;Checkpoint&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;You&lt;/td&gt;
&lt;td&gt;Hand over the task: issue link, goal, constraints; &lt;code&gt;AGENTS.md&lt;/code&gt; carries the project context&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;td&gt;Plan mode: produce &lt;code&gt;task-01.md&lt;/code&gt; with testable requirements and acceptance criteria&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;You&lt;/td&gt;
&lt;td&gt;Sign off on the spec&lt;/td&gt;
&lt;td&gt;yes, cheapest place to catch a mistake&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;td&gt;Turn the acceptance criteria into failing tests; you sign off and lock them&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;td&gt;Code until green, without write access to the tests&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;Agent&lt;/td&gt;
&lt;td&gt;Review loop, three rounds max; then you read the diff and merge&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Your attention goes to 3, 4 and 6. The test-editing problem lives between 4 and 5, and locking at step 4 is what makes step 5 safe to leave alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  The moment from the live demo that fits here
&lt;/h2&gt;

&lt;p&gt;The demo half ran in IntelliJ IDEA against &lt;code&gt;spring-petclinic-kotlin&lt;/code&gt;. Sergey asked the agent to write a Skill for controller tests, and the &lt;code&gt;SKILL.md&lt;/code&gt; it produced ended with an acceptance checklist: right test slice, observable behavior pinned, assertions left at full strength. A separate review agent went over the skill and reported nothing.&lt;/p&gt;

&lt;p&gt;Then he deleted &lt;code&gt;CrashControllerTest.kt&lt;/code&gt;, asked for tests for &lt;code&gt;CrashController.kt&lt;/code&gt;, and the agent wrote the test without using the skill. The cause, found live: the frontmatter had &lt;code&gt;agent: null&lt;/code&gt; and no &lt;code&gt;used-by&lt;/code&gt; field, so nothing told the agent the skill was available to it. One edit to the file and the second request showed &lt;code&gt;Used skill test-spring-controllers&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcx1g3jzw3so5hb77x292.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcx1g3jzw3so5hb77x292.jpg" alt="Demo: editing the SKILL.md frontmatter, on the right the Explyt panel with Session setup: 2 rules, 3 skills, memory on" width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We include this because the checklist inside that skill is the same defense as above, moved one layer earlier: the rule "do not weaken assertions" applied while the test is written, before any review runs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try this on Monday
&lt;/h2&gt;

&lt;p&gt;This block is our suggestion for your repository; the webinar did not show it. Add one rule to &lt;code&gt;AGENTS.md&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Tests&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Never modify, weaken, mock out or disable an existing test to make it pass.
&lt;span class="p"&gt;-&lt;/span&gt; If a test looks wrong, stop and report which test and why. Do not fix it silently.
&lt;span class="p"&gt;-&lt;/span&gt; Production code must change under src/main; tests under src/test are read-only during implementation.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then take one task from the backlog, write or generate the failing tests first, lock &lt;code&gt;src/test&lt;/code&gt; with Edit Scope or &lt;code&gt;.writeignore&lt;/code&gt;, and run the review prompt above. Compare the diff to what the agent used to hand you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;Edit Scope stops writes, so a bad test written before the lock stays bad; that is what the step 4 approval is for. The review agent is another model and can miss the same mock a human misses. Three rounds is a budget, and the exit report is the thing to read, since a clean round three means "nothing found", which is weaker than "verified". And none of this replaces reading the final diff yourself; it makes that read shorter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;→ &lt;a href="https://explyt.ai/en/blog/webinar-working-with-ai-tools-recap" rel="noopener noreferrer"&gt;Read the full recap on explyt.ai&lt;/a&gt;&lt;/strong&gt;: the other two anti-patterns with context-window numbers, the &lt;code&gt;task-01.md&lt;/code&gt; spec template, when subagents and git worktrees pay off, Rules, Skills, MCP and Memory Bank setup in 30 minutes, the full demo, and the free "Methodology for one task" PDF with the six steps as a printable checklist.&lt;/p&gt;

&lt;p&gt;Which of the three moves has an agent pulled on you: the weakened assert, the mock, or the &lt;code&gt;@Disabled&lt;/code&gt;? And what caught it, the review or production?&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/en/blog/webinar-working-with-ai-tools-recap" rel="noopener noreferrer"&gt;Green tests don't mean the task is done: notes from the Explyt webinar on working with AI agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/en/webinar-ai-tools-extended.html" rel="noopener noreferrer"&gt;Working with AI Tools at the User Level: Extended Guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Explyt documentation: &lt;a href="https://explyt.ai/docs/explyt-test/configuration/edit-scope" rel="noopener noreferrer"&gt;Edit Scope&lt;/a&gt;, &lt;a href="https://explyt.ai/docs/explyt-test/code-review/auto-review" rel="noopener noreferrer"&gt;Automatic code review&lt;/a&gt;, &lt;a href="https://explyt.ai/docs/explyt-test/configuration/skills" rel="noopener noreferrer"&gt;Skills&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>java</category>
      <category>jetbrains</category>
    </item>
    <item>
      <title>Every fix compiled and passed lint. Every fix broke the Kafka pipeline somewhere new. Read afterwards, the AI agent's session log showed all the defects violating one async-path rule. What does your agent call "fixed": a clean build, or a passing rerun?</title>
      <dc:creator>Viktoria</dc:creator>
      <pubDate>Wed, 23 Sep 2026 11:55:05 +0000</pubDate>
      <link>https://dev.to/marketing_explyt_a7b53da9/every-fix-compiled-and-passed-lint-every-fix-broke-the-kafka-pipeline-somewhere-new-read-2klc</link>
      <guid>https://dev.to/marketing_explyt_a7b53da9/every-fix-compiled-and-passed-lint-every-fix-broke-the-kafka-pipeline-somewhere-new-read-2klc</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/marketing_explyt_a7b53da9/green-build-broken-invariant-one-kafka-session-where-every-fix-passed-and-nothing-held-4k2n" class="crayons-story__hidden-navigation-link"&gt;Green build, broken invariant: one Kafka session where every fix passed and nothing held&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/marketing_explyt_a7b53da9" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4074488%2F591a7871-a646-4ad3-a738-9f5034840513.png" alt="marketing_explyt_a7b53da9 profile" class="crayons-avatar__image" width="96" height="96"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/marketing_explyt_a7b53da9" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Viktoria
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Viktoria
                
                
              
              &lt;div id="story-author-preview-content-4724727" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/marketing_explyt_a7b53da9" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4074488%2F591a7871-a646-4ad3-a738-9f5034840513.png" class="crayons-avatar__image" alt="" width="96" height="96"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Viktoria&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/marketing_explyt_a7b53da9/green-build-broken-invariant-one-kafka-session-where-every-fix-passed-and-nothing-held-4k2n" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 23&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/marketing_explyt_a7b53da9/green-build-broken-invariant-one-kafka-session-where-every-fix-passed-and-nothing-held-4k2n" id="article-link-4724727"&gt;
          Green build, broken invariant: one Kafka session where every fix passed and nothing held
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag crayons-tag--filled  " href="/t/news"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;news&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/machinelearning"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;machinelearning&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/github"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;github&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/marketing_explyt_a7b53da9/green-build-broken-invariant-one-kafka-session-where-every-fix-passed-and-nothing-held-4k2n" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;5&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/marketing_explyt_a7b53da9/green-build-broken-invariant-one-kafka-session-where-every-fix-passed-and-nothing-held-4k2n#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            8 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Green build, broken invariant: one Kafka session where every fix passed and nothing held</title>
      <dc:creator>Viktoria</dc:creator>
      <pubDate>Wed, 23 Sep 2026 11:48:01 +0000</pubDate>
      <link>https://dev.to/marketing_explyt_a7b53da9/green-build-broken-invariant-one-kafka-session-where-every-fix-passed-and-nothing-held-4k2n</link>
      <guid>https://dev.to/marketing_explyt_a7b53da9/green-build-broken-invariant-one-kafka-session-where-every-fix-passed-and-nothing-held-4k2n</guid>
      <description>&lt;p&gt;Every fix compiled and passed lint. Every fix was followed by a different breakage. All of them broke one sentence about asynchronous TID processing.&lt;/p&gt;

&lt;p&gt;You know this loop. A bug report comes in. You find something wrong, fix it, and the system fails again, somewhere else, in a new way. You fix that. It fails a third time. An hour later you have a diff that touches half a dozen files and a growing suspicion that the bugs were never separate.&lt;/p&gt;

&lt;p&gt;At Explyt, where we build an AI agent for JetBrains IDEs, we read session logs to see what debugging with an agent looks like outside a demo. This post is about one of those logs. Yurii Kostyukov, our product manager, went through it after the fact; he did not run the session, and neither did anyone else on our team. The developer worked in our agent, in plain chat mode, on a current frontier model. Debug mode stayed off. The codebase was Kotlin and Java on Spring Boot with Kafka, in a payments system.&lt;/p&gt;

&lt;p&gt;The longer narrative is on our blog as &lt;a href="https://explyt.ai/en/blog/one-symptom-ten-problems-kafka-invariant" rel="noopener noreferrer"&gt;One symptom, ten problems: a Kotlin + Spring Boot + Kafka debugging case&lt;/a&gt;. Treat the "ten" there as a working title; we could not verify a count from the log. This version is shorter, skips most of the story, and ends with something you can paste into your own repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A single symptom started the session. The developer and the agent fixed defect after defect, roughly half of them unrelated to that symptom, and every one of them violated the same rule about how a terminal ID moves through the asynchronous path.&lt;/li&gt;
&lt;li&gt;The agent helped with the small things: SQL to inspect state, risk warnings, local code explanations, no new debug logging. It never took over the investigation. The developer picked every next step.&lt;/li&gt;
&lt;li&gt;Fixes came before localization. The agent checked compilation and lint, then stopped. Whether the symptom went away was nobody's check, and several fixes did not survive the next run.&lt;/li&gt;
&lt;li&gt;Every runtime fact reached the agent through the chat window, pasted by hand. Its own terminal activity was limited to git.&lt;/li&gt;
&lt;li&gt;Debugging a distributed system means finding the invariant the defects share. We list four things an agent needs for that, which of them Explyt covers today, and which it does not.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The rule everything broke
&lt;/h2&gt;

&lt;p&gt;The developer summed up the scope of the session in one sentence:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Any operation on a TID must go through the single asynchronous path and must not let TMS and EMV drift apart.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;TID is a terminal ID. TMS and EMV are two downstream systems that have to agree about it. A rule like that lives between components; no single class owns it and the compiler does not check it. Each defect in the session was a place where the code stepped outside that rule.&lt;/p&gt;

&lt;p&gt;We are leaving the symptom, the defects and the code out on purpose, since they would identify the developer's project, and the shape of the session is enough to carry the lesson.&lt;/p&gt;

&lt;h2&gt;
  
  
  Helpful in the small, absent in the large
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Where the agent helped&lt;/th&gt;
&lt;th&gt;Where it left the developer alone&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Wrote SQL to look at the current rows in TMS and EMV&lt;/td&gt;
&lt;td&gt;Never proposed a theory of its own or an experiment to test it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flagged risks before touching code&lt;/td&gt;
&lt;td&gt;Presented verified facts and guesses in the same confident tone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Explained code correctly when asked&lt;/td&gt;
&lt;td&gt;Proposed fixes before the failure was localized&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Worked from existing log statements instead of adding new ones&lt;/td&gt;
&lt;td&gt;Never built a model of TID, TMS, EMV and what "consistent" means between them&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The left column is real value; a developer alone types all of that themselves. The right column is where the hours went.&lt;/p&gt;

&lt;p&gt;Take the third row. A fix arrives, it compiles, lint is clean, the developer applies it, and the next run breaks somewhere new. The log shows the agent verifying the build and stopping there. The symptom itself was never the acceptance criterion, for the agent or, in the moment, for the developer.&lt;/p&gt;

&lt;p&gt;The fourth row is the one we keep returning to. Without a model of the entities, every defect looks like a fresh bug with its own fix. With the model, they collapse into one rule violated in several places, and the search changes: the question stops being "why did this fail" and becomes "where else does TID state get written outside the async path".&lt;/p&gt;

&lt;h2&gt;
  
  
  Who ran the experiments
&lt;/h2&gt;

&lt;p&gt;Nobody but the developer. SQL snippets with their output, log lines from Spring Boot and from Kafka, HTTP responses, Kubernetes and Docker Compose output, configuration files: a human copied every one of them into the chat, while the agent's own terminal use stayed at git.&lt;/p&gt;

&lt;p&gt;Consider what that does to an agent's judgment. It sees the running system only through excerpts someone else chose, in an order someone else chose, with a framing already attached. From that position it cannot run an experiment; it can only comment on the experiments a human decided to run. The passivity in the table above follows directly from this, and no prompt wording changes it.&lt;/p&gt;

&lt;p&gt;The internal review of the session proposed two remedies. First, instruments: a browser and a terminal the agent operates itself, with instructions to use them. Second, a written debugging method: a journal of hypotheses, experiments and results, plus a running model of the data and its invariants. The references the review pointed to were &lt;a href="https://www.skills.sh/obra/superpowers/systematic-debugging" rel="noopener noreferrer"&gt;systematic-debugging&lt;/a&gt; from Superpowers and &lt;a href="https://www.skills.sh/wshobson/agents/parallel-debugging" rel="noopener noreferrer"&gt;parallel-debugging&lt;/a&gt; from wshobson/agents. Both remedies assume the agent can act and observe without a human relay. A method with no instruments is a checklist someone else has to fill in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four things a debugging agent needs
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;A written working theory outside the chat scroll, so that "verified" and "assumed" do not blur over a two-hour session.&lt;/li&gt;
&lt;li&gt;A way to execute the experiment behind the current hypothesis and read the outcome without a human in between.&lt;/li&gt;
&lt;li&gt;State observed first-hand: a variable value at a breakpoint, a row returned by a query the agent ran. A pasted excerpt is already somebody's interpretation.&lt;/li&gt;
&lt;li&gt;A standing question after the second defect in the same area: what do these two have in common?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The session had none of the four. The model was current; what it had to work with was a chat window, file access and git, pointed at an asynchronous payments pipeline. The harness is ours, and so are the gaps.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Explyt covers today
&lt;/h2&gt;

&lt;p&gt;We want to be precise here, because this is our product.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://explyt.ai/docs/explyt-test/tools/debugger" rel="noopener noreferrer"&gt;Debug mode&lt;/a&gt; addresses item 3 and part of item 2. It applies when the failure reproduces under a test, an application launch or one of the project's IDE run configurations. The agent confirms the cause in the debugger before editing: breakpoints, variable values, call stack. A minimal fix follows, and after it the same scenario runs again together with the related tests, which is the loop this session never entered.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5ssp2jqay8k30xijw084.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5ssp2jqay8k30xijw084.gif" alt="Explyt Debug mode on a Spring integration test: breakpoint, variable values, then the fix. Recorded on an earlier case, a different project from the Kafka session." width="640" height="360"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://explyt.ai/docs/explyt-test/tools/run-configurations" rel="noopener noreferrer"&gt;Run configurations&lt;/a&gt; address the rest of item 2. Builds and tests go through the IDE and come back as structured results: console output, test outcomes, compilation errors. The agent reads a red test instead of a green lint. The docs also say what a single debugger run does not prove: it confirms the fix for that scenario and does not stand in for the related tests.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcmu3md8d7blf40pfuga5.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcmu3md8d7blf40pfuga5.gif" alt="An Explyt session running a test through an IDE run configuration and reading the result back. Also recorded on a different project." width="760" height="321"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Items 1 and 4 stay open. Explyt's &lt;a href="https://explyt.ai/docs/explyt-test/configuration/memory-bank" rel="noopener noreferrer"&gt;Memory Bank&lt;/a&gt; persists facts and project agreements between chats; a per-bug hypothesis journal is a different artifact, and the product does not keep one for you. Reading Kafka topics, querying your database or pulling Kubernetes logs unaided is likewise absent from the documentation. That evidence still comes from the developer, or from an &lt;a href="https://explyt.ai/docs/explyt-test/tools" rel="noopener noreferrer"&gt;MCP server&lt;/a&gt; you attach, and each server decides what it exposes.&lt;/p&gt;

&lt;p&gt;So, had the failure in this log been reproducible in a test or a run configuration, Debug mode would have altered the middle of the loop: the state at the moment of the write would have been visible, and each fix would have been followed by a rerun instead of a lint pass. The edges would have stayed where they were. Someone still brings the infrastructure evidence, and someone still notices that the newest defect rhymes with the first one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try this on Monday
&lt;/h2&gt;

&lt;p&gt;Nothing in this section describes the session. It is what we would set up before the next one, with any agent that reads repository instructions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1. Give the working theory a file.&lt;/strong&gt; One per bug, outside the chat:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Working theory: &amp;lt;bug slug&amp;gt;&lt;/span&gt;

&lt;span class="gu"&gt;## Verified&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="nt"&gt;&amp;lt;fact&amp;gt;&lt;/span&gt; (how: breakpoint in X / query Y / test Z, when: &lt;span class="nt"&gt;&amp;lt;timestamp&amp;gt;&lt;/span&gt;)

&lt;span class="gu"&gt;## Assumed&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="nt"&gt;&amp;lt;claim&lt;/span&gt; &lt;span class="na"&gt;nobody&lt;/span&gt; &lt;span class="na"&gt;has&lt;/span&gt; &lt;span class="na"&gt;checked&lt;/span&gt; &lt;span class="na"&gt;yet&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;

&lt;span class="gu"&gt;## Current hypothesis&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;one&lt;/span&gt; &lt;span class="na"&gt;sentence&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
Experiment that would refute it: &lt;span class="nt"&gt;&amp;lt;command&lt;/span&gt; &lt;span class="err"&gt;/&lt;/span&gt; &lt;span class="na"&gt;test&lt;/span&gt; &lt;span class="err"&gt;/&lt;/span&gt; &lt;span class="na"&gt;breakpoint&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;

&lt;span class="gu"&gt;## Invariant candidates&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="nt"&gt;&amp;lt;a&lt;/span&gt; &lt;span class="na"&gt;rule&lt;/span&gt; &lt;span class="na"&gt;that&lt;/span&gt; &lt;span class="na"&gt;more&lt;/span&gt; &lt;span class="na"&gt;than&lt;/span&gt; &lt;span class="na"&gt;one&lt;/span&gt; &lt;span class="na"&gt;defect&lt;/span&gt; &lt;span class="na"&gt;has&lt;/span&gt; &lt;span class="na"&gt;violated&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 2. Make the agent keep it.&lt;/strong&gt; A rule for &lt;code&gt;AGENTS.md&lt;/code&gt; or your vendor's equivalent:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Debugging&lt;/span&gt;

Maintain .tasks/&lt;span class="nt"&gt;&amp;lt;bug&amp;gt;&lt;/span&gt;/theory.md and update it after each experiment.
Every entry under "Verified" names the command, test or breakpoint
that verified it. Everything else goes under "Assumed".

Do not propose a fix until the failure is reproduced in a test or a
run configuration and the cause is confirmed at a breakpoint or in
a query result you executed yourself.

Success means the original scenario passes on rerun. A clean build
and a clean linter are preconditions for a fix; they do not confirm it.

When a second defect appears in the same area, stop. Write down what
the two have in common before fixing either.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 3. Start under the debugger.&lt;/strong&gt; Pick one failing test, reproduce it with a breakpoint, and make the agent read the variables before it edits anything. If your agent has no debugger, the rule still helps: every time "Verified" stays empty, the missing instrument is visible in the journal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;One session, one developer, our own agent, reviewed after the fact by one person who did not run it. We are not publishing the defects or the code. The log also does not say whether a local test could reproduce the failure, which is what Debug mode requires. Treat this as one case and nothing more.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to go next
&lt;/h2&gt;

&lt;p&gt;This post keeps the story short on purpose. The full write-up, with the session walked through step by step and a closer look at each of the four failures, is on our blog:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;→ &lt;a href="https://explyt.ai/en/blog/one-symptom-ten-problems-kafka-invariant" rel="noopener noreferrer"&gt;Read the full case on explyt.ai: One symptom, ten problems&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Our earlier case, &lt;a href="https://explyt.ai/en/blog/superpowers-vs-debugger-explyt" rel="noopener noreferrer"&gt;Superpowers vs Plain Old Debugger in Explyt&lt;/a&gt;, covers the opposite shape: one bug, one cause, two agent runs.&lt;/p&gt;

&lt;p&gt;When your agent fixed a bug and something else broke right after, did you stop to ask what the two had in common? Tell us the invariant you found in the comments.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Site version of this case: &lt;a href="https://explyt.ai/en/blog/one-symptom-ten-problems-kafka-invariant" rel="noopener noreferrer"&gt;One symptom, ten problems: a Kotlin + Spring Boot + Kafka debugging case&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Earlier case: &lt;a href="https://explyt.ai/en/blog/superpowers-vs-debugger-explyt" rel="noopener noreferrer"&gt;Superpowers vs Plain Old Debugger in Explyt&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Explyt docs: &lt;a href="https://explyt.ai/docs/explyt-test/tools/debugger" rel="noopener noreferrer"&gt;Debug mode&lt;/a&gt;, &lt;a href="https://explyt.ai/docs/explyt-test/tools/run-configurations" rel="noopener noreferrer"&gt;Run configurations&lt;/a&gt;, &lt;a href="https://explyt.ai/docs/explyt-test/tools" rel="noopener noreferrer"&gt;Tools and integrations&lt;/a&gt;, &lt;a href="https://explyt.ai/docs/explyt-test/configuration/memory-bank" rel="noopener noreferrer"&gt;Memory Bank&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Debugging method references: &lt;a href="https://www.skills.sh/obra/superpowers/systematic-debugging" rel="noopener noreferrer"&gt;obra/superpowers systematic-debugging&lt;/a&gt;, &lt;a href="https://www.skills.sh/wshobson/agents/parallel-debugging" rel="noopener noreferrer"&gt;wshobson/agents parallel-debugging&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>machinelearning</category>
      <category>news</category>
      <category>github</category>
      <category>ai</category>
    </item>
    <item>
      <title>I gave two AI agents the same bug and the same model. One patched based on a guess from the stack trace. The other set a real breakpoint, ran the JetBrains debugger, and saw what actually broke.</title>
      <dc:creator>Viktoria</dc:creator>
      <pubDate>Fri, 18 Sep 2026 15:05:24 +0000</pubDate>
      <link>https://dev.to/marketing_explyt_a7b53da9/i-gave-two-ai-agents-the-same-bug-and-the-same-model-one-patched-based-on-a-guess-from-the-stack-58nn</link>
      <guid>https://dev.to/marketing_explyt_a7b53da9/i-gave-two-ai-agents-the-same-bug-and-the-same-model-one-patched-based-on-a-guess-from-the-stack-58nn</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/marketing_explyt_a7b53da9/superpowers-vs-plain-old-debugger-in-explyt-3271" class="crayons-story__hidden-navigation-link"&gt;Superpowers vs Plain Old Debugger in Explyt&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/marketing_explyt_a7b53da9" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4074488%2F591a7871-a646-4ad3-a738-9f5034840513.png" alt="marketing_explyt_a7b53da9 profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/marketing_explyt_a7b53da9" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Viktoria
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Viktoria
                
                
              
              &lt;div id="story-author-preview-content-4564899" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/marketing_explyt_a7b53da9" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4074488%2F591a7871-a646-4ad3-a738-9f5034840513.png" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Viktoria&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/marketing_explyt_a7b53da9/superpowers-vs-plain-old-debugger-in-explyt-3271" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 3&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/marketing_explyt_a7b53da9/superpowers-vs-plain-old-debugger-in-explyt-3271" id="article-link-4564899"&gt;
          Superpowers vs Plain Old Debugger in Explyt
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/programming"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;programming&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/community"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;community&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/machinelearning"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;machinelearning&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/marketing_explyt_a7b53da9/superpowers-vs-plain-old-debugger-in-explyt-3271" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;5&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/marketing_explyt_a7b53da9/superpowers-vs-plain-old-debugger-in-explyt-3271#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            6 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Ever watch your AI agent rename the wrong thing because it just grep'd for a string match? Your JetBrains IDE already knows better — real symbol resolution, real callers, real refactoring. Now Claude Code, Codex, Cursor &amp; OpenCode can tap into it, free.</title>
      <dc:creator>Viktoria</dc:creator>
      <pubDate>Fri, 18 Sep 2026 15:01:44 +0000</pubDate>
      <link>https://dev.to/marketing_explyt_a7b53da9/ever-watch-your-ai-agent-rename-the-wrong-thing-because-it-just-grepd-for-a-string-match-your-24pp</link>
      <guid>https://dev.to/marketing_explyt_a7b53da9/ever-watch-your-ai-agent-rename-the-wrong-thing-because-it-just-grepd-for-a-string-match-your-24pp</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/marketing_explyt_a7b53da9/give-claude-code-your-jetbrains-ide-tools-for-free-471" class="crayons-story__hidden-navigation-link"&gt;Give Claude Code your JetBrains IDE tools for free&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/marketing_explyt_a7b53da9" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4074488%2F591a7871-a646-4ad3-a738-9f5034840513.png" alt="marketing_explyt_a7b53da9 profile" class="crayons-avatar__image" width="" height=""&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/marketing_explyt_a7b53da9" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Viktoria
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Viktoria
                
                
              
              &lt;div id="story-author-preview-content-4444101" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/marketing_explyt_a7b53da9" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4074488%2F591a7871-a646-4ad3-a738-9f5034840513.png" class="crayons-avatar__image" alt="" width="96" height="96"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Viktoria&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/marketing_explyt_a7b53da9/give-claude-code-your-jetbrains-ide-tools-for-free-471" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Aug 20&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/marketing_explyt_a7b53da9/give-claude-code-your-jetbrains-ide-tools-for-free-471" id="article-link-4444101"&gt;
          Give Claude Code your JetBrains IDE tools for free
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/programming"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;programming&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/community"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;community&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/llm"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;llm&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/marketing_explyt_a7b53da9/give-claude-code-your-jetbrains-ide-tools-for-free-471" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/raised-hands-74b2099fd66a39f2d7eed9305ee0f4553df0eb7b4f11b01b6b1b499973048fe5.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;4&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/marketing_explyt_a7b53da9/give-claude-code-your-jetbrains-ide-tools-for-free-471#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            6 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Has this happened on your team? How do you check whether a developer still owns the code the agent wrote?</title>
      <dc:creator>Viktoria</dc:creator>
      <pubDate>Tue, 15 Sep 2026 13:47:20 +0000</pubDate>
      <link>https://dev.to/marketing_explyt_a7b53da9/has-this-happened-on-your-team-how-do-you-check-whether-a-developer-still-owns-the-code-the-agent-1emf</link>
      <guid>https://dev.to/marketing_explyt_a7b53da9/has-this-happened-on-your-team-how-do-you-check-whether-a-developer-still-owns-the-code-the-agent-1emf</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/marketing_explyt_a7b53da9/the-agent-wrote-the-code-the-tests-are-green-the-linter-is-clean-you-cant-explain-what-changed-53mn" class="crayons-story__hidden-navigation-link"&gt;The agent wrote the code. The tests are green. The linter is clean. You can't explain what changed.&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/marketing_explyt_a7b53da9" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4074488%2F591a7871-a646-4ad3-a738-9f5034840513.png" alt="marketing_explyt_a7b53da9 profile" class="crayons-avatar__image" width="96" height="96"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/marketing_explyt_a7b53da9" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Viktoria
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Viktoria
                
                
              
              &lt;div id="story-author-preview-content-4659957" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/marketing_explyt_a7b53da9" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4074488%2F591a7871-a646-4ad3-a738-9f5034840513.png" class="crayons-avatar__image" alt="" width="96" height="96"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Viktoria&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/marketing_explyt_a7b53da9/the-agent-wrote-the-code-the-tests-are-green-the-linter-is-clean-you-cant-explain-what-changed-53mn" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 15&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/marketing_explyt_a7b53da9/the-agent-wrote-the-code-the-tests-are-green-the-linter-is-clean-you-cant-explain-what-changed-53mn" id="article-link-4659957"&gt;
          The agent wrote the code. The tests are green. The linter is clean. You can't explain what changed.
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/programming"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;programming&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/javascript"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;javascript&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/python"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;python&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/marketing_explyt_a7b53da9/the-agent-wrote-the-code-the-tests-are-green-the-linter-is-clean-you-cant-explain-what-changed-53mn" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;5&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/marketing_explyt_a7b53da9/the-agent-wrote-the-code-the-tests-are-green-the-linter-is-clean-you-cant-explain-what-changed-53mn#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            8 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>The agent wrote the code. The tests are green. The linter is clean. You can't explain what changed.</title>
      <dc:creator>Viktoria</dc:creator>
      <pubDate>Tue, 15 Sep 2026 13:43:01 +0000</pubDate>
      <link>https://dev.to/marketing_explyt_a7b53da9/the-agent-wrote-the-code-the-tests-are-green-the-linter-is-clean-you-cant-explain-what-changed-53mn</link>
      <guid>https://dev.to/marketing_explyt_a7b53da9/the-agent-wrote-the-code-the-tests-are-green-the-linter-is-clean-you-cant-explain-what-changed-53mn</guid>
      <description>&lt;p&gt;There is a moment in a code review where a developer is asked why the code does what it does, and the answer is: "Claude wrote it." The code works. The tests pass. The reviewer approves. Nobody on the call can explain the architecture of the PR, including the person who submitted it.&lt;/p&gt;

&lt;p&gt;At Explyt, where we build an AI agent for JetBrains IDEs, two independent customer development interviews reported the same timeline: a month of agent-driven development is enough to lose the mental model of the code. The tests are green, the linter is clean, and the developer who shipped the PR cannot answer a question about it in review. The code works. The model is gone.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Two independent custdevs reported the same timeline: a month of agent-driven development, and the developer can no longer explain the code they are shipping.&lt;/li&gt;
&lt;li&gt;Compilation, green tests, and auto-LGTM verify the artifact. They do not verify that the author understands the invariant the code protects. A PR can pass all three on code nobody in the room can explain.&lt;/li&gt;
&lt;li&gt;The community calls this &lt;a href="https://www.reddit.com/r/webdev/comments/1rxznlw/comprehension_debt_the_silent_time_bomb_a_lot_of/" rel="noopener noreferrer"&gt;comprehension debt&lt;/a&gt;. It shows up before review metrics move. The first signal is a question the developer can't answer.&lt;/li&gt;
&lt;li&gt;The IDE keeps the running program. The chat transcript keeps the text. Understanding lives in the gap between them, and the debugger is where you close it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The one-month timeline
&lt;/h2&gt;

&lt;p&gt;Two customer development interviews, independent of each other, converged on the same number. A month of agent-driven development — where the agent writes the code and the developer reviews the diff — is enough to lose the mental model.&lt;/p&gt;

&lt;p&gt;The pattern is the same in both reports. The developer starts by reviewing every line. Within a week, they review the summary and skim the diff. Within two weeks, they approve based on the green tests and the agent's description. Within a month, they are asked a question about the code in a review and cannot answer it. The code works. The model is gone.&lt;/p&gt;

&lt;p&gt;What makes this worth reporting is the convergence: two people who never spoke to each other named the same timeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  What green tests don't prove
&lt;/h2&gt;

&lt;p&gt;Compilation, green tests, and auto-LGTM — the three signals most pipelines use to accept a change — can all pass on code nobody understands.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signal&lt;/th&gt;
&lt;th&gt;What it checks&lt;/th&gt;
&lt;th&gt;What it doesn't check&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Compiler&lt;/td&gt;
&lt;td&gt;Types match&lt;/td&gt;
&lt;td&gt;Intent is correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Test suite&lt;/td&gt;
&lt;td&gt;Code matches its own assumptions&lt;/td&gt;
&lt;td&gt;Assumptions are correct&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Auto-LGTM&lt;/td&gt;
&lt;td&gt;Pipeline ran&lt;/td&gt;
&lt;td&gt;Author can explain the result&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The compiler checks types, not intent. A method that compiles can still violate the invariant it was written to protect. The test suite checks the code against its own assumptions, and if the assumptions are wrong, the tests pass and the bug ships. Auto-LGTM checks that the pipeline ran, not that the author can explain the result.&lt;/p&gt;

&lt;p&gt;All three verify the artifact. None of them verify the author's model of the artifact. A developer who cannot explain why the code does what it does can still ship a PR that compiles, passes tests, and gets an auto-merge.&lt;/p&gt;

&lt;h2&gt;
  
  
  The day development stops
&lt;/h2&gt;

&lt;p&gt;There is a moment where the agent fails to fix the next bug, and the developer cannot step in.&lt;/p&gt;

&lt;p&gt;The developer who reviewed every line can step in. The developer who approved based on green tests for a month cannot — they don't know where the invariant lives, what the side effects are, or which module owns the state. Development stops because the human ran out of context, not because the agent ran out of capability. The tool is still there. The understanding is not.&lt;/p&gt;

&lt;p&gt;Reddit practitioners describe this as the limit of &lt;a href="https://www.reddit.com/r/ClaudeAI/comments/1uuf59z/im_a_vibe_coder_and_im_scared_that_i_have_no_idea/" rel="noopener noreferrer"&gt;"vibe coding"&lt;/a&gt; — the practice of accepting agent output without rebuilding a mental model. The practice works until the agent fails, and then the developer who was vibing is stuck. They can't debug what they don't understand, and they can't understand what they never read.&lt;/p&gt;

&lt;h2&gt;
  
  
  Explain it in your own words
&lt;/h2&gt;

&lt;p&gt;A practice reported from custdev: before accepting a diff, explain the change in your own words, without looking at the agent's summary.&lt;/p&gt;

&lt;p&gt;Compilation and green tests don't prove the author understands the invariant. A verbal explanation does. If you can close the diff, describe what changed, why it changed, and what breaks if it reverts — you own the change. If you can't, the diff owns you.&lt;/p&gt;

&lt;p&gt;The check takes a minute. Close the diff. Close the agent's summary. Open a blank document. Write: "This change does X because Y. If reverted, Z breaks." If you can't fill in X, Y, and Z from memory, you don't understand the change, and you shouldn't merge it.&lt;/p&gt;

&lt;p&gt;A &lt;a href="https://www.reddit.com/r/ExperiencedDevs/comments/1vg0cx8/what_do_you_do_when_a_developer_submits_ai/" rel="noopener noreferrer"&gt;Reddit team practice&lt;/a&gt; adds a social version: every PR needs a named engineer who can later explain the change from memory. LLM use is allowed, but one person must be able to answer for the result and the breakage. The rule does not ask the agent to explain itself — it asks the human to.&lt;/p&gt;

&lt;h2&gt;
  
  
  One day a week
&lt;/h2&gt;

&lt;p&gt;Another practice from custdev: one day a week, stop feature work. Spend the day simplifying architecture, reviewing code, and regaining understanding.&lt;/p&gt;

&lt;p&gt;The developer who reported this practice described it as maintenance — of the mental model, not of the code. The code is maintained by the agent. The model is maintained by the human. The model decays faster than the code, because the agent generates code faster than the human can rebuild the model, and the gap widens every week.&lt;/p&gt;

&lt;p&gt;One day a week is a practice one developer described, not a recommendation. What it tells you is that ownership is a recurring cost, not a one-time investment. You don't buy it by reading the code once. You pay it continuously, or you lose it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the IDE keeps that the chat doesn't
&lt;/h2&gt;

&lt;p&gt;The chat transcript keeps the text. The IDE keeps the running program. Understanding lives in the gap between them, and the debugger is where you close it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run configurations&lt;/strong&gt; let the agent run the test through the IDE and get back console output, test results, and compilation errors as a structured result. The developer sees which tests ran, which failed, and what the failure was — a structured result, not a line that says "tests passed."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvbatfl3oguq3nvaxtm7p.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvbatfl3oguq3nvaxtm7p.gif" alt="Run configurations: the IDE runs the test and returns the result. The developer sees what ran, not just that it passed." width="760" height="321"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Debug mode&lt;/strong&gt; asks the agent to set a breakpoint, read the variable values, inspect the call stack — before proposing a fix. The variable value at the breakpoint is an observation the chat cannot give you. It is also the moment where understanding gets rebuilt: you see the state, you see the decision point, you see why the code does what it does.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd5j4qtfhx9schujdv01z.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd5j4qtfhx9schujdv01z.gif" alt="Debug mode: breakpoint, variable values, call stack. The moment where the mental model gets rebuilt." width="640" height="360"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agent Changes&lt;/strong&gt; keeps the diff in a scoped panel separate from the chat. Each change is individually accepted or rejected. The diff is where the developer rebuilds the model — by reading the code that changed and asking whether it matches the intent, not by reading the agent's summary of the intent.&lt;/p&gt;

&lt;p&gt;None of these replace the human in the loop. They give the human something to loop on that is more than text. The chat says "the code does X." The IDE says "the variable at line 47 is null, and here is the call stack that got you there." The first is a claim. The second is a fact. Understanding starts from the second.&lt;/p&gt;

&lt;p&gt;Anthropic's own guidance points the same way. &lt;a href="https://code.claude.com/docs/en/best-practices" rel="noopener noreferrer"&gt;Claude Code best practices&lt;/a&gt;: "Claude stops when the work looks done. Without a check it can run, 'looks done' is the only signal available, and you become the verification loop." The check the IDE provides — a test result with names and failures, a debugger with variable values — is a check the chat cannot give you, and it is also where the mental model gets rebuilt. Verification and understanding come from the same instrument.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try this on your next PR
&lt;/h2&gt;

&lt;p&gt;Copy this rule into your &lt;code&gt;AGENTS.md&lt;/code&gt; or team checklist. Every line is a suggestion, not a claim — adapt it to your project.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Comprehension check (before merge)&lt;/span&gt;
&lt;span class="p"&gt;
1.&lt;/span&gt; Close the diff and the agent's summary.
&lt;span class="p"&gt;2.&lt;/span&gt; Open a blank document. Write:
   "This change does X because Y. If reverted, Z breaks."
&lt;span class="p"&gt;3.&lt;/span&gt; If you can't fill in X, Y, and Z from memory:
&lt;span class="p"&gt;   -&lt;/span&gt; Reopen the diff.
&lt;span class="p"&gt;   -&lt;/span&gt; Set a breakpoint at the decision point.
&lt;span class="p"&gt;   -&lt;/span&gt; Watch the variable. Read the call stack.
&lt;span class="p"&gt;   -&lt;/span&gt; Try again.
&lt;span class="p"&gt;4.&lt;/span&gt; If you still can't explain it after step 3, don't merge it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The check takes a minute. The cost of skipping it shows up in a month.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Isn't this just a skills problem? Get better at reading code and it goes away.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Reading code is the floor. The custdevs who reported the one-month timeline were experienced developers. The model decayed because the agent generates code faster than the human can read it, not because the human can't read. A faster reader buys you more time; the gap still widens.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If the agent can fix bugs on its own, why does my understanding matter?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because the agent can't fix every bug. The moment it fails — the edge case, the production-only issue, the interaction between three modules — you are the one who steps in. If you can't, development stops. Two custdevs described the same month-long window after which they could no longer step in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Won't better tests solve this?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Tests check the code against its own assumptions. If the assumptions are wrong, the tests pass and the bug ships. Tests verify the artifact, not the author's model of it. You can have 100% coverage and zero understanding.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is this against using AI agents for code?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The "explain it in your own words" check and the "one day a week" practice come from developers who use agents daily. They are not arguing against agents. They are describing what they do to keep owning the code the agent writes. The question is not whether to use the agent. The question is what you do alongside it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The agent writes the code. The tests are green. The linter is clean. The question is not whether the code works — the question is whether you can still explain it. If you can't, the code owns you, not the other way around.&lt;/p&gt;

&lt;p&gt;The next time the agent produces a diff, close the chat. Open the diff. Explain the change in your own words: what changed, why, and what breaks if it reverts. If you can't, set a breakpoint at the decision point and watch the variable. That is one fact the chat cannot give you, and it is where understanding starts.&lt;/p&gt;

&lt;p&gt;If you work in a JetBrains IDE, &lt;a href="https://plugins.jetbrains.com/plugin/27979-explyt-ai-agent" rel="noopener noreferrer"&gt;Explyt on the Marketplace&lt;/a&gt; gives the agent the IDE run and the debugger. The rest is on you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Related reading&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/en/blog/one-symptom-ten-problems-kafka-invariant" rel="noopener noreferrer"&gt;One symptom, ten problems: a Kotlin + Spring Boot + Kafka debugging case&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/en/blog/agent-commit-created-but-it-isnt" rel="noopener noreferrer"&gt;Commit created — but it isn't&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Explyt, Run configurations: &lt;a href="https://explyt.ai/docs/explyt-test/tools/run-configurations" rel="noopener noreferrer"&gt;https://explyt.ai/docs/explyt-test/tools/run-configurations&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Explyt, Debug mode: &lt;a href="https://explyt.ai/docs/explyt-test/tools/debugger" rel="noopener noreferrer"&gt;https://explyt.ai/docs/explyt-test/tools/debugger&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Explyt on JetBrains Marketplace: &lt;a href="https://plugins.jetbrains.com/plugin/27979-explyt-ai-agent" rel="noopener noreferrer"&gt;https://plugins.jetbrains.com/plugin/27979-explyt-ai-agent&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Explyt blog, One symptom, ten problems: &lt;a href="https://explyt.ai/en/blog/one-symptom-ten-problems-kafka-invariant" rel="noopener noreferrer"&gt;https://explyt.ai/en/blog/one-symptom-ten-problems-kafka-invariant&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Explyt blog, Commit created — but it isn't: &lt;a href="https://explyt.ai/en/blog/agent-commit-created-but-it-isnt" rel="noopener noreferrer"&gt;https://explyt.ai/en/blog/agent-commit-created-but-it-isnt&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Anthropic, Claude Code best practices: &lt;a href="https://code.claude.com/docs/en/best-practices" rel="noopener noreferrer"&gt;https://code.claude.com/docs/en/best-practices&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Anthropic, Building effective agents: &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;https://www.anthropic.com/engineering/building-effective-agents&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Reddit, Comprehension debt: the silent time bomb: &lt;a href="https://www.reddit.com/r/webdev/comments/1rxznlw/comprehension_debt_the_silent_time_bomb_a_lot_of/" rel="noopener noreferrer"&gt;https://www.reddit.com/r/webdev/comments/1rxznlw/comprehension_debt_the_silent_time_bomb_a_lot_of/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Reddit, What do you do when a developer submits AI generated code they clearly don't understand?: &lt;a href="https://www.reddit.com/r/ExperiencedDevs/comments/1vg0cx8/what_do_you_do_when_a_developer_submits_ai/" rel="noopener noreferrer"&gt;https://www.reddit.com/r/ExperiencedDevs/comments/1vg0cx8/what_do_you_do_when_a_developer_submits_ai/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Reddit, I'm a vibe coder and I'm scared that I have no idea what I'm doing anymore: &lt;a href="https://www.reddit.com/r/ClaudeAI/comments/1uuf59z/im_a_vibe_coder_and_im_scared_that_i_have_no_idea/" rel="noopener noreferrer"&gt;https://www.reddit.com/r/ClaudeAI/comments/1uuf59z/im_a_vibe_coder_and_im_scared_that_i_have_no_idea/&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>javascript</category>
      <category>python</category>
    </item>
    <item>
      <title>Why the agent needs access to the environment, not a human’s retelling of it</title>
      <dc:creator>Viktoria</dc:creator>
      <pubDate>Thu, 10 Sep 2026 18:50:00 +0000</pubDate>
      <link>https://dev.to/marketing_explyt_a7b53da9/why-the-agent-needs-access-to-the-environment-not-a-humans-retelling-of-it-4odp</link>
      <guid>https://dev.to/marketing_explyt_a7b53da9/why-the-agent-needs-access-to-the-environment-not-a-humans-retelling-of-it-4odp</guid>
      <description>&lt;p&gt;&lt;em&gt;A retold log is a filtered log. The agent should query the database, read the broker, hit the endpoint itself — and record what it saw.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;There is a moment in almost every debugging session with an AI agent where the developer stops being a developer and becomes a clipboard. The agent asks for the error. The developer runs the command, scrolls, selects the part that looks relevant, pastes it. The agent asks for the row in the database. The developer opens a SQL client, runs a query, pastes the result. The agent asks what the endpoint returned. Another terminal, another paste.&lt;/p&gt;

&lt;p&gt;Each paste is a decision about what matters, made by the person who does not yet know what matters.&lt;/p&gt;

&lt;p&gt;I am the product manager at Explyt, where we build an AI agent for JetBrains IDEs. I have written before about a session in our own agent where the developer pasted SQL results, Kafka and Spring logs, HTTP responses and Kubernetes messages by hand while the agent's terminal use was limited to git (&lt;a href="https://explyt.ai/en/blog/one-symptom-ten-problems-kafka-invariant" rel="noopener noreferrer"&gt;One symptom, ten problems&lt;/a&gt;). That session sent me looking for how common the pattern is outside our own logs. The issue trackers are full of it, and Anthropic's engineering posts describe the fix while the same company's prompt examples still say "[paste error]".&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Public issues on Claude Code and Copilot, plus Reddit threads about Cursor and Claude Code, document the same loop: the agent cannot see the terminal, the database or the CI logs, so the human copies output into the chat, and the agent reasons about the copy.&lt;/li&gt;
&lt;li&gt;The copy passes through several filters: what the human chose to run, what they chose to select, what the tooling truncated without saying so, and what the agent's own compaction later summarized away. Each filter removes the kind of detail that debugging turns on.&lt;/li&gt;
&lt;li&gt;When the agent does have access, it often leaves it unused. Issues show agents guessing column names with SSH access available, confirming a production config without checking, and asking the human to run a query the agent's own MCP server could run.&lt;/li&gt;
&lt;li&gt;Anthropic's engineering guidance says agents need "ground truth" from the environment at each step. The instruments to get it exist as MCP servers for databases, Kafka and Kubernetes, most with read-only modes.&lt;/li&gt;
&lt;li&gt;Direct access is half of the fix. The other half is a record of what the agent queried, what came back, and which conclusion rests on which observation.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The relay loop, in the words of the people stuck in it
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/anthropics/claude-code/issues/69360" rel="noopener noreferrer"&gt;anthropics/claude-code #69360&lt;/a&gt; (June 2026): "Claude Code (VS Code extension) cannot read the existing terminal output. When a user runs a command that produces errors, they must manually copy-paste the output into the chat. This adds friction to the debugging workflow and breaks the natural "run → see error → ask Claude" loop."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/anthropics/claude-code/issues/13865" rel="noopener noreferrer"&gt;anthropics/claude-code #13865&lt;/a&gt; (December 2025) spells out the steps and names the consequence: "Currently, if I want to debug a runtime issue, I have to: 1. Manually ask Claude to add print/log statements. 2. Manually run the app/script. 3. Copy/paste the terminal output back into Claude Code. 4. Ask Claude to analyze it. 5. Ask Claude to remove the logs afterwards. This manual "human-in-the-loop" data transfer is slow and error-prone, often leading to "shotgun debugging" where the model guesses the solution without concrete runtime evidence."&lt;/p&gt;

&lt;p&gt;The Copilot side has the same report from the other direction. In &lt;a href="https://github.com/microsoft/vscode-copilot-release/issues/10778" rel="noopener noreferrer"&gt;microsoft/vscode-copilot-release #10778&lt;/a&gt; (May 2025): "The agent fails to see the output from the terminal so it starts to run another command in a different way repeatedly failing to see the output. If I select the terminal output and select add to chat, it can see the output." An agent that cannot see the result of its own command reruns it and still sees nothing; in that thread the paste was the only channel that worked.&lt;/p&gt;

&lt;p&gt;Reddit adds the texture. A Cursor user in &lt;a href="https://www.reddit.com/r/ClaudeAI/comments/1psuh3t/can_claude_code_create_automation_loops_better/" rel="noopener noreferrer"&gt;r/ClaudeAI&lt;/a&gt; (December 2025): "Cursor also can't directly read and make use of the error logs in the PowerShell console. It feels broken for me to manually copy paste the logs to the Cursor agent." Another in &lt;a href="https://www.reddit.com/r/cursor/comments/1u3kvp6/open_in_terminal_pane_removal_is_breaking_the/" rel="noopener noreferrer"&gt;r/cursor&lt;/a&gt; (June 2026), after a tooling change: "Only option is copy the command and re-run it yourself, which kind of defeats the whole point of having an agent run things!"&lt;/p&gt;

&lt;p&gt;These are personal accounts with low vote counts, and I cite them as such. What they share with the GitHub issues is the shape: for everything outside the repository, the human carries the data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four filters between the environment and the agent
&lt;/h2&gt;

&lt;p&gt;A retold log is filtered whether or not anyone intends it.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. What to run
&lt;/h3&gt;

&lt;p&gt;The first filter is the choice of what to run. The developer decides which command, which query, which endpoint. If the hypothesis in their head is wrong, the evidence they collect is evidence for or against the wrong hypothesis, and the query that would have tested the right one never runs.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. What to select
&lt;/h3&gt;

&lt;p&gt;The second is the choice of what to select. Nobody pastes a 4,000-line log. They scroll to the stack trace or the last twenty lines and paste that. The warning three seconds earlier and the second exception in a different thread stay in the terminal.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Truncation by tooling
&lt;/h3&gt;

&lt;p&gt;The third is the one nobody chooses: truncation by tooling. A &lt;a href="https://www.reddit.com/r/ClaudeCode/comments/1spiy8t/token_optimizers_for_ai_coding_agents_are/" rel="noopener noreferrer"&gt;r/ClaudeCode thread&lt;/a&gt; (April 2026) tested a popular token-saving wrapper. On a package listing: "The remaining 314 packages are invisible. There is no truncation indicator. The output looks complete." In a separate example on log levels, the same author describes the downstream effect: "An AI agent doing incident triage reads the logs, sees one error (a transient retry), and concludes it's a minor blip. It applies a small fix and closes the investigation." Truncation is also the default in the agent itself: "For Claude Code, we restrict tool responses to 25,000 tokens by default", from Anthropic's &lt;a href="https://www.anthropic.com/engineering/writing-tools-for-agents" rel="noopener noreferrer"&gt;Writing effective tools for AI agents&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Compaction
&lt;/h3&gt;

&lt;p&gt;The fourth looks like memory. Anthropic's &lt;a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents" rel="noopener noreferrer"&gt;context engineering post&lt;/a&gt; describes compaction as "summarizing its contents, and reinitiating a new context window with the summary", where "the model preserves architectural decisions, unresolved bugs, and implementation details while discarding redundant tool outputs or messages." The word "redundant" is doing a lot of work there. In a debugging session, the tool output that looked redundant at the time of compaction is often the one you need two hours later.&lt;/p&gt;

&lt;p&gt;Stack the four and the agent is reasoning about a document written by a human with a hypothesis, cut to fit a chat box, truncated by a wrapper, and then summarized. It is a wonder it gets anything right.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwpko10t49qhlls0oh2vx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwpko10t49qhlls0oh2vx.png" alt="Four filters between the environment and the agent. Each one removes the kind of detail debugging turns on." width="800" height="448"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Four filters between the environment and the agent. Each one removes the kind of detail debugging turns on.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  When the agent has access and still asks you
&lt;/h2&gt;

&lt;p&gt;Direct access turns out to be necessary and, on its own, insufficient, because agents with access often leave it unused.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/anthropics/claude-code/issues/43387" rel="noopener noreferrer"&gt;anthropics/claude-code #43387&lt;/a&gt; (April 2026) is the sharpest example I found. The agent "Wrote 1,700 lines of PHP with 28 database actions using completely guessed column names — never ran SHOW COLUMNS or any test query against the real database." The reporter adds: "Claude has SSH access to the server and could run SHOW COLUMNS, curl endpoints, or test queries — but never does before saying "done"." Thirty of forty-six queries referenced columns that did not exist while the SSH session sat unused.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/anthropics/claude-code/issues/60107" rel="noopener noreferrer"&gt;anthropics/claude-code #60107&lt;/a&gt; (May 2026, a report the user had the agent itself write up) moves the same failure to production: "Claude Code confirmed a Postgres dialect option (&lt;code&gt;statement_timeout&lt;/code&gt;) as valid for a Heroku-hosted production app when I explicitly asked "are you sure this is a valid option?". The agent answered "yes" without verifying against the actual prod environment. The change passed all local tests (local Postgres in Docker has no pgbouncer), shipped to production, and crashed every dyno at boot." Local and production gave different answers to the same question, and the agent queried neither.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/anthropics/claude-code/issues/86718" rel="noopener noreferrer"&gt;anthropics/claude-code #86718&lt;/a&gt; (August 2026) is the near-miss: the agent "without verifying, concluded that a database record … had been overwritten/corrupted … and stated this as fact", and "had already suggested deleting the affected database record before I stopped it." The agent guessed that data was damaged and proposed deleting it before anyone had checked.&lt;/p&gt;

&lt;p&gt;And from the MCP side, a &lt;a href="https://www.reddit.com/r/cursor/comments/1loojez/serverpostgres_mcp_still_working_for_anyone/" rel="noopener noreferrer"&gt;r/cursor user&lt;/a&gt; (July 2025) with a Postgres MCP server connected: "Tried 4 different models to try to get it to query my dev database, all attempts fail with the model just asking me to run the query myself." A &lt;a href="https://www.reddit.com/r/ClaudeCode/comments/1ux854f/significant_drop_in_intelligence_today_on_opus_48/" rel="noopener noreferrer"&gt;r/ClaudeCode user&lt;/a&gt; (July 2026) on CI: "it tried a bunch of local builds before then asking me to paste in the logs from the failed CI deployment. It has access to the CI logs, and has used them in the past."&lt;/p&gt;

&lt;p&gt;So the instruments have to be there, and the agent has to be held to using them. The second part is a harness and prompting problem, and I come back to it below.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Anthropic says the agent should do
&lt;/h2&gt;

&lt;p&gt;Anthropic, &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;Building effective agents&lt;/a&gt;: "During execution, it's crucial for the agents to gain "ground truth" from the environment at each step (such as tool call results or code execution) to assess its progress."&lt;/p&gt;

&lt;p&gt;Anthropic, &lt;a href="https://code.claude.com/docs/en/best-practices" rel="noopener noreferrer"&gt;Claude Code best practices&lt;/a&gt;: "Claude stops when the work looks done. Without a check it can run, "looks done" is the only signal available, and you become the verification loop: every mistake waits for you to notice it." The same page defines the check as "anything that returns a signal Claude can read in the conversation: a test suite, a build exit code, a linter, a script that diffs output against a fixture, or a browser screenshot compared against a design".&lt;/p&gt;

&lt;p&gt;The same page also, in its prompt-example table, recommends writing "[paste error]" to the agent. I mention it because it shows how deep the relay habit runs: even the document that tells you to make the agent its own verification loop assumes the human will hand it the error.&lt;/p&gt;

&lt;p&gt;The context engineering post gives the alternative, and the alternative is many small retrievals made by the agent itself: "Letting agents navigate and retrieve data autonomously also enables progressive disclosure", and "The model can write targeted queries, store results, and leverage Bash commands like head and tail to analyze large volumes of data without ever loading the full data objects into context."&lt;/p&gt;

&lt;p&gt;The research on long contexts points the same way. &lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;Lost in the Middle&lt;/a&gt; (Liu et al., 2023) found that "performance can degrade significantly when changing the position of relevant information", and Chroma's &lt;a href="https://research.trychroma.com/context-rot" rel="noopener noreferrer"&gt;Context Rot&lt;/a&gt; report (2025), across 18 models, concludes that "models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows." A bigger paste makes the second problem worse while leaving the first four filters in place.&lt;/p&gt;

&lt;h2&gt;
  
  
  The instruments exist
&lt;/h2&gt;

&lt;p&gt;"Query the database, read the broker, hit the endpoint" sounds like a wish list. As of September 2026 it is a list of MCP servers, most with a read-only switch. Star counts and dates are from the GitHub API on 2026-09-08.&lt;/p&gt;

&lt;h3&gt;
  
  
  Databases
&lt;/h3&gt;

&lt;p&gt;For databases: &lt;a href="https://github.com/crystaldba/postgres-mcp" rel="noopener noreferrer"&gt;crystaldba/postgres-mcp&lt;/a&gt; (3.3k stars) has a "Restricted Mode" that "Limits operations to read-only transactions and imposes constraints on resource utilization" and is described as "suitable for production environments". &lt;a href="https://github.com/bytebase/dbhub" rel="noopener noreferrer"&gt;bytebase/dbhub&lt;/a&gt; (3.5k stars, release 2026-09-02) covers Postgres, MySQL, SQL Server, MariaDB and SQLite with "Read-only mode, row limiting, and query timeout to prevent runaway operations". Google's &lt;a href="https://github.com/googleapis/mcp-toolbox" rel="noopener noreferrer"&gt;MCP Toolbox for Databases&lt;/a&gt; (16k stars) ships "Prebuilt generic tools for instant data exploration (e.g., list_tables, execute_sql)". The original reference Postgres server from the Model Context Protocol project is archived; the current reference set is filesystem, git, fetch, memory and a few others.&lt;/p&gt;

&lt;h3&gt;
  
  
  Kafka
&lt;/h3&gt;

&lt;p&gt;For Kafka: &lt;a href="https://github.com/confluentinc/mcp-confluent" rel="noopener noreferrer"&gt;confluentinc/mcp-confluent&lt;/a&gt; (v1.5.0, July 2026) exposes &lt;code&gt;list-topics&lt;/code&gt;, &lt;code&gt;consume-messages&lt;/code&gt; and &lt;code&gt;list-consumer-groups&lt;/code&gt;, and also &lt;code&gt;produce-message&lt;/code&gt;, &lt;code&gt;create-topics&lt;/code&gt; and &lt;code&gt;delete-topics&lt;/code&gt;. It has no read-only flag, but its &lt;code&gt;--allow-tools&lt;/code&gt; and &lt;code&gt;--block-tools&lt;/code&gt; options let you start it with the write tools removed, which is how I would run it for debugging.&lt;/p&gt;

&lt;h3&gt;
  
  
  Kubernetes
&lt;/h3&gt;

&lt;p&gt;For Kubernetes: &lt;a href="https://github.com/containers/kubernetes-mcp-server" rel="noopener noreferrer"&gt;containers/kubernetes-mcp-server&lt;/a&gt; (2.1k stars, actively maintained) has &lt;code&gt;--read-only&lt;/code&gt;, which "will not allow any write operations (create, update, delete)", plus &lt;code&gt;--disable-destructive&lt;/code&gt;, and a &lt;code&gt;pods_log&lt;/code&gt; tool that does the obvious thing.&lt;/p&gt;

&lt;h3&gt;
  
  
  HTTP
&lt;/h3&gt;

&lt;p&gt;For HTTP: the reference &lt;a href="https://github.com/modelcontextprotocol/servers/tree/main/src/fetch" rel="noopener noreferrer"&gt;fetch server&lt;/a&gt; converts pages to markdown, so it fetches web pages and does no API work. For "hit the endpoint itself" you want the agent's terminal with &lt;code&gt;curl&lt;/code&gt;, or a purpose-built server.&lt;/p&gt;

&lt;p&gt;Each of these servers exposes access the developer already has, behind a permission model.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fss22npkr4m19dkos1xu5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fss22npkr4m19dkos1xu5.png" alt="Connecting an MCP server in Explyt (STDIO transport). The docs example is the sequential-thinking server; a database or Kafka server is configured the same way, with its read-only flags in the arguments." width="800" height="619"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Connecting an MCP server in Explyt (STDIO transport). The docs example is the sequential-thinking server; a database or Kafka server is configured the same way, with its read-only flags in the arguments.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The other half: record what it saw
&lt;/h2&gt;

&lt;p&gt;I disagree with the simplest version of my own thesis. Give the agent a terminal, a database and a broker, and you have replaced a filtered log with an unfiltered stream that nobody reads. The developer stops being the clipboard and becomes the person who has to trust a conclusion with no visible chain of evidence.&lt;/p&gt;

&lt;p&gt;The fix is the one that works for human engineers: a record. Which query ran, against which environment, what came back, and which conclusion rests on which result. Anthropic's tooling post recommends the reviewer's version of this: "Review the raw transcripts (including tool calls and tool responses) to catch any behavior not explicitly described in the agent's CoT." The agent's version is to keep that record itself, in the open, as it works.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdco4zsdf3q7oj7fwuhq3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdco4zsdf3q7oj7fwuhq3.png" alt="Every claim about the environment carries the observation it came from: the query, the environment and the timestamp next to the value." width="800" height="448"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Every claim about the environment carries the observation it came from: the query, the environment and the timestamp next to the value.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In practice I now hold any agent I use for debugging to a short set of rules.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Every claim about the environment carries the observation it came from: the query, the environment and the timestamp next to the value, so "the row has &lt;code&gt;status = PENDING&lt;/code&gt;" arrives with "query at 14:02 against staging".&lt;/li&gt;
&lt;li&gt;The environment is named every time, because #60107 above is what happens when local and production are treated as one.&lt;/li&gt;
&lt;li&gt;And the record survives compaction: if the agent summarizes its own context, the observations stay in a file that the summary links to.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The problem sits in the harness as much as in the model. #43387 and the two MCP stories above show models with access declining to use it. A rule in the repository's agent instructions that says "before stating anything about database, broker or runtime state, query it and quote the result" is blunt. In my own sessions it removed most of the guessed claims.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Explyt stands
&lt;/h2&gt;

&lt;p&gt;Explyt is our agent for JetBrains IDEs, and I want to state what its documentation does and does not cover on this topic.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://explyt.ai/docs/explyt-test/tools" rel="noopener noreferrer"&gt;tools documentation&lt;/a&gt; frames the goal in the same terms as this article: tools "give the agent facts about your project and let it perform actions without you manually copying long logs and search results into the chat."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4dugwkhoo61drw4qk6i8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4dugwkhoo61drw4qk6i8.png" alt="Once connected, each MCP server appears in the Tools menu as its own group next to the built-in IDE tools. Screenshot from the Explyt documentation." width="462" height="556"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Once connected, each MCP server appears in the Tools menu as its own group next to the built-in IDE tools. Screenshot from the Explyt documentation.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Two of those sources come from the IDE itself, not from anything a database or a broker exposes, and I start with them because they are the ones a paste cannot replace.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://explyt.ai/docs/explyt-test/tools/run-configurations" rel="noopener noreferrer"&gt;Run configurations&lt;/a&gt;: the agent runs a configured test, build or application through the IDE and "returns the result to the agent: console output, test results, compilation errors." That closes the "copy the terminal output back" loop from #13865 for anything the IDE knows how to run. The docs are explicit that in Rider the tool currently does not build the project and does not collect coverage.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs9z0ieygxj6hm4d7uooi.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs9z0ieygxj6hm4d7uooi.gif" alt="Run configurations: the IDE runs the test and returns console output, test results and compilation errors to the agent — the " width="760" height="321"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Run configurations: the IDE runs the test and returns console output, test results and compilation errors to the agent — the "copy the terminal output back" step from #13865 disappears for anything the IDE can run.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://explyt.ai/docs/explyt-test/tools/debugger" rel="noopener noreferrer"&gt;Debug mode documentation&lt;/a&gt; states the principle for the runtime case in one line: "The cause is confirmed by breakpoints, variable values, and the call stack, not just by a guess based on the code or the log." Replace "the log" with "the pasted log" and you have this article. A variable read at a breakpoint is an observation no log contains, pasted or not; an agent without a debugger is back in the relay loop for that class of bug, whatever else it can query.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpreeaxrsp00ukdgiz75r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpreeaxrsp00ukdgiz75r.png" alt="Frame from the official Explyt docs video (Debug mode); different project from the one in this article. The agent sets its own breakpoints and reads the debugger, instead of asking for a pasted log." width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Frame from the official Explyt docs video (Debug mode); different project from the one in this article. The agent sets its own breakpoints and reads the debugger, instead of asking for a pasted log.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The rest of the environment comes through the two channels any agent has. The terminal tools, per the &lt;a href="https://explyt.ai/docs/explyt-test/configuration/reference/tools-reference" rel="noopener noreferrer"&gt;tools reference&lt;/a&gt;: "The terminal tools run commands and manage processes. You need them when there is no suitable IDE configuration or when the task involves an external utility." The same page lists what to check before a command runs, because "A command can change files, a database, the environment or a remote system." This is the tool that runs &lt;code&gt;psql&lt;/code&gt; or &lt;code&gt;kubectl&lt;/code&gt; when you allow it.&lt;/p&gt;

&lt;p&gt;MCP: "Once an MCP server is connected, its tools appear in the Tools menu as a separate group. The names, parameters and consequences of a call are defined by the server itself." The &lt;a href="https://explyt.ai/docs/explyt-test/tools/mcp" rel="noopener noreferrer"&gt;MCP page&lt;/a&gt;, in its GitHub example, suggests to "Test the connection with a safe read operation first", and notes that each MCP result block "contains the outcome and diagnostic details of the call, which are useful when verifying a result or investigating an error."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frq0glw8ieklreq1rf7hm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frq0glw8ieklreq1rf7hm.png" alt="An MCP server call as it appears in the Explyt chat (Chrome MCP in the docs example). The result block keeps the outcome and diagnostics of each call in the transcript — the " width="799" height="327"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;An MCP server call as it appears in the Explyt chat (Chrome MCP in the docs example). The result block keeps the outcome and diagnostics of each call in the transcript — the "record what it saw" half of the argument, for MCP. Screenshot from the Explyt documentation.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;What the documentation does not describe is any built-in Kafka, SQL or Kubernetes access. Those come through the terminal or through an MCP server you connect and scope, and they land in the same session as the run configurations and the debugger, so one transcript holds the test result, the variable value and the query result side by side. And pasting remains supported: the &lt;a href="https://explyt.ai/docs/explyt-test/agent/chat" rel="noopener noreferrer"&gt;chat documentation&lt;/a&gt; says large pastes become attachments and "The agent receives the complete text." Pasting should stay available as the fallback; the agent's own query should be the default.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqpnnekuy57vytfgzvm5i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqpnnekuy57vytfgzvm5i.png" alt="Frame from the official Explyt docs video; different project. The agent's own reply — " width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Frame from the official Explyt docs video; different project. The agent's own reply — "Please paste the log output…" — is the relay habit this article is about. Pasting stays available as the fallback; the agent's own query should be the default.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;The GitHub issues above are individual reports; I have not measured how often the relay loop occurs or how often agents with access decline to use it. Reddit threads are personal accounts with low scores. MCP server statistics are a snapshot from one day and the read-only behavior comes from the READMEs; I have not audited the servers myself. The Explyt claims are limited to what the documentation says; the session I referenced is described in the linked article and its details are not published.&lt;/p&gt;

&lt;h2&gt;
  
  
  One thing to try this week
&lt;/h2&gt;

&lt;p&gt;Pick the next runtime bug where you would normally paste a log. Connect one read-only source the agent can query itself, a database MCP in restricted mode or the terminal with &lt;code&gt;kubectl logs&lt;/code&gt;, and add one line to your agent instructions: state nothing about the environment without quoting the query and the result. Then read the transcript and count how many of the agent's claims come with an observation attached. I expected the count to be low the first time I did this, and it was.&lt;/p&gt;

&lt;p&gt;If you work in a JetBrains IDE, &lt;a href="https://plugins.jetbrains.com/plugin/27979-explyt-ai-agent" rel="noopener noreferrer"&gt;Explyt on the Marketplace&lt;/a&gt; gives the agent the two sources a paste cannot replace, the IDE run and the debugger, and takes MCP servers for the rest.&lt;/p&gt;

&lt;p&gt;What did your agent last ask you to paste that it could have fetched itself, and what happened when you gave it the access? The comments are open; I am collecting the cases where the query the human never ran turned out to be the one that mattered.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Issue reports and threads&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/anthropics/claude-code/issues/69360" rel="noopener noreferrer"&gt;anthropics/claude-code #69360, terminal output not readable&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/anthropics/claude-code/issues/13865" rel="noopener noreferrer"&gt;anthropics/claude-code #13865, Debug Mode feature request&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/anthropics/claude-code/issues/43387" rel="noopener noreferrer"&gt;anthropics/claude-code #43387, guessed column names with SSH access available&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/anthropics/claude-code/issues/60107" rel="noopener noreferrer"&gt;anthropics/claude-code #60107, production config confirmed without checking&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/anthropics/claude-code/issues/86718" rel="noopener noreferrer"&gt;anthropics/claude-code #86718, guessed data corruption&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/microsoft/vscode-copilot-release/issues/10778" rel="noopener noreferrer"&gt;microsoft/vscode-copilot-release #10778, agent cannot see terminal output&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.reddit.com/r/ClaudeCode/comments/1spiy8t/token_optimizers_for_ai_coding_agents_are/" rel="noopener noreferrer"&gt;r/ClaudeCode, token optimizers drop lines silently&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.reddit.com/r/ClaudeCode/comments/1ux854f/significant_drop_in_intelligence_today_on_opus_48/" rel="noopener noreferrer"&gt;r/ClaudeCode, agent asks for CI logs it can access&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.reddit.com/r/cursor/comments/1loojez/serverpostgres_mcp_still_working_for_anyone/" rel="noopener noreferrer"&gt;r/cursor, Postgres MCP connected, agent asks human to run the query&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.reddit.com/r/cursor/comments/1u3kvp6/open_in_terminal_pane_removal_is_breaking_the/" rel="noopener noreferrer"&gt;r/cursor, terminal pane removal&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.reddit.com/r/ClaudeAI/comments/1psuh3t/can_claude_code_create_automation_loops_better/" rel="noopener noreferrer"&gt;r/ClaudeAI, copy-pasting PowerShell logs to Cursor&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Vendor and research&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;Anthropic, Building effective agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://code.claude.com/docs/en/best-practices" rel="noopener noreferrer"&gt;Anthropic, Claude Code best practices&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents" rel="noopener noreferrer"&gt;Anthropic, Effective context engineering for AI agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.anthropic.com/engineering/writing-tools-for-agents" rel="noopener noreferrer"&gt;Anthropic, Writing effective tools for AI agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;Liu et al., Lost in the Middle (2023)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://research.trychroma.com/context-rot" rel="noopener noreferrer"&gt;Chroma Research, Context Rot (2025)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;MCP servers&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/crystaldba/postgres-mcp" rel="noopener noreferrer"&gt;crystaldba/postgres-mcp&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/bytebase/dbhub" rel="noopener noreferrer"&gt;bytebase/dbhub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/googleapis/mcp-toolbox" rel="noopener noreferrer"&gt;googleapis/mcp-toolbox&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/confluentinc/mcp-confluent" rel="noopener noreferrer"&gt;confluentinc/mcp-confluent&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/containers/kubernetes-mcp-server" rel="noopener noreferrer"&gt;containers/kubernetes-mcp-server&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/modelcontextprotocol/servers/tree/main/src/fetch" rel="noopener noreferrer"&gt;modelcontextprotocol/servers, fetch&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Explyt documentation&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/docs/explyt-test/tools" rel="noopener noreferrer"&gt;Tools and integrations&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/docs/explyt-test/tools/run-configurations" rel="noopener noreferrer"&gt;Run configurations&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/docs/explyt-test/configuration/reference/tools-reference" rel="noopener noreferrer"&gt;Tools reference (terminal, MCP)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/docs/explyt-test/tools/mcp" rel="noopener noreferrer"&gt;MCP&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/docs/explyt-test/agent/chat" rel="noopener noreferrer"&gt;Chat (paste handling)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/docs/explyt-test/tools/debugger" rel="noopener noreferrer"&gt;Debugging with Debug mode&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/en/blog/one-symptom-ten-problems-kafka-invariant" rel="noopener noreferrer"&gt;Explyt blog, One symptom, ten problems&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://plugins.jetbrains.com/plugin/27979-explyt-ai-agent" rel="noopener noreferrer"&gt;Explyt on JetBrains Marketplace&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>javascript</category>
    </item>
    <item>
      <title>Explyt 5.19 adds an OpenAI-compatible gateway to the subscription. Here is the setup for Claude Code and Cursor.</title>
      <dc:creator>Viktoria</dc:creator>
      <pubDate>Wed, 09 Sep 2026 15:22:00 +0000</pubDate>
      <link>https://dev.to/marketing_explyt_a7b53da9/explyt-519-adds-an-openai-compatible-gateway-to-the-subscription-here-is-the-setup-for-claude-2lk</link>
      <guid>https://dev.to/marketing_explyt_a7b53da9/explyt-519-adds-an-openai-compatible-gateway-to-the-subscription-here-is-the-setup-for-claude-2lk</guid>
      <description>&lt;p&gt;If you use more than one coding agent, you already know the tax: a separate model account for each, a separate card on file, a separate place where the balance runs out. Explyt 5.19 removes that for a specific set of tools. The subscription that used to power only the Explyt plugin in JetBrains IDEs now gives you an OpenAI-compatible URL and an API key, and the key works in Claude Code, Claude Desktop, Cursor, Codex, Cline, OpenCode and Kilo Code.&lt;/p&gt;

&lt;p&gt;This post is the practical version: where the values are, three configs you can paste, which models answer, and the boundary that the release notes put in bold.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 0: get the two values
&lt;/h2&gt;

&lt;p&gt;Open your &lt;a href="https://app.explyt.ai/" rel="noopener noreferrer"&gt;Explyt account&lt;/a&gt;. With an active subscription or a positive Flex minutes balance you will see a block called &lt;strong&gt;Connect to Explyt&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8lgi7x6ih8al9u2ivmcu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8lgi7x6ih8al9u2ivmcu.png" alt="The Connect to Explyt block in the personal account" width="799" height="635"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It holds two things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The URL &lt;code&gt;https://api.explyt.ai/openai/api&lt;/code&gt;. Same for everyone.&lt;/li&gt;
&lt;li&gt;Your API key. &lt;strong&gt;Copy key&lt;/strong&gt; copies it; the eye icon reveals it.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If the block is not there, the account has neither a subscription nor Flex minutes, and there is nothing to connect yet.&lt;/p&gt;

&lt;p&gt;Keep the key like a password. It is one key for every tool; &lt;strong&gt;Reissue&lt;/strong&gt; in the account kills the old one instantly, and you then update every tool you configured.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Claude Code
&lt;/h2&gt;

&lt;p&gt;From &lt;a href="https://explyt.ai/docs/explyt/connect/claude-code" rel="noopener noreferrer"&gt;the Claude Code page&lt;/a&gt;. Open or create &lt;code&gt;~/.claude/settings.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"env"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_BASE_URL"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.explyt.ai/openai/api"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ANTHROPIC_API_KEY"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&amp;lt;YOUR_API_KEY&amp;gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Start Claude Code. When it asks &lt;em&gt;Do you want to use this API key?&lt;/em&gt;, answer yes. If it misbehaves later, &lt;code&gt;/logout&lt;/code&gt; in the chat and start again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Codex CLI
&lt;/h2&gt;

&lt;p&gt;From &lt;a href="https://explyt.ai/docs/explyt/connect/codex-cli" rel="noopener noreferrer"&gt;the Codex CLI page&lt;/a&gt;. Open or create &lt;code&gt;~/.codex/config.toml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="py"&gt;model&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"auto"&lt;/span&gt;
&lt;span class="py"&gt;model_provider&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"explyt"&lt;/span&gt;
&lt;span class="py"&gt;model_catalog_json&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"./model-catalog-full.json"&lt;/span&gt;

&lt;span class="nn"&gt;[model_providers.explyt]&lt;/span&gt;
&lt;span class="py"&gt;name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"explyt"&lt;/span&gt;
&lt;span class="py"&gt;base_url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"https://api.explyt.ai/openai/api"&lt;/span&gt;
&lt;span class="py"&gt;wire_api&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"responses"&lt;/span&gt;
&lt;span class="py"&gt;requires_openai_auth&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two more files sit next to it, both described on the docs page: &lt;code&gt;model-catalog-full.json&lt;/code&gt; with the list of Explyt models (a ready-made file is linked there) and &lt;code&gt;~/.codex/auth.json&lt;/code&gt; with &lt;code&gt;"auth_mode": "apikey"&lt;/code&gt; and your key. If you were signed in through a ChatGPT account before, run &lt;code&gt;codex logout&lt;/code&gt; first and then create &lt;code&gt;auth.json&lt;/code&gt;. One caveat from the same page: Codex CLI and the Codex app in ChatGPT share &lt;code&gt;config.toml&lt;/code&gt;, and the ChatGPT app can only use OpenAI models.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Cursor
&lt;/h2&gt;

&lt;p&gt;From &lt;a href="https://explyt.ai/docs/explyt/connect/cursor" rel="noopener noreferrer"&gt;the Cursor page&lt;/a&gt;. &lt;strong&gt;Settings → Models → API Keys&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Turn on &lt;strong&gt;OpenAI API Key&lt;/strong&gt; and paste the Explyt key.&lt;/li&gt;
&lt;li&gt;Turn on &lt;strong&gt;Override OpenAI Base URL&lt;/strong&gt; and paste &lt;code&gt;https://api.explyt.ai/openai/api&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fom86ha1xv6g3ov8vfld8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fom86ha1xv6g3ov8vfld8.png" alt="Cursor: the API Keys block with the Explyt key and URL" width="799" height="501"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Cursor allows custom keys only on a Pro plan or higher; the &lt;a href="https://explyt.ai/docs/explyt/connect/cursor" rel="noopener noreferrer"&gt;Cursor docs page&lt;/a&gt; spells this out, and the restriction comes from Cursor.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one rule that applies to every tool
&lt;/h2&gt;

&lt;p&gt;Do not also sign in to the tool's own account through ChatGPT or Anthropic. That sign-in sends requests around the Explyt gateway, and they get rejected.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everything else that connects
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Docs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Code&lt;/td&gt;
&lt;td&gt;CLI&lt;/td&gt;
&lt;td&gt;&lt;a href="https://explyt.ai/docs/explyt/connect/claude-code" rel="noopener noreferrer"&gt;page&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Codex CLI&lt;/td&gt;
&lt;td&gt;CLI&lt;/td&gt;
&lt;td&gt;&lt;a href="https://explyt.ai/docs/explyt/connect/codex-cli" rel="noopener noreferrer"&gt;page&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenCode&lt;/td&gt;
&lt;td&gt;CLI&lt;/td&gt;
&lt;td&gt;&lt;a href="https://explyt.ai/docs/explyt/connect/opencode" rel="noopener noreferrer"&gt;page&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Desktop&lt;/td&gt;
&lt;td&gt;Desktop app&lt;/td&gt;
&lt;td&gt;&lt;a href="https://explyt.ai/docs/explyt/connect/claude-desktop" rel="noopener noreferrer"&gt;page&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Codex in ChatGPT&lt;/td&gt;
&lt;td&gt;Desktop app&lt;/td&gt;
&lt;td&gt;&lt;a href="https://explyt.ai/docs/explyt/connect/codex-chatgpt-app" rel="noopener noreferrer"&gt;page&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cursor&lt;/td&gt;
&lt;td&gt;Editor&lt;/td&gt;
&lt;td&gt;&lt;a href="https://explyt.ai/docs/explyt/connect/cursor" rel="noopener noreferrer"&gt;page&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Code for VS Code&lt;/td&gt;
&lt;td&gt;Extension&lt;/td&gt;
&lt;td&gt;&lt;a href="https://explyt.ai/docs/explyt/connect/claude-code-vscode" rel="noopener noreferrer"&gt;page&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kilo Code&lt;/td&gt;
&lt;td&gt;Extension&lt;/td&gt;
&lt;td&gt;&lt;a href="https://explyt.ai/docs/explyt/connect/kilo-code" rel="noopener noreferrer"&gt;page&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cline&lt;/td&gt;
&lt;td&gt;Extension&lt;/td&gt;
&lt;td&gt;&lt;a href="https://explyt.ai/docs/explyt/connect/cline" rel="noopener noreferrer"&gt;page&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Explyt plugin (Community mode)&lt;/td&gt;
&lt;td&gt;JetBrains IDE plugin&lt;/td&gt;
&lt;td&gt;&lt;a href="https://explyt.ai/docs/explyt/connect/explyt" rel="noopener noreferrer"&gt;page&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Claude Desktop uses a dedicated gateway dialog: &lt;strong&gt;Configure third-party inference → Connection → Gateway&lt;/strong&gt;, URL into &lt;strong&gt;Gateway base URL&lt;/strong&gt;, key into &lt;strong&gt;Gateway API key&lt;/strong&gt;, scheme &lt;strong&gt;x-api-key&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7chfh74zm7autgwuppff.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7chfh74zm7autgwuppff.png" alt="Claude Desktop: configuring the Explyt gateway" width="800" height="649"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Anything not in the table that accepts an OpenAI-compatible base URL plus key will probably work; we have not verified it. Errors and their meanings are collected on the &lt;a href="https://explyt.ai/docs/explyt/connect/errors" rel="noopener noreferrer"&gt;error messages&lt;/a&gt; page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which models answer
&lt;/h2&gt;

&lt;p&gt;OpenAI and Anthropic models, plus &lt;code&gt;auto&lt;/code&gt;, which picks a model for you and has the highest availability. Many tools fetch the list from the gateway on their own. If yours does not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.explyt.ai/openai/api/models &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &amp;lt;YOUR_API_KEY&amp;gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The only restriction: the Codex app in ChatGPT works with OpenAI models and &lt;code&gt;auto&lt;/code&gt;. For Anthropic models in a Codex workflow use Codex CLI.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the key does not carry
&lt;/h2&gt;

&lt;p&gt;This is the part to read before you tell your team "we have Explyt in Cursor now".&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The key gives access to models and nothing else. In an external tool you will not get:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;IDE tools: project navigation, run configurations, the debugger, IDE-powered refactoring;&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;skills, modes, rules and checks of the Explyt agent;&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;subscription data inside the third-party tool's interface. Remaining minutes and the current chat's spend live in your personal account.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Paying for tokens through Explyt does not make Claude Code aware of your IntelliJ project's run configurations or let it stop on a breakpoint. Those capabilities live in the JetBrains IDE. The route for them is different and older: the &lt;a href="https://explyt.ai/docs/explyt-test/tools/external-agents-mcp" rel="noopener noreferrer"&gt;Explyt MCP server&lt;/a&gt; from 5.17, which exposes selected tool groups to an external agent, does no inference of its own, and needs no subscription.&lt;/p&gt;

&lt;p&gt;The key and the MCP server solve different problems and can work together: model through the gateway, IDE tools through MCP, both in one Claude Code session.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you also use the plugin
&lt;/h2&gt;

&lt;p&gt;The Explyt plugin in a JetBrains IDE has two ways to use the subscription. &lt;strong&gt;Personal&lt;/strong&gt; signs in from the plugin and knows about the subscription: remaining minutes, current chat spend, a live model list with parameters. &lt;strong&gt;Community&lt;/strong&gt; uses the same URL and key as any other tool and sees only an OpenAI-compatible provider. Minutes cost the same either way. Stay on Personal unless you have a reason not to.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two fixes in the same release
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Context indicator.&lt;/strong&gt; The chat's fill bar used to understate history size by about a third, because agent chats are mostly code and tool output, which tokenize more densely than the prose the estimate assumed. A chat at the model's limit looked half empty, auto-compression did not start, and the request failed at the provider. Now Explyt reconciles the estimate with the input token count the provider reports after each response, and keeps the correction with the chat across IDE restarts. Before the first response you still see the local estimate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MCP tokens.&lt;/strong&gt; Servers that authenticate through a browser hand out tokens that live about an hour. The first call after expiry failed. Explyt now renews ahead of expiry, one token per server. After the update, sign in once more for each such server. Manual-header and local servers are untouched.&lt;/p&gt;

&lt;p&gt;Also: fewer file-system checks for paths the agent mentioned but that do not exist in the project, and forks and subagents inherit the already compressed history rather than paying for the uncompressed one.&lt;/p&gt;

&lt;h2&gt;
  
  
  For teams
&lt;/h2&gt;

&lt;p&gt;One Explyt subscription per person covers every client that person uses. In Explyt Enterprise the gateway sits inside the company perimeter; the administrator issues the URL and key, and traffic from Cursor, Claude Code or Codex to the models goes through the corporate platform, not through personal vendor accounts. See the &lt;a href="https://explyt.ai/en/pricing" rel="noopener noreferrer"&gt;pricing page&lt;/a&gt; or &lt;a href="https://explyt.ai/en/contacts" rel="noopener noreferrer"&gt;contact sales&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it on Monday
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Open your &lt;a href="https://app.explyt.ai/" rel="noopener noreferrer"&gt;Explyt account&lt;/a&gt;, copy the key.&lt;/li&gt;
&lt;li&gt;Pick the one tool you open most and paste the config above.&lt;/li&gt;
&lt;li&gt;Send one request, then open the account and confirm the minutes moved.&lt;/li&gt;
&lt;li&gt;If the same tool also needs your IDE, wire the &lt;a href="https://explyt.ai/docs/explyt-test/tools/external-agents-mcp" rel="noopener noreferrer"&gt;MCP server&lt;/a&gt; next; it is a separate setup.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Which tool did you connect first, and did it pick up the model list on its own? Tell us in the comments.&lt;/p&gt;

&lt;p&gt;Full release notes: &lt;a href="https://explyt.ai/en/blog/release-explyt-5-19" rel="noopener noreferrer"&gt;Explyt 5.19 on explyt.ai&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>javascript</category>
      <category>python</category>
      <category>agents</category>
    </item>
    <item>
      <title>One symptom, ten problems: a Kotlin + Spring Boot + Kafka debugging case</title>
      <dc:creator>Viktoria</dc:creator>
      <pubDate>Tue, 08 Sep 2026 13:21:38 +0000</pubDate>
      <link>https://dev.to/marketing_explyt_a7b53da9/one-symptom-ten-problems-a-kotlin-spring-boot-kafka-debugging-case-k4j</link>
      <guid>https://dev.to/marketing_explyt_a7b53da9/one-symptom-ten-problems-a-kotlin-spring-boot-kafka-debugging-case-k4j</guid>
      <description>&lt;p&gt;&lt;em&gt;A single visible symptom led to many defects, half unrelated to it but all tied to one invariant about async TID processing. Debugging is invariant discovery.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;One visible symptom opened the session. By the end of it the developer and the agent had fixed many defects, about half of them unrelated to that symptom. All of them, though, sat inside the same scope, and the developer summed that scope up in one sentence:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Any operation on a TID must go through the single asynchronous path and must not let TMS and EMV drift apart.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I am the product manager at Explyt, where we build an AI agent for JetBrains IDEs. This session ran in our own agent, in regular chat mode, with a current frontier model, on a Kotlin and Java codebase with Spring Boot and Kafka in a payments system. I did not run it myself; I went through the log afterwards as part of our review of how people debug with an agent in practice. Debug mode was not used. What follows is a review of the shape of the work, with the session's own details left out. The defects are described as kinds, and the number in the title is the topic's working name; the log gives no count I could verify.&lt;/p&gt;

&lt;p&gt;Previously on the Explyt blog, we compared two agent runs on one JVM bug, one working from the log and one from a breakpoint: &lt;a href="https://explyt.ai/en/blog/superpowers-vs-debugger-explyt" rel="noopener noreferrer"&gt;Superpowers vs Plain Old Debugger in Explyt&lt;/a&gt;. That case was one bug with one cause. This one has the opposite shape, one symptom with many causes, and it stresses a different part of the debugging loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;One symptom, many defects, half of them unrelated to the symptom, all of them violating one invariant about asynchronous TID processing.&lt;/li&gt;
&lt;li&gt;Our agent was tactically helpful: it proposed SQL, warned about risks, explained code locally, and did not scatter new debug logging. It was strategically passive: the developer drove every step.&lt;/li&gt;
&lt;li&gt;It did not separate verified facts from hypotheses, and it proposed fixes before the problem was localized. Those fixes compiled and passed lint. Afterwards something else broke, differently.&lt;/li&gt;
&lt;li&gt;The developer was the agent's only sensor. Kafka and Spring logs, SQL results, HTTP responses, Kubernetes and Compose messages, config files: all pasted into the chat by hand. In the terminal the agent ran only git.&lt;/li&gt;
&lt;li&gt;Debugging a distributed system is a search for the invariant that all the defects violate. The agent needs instruments for that, and I say below which of them Explyt has today and which it does not.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What the session looked like
&lt;/h2&gt;

&lt;p&gt;I cannot show you the symptom or the defects; the log belongs to the developer's project. What I can describe is the shape.&lt;/p&gt;

&lt;p&gt;The developer came in with one observable problem. Working through it with the agent, they found a defect, fixed it, and something else broke, in a different way. Then another defect, another fix, another different breakage. About half of what they fixed had nothing to do with the original symptom. Every fix, however, touched the same scope: how the system handles a terminal ID (TID) across its asynchronous path and keeps two downstream systems, TMS and EMV, consistent. The sentence in bold above is how the developer described that scope.&lt;/p&gt;

&lt;p&gt;A few times the developer asked the agent to add something and later asked it to roll that back. That is the texture of the session: iterative, exploratory, with the agent along for the ride.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the agent did well
&lt;/h2&gt;

&lt;p&gt;The agent proposed SQL queries to inspect the current state in TMS and EMV, so the discussion started from actual rows. It warned about risks in code it was asked to change. It explained what a given piece of code did, when asked, and got it right. It did not add new debug logging; it worked from the log statements that already existed, which kept the diff clean.&lt;/p&gt;

&lt;p&gt;All of that is real help; a developer working alone would have done that typing themselves.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it fell apart
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Strategic passivity
&lt;/h3&gt;

&lt;p&gt;The first failure was strategic passivity. The agent did not run the investigation. It waited for the developer to bring the next piece of evidence, reacted to it, and waited again. At no point did it say: here is my current theory, here is what would falsify it, let me check. The developer decided what to look at next, every time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Facts and guesses in one voice
&lt;/h3&gt;

&lt;p&gt;The second: verified facts and guesses came out in the same confident voice. The developer had to keep track of which statements had been checked and which had not. In a session with many defects that bookkeeping is most of the job, and the agent was not doing it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Fixing before localizing
&lt;/h3&gt;

&lt;p&gt;Third, the agent proposed fixes before the problem was localized. The developer applied them; the log shows no rollback of those fixes, at least not in the chat. The agent checked that each change compiled and passed lint. It did not check that the fix removed the symptom. Several fixes turned out to be ineffective: something else broke afterwards, differently. Reading that part was uncomfortable. I have accepted a plausible fix at the end of a long day for the same reason: it compiled.&lt;/p&gt;

&lt;h3&gt;
  
  
  No model of the entities
&lt;/h3&gt;

&lt;p&gt;The fourth failure is the one that matters for tooling. The agent never built a model of how the entities relate: what TID is, what TMS and EMV each own, what "consistent" means between them. Without that model, every defect looked like a separate bug; with it, they are one invariant violated in several places.&lt;/p&gt;

&lt;h2&gt;
  
  
  The developer as the only sensor
&lt;/h2&gt;

&lt;p&gt;The developer pasted into the chat, by hand, fragments of SQL scripts and their results, Spring Boot and Kafka log excerpts, HTTP responses, messages from Kubernetes and Docker Compose, and configuration files. The agent's own use of the terminal was limited to git.&lt;/p&gt;

&lt;p&gt;So the agent's picture of the running system was whatever the developer chose to show it, in the order they chose, with their framing already applied. That is a poor position from which to form an independent theory, and it explains the passivity. This is the failure I keep coming back to, and I do not have a tidier answer than "give it instruments".&lt;/p&gt;

&lt;p&gt;The review proposed two fixes. Give the agent a browser and a terminal, and make it use them. And give it a debugging method with a written journal: hypotheses, experiments, results, and a working theory that includes the data model and its invariants, along the lines of the &lt;a href="https://www.skills.sh/obra/superpowers/systematic-debugging" rel="noopener noreferrer"&gt;Superpowers systematic-debugging skill&lt;/a&gt; and the &lt;a href="https://www.skills.sh/wshobson/agents/parallel-debugging" rel="noopener noreferrer"&gt;parallel-debugging&lt;/a&gt; variant. Both fixes assume the agent can run an experiment and read its result without a human in between. A method without instruments is a checklist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Debugging is invariant discovery
&lt;/h2&gt;

&lt;p&gt;The framing I took away from this log: debugging a distributed system is a search for the sentence that all the defects violate, and the individual fixes are almost a byproduct. A debugging agent therefore needs four things.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A working theory in writing, separate from the chat scroll, so that "checked" and "assumed" stay apart over a long session.&lt;/li&gt;
&lt;li&gt;The ability to run the experiment that would confirm or refute the current hypothesis, and to read the result itself.&lt;/li&gt;
&lt;li&gt;Direct observation of state, a variable at a breakpoint or a row from a query it ran, because pasted excerpts are already interpretations.&lt;/li&gt;
&lt;li&gt;After the second unrelated defect, the question of what the two have in common.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The session failed on all four. The model was a current frontier model. The harness gave it chat, file access and git, and asked it to debug an asynchronous payments pipeline. That is where the four failures come from, and it is our harness.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Explyt does about it today, and what it does not
&lt;/h2&gt;

&lt;p&gt;Explyt covers part of this loop today, and I want to be precise about which part.&lt;/p&gt;

&lt;p&gt;Explyt's documented &lt;a href="https://explyt.ai/docs/explyt-test/tools/debugger" rel="noopener noreferrer"&gt;Debug mode&lt;/a&gt; exists for the "observe state directly" requirement: for a bug you can reproduce by running a test, an application or an existing IDE run configuration, the agent is asked to confirm the cause under the debugger first, with breakpoints, variable values and the call stack, before it changes any code. Then a minimal fix, then a rerun of the original scenario and the related tests. That is the experiment loop this session never had.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5ssp2jqay8k30xijw084.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5ssp2jqay8k30xijw084.gif" alt="Explyt Debug mode on a Spring integration test: breakpoint, variable values, then the fix. From an earlier case; a different project from the Kafka session." width="640" height="360"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://explyt.ai/docs/explyt-test/tools/run-configurations" rel="noopener noreferrer"&gt;Run configurations&lt;/a&gt; close the other gap: the agent runs the build or the test through the IDE and gets back console output, test results and compilation errors as a structured tool result, so it has something to read beyond "compiles and lints". The docs are explicit that a single debugger run confirms the fix only for that scenario and does not replace the related tests.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcmu3md8d7blf40pfuga5.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcmu3md8d7blf40pfuga5.gif" alt="An Explyt session running a test through an IDE run configuration and reading the result back. Also a different project." width="760" height="321"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Nothing in the Explyt documentation reads Kafka topics, runs SQL against your database or pulls logs from Kubernetes on its own. Those sources still enter the session the way they did in this log, through the developer, or through an MCP server you connect yourself; the &lt;a href="https://explyt.ai/docs/explyt-test/tools" rel="noopener noreferrer"&gt;tools documentation&lt;/a&gt; describes MCP as the way to attach browsers, knowledge bases and other external services, and what a given server exposes is up to that server. Explyt's &lt;a href="https://explyt.ai/docs/explyt-test/configuration/memory-bank" rel="noopener noreferrer"&gt;Memory Bank&lt;/a&gt; keeps facts and project agreements across chats, which is useful, and it is a different thing from a per-session hypothesis journal. The working theory is still the developer's to keep.&lt;/p&gt;

&lt;p&gt;So if this symptom can be reproduced by a test or a run configuration, Debug mode would have changed the middle of the loop: the agent could have seen the actual state at the moment of the write, and rerun the scenario after each fix, with lint no longer the last check. It would not have changed the edges. Someone still has to bring the Kafka and infrastructure evidence, and someone still has to notice that the latest defect rhymes with the first one.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you want to try this on your own bug
&lt;/h2&gt;

&lt;p&gt;This part is a suggestion for your next session, and nothing here describes what happened in the one above.&lt;/p&gt;

&lt;p&gt;Whatever agent you use, put the invariant question into its instructions so it stops being your job to remember it. A repository-level rule along these lines works with most agents that read &lt;code&gt;AGENTS.md&lt;/code&gt; or a vendor equivalent:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Debugging&lt;/span&gt;

Keep a written working theory in a file under .tasks/, updated after
every experiment: what is verified (with the command or breakpoint that
verified it), what is assumed, and the current hypothesis.

Before proposing a fix, localize: reproduce the failure under a test or
run configuration and confirm the cause at a breakpoint or in a query
result you ran yourself.

After a fix, rerun the original scenario. "Compiles and lints" is not a result.

After the second defect in the same area, stop and write down what the
two have in common before fixing either.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then reproduce one failing test in the debugger and make the agent read the variables before it edits anything. If the agent has no debugger, the rule still helps; it just makes the missing instrument visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;The evidence here is one session log from one developer using our agent, reviewed by me after the fact. I did not run it and I am not publishing its details, defect list or code. Whether the symptom was reproducible under a local test, which is what Debug mode requires, I do not know from the log. Nothing here is a benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try the loop in a JetBrains IDE
&lt;/h2&gt;

&lt;p&gt;If you already work in a JetBrains IDE, &lt;a href="https://plugins.jetbrains.com/plugin/27979-explyt-ai-agent" rel="noopener noreferrer"&gt;install Explyt from the Marketplace&lt;/a&gt; and reproduce one failing test in Debug mode. Ask the agent to confirm the cause at a breakpoint before it edits anything, then make it rerun the scenario after the fix.&lt;/p&gt;

&lt;p&gt;When your agent fixed a bug and something else broke, did you go back and ask what the two had in common? I am collecting the invariants people found that way; the comments are open.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.skills.sh/obra/superpowers/systematic-debugging" rel="noopener noreferrer"&gt;Superpowers, systematic-debugging skill&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.skills.sh/wshobson/agents/parallel-debugging" rel="noopener noreferrer"&gt;wshobson/agents, parallel-debugging skill&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/docs/explyt-test/tools/debugger" rel="noopener noreferrer"&gt;Explyt documentation, Debugging with Debug mode&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/docs/explyt-test/tools/run-configurations" rel="noopener noreferrer"&gt;Explyt documentation, Run configurations&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/docs/explyt-test/tools" rel="noopener noreferrer"&gt;Explyt documentation, Tools and integrations&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/docs/explyt-test/configuration/memory-bank" rel="noopener noreferrer"&gt;Explyt documentation, Agent memory (Memory Bank)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://explyt.ai/en/blog/superpowers-vs-debugger-explyt" rel="noopener noreferrer"&gt;Explyt blog, Superpowers vs Plain Old Debugger in Explyt&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://plugins.jetbrains.com/plugin/27979-explyt-ai-agent" rel="noopener noreferrer"&gt;Explyt on JetBrains Marketplace&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
      <category>javascript</category>
    </item>
    <item>
      <title>Superpowers vs Plain Old Debugger in Explyt</title>
      <dc:creator>Viktoria</dc:creator>
      <pubDate>Thu, 03 Sep 2026 11:04:36 +0000</pubDate>
      <link>https://dev.to/marketing_explyt_a7b53da9/superpowers-vs-plain-old-debugger-in-explyt-3271</link>
      <guid>https://dev.to/marketing_explyt_a7b53da9/superpowers-vs-plain-old-debugger-in-explyt-3271</guid>
      <description>&lt;p&gt;An AI agent can follow a debugging checklist and still miss the bug.&lt;/p&gt;

&lt;p&gt;The question I kept coming back to was simple: who ran the experiment?&lt;/p&gt;

&lt;p&gt;I interviewed our users, mostly senior developers. In those interviews, debugging a hard problem came up most often. In practice, the investigation is still manual: a developer supplies logs, sets breakpoints, tests a theory, and pastes the result back into the chat.&lt;/p&gt;

&lt;p&gt;A systematic debugging method gives the work a clear order: gather facts, form a hypothesis, run an experiment, then change the code.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/obra/superpowers" rel="noopener noreferrer"&gt;Superpowers&lt;/a&gt; was created by &lt;a href="https://blog.fsck.com" rel="noopener noreferrer"&gt;Jesse Vincent&lt;/a&gt; and the &lt;a href="https://primeradiant.com" rel="noopener noreferrer"&gt;Prime Radiant&lt;/a&gt; team. Matt Pocock helped popularize it; the collection includes a &lt;a href="https://www.skills.sh/obra/superpowers/systematic-debugging" rel="noopener noreferrer"&gt;systematic-debugging skill&lt;/a&gt;. I ran Opus 5 on the same real bug and the same prompt twice. The only thing I changed was the debugging workflow: Superpowers in one run, Explyt's built-in Debug skill in the other.&lt;/p&gt;

&lt;p&gt;I ran one controlled case; it is not a benchmark of either product. One number makes the contrast plain: Explyt used 67k tokens to solve this case, while Superpowers used 132k. The transcripts explain the gap: the debugger supplied a runtime fact early, while the other run spent more context on speculative code, tests, and refactoring.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug: one protected Windows directory, one crashed coroutine
&lt;/h2&gt;

&lt;p&gt;The test case came from a Kotlin project on Windows. The IDE log contained:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;java.nio.file.AccessDeniedException: C:\Config.Msi
    at java.nio.file.Files.list(Files.java:3422)
    at FacadeMentionCandidates.listDir(FacadeMentionCandidates.kt:83)
    at FacadeMentionCandidates.listChildren(FacadeMentionCandidates.kt:69)
    at MentionSuggestionMapperKt$suggestionsFor$2.invokeSuspend(MentionSuggestionMapper.kt:36)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The relevant code first checked whether a path was a directory and then enumerated it with &lt;code&gt;Files.list&lt;/code&gt;. The exception came from a protected Windows Installer directory. &lt;code&gt;FacadeMentionCandidates&lt;/code&gt; makes filesystem roots available to the &lt;code&gt;@&lt;/code&gt; mention flow, so a user navigating through a path such as &lt;code&gt;@/C:/&lt;/code&gt; can eventually reach a directory that the IDE process cannot enumerate.&lt;/p&gt;

&lt;p&gt;The error was not handled at the flow boundary. It escaped the suggestion flow and was reported as an unhandled coroutine exception in &lt;code&gt;WorkspaceBootstrap&lt;/code&gt;, with the IDE blaming the plugin. The same listing path is also used while searching around the focused file, so this was a small function with more than one way to fail.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run one: a plausible patch from Superpowers
&lt;/h2&gt;

&lt;p&gt;The Superpowers run started well. It searched the log instead of reading the whole file, followed the stack into &lt;code&gt;FacadeMentionCandidates&lt;/code&gt;, found the unguarded &lt;code&gt;Files.list&lt;/code&gt;, and stated a plausible root-cause hypothesis.&lt;/p&gt;

&lt;p&gt;Then the investigation stopped before the experiment. The run never set a debugger breakpoint or inspected a value from the failing execution. The proposed fix was to wrap the directory listing in a &lt;code&gt;try/catch&lt;/code&gt;, log the exception, and return &lt;code&gt;emptyList()&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That choice made the failure disappear from the call stack, but it changed the meaning of the result. An empty directory and a directory denied by ACL became the same value; other I/O failures still had no user-facing outcome. The mention UI could not tell the developer what happened, and the option to mention the directory remained indistinguishable from a normal empty-folder case.&lt;/p&gt;

&lt;p&gt;The regression test followed the same assumption. It asserted that an unreadable directory produced no children. On Windows, the first version could not create the required POSIX permission state and was skipped. To force the exception, the agent then added a &lt;code&gt;readDirectory&lt;/code&gt; function to the production constructor so the test could inject a throwing fake. That seam existed for the test, not for the product behavior.&lt;/p&gt;

&lt;p&gt;The run finished with seven passing tests, but those tests established that the workaround returned an empty list. They did not establish that the user-facing behavior was correct, or that the error had been verified against a real paused execution. I discarded the patch: it hid the distinction we needed and left a test-only abstraction in the production API.&lt;/p&gt;

&lt;p&gt;That is how a checklist turns into a ritual. Every step appears in the transcript; the experiment never does. A plausible local edit stands in for evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run two: Explyt puts the hypothesis inside the IDE
&lt;/h2&gt;

&lt;p&gt;With Explyt's built-in debugging skill, the model started from the same log and source files, but it could use the JetBrains debugger as part of the investigation.&lt;/p&gt;

&lt;p&gt;It created a minimal reproducer, set a line breakpoint on the &lt;code&gt;Files.list&lt;/code&gt; call, and ran the test under the debugger. When execution stopped, the value was concrete:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Debugger evaluation&lt;/th&gt;
&lt;th&gt;Observed value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;dir&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;C:\Config.Msi&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Files.isDirectory(dir)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;true&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Files.isReadable(dir)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;false&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Files.list(dir)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;throws &lt;code&gt;java.nio.file.AccessDeniedException&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The call stack matched the production trace: &lt;code&gt;listDir&lt;/code&gt; → &lt;code&gt;listChildren&lt;/code&gt; → &lt;code&gt;MentionSuggestionMapper.suggestionsFor&lt;/code&gt;, inside the coroutine that builds mention suggestions. The debugger turned "a protected path probably causes this" into a verified statement about the running program:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;the directory check passes, the process cannot read the directory, and the unguarded enumeration kills the suggestion flow.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The evidence kept the fix at the filesystem boundary, with focused regression coverage and the normal listing path left intact. The mention protocol and runtime model stayed unchanged, and debug prints were unnecessary.&lt;/p&gt;

&lt;p&gt;The narrow fix passed review and is now in production. A &lt;code&gt;try/catch&lt;/code&gt; was enough here. What mattered was choosing it after observing the failing execution, instead of using it as the first explanation for the failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The practical difference
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Superpowers run&lt;/th&gt;
&lt;th&gt;Explyt run&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Starting point&lt;/td&gt;
&lt;td&gt;Same log and source&lt;/td&gt;
&lt;td&gt;Same log and source&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Investigation channel&lt;/td&gt;
&lt;td&gt;Text, stack trace, and tests&lt;/td&gt;
&lt;td&gt;JetBrains debugger plus source and tests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Key fact&lt;/td&gt;
&lt;td&gt;Inferred from &lt;code&gt;AccessDeniedException&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Observed at a breakpoint: directory, permissions, and throw&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Token use in this run&lt;/td&gt;
&lt;td&gt;132k&lt;/td&gt;
&lt;td&gt;67k&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;First implementation&lt;/td&gt;
&lt;td&gt;Swallow the error as &lt;code&gt;emptyList()&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Keep the fix at the proven failing boundary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Test signal&lt;/td&gt;
&lt;td&gt;Passing tests for the fallback&lt;/td&gt;
&lt;td&gt;Runtime evidence followed by a focused regression test&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Result&lt;/td&gt;
&lt;td&gt;Patch discarded as technical debt&lt;/td&gt;
&lt;td&gt;Reviewed minimal fix deployed to production&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I could see the cleanup cost in the diff. The discarded patch carried a workaround-driven test seam. The shipped patch stayed at the failing boundary and left the normal listing path intact.&lt;/p&gt;

&lt;p&gt;For a runtime bug, a coherent explanation from a log is still only a theory. I want one fact captured from the running program: a variable value, a call stack, an execution path, or the exact branch that contradicts the model's theory.&lt;/p&gt;

&lt;p&gt;Explyt is an AI agent for JetBrains IDEs. In this run, it used a supported JetBrains IDE and run configuration. Its Debug workflow started the code under the JetBrains debugger and inspected breakpoints, variable values, call stacks, and execution paths before editing. The IDE supplies the evidence; the developer still decides whether the hypothesis and the resulting diff are acceptable.&lt;/p&gt;

&lt;p&gt;What I can say from these transcripts is limited. Superpowers can be combined with a debugger, and Explyt's debugger depth depends on the IDE, language, and run configuration. Check the &lt;a href="https://explyt.ai/docs/explyt-test/overview/feature-matrix?_highlight=feature&amp;amp;_highlight=matrix" rel="noopener noreferrer"&gt;feature matrix&lt;/a&gt; for the supported combinations.&lt;/p&gt;

&lt;p&gt;The token count points in the same direction. Explyt solved this case in 67k tokens; Superpowers consumed 132k. A debugger did not merely add another tool to the menu. It gave the model the fact it needed before more speculative turns accumulated.&lt;/p&gt;

&lt;p&gt;JetBrains has been running paired A/B tests on "token-saving" skills, with a useful warning against reading too much into their README claims. In the &lt;a href="https://blog.jetbrains.com/ai/2026/07/speak-to-ai-agents-like-cavemen-tosave-tokens/" rel="noopener noreferrer"&gt;Caveman test&lt;/a&gt;, the advertised 65% reduction became 8.5% fewer output tokens on real agentic tasks, with activation forced. In the newer &lt;a href="https://blog.jetbrains.com/ai/2026/07/rtk-claude-code-token-savings/" rel="noopener noreferrer"&gt;rtk test&lt;/a&gt;, the with-rtk arm cost a median 7.6% more per task at low reasoning effort, and showed no saving at high effort. The latest &lt;a href="https://blog.jetbrains.com/ai/2026/07/ponytail-skill-claude-tested/" rel="noopener noreferrer"&gt;Ponytail test&lt;/a&gt; found a 10.3% cost reduction on its own benchmark. The results are mixed; the consistent lesson is to measure the whole agent run, not a skill's self-reported counter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try the first useful task
&lt;/h2&gt;

&lt;p&gt;If you already work in a JetBrains IDE, &lt;a href="https://plugins.jetbrains.com/plugin/27979-explyt-ai-agent" rel="noopener noreferrer"&gt;install Explyt from the Marketplace&lt;/a&gt;, open a reproducible failing test, and let the agent inspect the first runtime value that contradicts its theory. Then review the diff before accepting the fix.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>community</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>The first essay in this series argued that CMMI's core insight survives and its quiet assumption does not: process maturity still predicts delivery reliability, but the model was sized for code written by people at a human pace.</title>
      <dc:creator>Viktoria</dc:creator>
      <pubDate>Wed, 02 Sep 2026 18:11:31 +0000</pubDate>
      <link>https://dev.to/marketing_explyt_a7b53da9/the-first-essay-in-this-series-argued-that-cmmis-core-insight-survives-and-its-quiet-assumption-4chp</link>
      <guid>https://dev.to/marketing_explyt_a7b53da9/the-first-essay-in-this-series-argued-that-cmmis-core-insight-survives-and-its-quiet-assumption-4chp</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/marketing_explyt_a7b53da9/adlc-the-lifecycle-taking-shape-38cd" class="crayons-story__hidden-navigation-link"&gt;ADLC: The Lifecycle Taking Shape&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/marketing_explyt_a7b53da9" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4074488%2F591a7871-a646-4ad3-a738-9f5034840513.png" alt="marketing_explyt_a7b53da9 profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/marketing_explyt_a7b53da9" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Viktoria
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Viktoria
                
                
              
              &lt;div id="story-author-preview-content-4558390" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/marketing_explyt_a7b53da9" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4074488%2F591a7871-a646-4ad3-a738-9f5034840513.png" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Viktoria&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/marketing_explyt_a7b53da9/adlc-the-lifecycle-taking-shape-38cd" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 2&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/marketing_explyt_a7b53da9/adlc-the-lifecycle-taking-shape-38cd" id="article-link-4558390"&gt;
          ADLC: The Lifecycle Taking Shape
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/programming"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;programming&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/webdev"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;webdev&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/llm"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;llm&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/marketing_explyt_a7b53da9/adlc-the-lifecycle-taking-shape-38cd" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/raised-hands-74b2099fd66a39f2d7eed9305ee0f4553df0eb7b4f11b01b6b1b499973048fe5.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;4&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/marketing_explyt_a7b53da9/adlc-the-lifecycle-taking-shape-38cd#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            11 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
  </channel>
</rss>
