<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Manos Saratsis</title>
    <description>The latest articles on DEV Community by Manos Saratsis (@manos-saratsis).</description>
    <link>https://dev.to/manos-saratsis</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3802587%2Fac49f19b-dbd9-4ee3-b174-b1ad65285816.jpg</url>
      <title>DEV Community: Manos Saratsis</title>
      <link>https://dev.to/manos-saratsis</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/manos-saratsis"/>
    <language>en</language>
    <item>
      <title>ISO 42001 for Engineering Teams: What It Actually Asks You to Prove</title>
      <dc:creator>Manos Saratsis</dc:creator>
      <pubDate>Mon, 28 Sep 2026 14:39:19 +0000</pubDate>
      <link>https://dev.to/manos-saratsis/iso-42001-for-engineering-teams-what-it-actually-asks-you-to-prove-3npj</link>
      <guid>https://dev.to/manos-saratsis/iso-42001-for-engineering-teams-what-it-actually-asks-you-to-prove-3npj</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://dromeas.ai/blog/iso-42001-for-engineering-teams" rel="noopener noreferrer"&gt;Dromeas blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In short:&lt;/strong&gt; ISO 42001 is a voluntary, certifiable standard for managing AI systems. For engineering teams it's mostly about evidence: showing how AI-assisted changes are reviewed, tested and approved. It complements the EU AI Act and SOC 2 rather than replacing them.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Framework&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Question it answers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ISO 42001&lt;/td&gt;
&lt;td&gt;Voluntary certification (3 years, with surveillance audits)&lt;/td&gt;
&lt;td&gt;Is your AI governed?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;EU AI Act&lt;/td&gt;
&lt;td&gt;Law, with penalties&lt;/td&gt;
&lt;td&gt;Does your AI meet legal obligations in the EU?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SOC 2&lt;/td&gt;
&lt;td&gt;Attestation&lt;/td&gt;
&lt;td&gt;Do your security controls work?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;ISO 42001 keeps turning up in the same sentence as SOC 2 and the EU AI Act, usually on a slide that implies you need all three by next quarter. I wanted to know what it really asks of the people who write and ship code, so I sat down with it.&lt;/p&gt;

&lt;p&gt;Short version: it's less about paperwork than I expected, and a lot more about evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it is
&lt;/h2&gt;

&lt;p&gt;ISO/IEC 42001 is the international standard for an "AI management system." ISO published it in December 2023, and it's still the current edition. If you've been through ISO 27001, the shape will feel familiar: a set of clauses on how you run the system (the usual plan, do, check, act loop), plus an Annex A with 38 controls across nine areas. You pick the ones that apply and justify your choices in a Statement of Applicability.&lt;/p&gt;

&lt;p&gt;The important bit: it's not a spec for how your model should behave. It's about how your company governs the AI it builds, buys or uses. Who's accountable, how you assess risk, where your data comes from, how you watch things once they're live, and what happens when something goes wrong. It applies whether you're building AI products or just letting your engineers use AI tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who's going for it
&lt;/h2&gt;

&lt;p&gt;It's early. BCG announced in January that it was among the first 100 organizations certified worldwide. Pega got certified in February for Pega Cloud and its GenAI features. I couldn't find a reliable global count of certified companies, and I'd be a bit wary of anyone who quotes one.&lt;/p&gt;

&lt;p&gt;The trajectory looks a lot like SOC 2 a few years back. AI vendors selling into enterprise are the first to feel it, because it starts showing up in security questionnaires. If that's you, it's worth knowing what it'll ask before a customer asks first.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an auditor actually wants to see
&lt;/h2&gt;

&lt;p&gt;Certification happens in two stages. The first is a document review: your AI policy, your risk register, your Statement of Applicability. The second is the one that catches people out. The auditor checks that your controls actually run, which usually means interviews, watching processes, and pulling samples. They'll pick a feature or a date and ask you to show them the control working then.&lt;/p&gt;

&lt;p&gt;For an engineering team, the Annex A areas that land on your desk are roughly these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Lifecycle (A.6):&lt;/strong&gt; design decisions, what you verified before deploying, what changed in each release.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data (A.7):&lt;/strong&gt; where your training, fine-tuning or RAG data came from, how you checked its quality, and what you did when problems showed up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Impact assessment (A.5):&lt;/strong&gt; done before launch, not written up after an incident.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use (A.9):&lt;/strong&gt; evidence the system is used within its intended purpose.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Suppliers (A.10):&lt;/strong&gt; if you call a foundation model API, who's responsible for what.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The general rule auditors follow will sound familiar if you've done SOC 2. A record that was produced independently, timestamped and tied to a specific decision counts for a lot more than a policy saying the process happens. Policy says what should happen. Evidence shows it did.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it fits with the EU AI Act and SOC 2
&lt;/h2&gt;

&lt;p&gt;We've written about the &lt;a href="https://dromeas.ai/compliance/eu-ai-act" rel="noopener noreferrer"&gt;EU AI Act&lt;/a&gt; before. It's law, with real penalties, and it applies if you serve the EU market. ISO 42001 is voluntary. You get certified by a third party, the certificate lasts three years, and there are surveillance audits in between.&lt;/p&gt;

&lt;p&gt;They cover a lot of the same ground (data governance, risk, human oversight, transparency), but one doesn't satisfy the other. If you're building a high-risk system under the Act, a 42001 certificate is useful supporting evidence. It doesn't replace the legal obligation.&lt;/p&gt;

&lt;p&gt;SOC 2 is a different question again. It tells a customer your security controls work: access, encryption, incident response. It says nothing about how you govern the AI itself. Some analysts are already describing ISO 42001 plus SOC 2 Type II as the new baseline for AI vendors selling to enterprise. That's one view, not a rule, but it gives you a sense of where things are heading. Three audits, three different questions: is your AI governed, are your security controls working, is your infrastructure secure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Dromeas fits
&lt;/h2&gt;

&lt;p&gt;We don't certify anyone. That's accredited third-party work, and it should stay that way.&lt;/p&gt;

&lt;p&gt;What we do produce is the kind of record the Stage 2 auditor goes looking for. Every pull request and trunk commit gets reviewed. Every tagged release gets a verdict across six checks (quality, security, compliance, testing, docs, instrumentation), grounded in the actual diff and stamped with a date. If an auditor asks "show me what was verified before this release shipped," that's already sitting there. More on how we handle frameworks on our &lt;a href="https://dromeas.ai/compliance" rel="noopener noreferrer"&gt;compliance page&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;— Manos&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources:&lt;/strong&gt; &lt;a href="https://www.iso.org/standard/42001" rel="noopener noreferrer"&gt;ISO/IEC 42001:2023&lt;/a&gt; · &lt;a href="https://www.bcg.com/news/27january2026-bcg-certified-international-standard-ai-management-systems" rel="noopener noreferrer"&gt;BCG, Jan 27, 2026&lt;/a&gt; · &lt;a href="https://www.businesswire.com/news/home/20260203185475/en/" rel="noopener noreferrer"&gt;Pega via Businesswire, Feb 3, 2026&lt;/a&gt; · &lt;a href="https://www.konfirmity.com/blog/iso-42001-controls" rel="noopener noreferrer"&gt;Konfirmity, Annex A controls&lt;/a&gt; · &lt;a href="https://www.isms.online/iso-42001/vs-eu-ai-act/" rel="noopener noreferrer"&gt;ISMS.online, ISO 42001 vs EU AI Act&lt;/a&gt; · &lt;a href="https://www.vanta.com/collection/iso-42001/iso-42001-and-eu-ai-act" rel="noopener noreferrer"&gt;Vanta, EU AI Act &amp;amp; ISO 42001&lt;/a&gt; · &lt;a href="https://www.knowlee.ai/blog/iso-42001-vs-soc2-vs-iso-27001-comparison" rel="noopener noreferrer"&gt;Knowlee, ISO 42001 vs SOC 2 vs ISO 27001&lt;/a&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
      <category>compliance</category>
      <category>devops</category>
    </item>
    <item>
      <title>Behavioral Bugs Are Still Slipping Through AI Code Review. Here's What Actually Catches Them.</title>
      <dc:creator>Manos Saratsis</dc:creator>
      <pubDate>Thu, 17 Sep 2026 15:28:04 +0000</pubDate>
      <link>https://dev.to/manos-saratsis/behavioral-bugs-are-still-slipping-through-ai-code-review-heres-what-actually-catches-them-3geb</link>
      <guid>https://dev.to/manos-saratsis/behavioral-bugs-are-still-slipping-through-ai-code-review-heres-what-actually-catches-them-3geb</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://dromeas.ai/blog/behavioral-bugs-ai-code-static-analysis-bug-tracing" rel="noopener noreferrer"&gt;Dromeas blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The pull request looks clean, the linter's green, the tests pass — and the code still does the wrong thing the moment someone hands it an input nobody thought about. Not a crash, not a lint warning. Just the wrong answer, quietly.&lt;/p&gt;

&lt;p&gt;I've been digging into why that keeps happening even as AI writes more and more of our code, and it turns out there's decent data on it now — not just vibes. So let's go through what the numbers actually say, why the tools most of us already run can't really catch this category of bug, and what I think actually works.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the 2026 data says
&lt;/h2&gt;

&lt;p&gt;First: people don't fully trust this code, and it's specifically about correctness, not whether it runs. Sonar's 2026 State of Code Developer Survey found 96% of developers don't fully trust that AI-generated code is functionally correct, and 61% agree that "AI often produces code that looks correct but isn't reliable." Only 48% say they always verify AI-assisted code before committing it. Worth sitting with that for a second: the complaint isn't "it doesn't compile." It's "it runs, it looks fine, and I don't actually know if it's right."&lt;/p&gt;

&lt;p&gt;Second, when you break down what kind of bugs show up, it lines up with that. A 2026 empirical study ("Debt Behind the AI Boom") looked across five widely used AI coding tools — Copilot, Claude, Cursor, Gemini, Devin — and found code smells made up 89.3% of flagged issues (the stuff any linter catches fine). Correctness issues were a smaller slice, 6.0% — but the single most common one was "undefined variable or reference," almost 24,000 instances, which the researchers describe as code that "may look locally correct, but still fails to stay consistent with the surrounding context." That's a pretty good one-line description of the whole problem. More than 15% of commits from every single tool they studied introduced at least one issue, and 22.7% of the AI-introduced issues they tracked were still sitting in the repo, unnoticed, at the time of the study — some for nine months or longer.&lt;/p&gt;

&lt;p&gt;Third, when this reaches production, it's not staying theoretical. New Relic's 2026 State of AI Coding report found 82% of organizations had at least one major production failure caused by AI code in the past six months, and 78% report a measurable spike in incidents tied to AI code overall — with AI-generated code introducing roughly 1.7x more critical runtime issues than human-reviewed code. CloudBees' 2026 State of Code Abundance report found something similar from a different angle: 81% of enterprise leaders report increased production issues tied to AI code, and their own summary of it stuck with me — "writing code is no longer the primary bottleneck, governing it is." Same conclusion, three separate 2026 surveys.&lt;/p&gt;

&lt;p&gt;And the backdrop makes all of this harder to catch by hand. GitClear's 2026 research tracked eight quality signals across 623 million code changes from 2023 to 2026, and the trend lines aren't great: duplicated code blocks up 81% since 2023 (highest on record), copy-paste share of changed lines up from 9.4% to 15.7%, actual refactoring down about 70% over the same stretch. GitClear's own read on it: block duplication is the single risk signal most tied to defects and propagated bugs in the research they looked at. So more code is shipping, less of it is getting consolidated or re-examined, and reviewers have less bandwidth per line than ever.&lt;/p&gt;

&lt;p&gt;Put simply: it's not that AI writes "bad" code. It clears the bars we're good at checking automatically — style, structure, the obviously dangerous patterns — and it's specifically weaker on whether the logic holds up for a real input, in a way that's now showing up as real production failures, not just review comments. That's worth being precise about, because it tells you exactly why the tools most teams already run don't solve it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What these bugs actually look like
&lt;/h2&gt;

&lt;p&gt;"Behavioral bug" is the term I keep reaching for, but it's worth knowing it's not the only name people use for this — you'll see the same basic idea called logic bugs or logic errors (the plain-English default), semantic bugs (more of an academic/compiler-literature term — the code is syntactically fine but means the wrong thing), correctness bugs or correctness issues (the term the arxiv paper above uses), business logic bugs or business logic vulnerabilities (the AppSec framing), and functional bugs (QA/testing terminology). Some people just call them silent bugs or silent failures, which honestly might be the most useful name of the bunch — it points at the actual property that makes them dangerous: nothing crashes, nothing throws, nothing shows red in CI.&lt;/p&gt;

&lt;p&gt;In practice, most of what falls under that umbrella breaks down into a handful of recurring shapes:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Boolean logic bugs&lt;/strong&gt; — an inverted condition, the wrong operator, a guard clause that lets through exactly the case it was written to block. Classic version: a permission check written with &lt;code&gt;||&lt;/code&gt; where it needed &lt;code&gt;&amp;amp;&amp;amp;&lt;/code&gt;, so access gets granted if any one condition is met instead of requiring all of them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reference correctness bugs&lt;/strong&gt; — stale closures, the wrong variable captured, the wrong object mutated once two calls overlap. Classic version: a loop that captures its loop variable by reference in a callback, so every callback ends up pointing at the last value instead of the one that was current when it was created.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Boundary and indexing bugs&lt;/strong&gt; — the everyday off-by-one: loop bounds, first/last-element handling, pagination math. Classic version: a "load more" offset computed as &lt;code&gt;page * pageSize&lt;/code&gt; instead of &lt;code&gt;(page - 1) * pageSize&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Normalization-symmetry bugs&lt;/strong&gt; — a value gets normalized on write but not on read (case, whitespace, encoding), so a lookup silently misses. Classic version: emails lowercased at signup but not at login.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Nullability bugs&lt;/strong&gt; — a new code path where a value can legitimately be absent, and nothing downstream accounts for it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shared-state concurrency bugs&lt;/strong&gt; — races between concurrent writes, a rollback to a stale captured value, a cache write that clobbers something fresher.&lt;/p&gt;

&lt;p&gt;None of these are exotic. Every engineer has shipped at least one of each at some point. What's changed is the volume and the plausibility — code that reads as confidently correct, generated fast enough that nobody's tracing each new input by hand anymore.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why static analysis can't really get at this
&lt;/h2&gt;

&lt;p&gt;Tools like SonarQube, Semgrep, and CodeQL work by pattern matching — an AST shape, a taint path from untrusted input to a dangerous sink, something that matches a known CWE. Genuinely useful, and it's why these tools are worth running. But it only works if the bug matches a pattern someone already wrote a rule for.&lt;/p&gt;

&lt;p&gt;Most behavioral bugs don't work that way. Gecko Security put this well in a piece on why static analysis struggles with business logic: a lot of these bugs are about something being &lt;em&gt;missing&lt;/em&gt; — a missing auth check, a missing validation step — rather than something dangerous being &lt;em&gt;present&lt;/em&gt; that a rule can flag. Taint analysis can tell you untrusted data reaches a sensitive spot. It can't tell you whether the authorization logic guarding that spot is actually correct. And a lot of the bugs that cause real incidents span multiple files — their example is a real Cal.com auth-bypass bug that needed three separate issues chained together across different files, each looking fine on its own.&lt;/p&gt;

&lt;p&gt;So that's the real limitation: pattern matching can tell you something &lt;em&gt;looks like&lt;/em&gt; a shape it's seen before. It can't tell you that a specific input, actually walked through the logic, produces a different result than it should. That requires tracing behavior against intent — and it happens to be exactly where the data above says AI-generated code is weakest.&lt;/p&gt;

&lt;p&gt;It's also why I think "detect that AI wrote this, then review it harder" — which a few vendors have shipped, SonarQube's AI Code Assurance being one — is sorting on the wrong axis. It's routing on &lt;em&gt;who&lt;/em&gt; wrote the code, when the thing that actually matters is &lt;em&gt;what the code does&lt;/em&gt;. The question was never "who wrote this" — it's "does this do the right thing for a real input," and that needs tracing, not detection.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this leaves Bug Tracing
&lt;/h2&gt;

&lt;p&gt;This is basically the problem we built Bug Tracing to solve inside Dromeas — trace real inputs through the actual changed code and check what comes out, instead of scanning for patterns. The six categories above are exactly what it scores every change against, and it only reports a finding if it can name the concrete input that triggers it — no reproducible trace, no report.&lt;/p&gt;

&lt;p&gt;And because "review a diff" and "audit a whole system" are genuinely different jobs, it runs at whichever scope your team actually works at: on the pull request itself (changed lines plus their blast radius), on a branch or release after merge for teams doing trunk-based development, or as an on-demand sweep across a whole repository when you just want a second opinion on a system nobody's looked at closely in a while.&lt;/p&gt;

&lt;p&gt;If this is a problem you're running into, happy to show you how it works — the feature page is at &lt;a href="https://dromeas.ai/bug-tracing" rel="noopener noreferrer"&gt;dromeas.ai/bug-tracing&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>testing</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Shadow AI in Your Codebase: The Governance Gap Most CISOs Haven't Mapped Yet</title>
      <dc:creator>Manos Saratsis</dc:creator>
      <pubDate>Sun, 13 Sep 2026 13:27:37 +0000</pubDate>
      <link>https://dev.to/manos-saratsis/shadow-ai-in-your-codebase-the-governance-gap-most-cisos-havent-mapped-yet-eo9</link>
      <guid>https://dev.to/manos-saratsis/shadow-ai-in-your-codebase-the-governance-gap-most-cisos-havent-mapped-yet-eo9</guid>
      <description>&lt;p&gt;Your AI code policy covers approved tools. The harder question is whether every change faces the same evidence-backed review, regardless of what wrote it.&lt;/p&gt;

&lt;p&gt;Shadow AI is no longer a vague compliance concern. It describes employees using AI systems outside an organization's approved controls. The result can be both data exposure and software entering production without a consistent review record. That second problem matters even when the code looks ordinary.&lt;/p&gt;

&lt;p&gt;A note on sourcing: the figures below come from a mix of primary surveys and industry roundups, and each is attributed to the publication that states it. Shadow-AI measurement is still thin, so treat the percentages as directional signals rather than settled numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  The visibility gap is larger than most programs assume
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://airia.com/blog/shadow-ai-statistics-key-data-points-every-ciso-needs-in-2026/" rel="noopener noreferrer"&gt;Airia's 2026 overview&lt;/a&gt; collects findings from named third-party studies: source code represented 35% of sensitive-data events involving AI tools in Cyberhaven's 2024 data; 60% of security teams lacked visibility into which AI tools employees used, per Cisco's 2024 AI Readiness Index; and Gartner reported in 2025 that 43% could not audit or inventory the AI tools in use.&lt;/p&gt;

&lt;p&gt;A separate &lt;a href="https://jumpcloud.com/blog/11-stats-about-shadow-ai-in-2026" rel="noopener noreferrer"&gt;JumpCloud summary&lt;/a&gt; reports widespread unsanctioned assistant use. Survey percentages are useful warning signals, but they don't establish which tool authored a given line in your repository. That distinction matters when choosing a control.&lt;/p&gt;

&lt;h2&gt;
  
  
  Source code creates two different shadow-AI risks
&lt;/h2&gt;

&lt;p&gt;Input risk appears when proprietary code, secrets, customer data, or architecture details are sent to an unapproved service. That requires controls such as approved-tool policy, identity and endpoint management, DLP, vendor review, and data-retention rules.&lt;/p&gt;

&lt;p&gt;Output risk appears when generated code is accepted because it looks plausible, or enters through a workflow that never receives equivalent quality, security, and compliance checks. This is the part a repository and release review layer can address directly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Govern the output without pretending the tool is invisible
&lt;/h2&gt;

&lt;p&gt;Banning unsanctioned tools doesn't guarantee compliance, but governing only the output doesn't solve data leakage either. A durable program needs both: discover and control the tools where possible, then require the same review bar for every software change regardless of authorship.&lt;/p&gt;

&lt;p&gt;This avoids a fragile "detect AI code, then review it differently" path. Pull requests, direct commits to the default branch, and pre-PR local diffs can all enter through different routes. The control should follow the change.&lt;/p&gt;

&lt;h2&gt;
  
  
  What an AI-BOM can, and cannot, tell you
&lt;/h2&gt;

&lt;p&gt;Dromeas generates an AI Bill of Materials when a product actually invokes AI or machine-learning systems: a machine-readable CycloneDX 1.6 ML-BOM plus a human-readable report covering detected AI/ML components, data flows, risk classification, EU AI Act obligations, and source-backed provenance. For multi-repository products, the documentation process assembles evidence across the whole product rather than treating each repository as an isolated system.&lt;/p&gt;

&lt;p&gt;That's an inventory of AI/ML components used by the software. It is not a forensic attribution system for which coding assistant wrote each line, and it doesn't calculate a trustworthy percentage of AI-authored code. It complements workforce AI-tool inventory, DLP, endpoint controls, and an AI-use policy; it doesn't replace them.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical first step
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Inventory the services: use identity, network, endpoint, procurement, and expense data to find unapproved AI tools.&lt;/li&gt;
&lt;li&gt;Protect the inputs: define what code and data may be sent to each approved service, and enforce it with technical controls.&lt;/li&gt;
&lt;li&gt;Standardize the output gate: review PRs, trunk commits, and local agent diffs against one policy.&lt;/li&gt;
&lt;li&gt;Record what shipped: keep review and release evidence, and generate an AI-BOM for software that contains or invokes AI/ML components.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Shadow AI can't be solved by one inventory or one scanner. It becomes governable when tool controls and change-level evidence meet at the release boundary.&lt;/p&gt;




&lt;p&gt;Full piece, with more on Dromeas's changed-code and blast-radius review pipeline (independent Security, Quality, Compliance, and Bug Tracing analysts) plus AI-BOM documentation, originally published at dromeas.ai: &lt;a href="https://dromeas.ai/blog/shadow-ai-in-your-codebase" rel="noopener noreferrer"&gt;https://dromeas.ai/blog/shadow-ai-in-your-codebase&lt;/a&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
      <category>devops</category>
      <category>cybersecurity</category>
    </item>
    <item>
      <title>Which AI Model Writes the Most Secure Code? What the 2026 Data Actually Shows</title>
      <dc:creator>Manos Saratsis</dc:creator>
      <pubDate>Sun, 13 Sep 2026 13:26:40 +0000</pubDate>
      <link>https://dev.to/manos-saratsis/which-ai-model-writes-the-most-secure-code-what-the-2026-data-actually-shows-10io</link>
      <guid>https://dev.to/manos-saratsis/which-ai-model-writes-the-most-secure-code-what-the-2026-data-actually-shows-10io</guid>
      <description>&lt;p&gt;Ask five engineering leaders which AI coding model is "safe," and you'll get five confident, contradictory answers. &lt;a href="https://www.veracode.com/blog/2026-genai-code-security-report-ai-risk/" rel="noopener noreferrer"&gt;Veracode's 2026 GenAI Code Security Report&lt;/a&gt; gives the question a measured answer: across its benchmark tasks, AI-generated code passed security tests 56% of the time and introduced a risky vulnerability in the other 44%. That's effectively flat from the previous report's 55% pass rate, while Veracode estimates AI now authors roughly half of committed code.&lt;/p&gt;

&lt;p&gt;Flat security performance at higher volume isn't a wash. It's the same failure rate landing on more production code.&lt;/p&gt;

&lt;h2&gt;
  
  
  The plateau nobody's marketing slide mentions
&lt;/h2&gt;

&lt;p&gt;A separate &lt;a href="https://www.coderabbit.ai/blog/state-of-ai-vs-human-code-generation-report" rel="noopener noreferrer"&gt;CodeRabbit study&lt;/a&gt; reported 2.74x more vulnerabilities in AI-co-authored pull requests than in human-only pull requests, from 470 open-source PRs (320 AI-co-authored, 150 human-only), not from Veracode. CodeRabbit notes an important limitation: authorship was inferred from signals rather than confirmed ground truth.&lt;/p&gt;

&lt;p&gt;Different datasets point in the same direction: model capability and secure output don't rise on the same curve.&lt;/p&gt;

&lt;h2&gt;
  
  
  Vulnerability class matters more than the average
&lt;/h2&gt;

&lt;p&gt;Veracode's aggregate result hides a much sharper spread:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;SQL injection: 83% pass rate&lt;/li&gt;
&lt;li&gt;Cryptographic implementation: 87% pass rate&lt;/li&gt;
&lt;li&gt;Cross-site scripting: 15% pass rate&lt;/li&gt;
&lt;li&gt;Log injection: 12% pass rate&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The hard cases require context: where untrusted data entered, how it moved through calls, where it reached a sink. That's a dataflow problem, not a syntax problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model selection is a security decision, not a security control
&lt;/h2&gt;

&lt;p&gt;GPT-5.5 led Veracode's benchmark at 68%. More than half of tested models clustered at 50-53%. Coding-specialized models averaged 51%, general-purpose models 52%, and reasoning models 56% vs. 51% for non-reasoning variants. Model size showed little correlation with security outcomes.&lt;/p&gt;

&lt;p&gt;There's no universal "most secure" model. Even the benchmark leader failed nearly one test in three.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do with this data this quarter
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Ask for weakness-level evidence, not just an aggregate score.&lt;/li&gt;
&lt;li&gt;Review at the change boundary: local diffs, PRs, trunk commits, while context is fresh.&lt;/li&gt;
&lt;li&gt;Prioritize reachable risk over raw finding counts.&lt;/li&gt;
&lt;li&gt;Record the policy: which models, which checks, who signs off.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The 2026 data doesn't identify a shortcut past review. It makes the case for a review layer that stays useful when the model roster changes.&lt;/p&gt;




&lt;p&gt;Full piece, with more on Dromeas's multi-model review pipeline (independent Security, Quality, Compliance, and Bug Tracing analysts), originally published at dromeas.ai: &lt;a href="https://dromeas.ai/blog/which-ai-model-writes-most-secure-code" rel="noopener noreferrer"&gt;https://dromeas.ai/blog/which-ai-model-writes-most-secure-code&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>machinelearning</category>
      <category>webdev</category>
    </item>
    <item>
      <title>MCP for Code Review: What It Is, and How to Add a Review Layer to Claude Code or Cursor</title>
      <dc:creator>Manos Saratsis</dc:creator>
      <pubDate>Wed, 09 Sep 2026 13:24:23 +0000</pubDate>
      <link>https://dev.to/manos-saratsis/mcp-for-code-review-what-it-is-and-how-to-add-a-review-layer-to-claude-code-or-cursor-18p1</link>
      <guid>https://dev.to/manos-saratsis/mcp-for-code-review-what-it-is-and-how-to-add-a-review-layer-to-claude-code-or-cursor-18p1</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://dromeas.ai/blog/mcp-review-layer-guide" rel="noopener noreferrer"&gt;Dromeas blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Your coding agent writes code. MCP is how you give it a reviewer. Here's what the protocol actually does — and a working setup you can copy in minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  MCP, in one paragraph
&lt;/h2&gt;

&lt;p&gt;The Model Context Protocol is an open standard that lets an AI client — Claude Code, Cursor, Windsurf, Zed — call external tools through a uniform interface. Instead of pasting output between your terminal and a review dashboard, the agent invokes a tool like &lt;code&gt;review_pull_request&lt;/code&gt; directly and receives the result as structured data it can reason about. Tool support, not copy-paste, is what makes an agentic loop possible: the agent can act, observe the verdict, and repair — the loop we described in &lt;a href="https://dromeas.ai/blog/loop-engineering-with-dromeas" rel="noopener noreferrer"&gt;loop engineering&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a review layer belongs in the loop
&lt;/h2&gt;

&lt;p&gt;A coding agent with no review layer grades its own homework. It writes a diff, checks that it compiles, and declares victory. A review layer changes the economics: the same agent submits its diff, gets an independent multi-model verdict — security, quality, compliance — and fixes what it finds before a human spends a minute on it. Self-review before the PR is the single highest-leverage place to insert checking, because the cost of a fix there is one tool call, not a review round-trip.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup, step by step
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Connect your repositories.&lt;/strong&gt; Sign up at dromeas.ai and connect GitHub, GitLab or Bitbucket. Dromeas builds a typed code map of your repos so review verdicts come with real context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Get your MCP endpoint.&lt;/strong&gt; Open Settings &amp;gt; MCP in the Dromeas app and copy your workspace's MCP server URL and API key. One endpoint covers Claude Code, Cursor, Windsurf, Zed and any other MCP-capable client.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add the server to your client.&lt;/strong&gt; In Claude Code run &lt;code&gt;claude mcp add&lt;/code&gt; with the URL; in Cursor, add the server under Settings &amp;gt; MCP. The review tools appear in your agent's tool list immediately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ask your agent to self-review.&lt;/strong&gt; Before opening a PR, ask your coding agent to submit the diff for review. It calls &lt;code&gt;review_pull_request&lt;/code&gt; or &lt;code&gt;review_local_diff&lt;/code&gt;, gets a multi-model verdict, and can fix what it finds before a human ever looks.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What your agent can actually do once connected
&lt;/h2&gt;

&lt;p&gt;The Dromeas MCP server exposes the full review surface as tools: &lt;code&gt;review_pull_request&lt;/code&gt; and &lt;code&gt;review_local_diff&lt;/code&gt; for verdicts, &lt;code&gt;code_finder_search&lt;/code&gt; and the code-map tools for cheap context before edits, &lt;code&gt;get_findings&lt;/code&gt; and the fix tools for acting on results, and release tools like &lt;code&gt;describe_release_state&lt;/code&gt; for go/no-go decisions. The point is that review isn't a dashboard your agent can't see — it's a function it can call.&lt;/p&gt;

&lt;p&gt;Because every tool call is scoped to your workspace and logged, you keep the audit trail that matters when agents start merging on their own. That's the same provenance property the &lt;a href="https://dromeas.ai/blog/ciso-guide-ai-generated-code" rel="noopener noreferrer"&gt;CISO checklist&lt;/a&gt; depends on.&lt;/p&gt;




&lt;p&gt;Connect a repo, copy your MCP endpoint, and your coding agent gets a six-agent review council it can call from the terminal. &lt;a href="https://dromeas.ai/blog/loop-engineering-with-dromeas" rel="noopener noreferrer"&gt;Read about loop engineering&lt;/a&gt; for the fuller picture of what the loop looks like end to end.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>programming</category>
      <category>claude</category>
    </item>
    <item>
      <title>The CISO's Guide to AI-Generated Code at Scale</title>
      <dc:creator>Manos Saratsis</dc:creator>
      <pubDate>Wed, 09 Sep 2026 13:23:14 +0000</pubDate>
      <link>https://dev.to/manos-saratsis/the-cisos-guide-to-ai-generated-code-at-scale-3f6p</link>
      <guid>https://dev.to/manos-saratsis/the-cisos-guide-to-ai-generated-code-at-scale-3f6p</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://dromeas.ai/blog/ciso-guide-ai-generated-code" rel="noopener noreferrer"&gt;Dromeas blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;AI-generated code has 2.74x more vulnerabilities than human-written code. Copilot adoption is near-universal in the Fortune 100. And the EU AI Act's high-risk obligations now land on 2 December 2027. Here is the checklist that actually matters.&lt;/p&gt;

&lt;p&gt;The security conversation about AI-generated code has moved past "should we allow it" — your engineers already use it, and most of them use it daily. The question now is operational: how do you govern a codebase whose authorship is increasingly non-human, at a pace no review board can match?&lt;/p&gt;

&lt;p&gt;Three numbers frame the problem. AI-generated code carries roughly 2.74x more vulnerabilities than human-written code (&lt;a href="https://www.veracode.com/resources/analyst-reports/2025-genai-code-security-report/" rel="noopener noreferrer"&gt;Veracode 2025&lt;/a&gt;). GitHub Copilot is used by around 90% of the Fortune 100 (&lt;a href="https://github.blog/news-insights/company-news/github-copilot-the-agent-awakens/" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;), before you count Claude Code, Cursor, and the rest. And the EU AI Act's obligations for high-risk AI systems now apply from 2 December 2027 — pushed back from August 2026 by the &lt;a href="https://digital-strategy.ec.europa.eu/en/news/ai-omnibus-enters-force" rel="noopener noreferrer"&gt;AI Omnibus&lt;/a&gt;, which entered into force on 27 July 2026 (systems embedded in regulated products get until 2 August 2028). The extra runway does not change the expectation: documented provenance, an AI Bill of Materials, and evidence of how AI-assisted code was checked. If your governance plan is still "we'll review the PRs carefully," it is already behind the reality on your main branch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the risk actually concentrates
&lt;/h2&gt;

&lt;p&gt;Not all generated code carries the same risk. In practice the exposure clusters in three places, and each one has a different mitigation shape:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dependency and secret drift.&lt;/strong&gt; Agents reach for packages and patterns from training data — outdated versions, deprecated libraries, and occasionally hardcoded credentials in example-shaped code. Static dependency scanning catches the known-CVE slice; it misses the "correct-looking but wrong" slice, like an auth helper that works and is subtly unsafe.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The review blind spot.&lt;/strong&gt; Human reviewers read generated code more charitably than human-written code — it looks confident and idiomatic, so it gets skimmed. That is precisely the code the 2.74x number is describing. The blind spot isn't negligence; it's a predictable cognitive effect, which means it needs a systemic fix, not a reminder.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The provenance gap.&lt;/strong&gt; When an auditor (or your own incident review) asks "what wrote this line, under which policy, with which model," most teams today have no answer. That gap is what the AI Act's AI-BOM expectation is aimed at: not knowing the answer is becoming a compliance finding, not just an inconvenience.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a workable governance layer looks like
&lt;/h2&gt;

&lt;p&gt;The teams that get ahead of this do three things, and none of them involve slowing developers down:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;One bar, no detection step.&lt;/strong&gt; Apply the same security and quality checks to every change regardless of authorship. Architectures that first "detect AI code" and then route it somewhere stricter add a failure mode (the detection) and a governance seam (the routing). Sonar is already retiring the auto-detection half of its own feature — &lt;a href="https://dromeas.ai/blog/sonar-detects-ai-code-special-case" rel="noopener noreferrer"&gt;we wrote up why that's the wrong default&lt;/a&gt;. A single uniform gate can't be skipped by mislabeling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Post-merge coverage, not just PR gates.&lt;/strong&gt; Agents increasingly commit directly to trunk. If your security review only exists on the PR, a growing share of your production code passes no gate at all. Every commit — PR or trunk — should get the same automated pass.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Provenance as a byproduct, not a project.&lt;/strong&gt; Every review run should record what was checked, by which models, against which standard, with which verdict. Do that and the AI-BOM, the SOC 2 evidence, and the EU AI Act documentation all fall out of normal operation instead of becoming a quarterly scramble.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The questions to put in front of your team this quarter
&lt;/h2&gt;

&lt;p&gt;If you do nothing else, get concrete answers to these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which agents and assistants are committing code today, to which repos?&lt;/li&gt;
&lt;li&gt;What percentage of merges in the last 90 days had no human review?&lt;/li&gt;
&lt;li&gt;If an auditor asked for a bill of materials of the AI systems that touched production last month, could you produce one?&lt;/li&gt;
&lt;li&gt;Which checks run on trunk commits after merge — and who reads the results?&lt;/li&gt;
&lt;li&gt;When a generated-code incident happens, what is the provenance trail you would walk?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most organizations can answer one or two of these today. December 2027 sounds far away; the codebase you will be audited on is being written this quarter.&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
      <category>compliance</category>
      <category>programming</category>
    </item>
    <item>
      <title>PR Review vs. Trunk Review: A Practical Guide to Choosing (or Combining) Both</title>
      <dc:creator>Manos Saratsis</dc:creator>
      <pubDate>Wed, 09 Sep 2026 13:21:50 +0000</pubDate>
      <link>https://dev.to/manos-saratsis/pr-review-vs-trunk-review-a-practical-guide-to-choosing-or-combining-both-2gm6</link>
      <guid>https://dev.to/manos-saratsis/pr-review-vs-trunk-review-a-practical-guide-to-choosing-or-combining-both-2gm6</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://dromeas.ai/blog/pr-review-vs-trunk-review-guide" rel="noopener noreferrer"&gt;Dromeas blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;PR review and trunk-based review solve different problems, and most teams already run a hybrid without naming it. This guide is backed by data from &lt;a href="https://dromeas.ai/blog/state-of-ai-coding-2026" rel="noopener noreferrer"&gt;100,000+ real pull requests and 24,000 trunk commits across 500+ open-source repos&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;When we looked at that data, one finding kept surfacing in different forms: the PR-versus-trunk question isn't binary. Teams that think they've picked one are usually already running a hybrid, they just haven't named it. Here's what the data actually supports about when each one earns its keep.&lt;/p&gt;

&lt;h2&gt;
  
  
  What PR review is for
&lt;/h2&gt;

&lt;p&gt;PR review is built for the moments that benefit from a checkpoint: large or risky changes, anything that needs discussion before it merges. That matters because PR size in the wild is bimodal — plenty of small, quick changes, but a heavy tail of 1,000+ line PRs that are exactly where a pre-merge gate is worth the friction it adds. The PR is also the only place where a change is still cheap to reject: once it merges, every fix is a new change instead of a revision.&lt;/p&gt;

&lt;h2&gt;
  
  
  What trunk review is for
&lt;/h2&gt;

&lt;p&gt;Trunk review is built for everything else. Direct-to-trunk commits in our sample averaged 63x smaller than PRs — fast, low-friction, and increasingly the shape of how coding agents actually commit. Waiting for a full PR gate on every micro-change doesn't match agent-speed workflows, and it doesn't need to. The mistake is assuming "no PR" means "no review" — trunk commits can be reviewed after the fact against the same bar, without blocking the commit path.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tail is where the time goes
&lt;/h2&gt;

&lt;p&gt;The part that costs teams the most time isn't the average case in either lane — it's the tail. Rejected PRs took 6x longer to resolve than approved ones in our data. That's where review capacity actually gets consumed: not on the easy approvals, but on the changes that go back and forth. Any review strategy that doesn't have an answer for the tail — smaller initial diffs, automated first-pass review, faster feedback — will feel slow no matter which lane you picked.&lt;/p&gt;

&lt;h2&gt;
  
  
  A decision framework that matches reality
&lt;/h2&gt;

&lt;p&gt;Instead of "PRs for everything" or "trunk for everything," score each change on four axes and let the lane follow:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Signal&lt;/th&gt;
&lt;th&gt;Lean PR gate&lt;/th&gt;
&lt;th&gt;Lean trunk + post-merge review&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Change size&lt;/td&gt;
&lt;td&gt;Hundreds of lines and up; the heavy tail&lt;/td&gt;
&lt;td&gt;Tens of lines; single-purpose commits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blast radius&lt;/td&gt;
&lt;td&gt;Touches shared contracts, auth, billing, migrations&lt;/td&gt;
&lt;td&gt;Local, reversible, behind a flag&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Contributor type&lt;/td&gt;
&lt;td&gt;New contributor, unfamiliar area, cross-team change&lt;/td&gt;
&lt;td&gt;Owner of the area, or an agent doing a scoped fix&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Discussion value&lt;/td&gt;
&lt;td&gt;Design decisions others need to weigh in on&lt;/td&gt;
&lt;td&gt;Mechanical or already-agreed work&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Most teams discover their real policy is already this table, applied informally. Making it explicit does two things: it stops the arguments about whether a given change "deserved" a PR, and it makes the post-merge lane a first-class citizen instead of an unreviewed loophole.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why most teams will end up running both
&lt;/h2&gt;

&lt;p&gt;The forces pushing the two lanes apart are getting stronger, not weaker. Risky changes are getting larger (more generated code per PR), and routine changes are getting smaller and more frequent (agents committing at agent speed). A single gate tuned for one of those shapes fails the other: too much friction on the small stuff, too little scrutiny on the big stuff.&lt;/p&gt;

&lt;p&gt;The workable shape is both lanes with the same rigor on each: a real gate on the PR for the changes that benefit from one, and automatic, same-standard review on every trunk commit for the ones that don't. That's the model Dromeas implements — the same six-agent pipeline and multi-model council runs on PRs and on trunk commits alike, so the lane choice is a workflow decision, not a quality decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  The data behind this guide
&lt;/h2&gt;

&lt;p&gt;104,968 PRs and 23,964 trunk commits measured across 503 open-source repos — size distributions, review latency, and the slow-motion rejection problem in full.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dromeas.ai/blog/state-of-ai-coding-2026" rel="noopener noreferrer"&gt;Read the State of AI Coding 2026&lt;/a&gt; · &lt;a href="https://dromeas.ai/blog/pr-vs-trunk-what-code-review-actually-looks-like" rel="noopener noreferrer"&gt;See the charts&lt;/a&gt;&lt;/p&gt;

</description>
      <category>programming</category>
      <category>ai</category>
      <category>softwareengineering</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How to Evaluate an AI Code Review Vendor in 2026: A Practical Checklist</title>
      <dc:creator>Manos Saratsis</dc:creator>
      <pubDate>Wed, 09 Sep 2026 13:20:29 +0000</pubDate>
      <link>https://dev.to/manos-saratsis/how-to-evaluate-an-ai-code-review-vendor-in-2026-a-practical-checklist-3iem</link>
      <guid>https://dev.to/manos-saratsis/how-to-evaluate-an-ai-code-review-vendor-in-2026-a-practical-checklist-3iem</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://dromeas.ai/blog/ai-code-review-vendor-checklist" rel="noopener noreferrer"&gt;Dromeas blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Every AI code review vendor's homepage says the same three words: "AI-powered code review." That's not useful information anymore — of course it's AI-powered, it's 2026. The questions that actually separate these tools are more specific, and most comparison content in this category is either a vendor's own battlecard or a listicle nobody fact-checked. Here are six questions worth asking directly, in a demo or a trial, before you sign anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Does it review one point in the pipeline, or every point?
&lt;/h2&gt;

&lt;p&gt;A bot on the PR is the default shape of this category. Ask what happens to a commit that lands directly on trunk, or a change that ships six weeks after the PR that introduced the underlying risk. If the answer is "we don't look there," that's not disqualifying on its own — but it's a gap you're accepting, and you should know you're accepting it. In our own measurements of 100,000+ PRs across 500+ open-source repos, a meaningful share of changes never went through a PR at all — so "we review PRs" and "we review your code" are materially different claims. &lt;a href="https://dromeas.ai/blog/pr-vs-trunk-what-code-review-actually-looks-like" rel="noopener noreferrer"&gt;The data on that is public&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Is a finding one model's opinion, or several checking each other?
&lt;/h2&gt;

&lt;p&gt;Single-model review inherits that one model's blind spots. Ask whether the vendor runs more than one model against the same diff, and whether you can see where they disagreed. "We use GPT-5" is not the same claim as "three models reviewed this independently and here's where they split." The disagreement record is the useful part: a tool that hides it is asking you to trust a single point of failure with a friendly UI. (We wrote up &lt;a href="https://dromeas.ai/blog/why-one-model-isnt-enough-llm-council" rel="noopener noreferrer"&gt;how we built our own multi-model council&lt;/a&gt; if you want to see what the transparent version looks like.)&lt;/p&gt;

&lt;h2&gt;
  
  
  3. What happens after "no findings"?
&lt;/h2&gt;

&lt;p&gt;A clean PR review says nothing about test coverage, stale docs, or missing observability on the code that just shipped. Ask whether the tool's job ends at the diff, or whether it's also checking the five other things that determine whether a release is actually safe. "No findings" should be the start of a release conversation, not the end of a review.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Who checks compliance, and how?
&lt;/h2&gt;

&lt;p&gt;If you're regulated — SOC 2, HIPAA, PCI, GDPR — ask directly whether compliance is a first-class check or a slide in the sales deck. Most tools in this category are quality- or security-first and don't cover this at all; that's a fair trade-off if you don't need it, and a real gap if you do. The tell: ask to see a compliance finding in the demo, not a slide about one.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Per-seat, or usage-based?
&lt;/h2&gt;

&lt;p&gt;Per-seat pricing was built for a world where humans opened PRs. As agents start committing more often than people do, seat count stops tracking the thing that's actually driving cost or risk. Ask how pricing behaves as commit volume rises independent of headcount — and whether you'll be penalized for the agent-heavy workflow the same vendor's marketing encourages.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Can your coding agent actually talk to it?
&lt;/h2&gt;

&lt;p&gt;MCP support is now table stakes to claim, but "we have an MCP server" and "your agent can ask this tool for a verdict and act on it" are different claims. Ask what a coding agent can actually do through the integration — read-only context, or a working verify-and-fix loop. A quick test: can your agent submit its uncommitted diff for review and receive a verdict it can act on, without you leaving the terminal?&lt;/p&gt;

&lt;h2&gt;
  
  
  Ask everyone — including us
&lt;/h2&gt;

&lt;p&gt;None of these questions require you to already know what you're comparing against. Ask them of any vendor, including us — the answers are the actual differentiator, not the marketing copy above them.&lt;/p&gt;

&lt;p&gt;If you want the specific answers for the vendors you're already evaluating, the &lt;a href="https://dromeas.ai/comparison" rel="noopener noreferrer"&gt;comparison hub&lt;/a&gt; has the full feature-by-feature breakdowns — CodeRabbit, SonarQube, Copilot, Snyk and more, each with sources and a last-verified date.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>security</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Run code review and releases from a conversation</title>
      <dc:creator>Manos Saratsis</dc:creator>
      <pubDate>Tue, 01 Sep 2026 06:32:59 +0000</pubDate>
      <link>https://dev.to/manos-saratsis/run-code-review-and-releases-from-a-conversation-22m1</link>
      <guid>https://dev.to/manos-saratsis/run-code-review-and-releases-from-a-conversation-22m1</guid>
      <description>&lt;p&gt;We just shipped chat as a first-class interface into Dromeas.&lt;/p&gt;

&lt;p&gt;Instead of clicking through dashboards, you can now just ask: "review PR 247 on the payment service," "are we good to release v1.9.0?," "fix the top security finding and open a PR." Dromeas reads the diff, queries the Code Map for blast radius, runs the same quality/security/compliance agents that guard your trunk, and shows the result as a live status card — not a wall of text.&lt;/p&gt;

&lt;p&gt;A few things worth calling out for anyone building similar agentic UX:&lt;/p&gt;

&lt;p&gt;It's not a separate system. The chat calls the exact same MCP primitives (get_findings, code_map_search, run_finding_fix, approve_pull_request, etc.) that our IDE integrations for Claude, Cursor, and Copilot use. Start a release check in Cursor, see it finish in the chat.&lt;br&gt;
Autonomy is a dial, not a toggle. Every workspace sets a default — manual, observe, assist, or auto — and you can override it per conversation. Manual shows a confirm card before anything ships; auto acts inside caps you set (file budget, severity threshold, model cost) and reports back after.&lt;br&gt;
Structured over conversational-only. Long-running actions (a review, a fix, a doc run) return a live card with step, progress, and a deep link — updating in place instead of dumping another paragraph into the thread.&lt;/p&gt;

&lt;p&gt;Video walkthrough: &lt;a href="https://youtu.be/ypCt0d8sGis" rel="noopener noreferrer"&gt;https://youtu.be/ypCt0d8sGis&lt;/a&gt;&lt;br&gt;
Try it: &lt;a href="https://dromeas.ai/chat" rel="noopener noreferrer"&gt;https://dromeas.ai/chat&lt;/a&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>automation</category>
      <category>mcp</category>
    </item>
    <item>
      <title>SonarQube Flags AI-Generated Code as a Special Case. That's the Wrong Default.</title>
      <dc:creator>Manos Saratsis</dc:creator>
      <pubDate>Tue, 01 Sep 2026 06:31:56 +0000</pubDate>
      <link>https://dev.to/manos-saratsis/sonarqube-flags-ai-generated-code-as-a-special-case-thats-the-wrong-default-2c0g</link>
      <guid>https://dev.to/manos-saratsis/sonarqube-flags-ai-generated-code-as-a-special-case-thats-the-wrong-default-2c0g</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://dromeas.ai/blog/sonar-detects-ai-code-special-case" rel="noopener noreferrer"&gt;Dromeas blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;SonarQube has a feature called AI Code Assurance. When it detects that a project uses GitHub Copilot — checked via the GitHub Copilot Business org settings — it tags the project &lt;code&gt;CONTAINS AI CODE&lt;/code&gt; and routes it through a dedicated quality gate built specifically for AI output, instead of the standard one (&lt;a href="https://www.sonarsource.com/blog/auto-detect-and-review-ai-generated-code-from-github-copilot/" rel="noopener noreferrer"&gt;Sonar&lt;/a&gt;). More recently, they shipped a plugin that goes further: inside the GitHub Copilot CLI itself, an agent's generated code now gets run through an automatic verify-fix-reanalyze loop before it ever reaches a pull request (&lt;a href="https://www.sonarsource.com/blog/now-available-sonarqube-plugin-for-github-copilot-cli/" rel="noopener noreferrer"&gt;Sonar&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;Both are sensible responses to a real problem — AI-generated code has roughly 2.74x more vulnerabilities than human-written code (&lt;a href="https://www.veracode.com/resources/analyst-reports/2025-genai-code-security-report/" rel="noopener noreferrer"&gt;Veracode 2025&lt;/a&gt;), so scrutinizing it more makes sense on paper. But look at the mechanism underneath both features: first you have to detect that a human didn't write this, then you route it somewhere stricter. AI-authored code is the exception you build a special lane for.&lt;/p&gt;

&lt;p&gt;That's worth questioning, and not because the detection is badly built. It's because "detect, then route differently" only works as well as the detection does — and detection is inherently one step behind.&lt;/p&gt;

&lt;h2&gt;
  
  
  What AI Code Assurance actually checks — and what it misses
&lt;/h2&gt;

&lt;p&gt;The detection step is worth being precise about, because its scope defines the size of the hole. SonarQube's autodetect mechanism evaluates Copilot usage patterns and code-contribution data through the GitHub Copilot Business organization API. Two consequences follow.&lt;/p&gt;

&lt;p&gt;First, the wire it trips is GitHub Copilot-specific. A team running Claude Code, Cursor, or a local agent alongside (or instead of) Copilot doesn't necessarily trip that wire the same way. The special lane exists, but not every AI-authored line of code is guaranteed to be in it.&lt;/p&gt;

&lt;p&gt;Second — and this is the part that makes the architecture question concrete rather than philosophical — Sonar's own docs now flag autodetect as &lt;strong&gt;deprecated&lt;/strong&gt; in SonarQube Server 2026.1 LTA, with removal planned and manual project labeling as the remaining path (&lt;a href="https://docs.sonarsource.com/sonarqube-server/2026.1/quality-standards-administration/ai-code-assurance/overview" rel="noopener noreferrer"&gt;docs.sonarsource.com&lt;/a&gt;). Read that slowly: the vendor that built "detect AI code, route it to a stricter gate" is retiring the detection half of the design and asking humans to self-declare instead. Manual labeling is detection too — it's just detection delegated to the person least incentivized to do it carefully, at exactly the moment agent-written code is becoming the majority of new lines.&lt;/p&gt;

&lt;p&gt;None of this is a knock on the feature's implementation. It's the predictable end-state of any architecture whose first step is "figure out whether a human wrote this."&lt;/p&gt;

&lt;h2&gt;
  
  
  The CLI plugin is the more interesting move
&lt;/h2&gt;

&lt;p&gt;To be fair to Sonar, the GitHub Copilot CLI plugin (June 2026) is the smarter of the two ideas. It doesn't wait for detection at the project level — it runs an agentic loop right in the terminal: analyze the agent's output, fix what it finds, re-analyze (&lt;code&gt;sonar analyze agentic&lt;/code&gt;) before the code reaches a PR. That's the right instinct: meet the agent where it works, and verify before merge rather than audit after.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "same rigor, no detection step" scales better
&lt;/h2&gt;

&lt;p&gt;The alternative to a better detector is not needing one.&lt;/p&gt;

&lt;p&gt;Every PR and every trunk commit that goes through &lt;a href="https://dromeas.ai" rel="noopener noreferrer"&gt;Dromeas&lt;/a&gt; runs the same six-agent pipeline — quality, security, compliance, testing, docs, instrumentation — reviewed by the same multi-model council, regardless of who or what wrote it. There's no &lt;code&gt;CONTAINS AI CODE&lt;/code&gt; badge, because there's no separate gate to route into. A junior engineer's Tuesday-afternoon commit and an autonomous agent's 2am commit get the identical bar.&lt;/p&gt;

&lt;p&gt;That's not a philosophical stance so much as a practical one: as the share of AI-authored code climbs toward the 65% Sonar's own 2026 developer survey projects for 2027, "detect it, then scrutinize it more" is a rule that has to run correctly on a shrinking minority of code to matter, while "scrutinize everything the same way" doesn't have that failure mode at all.&lt;/p&gt;

&lt;p&gt;There's a second-order benefit too. When there is no special lane, there is no lane-splitting argument — no "this was mostly agent-written, so the gate should have caught it" postmortem, and no quiet drift where one class of code quietly gets less review because nobody remembered to label the project. The bar is the bar.&lt;/p&gt;

&lt;p&gt;If you're weighing the two approaches side by side, the &lt;a href="https://dromeas.ai/comparison/dromeas-vs-sonarqube" rel="noopener noreferrer"&gt;full feature-by-feature breakdown&lt;/a&gt; — including AI Code Assurance and the new CLI plugin — is on the Dromeas vs SonarQube comparison page.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>programming</category>
    </item>
    <item>
      <title>Claude Code's ultrareview vs Dromeas Code Review with LLM council</title>
      <dc:creator>Manos Saratsis</dc:creator>
      <pubDate>Thu, 20 Aug 2026 14:51:07 +0000</pubDate>
      <link>https://dev.to/manos-saratsis/claude-codes-ultrareview-vs-dromeas-code-review-with-llm-council-4jji</link>
      <guid>https://dev.to/manos-saratsis/claude-codes-ultrareview-vs-dromeas-code-review-with-llm-council-4jji</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published at &lt;a href="https://dromeas.ai/blog/claude-ultrareview-vs-dromeas-code-review" rel="noopener noreferrer"&gt;dromeas.ai&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;We heard about Claude Code's ultrareview and got excited — a cloud-run, multi-agent deep review sounded like exactly the kind of thing worth building a workflow around.&lt;/p&gt;

&lt;p&gt;So we pointed it at changes in our own repo and compared it against Dromeas code review: three analyzers (quality, security, compliance) cross-checked by an LLM council. Dromeas held up well in that first pass.&lt;/p&gt;

&lt;p&gt;That result was interesting enough that we wanted a harder, more neutral test: a large, real, independently-approved pull request from a codebase neither tool had any stake in. So we picked openclaw/openclaw — a public, actively-developed agentic coding tool — and went looking for its biggest recently-merged, genuinely-reviewed PR. That led us to openclaw#124250, 31 files changed, approved by a human reviewer, and we ran the same head-to-head again.&lt;/p&gt;

&lt;h2&gt;
  
  
  The PR
&lt;/h2&gt;

&lt;p&gt;"Preserve ClawHub external source identity and expose only supported actions" — merged, approved by a human reviewer (not a bot self-merge), XL size: 31 files changed, +1,064/−116 lines, spanning the Control UI, macOS, iOS, and Android clients plus the backend that serves them.&lt;/p&gt;

&lt;p&gt;The bug it fixes: ClawHub's search API returns each result's source under a nested &lt;code&gt;install.reference&lt;/code&gt; field, but the client code expected a flat &lt;code&gt;installRef&lt;/code&gt;. Every external search result silently fell through to a synthesized &lt;code&gt;@owner/slug&lt;/code&gt; reference — quietly pointing installs at a different publisher's skill than the one the operator actually picked. An identity-spoofing bug in a skill-installation flow, fixed across five client surfaces.&lt;/p&gt;

&lt;h2&gt;
  
  
  What each tool found
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;ultrareview:&lt;/strong&gt; 1 finding, nit severity — a duplicate test assertion in an Android test file, unrelated to the identity-spoofing bug the PR exists to fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dromeas's LLM council:&lt;/strong&gt; 29 candidate findings raised, 17 kept after cross-verification. Three models (Opus 5, DeepSeek V4 Pro, GPT-5.6 Terra) independently analyzed the diff, then a decider cross-checked each finding. All 12 quality findings and all 5 security findings held up; 12 compliance findings were flagged as duplicates of already-caught security issues or dropped outright, with the report explaining why for each.&lt;/p&gt;

&lt;p&gt;None of Dromeas's 17 kept findings overlap with ultrareview's one — not because ultrareview did a bad job reading the diff, but because questions like "is this credential field masked" or "does this action get an audit trail" were never in its scope. Full breakdown, cost comparison (~$5 for the full council run vs. $5–25 typical for ultrareview), and the four findings flagged for manual triage are in the full post →&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dromeas.ai/blog/claude-ultrareview-vs-dromeas-code-review" rel="noopener noreferrer"&gt;Read the full comparison&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>codereview</category>
      <category>llm</category>
    </item>
    <item>
      <title>Loop Engineering: How to Actually Close the Loop When You're Coding With AI</title>
      <dc:creator>Manos Saratsis</dc:creator>
      <pubDate>Mon, 10 Aug 2026 12:09:47 +0000</pubDate>
      <link>https://dev.to/manos-saratsis/loop-engineering-how-to-actually-close-the-loop-when-youre-coding-with-ai-1cmi</link>
      <guid>https://dev.to/manos-saratsis/loop-engineering-how-to-actually-close-the-loop-when-youre-coding-with-ai-1cmi</guid>
      <description>&lt;p&gt;When we started experimenting with models coding we were looking at the right prompt, later at prompt chaining, then graphs. The higher the autonomy is we see there is the need for loops, both when coding but also when reviewing and releasing code.&lt;/p&gt;

&lt;p&gt;Original blog &lt;a href="https://dromeas.ai/blog/loop-engineering-with-dromeas" rel="noopener noreferrer"&gt;here&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fri0znxw1m04n1duem50g.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fri0znxw1m04n1duem50g.jpg" alt=" " width="800" height="456"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Loop diagram with four nodes — act, observe, verify, repair — connected in a cycle&lt;br&gt;
The loop. The interesting engineering lives in the bottom half.&lt;br&gt;
Now multiple models are great on writing code at the act step. What requires further improvement is how the code is self healing, improving performance, take solid architecture decisions. Get real inputs to keep going, better and better each time.&lt;/p&gt;

&lt;p&gt;That's the half of the loop Dromeas is built for. Not "write my code for me" — you already have something for that. Verify and repair. We run an amazing experiment: 16 feedback rounds with the same coding agent and a small React/TypeScript repo, building real features with and without the Dromeas MCP help. Btw this was the first time when the actual user is a machine and I have to admit that machines give more structured product feedback allowing a fun iteration until we got something useful.&lt;/p&gt;

&lt;p&gt;Below is what actually held up. Including the parts that didn't.&lt;/p&gt;

&lt;p&gt;What loop engineering looks like while you're coding&lt;br&gt;
The loop while building a feature is four moves, and only two of them are the fun ones.&lt;/p&gt;

&lt;p&gt;Orient. start_task({ repo_full_name, task }) in one call: where the code lives (paths, line ranges, symbols, callers), what's already known-broken in those files, and a token-budgeted read plan.&lt;br&gt;
Act. Read the ranges from disk, make the change, keep the diff scoped.&lt;br&gt;
Verify. Typecheck and lint first — they're free. Then verify_change with the full post-change content of every changed file. It analyses your uncommitted work; nothing needs to be committed or pushed.&lt;br&gt;
Repair. get_findings({ trunk_review_id }) to read the blockers, then preview_fix for a diff or run_finding_fix to push one. Then back to step 2.&lt;br&gt;
The bugs it caught that the agent didn't&lt;br&gt;
This was the most consistent benefit across all 16 rounds, and it's the one worth the money. Not lint noise, not style nits — actual logic errors in code written minutes earlier.&lt;/p&gt;

&lt;p&gt;While building a radius-select tool, a new commit() function in useRadiusSelect.ts didn't guard against a null draft. A spurious commit() with draft === null would silently wipe an already-committed selection. get_findings flagged it; the fix was one line — if (!draft) return; — and the agent's own log said it "would probably not have caught that on my own re-read."&lt;/p&gt;

&lt;p&gt;Better one: a handlePointerUp click-vs-drag detector whose moved flag was set to true at pointer-down time. Meaning it never measured movement at all, so the comment right above it ("treat as a click if the pointer barely moved") was lying about what the code did. No linter or typechecker catches that. It takes reading code against its own comment — and two independent runs in two different rounds both caught it.&lt;/p&gt;

&lt;p&gt;Knowing when to stop searching&lt;br&gt;
Search tools have a failure mode where the agent keeps searching because searching feels like progress. Every code_finder_search and start_task response carries a value_signal, and it's honestly calibrated rather than self-flattering: tested side by side against a 22-symbol repo and a ~3,000-node repo, it rated the small repo's search value low (flat scores, most of the repo returned) and the large one medium. On low signal it returns an empty next_calls list, specifically so the agent doesn't reflexively chain another query.&lt;/p&gt;

&lt;p&gt;On the star-map repo, getting_started reported first_move: "read_files" with the reason spelled out: "only 22 indexed symbols — reading the few source files end-to-end beats any search here." Every agent that followed it stopped after one or two orientation calls. That's real credit and context-window savings, from a tool telling you not to use it.&lt;/p&gt;

&lt;p&gt;"What else touches this?" in one call&lt;br&gt;
Search results inline the caller/callee graph and blast radius for top hits, not just a path. Searching worldToScreen came back with its file, its line range, its 3 callers (StarMapCanvas, hitTestStar, and the containing file) and a blast radius of 3 — enough to know a signature change ripples into exactly those three places, without opening any of them first. That's a manual grep chase replaced by one response.&lt;/p&gt;

&lt;p&gt;Experiment rounds&lt;br&gt;
16&lt;br&gt;
same agent, same repo, with and without the MCP&lt;br&gt;
Real logic bugs&lt;br&gt;
2&lt;br&gt;
caught in freshly written code, pre-commit&lt;br&gt;
Blast radius&lt;br&gt;
1 call&lt;br&gt;
callers + impact inlined with search hits&lt;br&gt;
Findings surfaced at bootstrap&lt;br&gt;
222&lt;br&gt;
60 critical, across two repos&lt;br&gt;
We will cover in a different blog how loop engineering works when reviewing code or releasing products.&lt;/p&gt;

&lt;p&gt;New Dromeas Skills&lt;br&gt;
To get started with this we shipped three agent skills you can install straight into your coding tool from Workspace management → Agent instructions &amp;amp; skills&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxl42b2oyvoos7xfgkz1l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxl42b2oyvoos7xfgkz1l.png" alt=" " width="800" height="513"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The honest downsides&lt;br&gt;
While our early Loop engineering has been great so far, there are some downsides that you need to be aware.&lt;/p&gt;

&lt;p&gt;It costs tokens, and on small codebases the math is worse. Orientation calls, verification payloads and findings all land in the context window. On a 22-symbol repo the agent can just read every file — and, to its credit, our own value_signal and first_move hints say exactly that. The loop earns its keep as coupling and history grow; on a toy repo it's overhead.&lt;br&gt;
It makes tasks take longer. Verification is a real analysis pass, not a lint run — typically 60–120 seconds of polling per loop iteration. If you run it after every micro-edit, you'll feel it. Batch your edits, verify once per meaningful change. We're actively working on cutting that wall-clock time (quick: true already drops compliance for a materially shorter security + quality loop).&lt;br&gt;
It's a loop, which means discipline. The value shows up when you actually read the findings and go back to step 2. An agent that dispatches a verification and then declares victory without reading the verdict has gained nothing. Notably, verdict: "unknown" with analyzed: false, retryable: true is not a pass — and yes, we had to write that in bold in the skill files.&lt;br&gt;
Should I get started?&lt;br&gt;
Actually, yes — especially if you build something robust and you have a sizeable codebase you will get real code context around dependencies, code scope and issues identified. Then your agent will self-heal your code every time it touches components with issues, always with some cost on time&amp;amp;token per task. Just add the new skills and get going.&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
