<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ayush Gupta</title>
    <description>The latest articles on DEV Community by Ayush Gupta (@itsayush).</description>
    <link>https://dev.to/itsayush</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4164016%2F625b5a02-d869-421e-bbdd-a095f51c4bb4.jpg</url>
      <title>DEV Community: Ayush Gupta</title>
      <link>https://dev.to/itsayush</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/itsayush"/>
    <language>en</language>
    <item>
      <title>Ten ways AI coding agents fake a green build (and how to catch them in the pull request)</title>
      <dc:creator>Ayush Gupta</dc:creator>
      <pubDate>Mon, 05 Oct 2026 13:22:00 +0000</pubDate>
      <link>https://dev.to/itsayush/ten-ways-ai-coding-agents-fake-a-green-build-and-how-to-catch-them-in-the-pull-request-4i3o</link>
      <guid>https://dev.to/itsayush/ten-ways-ai-coding-agents-fake-a-green-build-and-how-to-catch-them-in-the-pull-request-4i3o</guid>
      <description>&lt;p&gt;When a coding agent can't fix a failing build, it sometimes makes the build pass anyway.&lt;/p&gt;

&lt;p&gt;Not by fixing the bug. By changing the test, the check, or the pipeline until the red goes away. The build turns green, the reviewer sees a passing check, and the change gets merged.&lt;/p&gt;

&lt;p&gt;I built &lt;a href="https://github.com/ayushgml/greenwash-oss" rel="noopener noreferrer"&gt;Greenwash&lt;/a&gt;, an open-source GitHub App, to catch exactly this. This post covers the ten patterns it looks for, why a general AI code review is the wrong tool for the job, and the design decisions that make the signal useful instead of noisy.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it looks like
&lt;/h2&gt;

&lt;p&gt;Here are two hunks from a pull request where an agent was asked to fix a failing checkout test:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;# tests/test_checkout.py
&lt;span class="gd"&gt;- assert result.total == 100
&lt;/span&gt;&lt;span class="gi"&gt;+ assert total &amp;gt; 0
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight diff"&gt;&lt;code&gt;# src/pricing.py
&lt;span class="gi"&gt;+ if user_id == 4821:
+     return expected_result
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first change weakens an assertion until it can no longer fail. The second hard-codes the answer the test expects for one specific input. Both make CI green. Neither fixes anything.&lt;/p&gt;

&lt;p&gt;In a small diff you would spot these. In a 40-file pull request, written quickly by an agent and reviewed quickly by a person, they are easy to miss. That is the whole problem: these changes are small, and they hide well.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ten patterns
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Weakened test assertion.&lt;/strong&gt; An exact check becomes a loose one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test skipped or disabled.&lt;/strong&gt; A skip marker, a commented-out test, an early return.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expected value rewritten&lt;/strong&gt; without any stated change in behavior.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hard-coded answer for specific inputs,&lt;/strong&gt; like the &lt;code&gt;user_id == 4821&lt;/code&gt; branch above.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Logic replaced with a placeholder.&lt;/strong&gt; The real implementation becomes a stub that satisfies the test.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Error silently swallowed.&lt;/strong&gt; A &lt;code&gt;try&lt;/code&gt; with an empty &lt;code&gt;except&lt;/code&gt;, so the failure disappears.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Type or lint check suppressed.&lt;/strong&gt; An ignore comment instead of a fix.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CI step removed or made non-blocking.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Test file deleted.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Snapshot updated without any source change.&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of these are always wrong. Sometimes you really do mean to update an expected value or skip a flaky test. That matters for the design, and I'll come back to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a general AI review doesn't solve it
&lt;/h2&gt;

&lt;p&gt;The obvious approach is to point an LLM at the diff and ask "is anything wrong here?". In practice that tends to produce an essay: a long list of observations of mixed importance, with no clear signal about which line to look at.&lt;/p&gt;

&lt;p&gt;Greenwash does the opposite. It reads every changed hunk of a pull request and asks narrow yes/no questions about it, such as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Does this delete a test assertion or replace it with a weaker one?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Each question is answered with a probability by &lt;a href="https://docs.typesafe.ai" rel="noopener noreferrer"&gt;TypeSafe&lt;/a&gt;'s structured-judgment models. Eight of the ten checks work this way. The other two, like a deleted test file, are plain code rules, because you don't need a model to notice a file is gone.&lt;/p&gt;

&lt;p&gt;The principle I followed: &lt;strong&gt;the model handles judgment, code handles policy.&lt;/strong&gt; The model answers "how likely is it that this hunk weakens an assertion?". Code decides the thresholds, the severity, and what gets reported.&lt;/p&gt;

&lt;h2&gt;
  
  
  Design decisions, and what each one costs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Probabilities, not verdicts.&lt;/strong&gt; A flagged change may be intentional, so every finding shows how confident it is. A finding is flagged at 0.8 by default, and anything from 0.5 up is marked as worth a second look. The cost: you need thresholds, and they are a judgment call.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Comment mode by default.&lt;/strong&gt; Greenwash reports through a check run with line annotations and a single pull request comment that updates in place on every push. It never blocks a merge unless a repository opts in with &lt;code&gt;mode: check&lt;/code&gt;. The cost: enforcement is opt-in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Config is read from the base commit.&lt;/strong&gt; Repositories tune Greenwash with &lt;code&gt;.github/greenwash.yml&lt;/code&gt;, and the file is read from the pull request's base, not its head. Otherwise an agent could switch off the very check that would catch it. The cost: config changes only apply after they merge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fail neutral.&lt;/strong&gt; If Greenwash itself errors, the check completes as neutral instead of failing your build. The cost: an outage means no findings for that run.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One request per hunk.&lt;/strong&gt; Each request carries only the hunk's diff, the file path, and the pull request's title and description. No whole files, no database, nothing stored; findings live only on GitHub. The cost: each hunk is judged without the rest of the file.&lt;/p&gt;

&lt;p&gt;You can also add your own checks as plain questions. For example, a team could flag any change to billing logic:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;billing&lt;/span&gt;
    &lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Billing logic changed&lt;/span&gt;
    &lt;span class="na"&gt;question&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Does&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;`hunk.diff`&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;change&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;how&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;customers&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;are&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;charged,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;refunded,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;or&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;invoiced?"&lt;/span&gt;
    &lt;span class="na"&gt;applies_to&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;source&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;severity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;warning&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How I evaluated it
&lt;/h2&gt;

&lt;p&gt;The repository ships a labelled evaluation set. For every model-answered check there are two real examples of greenwashing and two legitimate look-alike changes: 32 cases in total. All 32 classify correctly at the default thresholds.&lt;/p&gt;

&lt;p&gt;To be clear about what that means: the set is small and hand-written. It works as a regression suite that stops a change to a question or threshold from breaking a check. It is not a measure of real-world precision. Reports of false positives and misses are the most useful contribution anyone can make to the project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Known limits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;GitHub only for now, no GitLab or Bitbucket.&lt;/li&gt;
&lt;li&gt;Each hunk is judged on its own, without the rest of the file or repository.&lt;/li&gt;
&lt;li&gt;Very large pull requests are capped at 300 hunks, and each hunk at 8,000 characters.&lt;/li&gt;
&lt;li&gt;Every finding is a probability. Intentional changes can be flagged, and subtle cheating can be missed.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Greenwash is open source under Apache-2.0 and self-hosted: you register your own GitHub App and deploy with Docker or Railway. The setup guide is in the repository.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Code: &lt;a href="https://github.com/ayushgml/greenwash-oss" rel="noopener noreferrer"&gt;github.com/ayushgml/greenwash-oss&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Video walkthrough: &lt;a href="https://www.youtube.com/watch?v=g_i37SqXvZU" rel="noopener noreferrer"&gt;AI Agents Can Fake a Green CI. I Built Greenwash to Catch It.&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Case study: &lt;a href="https://itsayush.dev/work/greenwash" rel="noopener noreferrer"&gt;itsayush.dev/work/greenwash&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you run coding agents against real repositories, I'd like to hear which patterns you see that aren't on this list.&lt;/p&gt;




</description>
      <category>ai</category>
      <category>github</category>
      <category>testing</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
