<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Jeff Lenamon</title>
    <description>The latest articles on DEV Community by Jeff Lenamon (@lenamonj).</description>
    <link>https://dev.to/lenamonj</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4056670%2F1c71edb8-3b44-4dc2-8847-53aea265d86b.jpg</url>
      <title>DEV Community: Jeff Lenamon</title>
      <link>https://dev.to/lenamonj</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/lenamonj"/>
    <language>en</language>
    <item>
      <title>Alone at Midnight, I Sent an AI to Fix the Software the World Runs On</title>
      <dc:creator>Jeff Lenamon</dc:creator>
      <pubDate>Sun, 27 Sep 2026 14:44:54 +0000</pubDate>
      <link>https://dev.to/lenamonj/alone-at-midnight-i-sent-an-ai-to-fix-the-software-the-world-runs-on-c38</link>
      <guid>https://dev.to/lenamonj/alone-at-midnight-i-sent-an-ai-to-fix-the-software-the-world-runs-on-c38</guid>
      <description>&lt;p&gt;At 12:03 AM on September 6, I filed a pull request against Google's snappy compression library. The bug inside it was almost absurd. Compress a 4 GiB file with any release build, and the header it wrote swore the file held 0 bytes. A machine I built called Jeffy Loop had found it that night. I was about to ask Google to believe my machine.&lt;/p&gt;

&lt;p&gt;The maintainer merged it the next day. The loop's patches have now been merged into code owned by NVIDIA, Meta, Tesla, Google, Apple, Microsoft, Netflix, Apache, Oracle, IBM, Cisco, Square, Cloudflare, JetBrains and Canonical, and into the URL parser inside Node.js. That comes to 57 patches across 45 projects, every one written by the loop and every one accepted by a stranger who owed me nothing. It was one person, at night, betting that a machine held to evidence could find what busy people miss.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trust nobody had earned
&lt;/h2&gt;

&lt;p&gt;Think about what sits underneath every app you touched today. Compression libraries. Memory allocators. The URL parser inside Node.js. A Rust UUID library with 192 million downloads in the last 90 days. Much of it is kept alive by small teams of software engineers working together to maintain the integrity of the repo.&lt;/p&gt;

&lt;p&gt;Imagine AI that could be trusted to find real defects in that layer and fix them properly. Every piece of software stacked on top would get better at once. That would be one of the largest quiet upgrades in the history of computing.&lt;/p&gt;

&lt;p&gt;Nobody had earned that trust. AI agents made it cheap to break every rule of a good pull request at once, on a hundred repositories, before breakfast. So maintainers learned to close them without reading them, and they were right to.&lt;/p&gt;

&lt;p&gt;I wanted to earn it anyway, one merge at a time, from people with every reason to say no to someone like me.&lt;/p&gt;

&lt;h2&gt;
  
  
  A machine that cannot crown itself
&lt;/h2&gt;

&lt;p&gt;I built Jeffy Loop as an open-source control loop for Claude Code. It audits a codebase, attacks what it finds, fixes it, and proves the fix with a check that can fail. Every iteration is a local commit. A broken verify gets reverted. Nothing is ever pushed.&lt;/p&gt;

&lt;p&gt;The heart of it is that the agent never gets the last word. An adversarial evaluator and a shell gate re-check every claim that the work is done. On one build that evaluator was called 8 times and rejected 7. And the final verdict sits entirely outside my reach. A merged pull request is the one result the loop cannot award itself.&lt;/p&gt;

&lt;p&gt;Then I pointed it at the world. 132 open-source projects with no connection to me, each judged by its own test suite. 103 of them run to convergence across 13 languages, with no language-specific analyzer anywhere in the engine.&lt;/p&gt;

&lt;h2&gt;
  
  
  The door
&lt;/h2&gt;

&lt;p&gt;Every patch leaves my machine through a set of rules I built called housebroken, published on PyPI as housebroken-cli. Every rule in it came from a maintainer's verdict on a real pull request.&lt;/p&gt;

&lt;p&gt;One maintainer asked the question that every AI patch should have to survive: "Is there any actual bug here? Either a security issue where we read out-of-bounds, or a case where we claimed to succeed but returned incorrect output?" I had a measurement and no broken output. I conceded. He closed it with one line: "Closing - no user-visible bug." He was right, and from that day a finding backed only by a measurement never leaves my machine.&lt;/p&gt;

&lt;p&gt;NVIDIA's contributing guide says: "You must use your real name (sorry, no pseudonyms or anonymous contributions)." The loop had found a buffer allocated with zero bytes that the library then wrote a directory path into, crashing the process. A machine wrote that fix. NVIDIA wanted a human name standing behind it, and a signed commit to prove it. I put mine there. They merged it 12 days after I filed.&lt;/p&gt;

&lt;h2&gt;
  
  
  "Your AI is imagining things"
&lt;/h2&gt;

&lt;p&gt;Then came the review that could have ended it.&lt;/p&gt;

&lt;p&gt;An Apache Commons maintainer, reading a reflection fix the loop wrote for commons-lang, converted it to a draft and wrote: "I think this PR creates 2 bugs so it looks like we are missing some tests since the build was green." Then: "Your AI is imagining things when it talks about PR #1427 because that PR was closed without being merged."&lt;/p&gt;

&lt;p&gt;He was right on both counts. The fix had introduced two bugs the green build never caught. The description had borrowed its history from a summary table instead of the code's own record.&lt;/p&gt;

&lt;p&gt;Every skeptic of AI coding would have been vindicated in that one comment. I went to work. Each case he raised was reproduced, fixed and covered by a test. He raised a third. It got the same treatment. On September 9 he merged the second rework. And the mistake became machinery: history in a pull request now comes from git blame, never from a table.&lt;/p&gt;

&lt;p&gt;That is the engine of the whole project. Every closure that was my mistake was a mistake of reading, and every one became a gate that reads for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the strangers said
&lt;/h2&gt;

&lt;p&gt;At NVIDIA, an empty setting in a config file passed validation, and a container started with no GPU access and no error. At Netflix, every query parameter with an uppercase letter in its name came back empty. At Tesla, a proxy read every request body with no size limit. At Meta, StyleX emitted every declaration twice. At Microsoft, an allocator that promised zeroed memory handed back uninitialized memory, and the library's own author merged the fix.&lt;/p&gt;

&lt;p&gt;The URL parser inside Node.js took the loop's fix twelve minutes after it was opened. Microsoft's snmalloc went from filed at 2:12 PM to merged at 4:05 PM, with two words. An Apache maintainer asked for four more test cases, and sixteen seconds after CI passed on them he wrote "looks good, merged."&lt;/p&gt;

&lt;p&gt;And I published every loss beside those wins. 28 projects failed. mruby took 10 runs and 113 iterations and never converged. It sits in the same public scorecard as the merges, because a record that hides its losses deserves no one's trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  Come build this way
&lt;/h2&gt;

&lt;p&gt;I believe this is where software is going. Agents will write a growing share of the world's code. The ones worth trusting will carry their evidence with them, prove their fixes red then green, learn from every rejection, and hand the final verdict to someone who owes them nothing.&lt;/p&gt;

&lt;p&gt;It is buildable now. One person proved it at night in their spare time. Jeffy Loop is open source under the MIT license, and it installs with one command.&lt;/p&gt;

&lt;p&gt;Come knock on that door with me. Bring evidence.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Jeffy Loop:&lt;/strong&gt; &lt;a href="https://github.com/lenamonj/jeffy-loop" rel="noopener noreferrer"&gt;github.com/lenamonj/jeffy-loop&lt;/a&gt;, &lt;code&gt;pip install jeffy-loop&lt;/code&gt;. The receipts, every merged patch with its link, are on &lt;a href="https://github.com/lenamonj/jeffy-loop/blob/main/evals/README.md" rel="noopener noreferrer"&gt;the receipts page&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;housebroken:&lt;/strong&gt; &lt;a href="https://github.com/lenamonj/housebroken" rel="noopener noreferrer"&gt;github.com/lenamonj/housebroken&lt;/a&gt;, &lt;code&gt;pip install housebroken-cli&lt;/code&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>programming</category>
      <category>claudecode</category>
    </item>
    <item>
      <title>Jeffy Loop: The Coding Agent That Won’t Let Itself Lie</title>
      <dc:creator>Jeff Lenamon</dc:creator>
      <pubDate>Sun, 30 Aug 2026 00:58:50 +0000</pubDate>
      <link>https://dev.to/lenamonj/jeffy-loop-the-coding-agent-that-wont-let-itself-lie-3907</link>
      <guid>https://dev.to/lenamonj/jeffy-loop-the-coding-agent-that-wont-let-itself-lie-3907</guid>
      <description>&lt;p&gt;Most autonomous coding agents are optimistic. They audit, fix, declare victory, and sometimes the victory is mostly vibes.&lt;/p&gt;

&lt;p&gt;Jeffy Loop takes the opposite stance. Built for Claude Code, it forces the model to act like a disciplined principal engineer: audit first, fix one verified task at a time, checkpoint everything, and refuse to stop until “done” survives independent checks.&lt;/p&gt;

&lt;p&gt;It starts from Geoffrey Huntley’s Ralph technique (re-feeding one prompt in a loop) and wraps it in real engineering method.&lt;/p&gt;

&lt;p&gt;Type /jeffy 10 and walk away.&lt;/p&gt;

&lt;p&gt;Map the public surface and run a breadth-first audit. Turn every finding into a backlog item with a runnable acceptance check. Execute one verified, checkpointed task per iteration.&lt;/p&gt;

&lt;p&gt;Revert any change that breaks the project’s own tests via a verify gate.&lt;br&gt;
Only stop when a fresh audit is clean, an adversarial evaluator countersigns, and a shell script re-checks the claim.&lt;/p&gt;

&lt;p&gt;“Done” is machine-enforced, not the model’s opinion.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real receipts&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;52 open-source projects converged across 13 languages&lt;br&gt;
31 non-convergences fully documented&lt;br&gt;
Upstream fixes accepted in bat, fasthttp, jsoncpp, PapaParse, and chalk&lt;/p&gt;

&lt;p&gt;These include real High-severity bugs found behind green test suites in popular projects.&lt;/p&gt;

&lt;p&gt;Why it stands out&lt;/p&gt;

&lt;p&gt;It cannot claim it looked at code it never examined. Lessons become permanent rules. Progress means actual code movement. The engine itself is held to hundreds of behavioural checks.&lt;/p&gt;

&lt;p&gt;Prefer short, fresh-context runs over long sessions. Context pressure, stalls, and thrashing are all measured and cut short.&lt;/p&gt;

&lt;p&gt;If you’ve been burned by agents that rewrite working code or quietly delete failing tests, this is a different bet: process-heavy, transparent, and forced to prove its claims.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/lenamonj/jeffy-loop" rel="noopener noreferrer"&gt;https://github.com/lenamonj/jeffy-loop&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>webdev</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Four High-severity bugs were hiding behind a green test suite in a 7k-star library</title>
      <dc:creator>Jeff Lenamon</dc:creator>
      <pubDate>Fri, 31 Jul 2026 13:23:34 +0000</pubDate>
      <link>https://dev.to/lenamonj/four-high-severity-bugs-were-hiding-behind-a-green-test-suite-in-a-7k-star-library-57a8</link>
      <guid>https://dev.to/lenamonj/four-high-severity-bugs-were-hiding-behind-a-green-test-suite-in-a-7k-star-library-57a8</guid>
      <description>&lt;p&gt;Run &lt;code&gt;pytest&lt;/code&gt; on &lt;a href="https://github.com/kennethreitz/records" rel="noopener noreferrer"&gt;kennethreitz/records&lt;/a&gt; at upstream HEAD and you get the answer every maintainer wants: 31 passed. Green across the board.&lt;/p&gt;

&lt;p&gt;At that exact same commit, all of the following are true:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;db.query("INSERT ...")&lt;/code&gt; silently loses your data. The rows vanish when the connection closes. No error, no warning, nothing.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;db.bulk_query(...)&lt;/code&gt; loses data the same way.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;db.transaction()&lt;/code&gt; swallows every exception raised inside it. Your failed transaction reports success.&lt;/li&gt;
&lt;li&gt;Every single &lt;code&gt;query()&lt;/code&gt; call leaks a pooled connection. With a default-sized pool, the third query can hang your process.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Four High-severity bugs. One green test suite. I want to walk through how that is possible, because the mechanism is more interesting than any single bug, and it is probably present in more codebases than we would like to admit.&lt;/p&gt;

&lt;p&gt;Full disclosure before we start: I found these with an autonomous audit loop I built, and I will say a little about it at the end. Everything in this post is independently checkable. The repro script runs against upstream HEAD, and all four bugs were disclosed upstream with a PR offer in &lt;a href="https://github.com/kennethreitz/records/issues/236" rel="noopener noreferrer"&gt;records#236&lt;/a&gt;. None of this is a criticism of records or its maintainers. It is a beloved library with a beautiful API, and what happened to it is a story about dependency semantics, not carelessness.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually broke
&lt;/h2&gt;

&lt;p&gt;records was written in the SQLAlchemy 1.x era, and its &lt;code&gt;Connection&lt;/code&gt; class leans on two 1.x behaviors that SQLAlchemy 2.0 removed.&lt;/p&gt;

&lt;p&gt;First, autocommit. Under 1.x, executing a bare &lt;code&gt;INSERT&lt;/code&gt; through a connection could commit implicitly. Under 2.x, autocommit is gone: if nobody calls &lt;code&gt;commit()&lt;/code&gt;, the transaction rolls back when the connection closes. records never calls &lt;code&gt;commit()&lt;/code&gt; for plain queries, because it never had to. So on SQLAlchemy 2.x, &lt;code&gt;db.query("INSERT INTO users ...")&lt;/code&gt; executes, appears to succeed, and the row quietly disappears when the connection goes away. Same story for &lt;code&gt;bulk_query&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Second, &lt;code&gt;close_with_result&lt;/code&gt;. Under 1.x, this flag meant the connection would close itself once the result set was exhausted, which is how records returned connections to the pool. Under 2.x the flag is accepted and ignored. Every &lt;code&gt;query()&lt;/code&gt; checks out a pooled connection and never returns it. A pool of size 1 with 1 overflow dies on the third query.&lt;/p&gt;

&lt;p&gt;The third bug is older and more human. &lt;code&gt;db.transaction()&lt;/code&gt; wraps its body in a bare &lt;code&gt;except:&lt;/code&gt; that rolls back and then does not re-raise. Callers cannot tell a failed transaction from a successful one. Here is the detail that stopped me: upstream added the missing &lt;code&gt;raise&lt;/code&gt; on 2026-02-08 and reverted it the same day. The fix existed for a few hours. Something in the test suite presumably objected, and the fix lost.&lt;/p&gt;

&lt;p&gt;Which brings us to the real subject.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the suite stayed green
&lt;/h2&gt;

&lt;p&gt;The records test suite runs against SQLite, in memory. That choice is fast, hermetic, and completely blind to all four bugs.&lt;/p&gt;

&lt;p&gt;An in-memory SQLite database lives on a single connection. There is no separate reader that would fail to see uncommitted rows, so the missing commits are invisible: within that one shared connection, uncommitted state looks exactly like committed state. There is no real pool boundary to exhaust, so the leaked connections are invisible too. The suite cannot observe persistence, cannot observe pool behavior, and cannot observe commit semantics, because in its world those concepts do not exist.&lt;/p&gt;

&lt;p&gt;And the transaction bug? One test actively asserted the broken behavior. The suite did not just miss the bug. It defended it. That is almost certainly why the one-day fix got reverted: the fix made a test fail, and the test was wrong.&lt;/p&gt;

&lt;p&gt;This is the generalizable lesson, and it has nothing to do with AI or with records specifically. A test gate is a measurement instrument, and an instrument can be precise, fast, reliable, and pointed at the wrong world. records' gate measured "does this code work against a single shared in-memory connection," and the answer was honestly yes. Nobody was asking "does this code work against a database," which is the only question users care about.&lt;/p&gt;

&lt;p&gt;Green is not a property of the code. Green is a property of the code and the gate together.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the bugs were found
&lt;/h2&gt;

&lt;p&gt;I ran an autonomous improvement loop against a local clone. The loop's first iteration is always an audit, and the audit has one hard rule: a finding exists only if it can be reproduced on the spot. No "this looks wrong," no linting vibes. Every claim above began life as a small script the audit ran against upstream HEAD and watched fail.&lt;/p&gt;

&lt;p&gt;The audit also noticed something structural: three of the four Highs are the same bug wearing different costumes. Missing commits, ignored &lt;code&gt;close_with_result&lt;/code&gt;, and the varargs breakage all trace to one root cause, a &lt;code&gt;Connection&lt;/code&gt; class written against SQLAlchemy 1.x execution semantics. The loop has a three-strike rule for exactly this pattern: the third finding sharing one root cause forces a single structural fix, never a third spot patch.&lt;/p&gt;

&lt;p&gt;So the fix was structural. Commit-and-finalize when a result set is exhausted, immediate commit for DML statements, single-result connections actually returned to the pool, and a transaction wrapper that guarantees auto-commit never fires inside an explicit transaction. That last clause matters: the naive spot patch (just commit after every statement) silently breaks rollback, because a rollback that follows an auto-commit rolls back nothing. The structural fix proves rollback still works, with a test.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;transaction()&lt;/code&gt; fix restores the change upstream made and reverted: rollback, then re-raise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Giving the gate teeth
&lt;/h2&gt;

&lt;p&gt;Fixing code under a blind gate is pointless, so the loop's next task rebuilt the instrument. The test that asserted the swallow bug was corrected. A file-backed regression suite was added: persistence through both query paths and both API forms, pool exhaustion, exception propagation, explicit rollback.&lt;/p&gt;

&lt;p&gt;Then the honesty check, which is my favorite artifact of the whole run: the new regression suite was executed against the pre-fix code from the audit checkpoint, and 5 of its 6 tests fail there. A regression test that never failed on the broken code proves nothing. This one is proven to detect the world it claims to detect. After the fixes, the full suite is 37 passing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it declined to do
&lt;/h2&gt;

&lt;p&gt;Two things did not happen, and they matter as much as the fixes.&lt;/p&gt;

&lt;p&gt;The loop declined to impose a lint gate, with a written reason: records declares no lint configuration, so a lint pass would be cosmetic churn on a dormant project. And two decisions it had no right to make were filed for the owner instead of taken: whether to replace the unmaintained &lt;code&gt;docopt&lt;/code&gt; dependency, and whether to restore the multi-backend CI matrix. An autonomous tool that seizes owner decisions is worse than no tool.&lt;/p&gt;

&lt;p&gt;Everything shipped as artifacts: a &lt;code&gt;repro.py&lt;/code&gt; that demonstrates all four bugs on upstream HEAD and their absence after the patch, a &lt;code&gt;fixes.patch&lt;/code&gt; of about 180 lines covering code, tests, and packaging, and a journal of every iteration including the operator mistakes the method itself caught (a premature "passed" claim, an invalid negative-path check, and one mistyped commit hash in a ledger, each caught and corrected by the loop's own rules).&lt;/p&gt;

&lt;p&gt;All four findings went upstream with the repros and a PR offer. Merging anything is, as it should be, the maintainers' call.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;A green suite certifies the gate, not the code. Ask what world your gate can actually observe. In-memory databases, mocked networks, and frozen clocks are all worlds with physics different from production.&lt;/li&gt;
&lt;li&gt;When a correct fix makes a test fail, suspect the test. The one-day life of the &lt;code&gt;raise&lt;/code&gt; fix is what it looks like when a wrong test wins.&lt;/li&gt;
&lt;li&gt;Regression tests should be proven against the broken code. If you never watched them fail, you do not know what they detect.&lt;/li&gt;
&lt;li&gt;The third bug with the same root cause is not a bug, it is architecture. Fix the boundary, not the symptom.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The tool, briefly
&lt;/h2&gt;

&lt;p&gt;The loop is called Jeffy Loop. It is a free, MIT-licensed autonomous improvement loop for Claude Code: audit, backlog with a runnable acceptance check per task, one verified task per iteration behind local git checkpoints and a verify gate, convergence only when a fresh audit comes back clean. The records run used 7 of its 8-iteration budget. The whole engine is one shell script of about 100 lines, and the eval above, alongside the rest of the eval set across eight languages including deliberate control cases that came back clean, lives in the repo with every artifact mentioned here: &lt;a href="https://github.com/lenamonj/jeffy-loop" rel="noopener noreferrer"&gt;https://github.com/lenamonj/jeffy-loop&lt;/a&gt;&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>ai</category>
      <category>testing</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
