<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Necip Reis</title>
    <description>The latest articles on DEV Community by Necip Reis (@necip_reis_bf4457166f77d6).</description>
    <link>https://dev.to/necip_reis_bf4457166f77d6</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4169927%2F7710a339-9b54-478d-ace8-5b09a5f4e7e2.webp</url>
      <title>DEV Community: Necip Reis</title>
      <link>https://dev.to/necip_reis_bf4457166f77d6</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/necip_reis_bf4457166f77d6"/>
    <language>en</language>
    <item>
      <title>When an independent audit corrected a coding-agent test report</title>
      <dc:creator>Necip Reis</dc:creator>
      <pubDate>Wed, 07 Oct 2026 23:19:39 +0000</pubDate>
      <link>https://dev.to/necip_reis_bf4457166f77d6/when-an-independent-audit-corrected-a-coding-agent-test-report-351h</link>
      <guid>https://dev.to/necip_reis_bf4457166f77d6/when-an-independent-audit-corrected-a-coding-agent-test-report-351h</guid>
      <description>&lt;p&gt;I am the founder of AiOrch, the commercially licensed engineering orchestration platform discussed here.&lt;/p&gt;

&lt;p&gt;A coding-agent workflow can produce several different kinds of evidence: an implementation, a review verdict, an integrated branch, a test result and a pull request. Those records can disagree.&lt;/p&gt;

&lt;p&gt;In a recorded AiOrch Python TODO CLI demo, the final development report claimed 36 tests. A separate audit reran &lt;code&gt;python -m pytest tests/ -v&lt;/code&gt; and recorded 35 passed with exit status zero.&lt;/p&gt;

&lt;p&gt;That is a small discrepancy. It is also a useful example of why the generated report needs its own verification boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the run actually recorded
&lt;/h2&gt;

&lt;p&gt;The work was assigned to three implementation agents covering core logic, CLI work and packaging. Dependent work started after its prerequisite was approved. All three implementations received change requests before their final approval. The changes addressed scope cleanup, CLI error handling and packaging/documentation corrections.&lt;/p&gt;

&lt;p&gt;During integration, the record captured a conflict on the CLI branch and its later resolution. All three branches were marked merged inside the session.&lt;/p&gt;

&lt;p&gt;A separate audit session used four audit/follow-up agents. In addition to correcting the test count, it identified specification/report discrepancies and prepared follow-up fixes. The inspected record did not show a merged audit result. An approved audit session therefore does not establish that all of its follow-up changes reached the public PR.&lt;/p&gt;

&lt;p&gt;The case study links to demo-projects PR #5. It was still open when checked on 5 October 2026. Branch integration inside the session and a GitHub PR merge are different states.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five checks worth adding to an agent handoff
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Preserve the approved task specification. In this historical run, an original task file was removed and the audit reconstructed its specification from saved configuration. A durable copy gives later reviewers a stable reference.&lt;/li&gt;
&lt;li&gt;Record the exact command and exit status for each claimed validation result. A successful result without a command, revision and captured output is difficult to reproduce.&lt;/li&gt;
&lt;li&gt;Distinguish evidence from the agent's interpretation of it. Keep test output separately from the prose summary so a count or failure classification can be checked.&lt;/li&gt;
&lt;li&gt;Show integration state explicitly. A reviewed branch, an integrated branch and a delivered PR should each have their own status. If audit fixes have not been integrated, say so.&lt;/li&gt;
&lt;li&gt;Preserve unresolved gates. Missing integration evidence should remain visible in the handoff instead of being silently translated into a pass.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Why parallelism changes the handoff
&lt;/h2&gt;

&lt;p&gt;Separate git worktrees make it possible to assign independent implementation work without agents overwriting a shared checkout. They do not prove that the combined result satisfies the task. Dependencies, changed interfaces and merge conflicts still require integration work.&lt;/p&gt;

&lt;p&gt;Review verdicts and test commands answer different questions. A reviewer can identify scope or behavior problems; a test command can establish what passed for a particular checkout. Keeping both records makes the final result easier to inspect.&lt;/p&gt;

&lt;h2&gt;
  
  
  The scope of this example
&lt;/h2&gt;

&lt;p&gt;This is a small recorded Python demo, not a tool-comparison benchmark or a claim of production reliability. The run contains resume events, and the record does not identify the actor for every resume. It does not establish a run without operator intervention.&lt;/p&gt;

&lt;p&gt;Usage data was missing for several agents, so the dashboard's displayed zero cost cannot support a claim that the run was free. The published case study inspected historical results; it did not rerun them during publication.&lt;/p&gt;

&lt;p&gt;AiOrch runs its orchestration and worktrees on your infrastructure. Configured cloud inference providers still receive the context sent to them. Self-hosting the coordinator should not be described as proof that all inference remains local.&lt;/p&gt;

&lt;p&gt;The useful engineering result here is an inspectable chain from task to implementation, review, integration, audit and handoff—with disagreements retained in the record.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://aiorch.ai/docs/case-studies/todo-cli-review-audit/" rel="noopener noreferrer"&gt;Full case record, limits and public PR&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://aiorch.ai/guides/parallel-coding-agents/" rel="noopener noreferrer"&gt;Parallel orchestration guide&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>git</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
