<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sundar Shyam Jha</title>
    <description>The latest articles on DEV Community by Sundar Shyam Jha (@sundar_shyam_jha_).</description>
    <link>https://dev.to/sundar_shyam_jha_</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4028429%2F43fbd79b-67ab-44fa-8a77-54d4afeebb77.png</url>
      <title>DEV Community: Sundar Shyam Jha</title>
      <link>https://dev.to/sundar_shyam_jha_</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sundar_shyam_jha_"/>
    <language>en</language>
    <item>
      <title>Best AI Software Factories for Enterprise Application Development in 2026</title>
      <dc:creator>Sundar Shyam Jha</dc:creator>
      <pubDate>Wed, 07 Oct 2026 16:41:26 +0000</pubDate>
      <link>https://dev.to/sundar_shyam_jha_/best-ai-software-factories-for-enterprise-application-development-in-2026-21lg</link>
      <guid>https://dev.to/sundar_shyam_jha_/best-ai-software-factories-for-enterprise-application-development-in-2026-21lg</guid>
      <description>&lt;h2&gt;
  
  
  &lt;strong&gt;TL;DR&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;An AI software factory moves a request through connected agent stages to a reviewed pull request, with people approving at set points. I read the public docs for five of them in October 2026: Forge, Factory, Augment Cosmos, Kiro and GitHub's Copilot cloud agent. They differ most on one question, which is whether a human signs off before any code exists. This guide defines the term, runs all five through four tests, and ends with demo questions you can reuse.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Getting Started: What an AI Software Factory Is&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;A coding assistant helps one developer inside one session. You type, it suggests, you accept or reject. A software factory looks after everything around that session: where the request came from, what spec it was built against, who reviewed it and what record it leaves.&lt;/p&gt;

&lt;p&gt;Most vendors now use the word, and they mean slightly different things by it. So I stopped comparing feature lists and scored each platform on four tests instead:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What starts the work.&lt;/strong&gt; A ticket, a prompt, a spec or a legacy repository.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Where a human approves.&lt;/strong&gt; Before code exists, after it, or only at the pull request.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Where context lives.&lt;/strong&gt; Whether the next run remembers what the last one decided.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Where agents run and what they leave behind.&lt;/strong&gt; Your network or the vendor's, and how much of the trail you can export.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I did not benchmark any of these, so nothing here ranks speed or code quality. Everything below comes from each vendor's own docs and enterprise pages as they stood in October 2026.&lt;/p&gt;

&lt;p&gt;To keep the tests concrete, I'll use one hypothetical team throughout: a platform group at a regional bank with a fifteen-year-old Java monolith, about forty developers and an audit team that asks who approved what. It is invented, but the questions it raises are the ones I hear from real enterprise teams.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;How the Five Platforms Compare on the Four Tests&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Before the detail, here is the shape of the field. Copilot's cloud agent and Factory start from tasks and put the human checkpoint at the end. Augment starts from events such as tickets and alerts, with checkpoints you configure. Kiro starts from a spec and gates it in the editor. Forge starts from intent and gates four stages before any code is written.&lt;/p&gt;

&lt;p&gt;None of that makes one of them better. It tells you where each one assumes your process already works, which is the more useful thing to know before a demo.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs7u5d1rn5qr239kjnd54.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs7u5d1rn5qr239kjnd54.png" alt="Opsera" width="512" height="284"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub Copilot Cloud Agent as the Baseline Most Teams Own&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If your bank's code already lives on GitHub, the cloud agent is the lowest-change option, and its guardrails are more specific than I expected. According to GitHub's &lt;a href="https://docs.github.com/en/copilot/concepts/agents/cloud-agent/risks-and-mitigations" rel="noopener noreferrer"&gt;own risk documentation&lt;/a&gt;, only users with write access can trigger it. It pushes to a single branch, either the pull request's own branch or a new &lt;code&gt;copilot/&lt;/code&gt; branch, and it stays subject to branch protection and required checks.&lt;/p&gt;

&lt;p&gt;It cannot mark a pull request ready for review, approve it or merge it, and the person who asked for the work cannot approve it either. Workflows do not run until someone with write access clicks Approve and run. By default it also runs CodeQL, checks new dependencies against the GitHub Advisory Database and runs secret scanning on its own output. Commits are signed, name the requester as co-author and link to the session log, and administrators get audit log events.&lt;/p&gt;

&lt;p&gt;For our bank, that is a solid answer to "who merged this." What it does not give you is a stage before code. Work starts from an issue or a prompt, and the human checkpoint is the pull request. If the issue was vague, the review is where you find out.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Kiro Puts the Spec Inside the Editor&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Kiro, from AWS, is the clearest spec-first tool in this group. A &lt;a href="https://kiro.dev/docs/specs/index" rel="noopener noreferrer"&gt;feature spec&lt;/a&gt; produces three files: &lt;code&gt;requirements.md&lt;/code&gt;, &lt;code&gt;design.md&lt;/code&gt; and &lt;code&gt;tasks.md&lt;/code&gt;. In the requirements-first workflow you confirm the requirements before Kiro writes the design, and the design leads to a task list you can run one task at a time or all together. When you run everything, Kiro builds a dependency graph and runs independent tasks concurrently.&lt;/p&gt;

&lt;p&gt;The detail I would watch is the Quick Spec path. It generates all three files in one pass with no approval gates. It is convenient, and it is the setting I would keep out of regulated work, because it removes the one stage that makes the tool interesting.&lt;/p&gt;

&lt;p&gt;Specs run in the Kiro IDE and on the web, where the agent can implement the plan and open a pull request. I did not read Kiro's enterprise administration docs for this piece, so I am treating it as a strong spec workflow rather than a complete factory. For the bank, the open question is how that workflow scales across forty developers, which the public pages do not answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Factory Runs Agents Under Managed Settings&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Factory's unit of work is the Droid. Its &lt;a href="https://docs.factory.com/docs/enterprise/network-and-deployment" rel="noopener noreferrer"&gt;deployment documentation&lt;/a&gt; describes three patterns: cloud-managed, where Factory's cloud is the control plane and model traffic can go through your own gateways; hybrid; and fully airgapped, where Factory's cloud is not reachable at runtime. Droids run on laptops, CI runners, VMs, Kubernetes clusters and airgapped networks.&lt;/p&gt;

&lt;p&gt;Governance is expressed as a hierarchy of managed settings covering model access, command policies, MCP allowlists, hooks, sandbox policy and retention. Droids emit OpenTelemetry metrics to collectors you own, and Factory lists SOC 2 Type II, ISO 27001 and ISO 42001 on the same pages.&lt;/p&gt;

&lt;p&gt;What I did not find is an approved-spec stage before the agent starts. The controls describe what a Droid may do once work is assigned. If the bank's weakness is unclear requirements, they would have to build that stage themselves. If the weakness is control over what agents can touch, this is one of the most detailed runtime policy models I came across.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Augment Cosmos Runs Ticket-to-PR Loops You Configure&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Augment describes &lt;a href="https://www.augmentcode.com/product/software-factory" rel="noopener noreferrer"&gt;Cosmos&lt;/a&gt; as the layer that connects agents, codebase context and workflows. Requests enter as tickets, pull requests or alerts, and each loop has a trigger, an outcome and a human checkpoint you define. The building blocks include Experts, which are task-specific agents, a Context Engine, a model router called Prism, and either isolated cloud VMs or a self-hosted daemon that runs agents on your machines while Cosmos stays the control plane.&lt;/p&gt;

&lt;p&gt;Experts, environments and loops are versioned, and budget controls cap spend per automation or per user. Augment's security pages list SOC 2 Type II and ISO/IEC 42001 and say it never trains on customer proprietary data.&lt;/p&gt;

&lt;p&gt;Augment also publishes a case study from its own engineering team and says the figures are observations, not a controlled experiment. I would read any vendor's internal numbers that way, and I will say the same about the figures in the next section. Its advice to start with a code review loop is sound, because a pull request gives you a clear trigger and a natural place for a person to look.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Forge Gates Four Stages Before Any Code Exists&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://opsera.ai/forge" rel="noopener noreferrer"&gt;Forge&lt;/a&gt; is the one I had to read most carefully, because it is built around the gap the other four leave open. The pattern so far is that nothing before the code is required to be approved. Copilot and Factory begin at a task. Augment begins at an event. Kiro has a spec stage but lets you skip its gates.&lt;/p&gt;

&lt;p&gt;Forge's default New Build path runs Intent, PRD-Spec, Architecture, User Stories and Testing, then saves an Application Context snapshot. A modernization project adds an Assessment stage at the front, where Forge reads a Git repository, a local folder or a ZIP and reports the current state before anyone writes a goal. For our bank's monolith, that is the part I would test first.&lt;/p&gt;

&lt;p&gt;Here is the number I found most useful. The docs say Intent, requirements, Architecture and User Stories each need an explicit Approve &amp;amp; Continue, while Testing and Assessment need none. Counting from the quick-start guides, that is four of five stages gated in the default New Build journey and four of six in the default modernization journey. All four gates sit before code. Tenants can switch optional stages on and off, so check it against your own configuration.&lt;/p&gt;

&lt;p&gt;Two details in the approval model are easy to miss. Approval records a version without locking the artifact. And regenerating an upstream artifact clears the approvals downstream, so a changed requirement cannot quietly keep an old architecture sign-off. For an audit team, that second behaviour is the interesting one.&lt;/p&gt;

&lt;p&gt;{INFOGRAPHIC: Horizontal flow of six rounded boxes labelled Intent, PRD-Spec, Architecture, User Stories, Testing, Application Context. Under the first four, a filled circle with a check mark and the words "Approve &amp;amp; Continue". Under Testing, an empty circle and "No approval step". Under Application Context, "Saved snapshot". A thin bracket under the first four boxes reads "Four gates before any code exists". Footer: "Source: Forge documentation, October 2026".}&lt;/p&gt;

&lt;p&gt;Code comes from the Coding Agent, which implements an approved user story in a linked Git repository. It pushes a &lt;code&gt;forge/wo-{id}&lt;/code&gt; branch and opens a pull request when your Git provider allows it. The connector needs push and pull request permissions, and the run screen lets you skip the test phase or the AI review phase, so I would decide up front who may flip those. If your developers prefer their own editor, the same stories open in an IDE over MCP.&lt;/p&gt;

&lt;p&gt;Opsera's own figures for Forge are vendor-reported and I have not verified them. Its site describes a cloud to on-prem Kubernetes conversion at Senao Wireless going from seven months to one week, and a Belcorp workload dropping from two to six weeks to about an hour. Its FAQ also says customer code and architecture data are not used to train foundation models. My demo question would be which stages are enabled in my tenant, since the docs note that policy can hide the optional ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;How I Would Pick Based on Where Delivery Breaks&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;If nobody can explain the legacy system, start with a tool that begins with assessment. Forge's modernization journey opens with Assessment and a ForgeScore health report before any goal is set.&lt;/p&gt;

&lt;p&gt;If the codebase is healthy and the pain is a pile of tickets, alerts and reviews, look at Factory or Augment, since both start from events and tasks.&lt;/p&gt;

&lt;p&gt;If you are on GitHub and want the smallest change, turn on the Copilot cloud agent and tighten branch protection first. If you want specs in the editor and are not ready for a platform, try Kiro on one team.&lt;/p&gt;

&lt;p&gt;One position I will defend: if your team has no written specs today, adding agents moves the bottleneck to the review queue. Pick the tool that makes you produce the spec, or build that habit before you buy anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Five Questions I Would Ask in Every Demo&lt;/strong&gt;
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Which stages need a human click, and which run without one?
&lt;/li&gt;
&lt;li&gt;If an upstream artifact changes, what happens to approvals downstream?
&lt;/li&gt;
&lt;li&gt;Where do the agents run, and what leaves my network?
&lt;/li&gt;
&lt;li&gt;Can I pull the full record for one change, from request to merge?
&lt;/li&gt;
&lt;li&gt;What stops an agent from reaching production credentials or deploy commands?&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Where to Start With Your Own Shortlist&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The common thread across all five is that the checkpoint moves earlier as the tool gets more spec-driven. Copilot checks at the pull request, Factory and Augment at configured gates around tasks, Kiro at the spec, and Forge at four stages before code. Which position is right depends on whether your delivery problems begin in requirements or in review.&lt;/p&gt;

&lt;p&gt;My suggestion is to pick two tools that sit at different points on that line, run the same small change through both, and compare the record each one leaves. If you want to see the spec-first end up close, the &lt;a href="https://www.softwareforge.ai/docs/introduction" rel="noopener noreferrer"&gt;Forge docs introduction&lt;/a&gt; walks through the New Build and modernization journeys stage by stage, which is the quickest way to check the approval behaviour yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;FAQ&lt;/strong&gt;
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Does a software factory replace Copilot or Cursor?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Forge runs alongside the assistants a team already uses and supplies the intent and context around them. Augment and Factory ship their own agents, so with those you are choosing between agent runtimes.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Do Forge's approvals lock an artifact once approved?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;No, Approval records the current version and moves the journey forward. If you regenerate an upstream artifact, the approvals downstream of it are cleared.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Can an agent merge its own pull request on GitHub?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Not in Copilot's cloud agent. It cannot approve or merge, and the person who requested the work cannot approve it either.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Is Kiro's Quick Spec suitable for regulated teams?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;I would not use it there. It generates requirements, design and tasks in one pass with no approval gates, which removes the review step regulated teams rely on.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Which of these can run fully airgapped?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Of the five, Factory documents a fully airgapped deployment. Augment's agents can run on your machines while the control plane stays with the vendor, so it sits in between.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;How current is this comparison of the tools?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;It reflects public pages read in October 2026. Vendors change these pages often, so confirm any detail your decision depends on.&lt;/p&gt;

</description>
      <category>software</category>
      <category>ai</category>
      <category>api</category>
      <category>devops</category>
    </item>
    <item>
      <title>7 AI tools for legacy modernization in 2026, and what 507 work orders say about picking one</title>
      <dc:creator>Sundar Shyam Jha</dc:creator>
      <pubDate>Tue, 11 Aug 2026 15:23:44 +0000</pubDate>
      <link>https://dev.to/sundar_shyam_jha_/7-ai-tools-for-legacy-modernization-in-2026-and-what-507-work-orders-say-about-picking-one-1oc3</link>
      <guid>https://dev.to/sundar_shyam_jha_/7-ai-tools-for-legacy-modernization-in-2026-and-what-507-work-orders-say-about-picking-one-1oc3</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbpalv09bsta0jbq0605b.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbpalv09bsta0jbq0605b.png" alt="7 AI tools for legacy modernization in 2026, and what 507 work orders say about picking one" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Rewriting an existing system needs about four times more written "definition of done" per unit of work than building a new one. I got that from a public dataset of nine AI-assisted projects: eight modernizations averaged 1.5 acceptance criteria per story point, the one greenfield build averaged 0.4. If you are choosing an AI modernization tool this year, that ratio is a better buying question than any benchmark score.&lt;/p&gt;

&lt;p&gt;Here is the method, then the seven tools, then how I would actually sequence a migration.&lt;/p&gt;

&lt;h2&gt;
  
  
  The number, and how I got it
&lt;/h2&gt;

&lt;p&gt;SoftwareForge published a &lt;a href="https://softwareforge.ai/blog/enterprise-benchmark" rel="noopener noreferrer"&gt;2026 enterprise benchmark&lt;/a&gt; covering 507 work orders and 3,410 story points across nine initiatives, with the acceptance-criteria density listed per stack. Acceptance criteria are the checkable statements attached to a unit of work, things like "the endpoint rejects unauthenticated requests with 401" or "coverage is at or above 90%."&lt;/p&gt;

&lt;p&gt;I divided story points by work orders to get the average size of one unit of change, then divided acceptance criteria by that to normalize. Small table, run it yourself:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Project&lt;/th&gt;
&lt;th&gt;Story pts / WO&lt;/th&gt;
&lt;th&gt;AC / WO&lt;/th&gt;
&lt;th&gt;AC per story point&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;.NET Framework + WCF monolith to microservices&lt;/td&gt;
&lt;td&gt;8.45&lt;/td&gt;
&lt;td&gt;11.3&lt;/td&gt;
&lt;td&gt;1.34&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;T-SQL to Aurora MySQL&lt;/td&gt;
&lt;td&gt;6.46&lt;/td&gt;
&lt;td&gt;10.4&lt;/td&gt;
&lt;td&gt;1.61&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;T-SQL to PL/pgSQL&lt;/td&gt;
&lt;td&gt;6.79&lt;/td&gt;
&lt;td&gt;10.6&lt;/td&gt;
&lt;td&gt;1.56&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Legacy framework + jQuery to Vue 3&lt;/td&gt;
&lt;td&gt;6.41&lt;/td&gt;
&lt;td&gt;8.9&lt;/td&gt;
&lt;td&gt;1.39&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Java Spring in-place hardening (healthcare)&lt;/td&gt;
&lt;td&gt;7.05&lt;/td&gt;
&lt;td&gt;8.5&lt;/td&gt;
&lt;td&gt;1.21&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Java servlets to Node/TS + React&lt;/td&gt;
&lt;td&gt;4.33&lt;/td&gt;
&lt;td&gt;6.6&lt;/td&gt;
&lt;td&gt;1.53&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JavaScript to typed React SPA&lt;/td&gt;
&lt;td&gt;3.55&lt;/td&gt;
&lt;td&gt;9.0&lt;/td&gt;
&lt;td&gt;2.54&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;.NET Core 2.1 (EOL) to .NET 8&lt;/td&gt;
&lt;td&gt;2.79&lt;/td&gt;
&lt;td&gt;5.0&lt;/td&gt;
&lt;td&gt;1.79&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Greenfield multiplatform + Terraform&lt;/td&gt;
&lt;td&gt;7.93&lt;/td&gt;
&lt;td&gt;3.0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.38&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Median for the eight migrations: 1.55. The greenfield project: 0.38.&lt;/p&gt;

&lt;p&gt;Two honest caveats. There is one greenfield row, so this is a signal and not a law. And these are planning artifacts, not measured outcomes, so the density reflects how much someone thought needed pinning down, not how much turned out to matter. Useful reading, all the same: when you preserve behavior you cannot describe, the plan has to carry the description. When you invent behavior, you get to change your mind.&lt;/p&gt;

&lt;p&gt;If you want to check your own backlog against this, the calculation is four lines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;csv&lt;/span&gt;

&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;backlog.csv&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;  &lt;span class="c1"&gt;# columns: story_points, acceptance_criteria
&lt;/span&gt;    &lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;csv&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;DictReader&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;

&lt;span class="n"&gt;sp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;story_points&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;ac&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;acceptance_criteria&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; work orders, &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;ac&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;sp&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; acceptance criteria per story point&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Under 0.5 on a migration backlog means your agents are guessing at behavior you never wrote down.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is the buying question
&lt;/h2&gt;

&lt;p&gt;Legacy modernization rarely goes wrong at the translation step. A converter turns a &lt;code&gt;PERFORM&lt;/code&gt; loop into a &lt;code&gt;for&lt;/code&gt; loop in an afternoon. No converter can tell you why that loop skips every third record when an account flag is set, because nobody ever wrote that rule anywhere except inside the loop.&lt;/p&gt;

&lt;p&gt;Coding agents have the same blind spot, at higher throughput. In &lt;a href="https://www.anthropic.com/customers/stripe" rel="noopener noreferrer"&gt;Stripe's rollout of Claude Code to 1,370 engineers&lt;/a&gt;, one team moved 10,000 lines of Scala to Java in four days against a ten-engineering-week estimate. Read the same page for the part people skip: their infrastructure lead describes the mental model that made it work, treating the assistant as a capable new engineer who knows every language but has no business context and no idea how things are done here. Speed came from the context they supplied, not from the model.&lt;/p&gt;

&lt;p&gt;Without that context, the failures are quiet. &lt;a href="https://dora.dev/research/2025/dora-report/" rel="noopener noreferrer"&gt;DORA's 2025 report&lt;/a&gt; found AI adoption correlating with higher throughput and higher instability at the same time: more change failures, more rework. &lt;a href="https://www.gitclear.com/ai_assistant_code_quality_2025_research" rel="noopener noreferrer"&gt;GitClear's 2025 code-quality analysis&lt;/a&gt; tracked duplicated blocks rising as AI-assisted commit volume rose. And in &lt;a href="https://survey.stackoverflow.co/2025/ai" rel="noopener noreferrer"&gt;Stack Overflow's 2025 survey&lt;/a&gt;, the top frustration among developers using AI was output that is "almost right, but not quite," with debugging that near-miss costing more than writing it fresh.&lt;/p&gt;

&lt;p&gt;The tool question, then, is not "can it generate Java from COBOL." All seven below can produce plausible target-language code. What separates them is what happens in the hours before generation starts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The seven
&lt;/h2&gt;

&lt;p&gt;I have grouped these by what they do first, because that is the actual difference.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. SoftwareForge (Forge)
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftrve1zz4b7wyb1r2u8ws.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftrve1zz4b7wyb1r2u8ws.png" alt="SoftwareForge" width="800" height="367"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Disclosure up front: this list came out of research I did for SoftwareForge, and the dataset above is theirs. Judge the number on its method, not on my byline.&lt;/p&gt;

&lt;p&gt;Forge starts with a repository scan called ForgeScore, an eight-dimension health assessment across security, architecture, performance and AI adaptability, run before any migration plan gets written. The output becomes a Living Specification, a versioned machine-readable document holding business objectives, functional requirements, security obligations and an explicit out-of-scope boundary. Agents then execute Work Orders generated from that spec, each carrying its own acceptance criteria and a link back to the requirement that authorized it.&lt;/p&gt;

&lt;p&gt;Forge bets that the expensive discovery work should happen once, deterministically, rather than being repeated by every agent on every task. Their own writeup on &lt;a href="https://softwareforge.ai/blog/legacy-code-modernization" rel="noopener noreferrer"&gt;modernizing without breaking business logic&lt;/a&gt; is more useful than the product pages if you want the reasoning.&lt;/p&gt;

&lt;p&gt;Where it fits: regulated portfolios, multi-quarter migrations, anywhere someone will eventually ask which requirement authorized a specific line in production. Where it does not: a two-week framework upgrade. The planning layer is overhead you will not recover on small work.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. GitHub Copilot app modernization
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd83wy9hpeuafvkdpo6l0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd83wy9hpeuafvkdpo6l0.png" alt="GitHub Copilot app modernization" width="800" height="366"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Microsoft ships modernization paths for Java and .NET as guided upgrade flows: assess the project, apply the upgrade, fix what breaks, iterate. The &lt;a href="https://learn.microsoft.com/en-us/dotnet/core/porting/github-copilot-app-modernization-overview" rel="noopener noreferrer"&gt;.NET modernization overview&lt;/a&gt; and the &lt;a href="https://learn.microsoft.com/en-us/java/upgrade/overview" rel="noopener noreferrer"&gt;Java upgrade docs&lt;/a&gt; show the actual shape, and GitHub's own &lt;a href="https://docs.github.com/en/copilot/tutorials/modernize-legacy-code" rel="noopener noreferrer"&gt;modernize legacy code tutorial&lt;/a&gt; covers the general pattern.&lt;/p&gt;

&lt;p&gt;Strongest when the target is well-known and the gap is version drift rather than architecture. Framework and runtime upgrades, dependency remediation, Azure-bound refactors. This path assumes you already know what the system is supposed to do.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Amazon Q Developer transformation
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz6qbbauhz6cmj75c92k5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz6qbbauhz6cmj75c92k5.png" alt="Amazon Q Developer transformation" width="799" height="322"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;AWS runs transformation as a defined job type rather than a chat: you point it at a Java or .NET codebase and it produces an upgrade plan and applies it, documented in the &lt;a href="https://docs.aws.amazon.com/amazonq/latest/qdeveloper-ug/code-transformation.html" rel="noopener noreferrer"&gt;code transformation guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Same trade as the Microsoft option, on the other cloud. If your estate is already on AWS and the work is Java 8 or 11 moving forward, this is the shortest path. This gets narrower than it looks the moment your migration involves a database dialect change, which is exactly where the acceptance-criteria density in that table went highest.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. IBM watsonx Code Assistant for Z
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvjtoibi4570g425bpush.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvjtoibi4570g425bpush.png" alt="IBM watsonx Code Assistant for Z" width="800" height="338"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The mainframe case. &lt;a href="https://www.ibm.com/products/watsonx-code-assistant-z" rel="noopener noreferrer"&gt;IBM's product&lt;/a&gt; works COBOL through an explicit sequence: understand the application, refactor the business services, then transform to Java, with the option to stop after any stage.&lt;/p&gt;

&lt;p&gt;Staging like that is the interesting part. Stopping after "understand" and shipping nothing is a feature, not a limitation, and it is the honest answer for estates where nobody currently knows what the system does.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Moderne and OpenRewrite
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F00mbfuacj0tkx6k8obis.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F00mbfuacj0tkx6k8obis.png" alt="Moderne and OpenRewrite" width="800" height="318"&gt;&lt;/a&gt;&lt;br&gt;
The deterministic option. &lt;a href="https://docs.openrewrite.org/" rel="noopener noreferrer"&gt;OpenRewrite&lt;/a&gt; parses source into a lossless semantic tree and applies recipes, which are code, not prompts. Same input, same output, every run. Moderne is the commercial platform that runs recipes across many repositories at once.&lt;/p&gt;

&lt;p&gt;I would put this first for anything mechanical and repeated: framework migrations, API deprecations, dependency bumps across 300 services. Recipes cannot infer intent, and Moderne does not pretend otherwise. When a recipe exists for your change, using an LLM instead is a worse trade.&lt;/p&gt;
&lt;h3&gt;
  
  
  6. Devin
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqs9olc40582m13hcp5tr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqs9olc40582m13hcp5tr.png" alt="Devin" width="800" height="379"&gt;&lt;/a&gt;&lt;br&gt;
An autonomous agent that takes a task and runs it to completion with little supervision. Independence is the whole value and the whole risk here, because a vague goal produces a confident implementation of the wrong thing.&lt;/p&gt;

&lt;p&gt;Devin gets much better when it is fed bounded work with explicit in-scope and out-of-scope boundaries. Same claim as the rest of this post, from a different direction: the agent is not the variable, the brief is.&lt;/p&gt;
&lt;h3&gt;
  
  
  7. Mechanical Orchard
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc54rtxputi7zvlmu7dc2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc54rtxputi7zvlmu7dc2.png" alt="Mechanical Orchard" width="800" height="359"&gt;&lt;/a&gt;&lt;br&gt;
Behavior-first mainframe replacement. Rather than translating source, the approach observes what the running system actually does and builds a modern implementation to match that observed behavior.&lt;/p&gt;

&lt;p&gt;Of everything here this is the slowest and most expensive, and for a 40-year-old system with no surviving documentation and no surviving authors, it is sometimes the only defensible one.&lt;/p&gt;
&lt;h2&gt;
  
  
  How I would actually sequence this
&lt;/h2&gt;

&lt;p&gt;Nothing above works if you skip the step everyone skips.&lt;/p&gt;

&lt;p&gt;Before any tool touches the code, write characterization tests. Michael Feathers named these in &lt;em&gt;Working Effectively with Legacy Code&lt;/em&gt;, and the definition matters: a characterization test does not check that code does what it should. It records what the code currently does for real inputs, so every later change is measured against observed behavior rather than assumed intent. Writing one for an undocumented function is usually the fastest way to find a business rule that no design document mentions.&lt;/p&gt;

&lt;p&gt;Then the order I would run:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Scan and score, so priority is defensible to a reviewer rather than convenient for whoever is free&lt;/li&gt;
&lt;li&gt;Characterize the twenty modules carrying the most transaction volume&lt;/li&gt;
&lt;li&gt;Take the deterministic wins first (recipes over prompts wherever a recipe exists)&lt;/li&gt;
&lt;li&gt;Scope the rest into work orders small enough that one person can verify the acceptance criteria without re-reading the diff cold&lt;/li&gt;
&lt;li&gt;Route each work order to the cheapest model that clears it&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And be careful about the dependency layer, because it fails silently. There is no single command that audits a monorepo of independently managed services:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;services/paymentservice
mvn dependency:tree &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; dep-tree-payment.txt
&lt;span class="nb"&gt;cd &lt;/span&gt;services/inventoryservice
mvn dependency:tree &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; dep-tree-inventory.txt
&lt;span class="c"&gt;# repeat per service, then cross-reference each output against a CVE database&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A single service tree runs several hundred lines. An agent adding a legitimate payment SDK can pull a vulnerable auth library four levels down and nothing in the pull request will say so.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part the vendor decks leave out
&lt;/h2&gt;

&lt;p&gt;You should expect less speedup than the demos suggest, at least at first.&lt;/p&gt;

&lt;p&gt;METR ran a randomized controlled trial with experienced open-source developers on their own repositories and found &lt;a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/" rel="noopener noreferrer"&gt;they took 19% longer with AI tools&lt;/a&gt; while believing they had been sped up. Plenty of people quote that result as a permanent verdict, which is not fair to it. METR &lt;a href="https://metr.org/blog/2026-02-24-uplift-update/" rel="noopener noreferrer"&gt;revisited the design in February 2026&lt;/a&gt; and said plainly that developers are likely faster now, but that their newer data has selection problems severe enough that they will not put a number on it.&lt;/p&gt;

&lt;p&gt;Carry both halves of that. Familiarity with the codebase reduces the gain, self-reported speedup is unreliable, and the honest current answer on magnitude is that nobody has a clean measurement.&lt;/p&gt;

&lt;p&gt;Which brings it back to the ratio. The gain is not in the generation. Gains come from how much of the system's behavior you managed to write down before generation started. You can measure it on your own backlog this afternoon, and it does not require buying anything.&lt;/p&gt;

&lt;p&gt;If your migration backlog is sitting under 0.5 acceptance criteria per story point, no tool on this list will save you.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Agentic software development: how autonomous agents are changing the SDLC</title>
      <dc:creator>Sundar Shyam Jha</dc:creator>
      <pubDate>Tue, 21 Jul 2026 10:18:30 +0000</pubDate>
      <link>https://dev.to/sundar_shyam_jha_/agentic-software-development-how-autonomous-agents-are-changing-the-sdlc-3ccj</link>
      <guid>https://dev.to/sundar_shyam_jha_/agentic-software-development-how-autonomous-agents-are-changing-the-sdlc-3ccj</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv5gl1c37ym6r2v6kpcs8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv5gl1c37ym6r2v6kpcs8.png" alt="Agentic software development" width="800" height="397"&gt;&lt;/a&gt;&lt;br&gt;
In July 2025, an AI coding agent at Replit deleted a production database during a code freeze. It wiped &lt;a href="https://dev.to/jedrzejdocs/ai-coding-agents-arent-production-ready-heres-whats-actually-breaking-4oo7"&gt;1,206 executive records and 1,196 company records&lt;/a&gt;. The agent had been told, in plain language, not to touch production. It did anyway, then generated fake data to cover the gap.&lt;/p&gt;

&lt;p&gt;That incident is the whole debate about agentic software development in one story. The capability is real. The judgment is not. And the gap between the two is where enterprise engineering leaders are now spending their attention.&lt;/p&gt;

&lt;p&gt;This piece is for people who own delivery. I want to walk through what actually changes when agents stop being autocomplete and start participating across the lifecycle, where the real failure modes sit, and what a governed version of this looks like in practice.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;From assistant to participant&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;For two years, "AI in the SDLC" mostly meant a smarter autocomplete inside the editor. That era is closing. TuringBots are becoming agentic, with autonomous agents collaborating across the full lifecycle toward end-to-end automation rather than sitting inside one tool.&lt;/p&gt;

&lt;p&gt;The adoption curve backs this up. Around 84% of developers now use or plan to use AI tools, and more than half use them daily. On the benchmark side, agent performance on SWE-bench Verified rose from &lt;a href="https://arxiv.org/pdf/2604.26275" rel="noopener noreferrer"&gt;1.96% in October 2023 to 78.4% by April 2026&lt;/a&gt;. The machines got good at closing well-scoped tickets fast.&lt;/p&gt;

&lt;p&gt;Here is the part the benchmark charts miss. Getting good at writing code was never the bottleneck in enterprise delivery. The bottleneck was coordination, ambiguity, and lost knowledge. Agents that write code faster do not fix any of those. In some cases they make them worse.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;What agents change at each stage&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The best way to describe agentic development is not "agents do the work now." It is "the sequence and ownership of work move around." A useful frame is where an agent enters the lifecycle and what it receives when it gets there.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Requirements&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Instead of a human writing tickets after the fact, an agent parses business intent into structured requirements with acceptance criteria attached. The value is not speed. It is that ambiguity gets resolved at the specification stage, where a fix costs minutes, instead of surfacing during code review, where it costs a sprint.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Architecture&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Design constraints get captured before code starts, as machine-readable blueprints rather than a diagram in a wiki that goes stale by the second sprint.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Development&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;A coding agent receives a work order that already carries the requirement ID, the relevant architecture, security constraints, and acceptance criteria. The difference between a prompt and a work order is the difference between a prototype and a production workflow.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Testing&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Agents generate test suites from the specification in parallel with the code, rather than a QA engineer reconstructing intent weeks later.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Governance&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Policy checks and audit trails run continuously, built into every artifact, instead of being reconstructed at a release gate under pressure.&lt;/p&gt;

&lt;p&gt;Notice what ties those together. It is not the model. It is whether each stage hands the next stage enough structured information to act correctly. That single property decides whether agentic development helps you or buries you.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Where it actually breaks&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;I want to be specific here, because most coverage of agentic software development is either hype or fear. The real failure modes are well documented, and they aren’t about model intelligence.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Context is the only control surface, and it degrades&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Large language models are stateless and non-deterministic. Every decision they make comes from the tokens currently in the window. Practitioners have named the point where quality falls off: the &lt;a href="https://dev.to/ametel01/advanced-context-engineering-for-coding-agents-11p7"&gt;"dumb zone,"&lt;/a&gt; which often begins around 40% of context usage, well before the window is "full." More tokens don’t mean better outcomes. Better tokens do. When an agent loses the thread of the original intent halfway through a build, you get code that compiles and misses the requirement.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Productivity is not progress&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;This is the finding senior engineers feel in their bones. Across large developer surveys, AI increases the amount of code shipped, but &lt;a href="https://dev.to/ametel01/advanced-context-engineering-for-coding-agents-11p7"&gt;code churn increases even more&lt;/a&gt;, and teams rework agent output repeatedly. Brownfield codebases suffer worst. One research team put the mechanism plainly: agents are biased toward producing more code because generation is cheap, and toward local fixes because global redesign is expensive in tokens. That combination is a &lt;a href="https://arxiv.org/pdf/2604.26275" rel="noopener noreferrer"&gt;tech-debt factory&lt;/a&gt; if you let it run unsupervised.&lt;/p&gt;

&lt;h3&gt;
  
  
  Human review becomes the bottleneck
&lt;/h3&gt;

&lt;p&gt;If an agent can produce ten plausible patches an hour, the &lt;a href="https://arxiv.org/pdf/2604.26275" rel="noopener noreferrer"&gt;rate-limiting resource is human attention&lt;/a&gt;. It gets worse. AI review comments run about &lt;a href="https://arxiv.org/pdf/2603.15911" rel="noopener noreferrer"&gt;29.6 tokens per line of code versus 4.1 for human reviewers&lt;/a&gt;, because agents explain from first principles every time instead of leaning on shared project knowledge. You have not removed the review burden. You have multiplied it.&lt;/p&gt;

&lt;p&gt;The Stack Overflow engineering team reached the same practical conclusion in &lt;a href="https://stackoverflow.blog/2026/01/28/are-bugs-and-incidents-inevitable-with-ai-coding-agents/" rel="noopener noreferrer"&gt;January 2026&lt;/a&gt;: break tasks into the smallest possible chunks, keep commits small enough to actually review, and treat "long-running autonomous agent" claims as a sales tactic. &lt;/p&gt;

&lt;p&gt;Their phrasing was clear-eyed. Review AI-assisted code the way you would review any human commit, knowing there will be more issues in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The number that should shape your strategy&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Gartner projects that more than &lt;a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027" rel="noopener noreferrer"&gt;40% of agentic AI projects will be canceled by the end of 2027&lt;/a&gt;, mostly due to escalating cost, unclear ROI, or weak risk controls. &lt;/p&gt;

&lt;p&gt;Read alongside the fact that only around a quarter of organizations have scaled an agentic system in production, the message is not "avoid this." The message is that the teams who survive the cull will be the ones treating agents as engineering infrastructure with governance attached, not as a feature they switched on.&lt;/p&gt;

&lt;p&gt;There is a matching upside number worth holding next to it. Gartner also finds roughly &lt;a href="https://www.gartner.com/en/articles/ai-in-software-engineering" rel="noopener noreferrer"&gt;25 to 30% productivity improvement&lt;/a&gt; when AI is applied across the full lifecycle, versus about 10% when it is confined to code generation. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.pwc.com/m1/en/publications/2026/docs/gen-ai-survey.pdf" rel="noopener noreferrer"&gt;PwC's research on 377 technology leaders&lt;/a&gt; found that teams using GenAI across six or more SDLC stages release nearly twice as often. The return lives in breadth and coordination, not in typing speed.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;What a governed version looks like&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;So the practical question is not whether to use agents. It is how to give them enough structure that they help without shipping debt or deleting your database.&lt;/p&gt;

&lt;p&gt;This is the design problem SoftwareForge's &lt;a href="https://softwareforge.ai/" rel="noopener noreferrer"&gt;Forge platform&lt;/a&gt; was built around, and it is worth using as a concrete example because the mechanics are specific rather than aspirational.&lt;/p&gt;

&lt;p&gt;Forge treats the specification as the system of record. Business intent becomes a &lt;a href="https://softwareforge.ai/blog/ai-sdlc" rel="noopener noreferrer"&gt;Living Specification&lt;/a&gt;: a versioned, machine-readable document that carries objectives, requirements, security obligations, and compliance constraints forward through every handoff. &lt;/p&gt;

&lt;p&gt;Agents inherit that full context instead of starting from a blank prompt, which is the direct fix for the "dumb zone" problem. &lt;/p&gt;

&lt;p&gt;Forge calls the failure mode context rot, and the persistent context layer exists specifically to stop downstream agents from drifting from decisions made earlier in the build.&lt;/p&gt;

&lt;p&gt;Execution runs through Work Orders. Each one is human-auditable and carries the intent, constraints, and acceptance criteria from the specification, so an agent receives structured context rather than an isolated instruction. No agent action reaches production without passing through a Work Order a human can authorize. That is the mechanism that would have stopped the Replit scenario: the destructive action needed an approval it never had.&lt;/p&gt;

&lt;p&gt;If you run regulated systems, this is where the argument turns financial. Forge maps every agentic action to policy and keeps an &lt;a href="https://softwareforge.ai/blog/ai-sdlc" rel="noopener noreferrer"&gt;immutable audit trail retained for seven years&lt;/a&gt;. &lt;/p&gt;

&lt;p&gt;When an auditor asks how a specific line of code traces back to a requirement, you have a decision trail instead of a Slack archaeology project. &lt;/p&gt;

&lt;p&gt;The platform reports 66% of rework eliminated, 83% faster delivery, and zero architectural drift, all traceable to the intent captured at the start.&lt;/p&gt;

&lt;p&gt;If your teams are already generating code faster than they can safely review it, the missing layer is structured context and governed execution, not a better model. You can &lt;a href="https://app.softwareforge.ai/docs/introduction" rel="noopener noreferrer"&gt;see how Forge structures that here&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Where this leaves you&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Agentic software development is real, and it is not magic. The agents are genuinely capable of closing scoped work fast. They are also stateless, non-deterministic, and blind to your production environment unless you feed them the structure to see it.&lt;/p&gt;

&lt;p&gt;The teams getting durable value are not the ones with the cleverest prompts. They are the ones who decided early that speed without governance just ships debt faster, and who built the specification, context, and audit layer to match. &lt;/p&gt;

&lt;p&gt;That is the difference between an agent that clears your backlog and one that quietly refills it.&lt;/p&gt;

&lt;p&gt;If you are moving from experimentation to real delivery this year, start by asking one question of any agentic setup you evaluate: what does the agent receive when it starts, and who approves what it does before it reaches production. If the answer is "a prompt" and "no one," you already know how that ends. If you want to see the governed version in practice, &lt;a href="https://softwareforge.ai/contact-us" rel="noopener noreferrer"&gt;book a walkthrough with SoftwareForge&lt;/a&gt; and bring your messiest legacy repo.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>AI-SDLC Explained: The AI-Driven Software Development Lifecycle</title>
      <dc:creator>Sundar Shyam Jha</dc:creator>
      <pubDate>Tue, 14 Jul 2026 11:15:58 +0000</pubDate>
      <link>https://dev.to/sundar_shyam_jha_/ai-sdlc-explained-the-ai-driven-software-development-lifecycle-1o0</link>
      <guid>https://dev.to/sundar_shyam_jha_/ai-sdlc-explained-the-ai-driven-software-development-lifecycle-1o0</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffomj2udzo2s67dpsjrgo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffomj2udzo2s67dpsjrgo.png" alt="AI-SDLC Explained" width="800" height="400"&gt;&lt;/a&gt;&lt;br&gt;
Ask ten engineering leaders what "AI-SDLC" means and you'll get ten slightly different answers, and most of them will be describing an AI assistant that autocompletes code faster. &lt;/p&gt;

&lt;p&gt;That's part of it, but it's the smallest part. The real shift is that the stages themselves are starting to reorganize around what AI can and can't be trusted to do without a human checking its work.&lt;/p&gt;

&lt;p&gt;An AI-driven SDLC, in the sense the term is actually used by teams running it in production, means AI participates across planning, coding, testing, review, and deployment, but every one of those participations is governed, meaning it's authorized, logged, and traceable back to a specific business requirement. &lt;/p&gt;

&lt;p&gt;Take away the governance and you don't have an AI-SDLC. You have a fast way to generate code that a human still has to fully re-verify before it ships, which often cancels out the time savings.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Why the old SDLC model breaks under AI-generated code volume&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The traditional software development lifecycle assumes a roughly constant, human-paced rate of code production. &lt;/p&gt;

&lt;p&gt;Requirements gathering, design, coding, testing, deployment, maintenance, each stage sized around how fast a person can reasonably write and think through code. &lt;/p&gt;

&lt;p&gt;AI breaks that assumption at the second stage. A developer using an AI coding assistant can produce in an hour what used to take a day. Everything downstream, the review, the testing, the security scan, the audit trail, was never built to absorb that volume.&lt;/p&gt;

&lt;p&gt;Katie Norton at IDC named this directly: “as AI increases the speed and scale of code generation, organizations need stronger ways to validate what is being built, what risks it introduces, and whether it meets policy requirements before code progresses.” &lt;/p&gt;

&lt;p&gt;That's the exact bottleneck teams hit six months into adopting Copilot or Cursor at scale. &lt;/p&gt;

&lt;p&gt;Code generation stopped being the constraint. &lt;/p&gt;

&lt;p&gt;Verification became the constraint. And most teams didn't rebuild their SDLC to reflect that, they just added more AI on top of a review process designed for a slower era.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;What actually changes at each stage of the lifecycle&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Here's how each phase evolves when AI moves from assistant to active contributor.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Requirements and planning
&lt;/h3&gt;

&lt;p&gt;This is the stage most AI-SDLC pitches gloss over, and it's the one that determines whether everything downstream works. &lt;/p&gt;

&lt;p&gt;If the requirement lives only in someone's head or a Jira ticket with three lines of context, an AI agent generating code against it is guessing. &lt;/p&gt;

&lt;p&gt;The pattern that actually works is capturing intent as a structured, versioned artifact before generation starts, something a PRD, architecture doc, and compliance constraints can all attach to. &lt;/p&gt;

&lt;p&gt;SoftwareForge frames this as a Living Spec, described as a versioned, machine-readable specification for every mission, so agents inherit full intent, context, and compliance instead of starting from a blank slate.&lt;/p&gt;

&lt;p&gt;Whatever tool you use, this is the step that determines whether AI output actually matches what the business needs, or just what a vague prompt produced.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Coding
&lt;/h3&gt;

&lt;p&gt;This is where most people's mental model of "AI-SDLC" stops, and it's genuinely the least interesting part now. &lt;/p&gt;

&lt;p&gt;Every major coding assistant handles this reasonably well. The differentiator isn't which model writes better code anymore. It's whether that code inherits context from the previous stage or starts fresh every session, which is why context loss is the single most common complaint from engineers who've used AI assistants on large, unfamiliar codebases for more than a few weeks.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Review and compliance
&lt;/h3&gt;

&lt;p&gt;This is the stage that determines whether an AI-SDLC is real or just marketing. &lt;/p&gt;

&lt;p&gt;Every AI-generated change needs a record of what changed, why, and whether it clears policy before it merges, not after. Katie Norton described the mature version of this as bringing architecture analysis, security scanning, and compliance checks into a coordinated workflow that supports more autonomous execution while preserving auditability and governance across the software lifecycle. &lt;/p&gt;

&lt;p&gt;SoftwareForge's version of this concept, Work Orders, ties every agentic action to a human-auditable record, which matters most in exactly the moment nobody plans for: the incident review six months later when someone asks why a specific line of production code exists.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Testing and deployment
&lt;/h3&gt;

&lt;p&gt;AI can generate tests faster than a human, but the harder problem is knowing whether the tests actually cover what the requirement demanded, not just what the code happens to do. This loops back to the requirements stage. If intent was captured properly at the start, test coverage can be checked against it. If it wasn't, you're testing that the code does what it does, which tells you nothing about whether it does what it should.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Maintenance and modernization
&lt;/h3&gt;

&lt;p&gt;This stage gets almost no attention in AI-SDLC discussions, and it's arguably where the biggest time savings actually show up. Most enterprise engineering time doesn't go into net-new features. It goes into safely changing systems built years ago by people who've since left the company. &lt;/p&gt;

&lt;p&gt;Roman Vorel, a Fortune 100 technology executive, described what's actually lost in these systems: legacy software turns into a black box, and restoring the context around how it works is what turns it back into a usable enterprise asset instead of a growing liability. &lt;/p&gt;

&lt;p&gt;An AI-SDLC that only covers new code and ignores the legacy estate is solving maybe a third of the actual problem enterprise teams have.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh35y1el03nxkcebaq2ul.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh35y1el03nxkcebaq2ul.png" alt="What actually changes at each stage of the lifecycle" width="799" height="581"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;A concrete example of how this plays out on a real team&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Take a mid-size healthcare software company adding a new patient scheduling feature while also carrying HIPAA compliance obligations across an older claims processing system. Under a traditional setup, the new feature gets built fast with an AI assistant, then sits in a security queue for two weeks because nobody can quickly confirm it doesn't touch protected health data in a way that violates policy. &lt;/p&gt;

&lt;p&gt;The claims system, meanwhile, barely gets touched because the one engineer who understood it left months ago, and every change carries unknown risk.&lt;/p&gt;

&lt;p&gt;Under a governed AI-SDLC, the scheduling feature gets built against a spec that already encodes the HIPAA constraints, so the compliance check is confirming something the system flagged during generation, not discovering it cold at review time. &lt;/p&gt;

&lt;p&gt;The claims system gets scanned and scored first, so the team knows exactly which functions touch sensitive data before an AI agent or a human ever proposes a change to them. Venkat Gopalan at Belcorp described wanting exactly this outcome from an AI-driven SDLC: “acceleration to market with measurable business outcomes, freeing people up from grunt work for higher-value work, turning weeks of effort into minutes with an enterprise-grade approach.” The point is that the whole cycle, from requirement to production, stops losing time at the handoffs.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The gap between "AI writes code fast" and "AI-SDLC delivers outcomes"&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Srikrishnan Ganesan, CEO of RocketLane, put this gap in plain terms: AI can generate code in minutes, but turning it into real enterprise outcomes still takes weeks, and closing that gap means anchoring every artifact to intent and keeping governance inline instead of added on afterward.”&lt;/p&gt;

&lt;p&gt;That gap, weeks of translation work between "code exists" and "code is a shipped, compliant, business outcome," is the actual thing an AI-SDLC is supposed to close. If your current setup only speeds up the typing part, you haven't adopted an AI-SDLC. You've adopted a faster typewriter.&lt;/p&gt;

&lt;p&gt;AI has given every developer superpowers, but code generation without enterprise-grade governance limits how much of that power can actually be used, which is why making AI-generated code secure, compliant, and auditable by design matters as much as the generation itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;What to check before you call your setup an AI-SDLC&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Three honest questions, before anyone on the team starts using the term. &lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Can you trace any AI-generated production change back to the requirement that justified it, six months later, without digging through Slack history?
&lt;/li&gt;
&lt;li&gt;Does compliance checking happen during generation, or only after code is already written and someone has to catch problems manually?
&lt;/li&gt;
&lt;li&gt;And does your legacy codebase get the same governance as new code, or is AI only touching greenfield work while the oldest, riskiest systems stay untouched because nobody trusts an AI agent near them yet?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If the honest answer to any of those is no, you're running fast code generation with a traditional SDLC wrapped around it, which is a perfectly reasonable place to be. It's just not the same thing as an AI-SDLC, and knowing the difference matters when you're deciding what to fix next.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Where this leaves teams evaluating an AI-SDLC platform&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;This category is real, but the term gets used loosely enough that it's worth pressure-testing any vendor's claim against the five stages above rather than taking a demo at face value. &lt;/p&gt;

&lt;p&gt;SoftwareForge built its Forge platform specifically around carrying intent, context, and compliance from a first prompt through production, reporting 66% of rework eliminated and delivery running 83% faster with zero architectural drift across customer deployments, figures worth validating against a pilot on your own codebase rather than accepting outright, but directionally aligned with what happens when governance stops being rebuilt from scratch at every stage handoff.&lt;/p&gt;

&lt;p&gt;If you're trying to figure out where your own team's biggest time loss actually sits, requirements gathering, review bottlenecks, or the legacy code nobody wants to touch, that's worth measuring before picking a tool. &lt;/p&gt;

&lt;p&gt;SoftwareForge offers a free scan that scores your existing codebase across eight dimensions of technical debt, which is a reasonable first step regardless of which platform you end up choosing. You can &lt;a href="https://app.softwareforge.ai/" rel="noopener noreferrer"&gt;get your tech debt score here&lt;/a&gt; and see where your own AI-SDLC gap actually is.&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
