DEV Community

Cover image for Stop Writing Code, Start Directing AI: The Future of Software Engineering
CapeStart
CapeStart

Posted on Originally published at capestart.com AI-assisted

Stop Writing Code, Start Directing AI: The Future of Software Engineering

AI Coding Agents

A few years ago, the exciting part of AI in my work was watching it finish my lines of code. That already feels dated. For much of my work in 2026, I do not write the code at all. We direct AI Coding Agents that write it, then check what comes back. The interesting part is not the speed. It is that the hard part of the job has moved.

Anthropic’s 2026 Agentic Coding Trends Report describes the same shift and puts a useful limit on it. Engineers report using AI in roughly 60 percent of their work, yet say they can fully delegate only 0 to 20 percent of tasks. That gap is the whole story: constant use, limited handover, and human judgment at every boundary. Here is where things stand.

The Shift from Autocomplete To Agent Teams

For years the headline feature was the inline suggestion. Useful, but bounded. Today that is the floor, not the ceiling. The tools worth discussing take a goal, break it into steps, edit many files, run the tests, read the errors, and try again.

The comparison below reflects my own use and the consensus among developers I work with. It is not benchmark data, so treat it as a starting point, not a ranking.

The unit of work is no longer the keystroke. It is a task, or a whole feature.

The Agent Harness: Where Reliability Comes From

It is easy to give the model all the credit, but the model is half the story. On its own, a language model can only propose the next step. It cannot open a file, run your tests, or read the error that comes back. What turns a suggestion engine into something that finishes work is the layer around it, usually called the agent harness.

The harness is the support system that lets an agent act. It supplies tools: reading and writing files, running shell commands, searching the web, querying a database. It executes what the agent writes, ideally in a sandbox, so a bad command cannot damage your machine.

Most importantly, it runs the loop that makes multi-step work possible. The agent picks a step, runs it, checks the result, and decides what comes next, until the task is done. Without that loop, you do not have an agent. You have autocomplete with extra steps.

This layer is also where reliability comes from. Permissions decide what the agent may touch. Memory and checkpoints hold earlier decisions across a long task. Observability lets you audit what happened. None of it is exciting, and all of it separates a demo from something you would point at production code.

Every tool in the previous table sits on a harness like this, whether homegrown or borrowed. When developers say one agent is more reliable than another on multi-step work, they usually mean the quality of this layer, not the model. OpenAI makes the same argument in its Agents SDK harness and sandbox release, and the SDK reference lists the primitives.

Spec-Driven Development: Why Specifications Came Back

Anyone who has handed a loose prompt to an agent knows the failure mode. You get code that compiles, looks right, and quietly misses the point. Three problems drove teams to fix this. Prompts were vague, so the agent guessed and drifted from intent. Codebases grew, and the agent forgot decisions it made many files ago. And there was rarely a clear test of whether the output was correct.

The answer is an old idea in new packaging. Write the specification first, treat it as the source of truth, and generate the code and tests from it. When a requirement changes, update the spec and regenerate instead of patching code by hand. Of everything here, this is the change I expect to outlast the current tooling.

Two practices make it work. EARS notation, the Easy Approach to Requirements Syntax developed at Rolls-Royce, turns requirements into constrained, testable patterns: WHEN a user submits a form with invalid data THE SYSTEM SHALL display validation errors next to the relevant fields. That structure removes ambiguity an agent would otherwise fill with a guess. The second is a project constitution, a stable document fixing language choices, testing standards, and dependency rules, so the same arguments are not relitigated every feature.

Try the open tools: Spec Kit, cc-sdd, or Kiro’s spec best practices. But spec-first development comes with trade-offs that vendor blogs often overlook. Specifications are artifacts, and artifacts can drift. Once a spec falls out of sync with the code, it can be worse than having no spec at all because developers and agents may keep trusting outdated requirements.

Spec-first development is also a poor fit for early exploratory prototyping. If the requirements are still taking shape, forcing them into EARS notation can create a false sense of certainty rather than useful structure.

The evidence deserves similar caution. Published results for these tools are largely vendor case studies rather than controlled trials. Before citing productivity claims, check the underlying methodology, especially widely repeated figures such as features estimated at 40 hours being delivered in under eight.

MCP: The Integration Layer For AI Coding Agents

If the harness lets an agent act, the next question is how it reaches the tools and data it acts on. That is the Model Context Protocol, an open standard for connecting AI systems to tools and data, and it is now the default answer. Instead of every vendor building a private bridge to your database or ticket tracker, MCP gives them one way in. The current revision, published on 28 July 2026, is the largest since launch: a stateless core that scales on ordinary HTTP infrastructure, a formal extensions framework, server-rendered UIs through MCP Apps, long-running work through the Tasks extension, and tightened authorization. The 2026 roadmap is worth a read before you commit to a deployment shape.

One caveat deserves more attention than it gets. Every MCP server you connect is a new attack surface. Tool descriptions and returned content both enter the model context, which makes prompt injection a supply chain problem rather than a curiosity. Scope credentials per server, keep an inventory of what your agents can reach, and treat a new connector as a new dependency.

This leads to a skill I did not expect to matter so much: context engineering. What separates a good AI Coding Agent from a frustrating one is rarely model size. It is whether the tool indexes your repository, tracks dependencies, and reasons across the project rather than one file at a time. Supplying the right context and withholding the noise is the work. For more insights on context management and building reliable systems, check out From Prompt to Production: Building Enterprise-Grade AI Systems Without Fine-Tuning on the CapeStart AI & Technology Blog.

Code Review and Testing Move Into the Agent Loop

Once agents write a large share of the code, the bottleneck moves to checking it. AI review tools such as Greptile and CodeRabbit read a pull request against the whole codebase and surface cross-file issues a human skimming a diff would miss. Neither replaces a reviewer who understands the product, and both add noise on large diffs, so budget time to tune them. Test generation is moving the same way, and works best where tests derive from the spec.

Cost belongs here too. Developers now discuss token efficiency almost as much as capability, and prefer an agent that is right first time over one that retries its way there.

What the Productivity Evidence Actually Shows

Adoption is not in doubt. In Stack Overflow’s 2025 Developer Survey of more than 49,000 developers, 84 percent used or planned to use AI tools. Trust is a different matter: 33 percent said they trust the accuracy of AI output while 46 percent actively distrust it, and 45 percent named debugging AI-generated code as a specific frustration. Those figures are from 2025, the most recent broad baseline rather than today’s reading.

The number that should give everyone pause comes from METR’s randomized controlled trial, run between February and June 2025. Sixteen experienced open-source developers completed 246 real tasks in repositories they knew well, mostly using Cursor Pro with Claude 3.5 or 3.7 Sonnet. They predicted AI would make them 24 percent faster. Afterwards, they believed it had made them 20 percent faster. Measured, they were 19 percent slower.

Read that carefully. Sixteen developers is a small sample; the setting was mature codebases they already understood, and the tools have moved on. METR’s own 2026 follow-up ran into selection effects that made its central estimate unreliable. The durable finding is not the 19 percent. It is that perceived speedup is not evidence of speedup.

One more note on numbers, since you will meet this one everywhere. The claim that 41 percent of code is AI-generated traces back to a 2024 analysis and gets recycled without its date. Other measurements land closer to 27 to 30 percent. If you cannot source a figure to a primary study with a stated method, leave it out.

What this Means for Engineering Teams

The tools got good enough to hand work to. That is the whole change, and it matters more than any productivity slide, because it moves the hard part of the job from typing to judgment.

None of the evidence says put the tools down. It says the gains are real, uneven, and dependent on how you work. Agents pay off on greenfield code, clear tasks, and well-specified changes. They punish you on mature systems you already know well, which is where the METR trial landed. Know which one you are in before you delegate.

Three moves for this week:

  1. Delegate one whole task. Not a function, a task. Give your agent the goal, then read the diff line by line. You are calibrating your own trust, not testing the tool.
  2. Write one real specification. Use EARS phrasing for the acceptance criteria, commit it beside the code, and update it when requirements move. If you do only one thing on this list, do this one.
  3. Audit what your agents can reach. List every MCP server, the credentials it holds, and who approved it. A connector is a dependency.

And one thing to stop: quoting how much faster AI makes you. You do not know. Neither did the developers who felt 20 percent faster while measuring 19 percent slower. Track delivery, not feeling.

The engineers pulling ahead this year are not the ones typing the most. They are the ones who learned to direct the machines that type, and who still read the diff.

Author’s Note: This article was supported by AI-based research and writing, with Claude 5 assisting in the creation of text and images.

Top comments (0)