DEV Community

qnbs
qnbs

Posted on Fully Autonomous

Measuring Agentic Engineering: Count Review, Rework, and Value

“The agent finished in five minutes” is a timing observation, not a productivity result. Someone may have spent an hour shaping the request, reviewing a large patch, repairing a regression, or waiting while doing other work. Conversely, a task that took the same time may now be reliable enough to attempt when it was previously uneconomic.

Measuring agentic engineering well means measuring more than output volume or elapsed time. The useful question is whether a workflow produces more acceptable engineering value for the total human and system effort, without silently lowering quality or shifting cost to another stage.

Do not compress unlike outcomes into one score

Use a small dashboard rather than one magic multiplier:

Dimension What to capture Why it matters
Outcome Accepted tasks or capabilities, with a fixed definition of acceptance More code or more closed tickets can still mean less useful work
Elapsed time Start-to-accepted wall-clock time Shows user-visible lead time, but can hide parallel waiting
Human attention Active minutes for task framing, supervision, review, and recovery Agent runtime is not the same as human effort
Rework Returned patches, retries that changed nothing, reopened tasks, and reversions A fast first pass may create a slower overall path
Quality The same tests, defect checks, maintainability review, and security requirements A speed result that relaxes the quality bar is not comparable
System cost Model/tool usage or infrastructure cost when available Helps explain whether the workflow’s benefit is economical
New work enabled Previously deferred tasks that were completed, tracked separately Captures expanded capability without pretending there was an identical baseline

These measures answer different questions. Do not add minutes, dollars, defect counts, and newly enabled work into a single total unless a defensible value model exists and stakeholders agree with its assumptions.

Count the human work around the agent

Record at least four kinds of time:

  1. Framing: understanding the task, defining scope, preparing context.
  2. Supervision: answering questions, correcting direction, and resolving blocked states.
  3. Review: examining the patch, evidence, and remaining risk.
  4. Recovery: rework, repair, rollback, or incident response attributable to the attempt.

Also record agent runtime and idle waiting, but keep them separate. If a developer launches two independent agents and works on another task while they run, adding both agent runtimes to wall-clock time would overstate elapsed time. Counting only the five-minute wait would understate the person’s total effort if they also spent twenty minutes reviewing both results.

Agree on the accounting rule before gathering data. A useful policy might be: wall-clock time is measured from task start to accepted completion; active human minutes are the sum of time spent on this task even if it occurred in short intervals; agent runtime is recorded separately; overlapping agent runtime is never added to human time.

Compare like with like

AI-assisted and non-AI tasks should be comparable in scope, task class, repository familiarity, and acceptance bar. If people select only easy tasks for the agent and keep difficult work for themselves, the observed times do not estimate a general effect. If the workflow changes halfway through the measurement window, the before-and-after comparison also includes that change.

A practical internal pilot can:

  1. Choose a recurring task family, such as test updates or small dependency migrations.
  2. Record difficulty, repository familiarity, and expected risk before work starts.
  3. Use matched tasks or a randomized assignment where practical and acceptable.
  4. Freeze the same definition of “accepted” for both paths.
  5. Record every eligible task, including abandoned, blocked, or manually completed ones.
  6. Review the results by task class and experience level before calculating an overall figure.

A small pilot can identify workflow problems. It usually cannot establish a universal productivity effect. Report the number and kind of tasks, the spread of results, changes in the workflow, and what was excluded.

Design for the measurement problems already visible

METR’s randomized 2025 study of experienced open-source developers and familiar repositories reported a slowdown for the early-2025 tools it tested. Its 2026 update said the follow-on measurements were difficult to interpret: some developers avoided tasks assigned to the no-AI condition, and some found it hard to report time when several agents ran concurrently and they worked on unrelated tasks while waiting.

The lesson is methodological, not a forecast that agents slow developers down. The same workflow can change who accepts a task, which tasks are attempted, and how time is spent. A credible measurement plan should record those shifts instead of treating the task list and clock as neutral.

For example, report task substitution separately:

  • Existing-task effect: did the workflow change effort or quality on work the team already did?
  • New-task effect: did it make a new kind of work feasible?
  • Value effect: was the new work worth doing to the people who use it?

An agent can increase capability without shortening every task. It can also shorten a task while increasing review burden or defect risk. Keep these outcomes visible instead of forcing them into a single percentage.

Use a quality gate, not a speed target

Set the acceptance bar before the experiment. Specify what tests, review conditions, security controls, and operational evidence each task class requires. Decide how failed checks, reopened changes, or escapes affect the outcome.

For a pilot, ask a reviewer to classify defects without being told whether AI was used where feasible. Blinding may be impractical, but a consistent rubric still helps. Include time for review in the measurement; otherwise the team may appear faster simply because the review has not happened yet.

Do not reward an agent for closing more tasks if it broadens scope, weakens tests, or leaves work for someone else. A useful headline can be as simple as: “Across this task family, median accepted completion time changed by X; review effort changed by Y; defect/rework observations were Z; the sample is small and the task distribution changed in these ways.” Only calculate X, Y, or Z from actual observations.

A one-page measurement card

Before a pilot, agree on:

  • Question: What change are we trying to estimate?
  • Task set: Which tasks count, and how will their difficulty be recorded?
  • Comparison: How are assisted and comparison tasks selected?
  • Acceptance: Which checks and review conditions must both paths pass?
  • Time rule: How are framing, supervision, review, waiting, and overlap counted?
  • Quality signals: Which defects, reopens, or reversions will be tracked?
  • Analysis: Which groups or task types will be reported separately?
  • Stop rule: What quality, safety, or cost result would pause the pilot?

The stop rule matters. Averages should not hide a workflow that is faster on routine changes but unsafe on migration or security work.

Optimize for useful capability per unit of attention

The most valuable result may be fewer repetitive steps, faster feedback, less context switching, or the ability to attempt a previously deferred improvement. Those benefits matter, but they should be described accurately: a measured time reduction, a measured quality change, or a new capability observed in a specific workflow.

Agentic engineering is not measured well by how many tokens it generates, how many tools it calls, or how confidently it reports success. Measure outcomes people accept, all the effort required to reach them, and the quality of the resulting system. Then preserve uncertainty in the report.

References

Source, license, and AI assistance

This article develops the conditional-productivity and measurement questions in Part II of From Vibe Coding to Agentic Software Engineering. The source record credits ChatGPT as preparer and identifies CC BY-NC-SA 4.0. This version is substantially rewritten and expanded with a measurement dashboard, pilot design, and updated research references, and is shared under the same license: CC BY-NC-SA 4.0.

AI disclosure: The article text was generated primarily by AI. A human publisher supplied the topic, source material, and editorial direction, and remains responsible for checking claims and examples before publication.

Top comments (0)