Choosing an AI coding agent by one overall accuracy score can hide the work that matters most to your team. An agent may be effective for documentation while creating more review effort for feature development or bug fixes.
A safer evaluation separates tasks by type, measures acceptance over time, and connects generated changes to requirements, tests, reviews, and project context.
What the study measured
The paper Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance analyzed 7,156 pull requests from the AIDev dataset. It compared five coding agents across nine task categories and tracked acceptance rates over 32 weeks.
The primary metric was pull request acceptance. Acceptance is useful for determining whether submitted changes were accepted, but it is not a complete measure of code quality, security, maintainability, or developer productivity.
Why task type matters
The study found that the type of work often had a larger practical effect than the identity of the agent. Documentation tasks reached an 82.1% acceptance rate, compared with 66.1% for new-feature tasks—a difference of approximately 16 percentage points.
An overall average can therefore conceal where an agent is useful and where it creates additional review or rework.
Documentation has narrower review criteria
Documentation changes usually have a smaller scope and more visible acceptance criteria than new functionality. In the study, Claude Code led the documentation category with a 92.3% acceptance rate.
That result should not be generalized to architectural changes, integrations, or feature development. Keep these use cases separate when designing an internal benchmark.
Feature work needs stronger controls
New features had the lowest acceptance rate among the examples highlighted in the paper, at 66.1%. Feature work requires interpreting requirements, choosing dependencies, writing tests, and coordinating across contributors. A plausible patch can still miss the intended product behavior.
For this category, agent selection is only one control. Teams also need explicit requirements, linked work items, review ownership, test evidence, and a defined process for revising rejected changes.
There is no universal winner
OpenAI Codex produced consistently strong results across all nine task categories, with acceptance rates ranging from 59.6% to 88.6%. The paper also reports statistically significant advantages over other agents in several categories.
Category leaders still differed. Claude Code performed best for documentation and feature tasks in the reported results, while Cursor led the fix category with an 80.4% acceptance rate.
The study also found that Devin was the only agent with a consistent positive acceptance trend over the 32-week period, increasing by 0.77% per week. The other agents were described as largely stable during the observed period.
These are historical results from one dataset and time window. Treat them as evaluation guidance, not as a permanent ranking. Your repositories, prompts, review standards, and agent versions may produce different results.
Build a task-stratified benchmark
1. Define representative task categories
Create separate evaluation sets for the work your team actually performs. Possible categories include:
- Documentation updates
- Bug fixes
- New features
- Refactoring
- Test additions and maintenance
Record acceptance rates by category instead of combining every pull request into one score.
2. Define acceptance before execution
Give each task a clear objective, boundaries, relevant files or services, expected behavior, and validation requirements. Keep the requirement connected to the implementation and test evidence so reviewers can assess the change against its original intent.
For feature work, specify behavior and constraints before asking an agent to modify code. For fixes, include the failing behavior, reproduction conditions, and the expected regression test where applicable.
3. Measure more than acceptance
A useful benchmark should also capture:
- Review time
- Number of revisions
- Human editing required before acceptance
- Test failures
- Rejected pull requests
- Defects discovered after acceptance
Acceptance is an important outcome, but it does not describe the total delivery effort or the quality of the resulting software.
4. Repeat the evaluation
Agent performance can change as models, prompts, repository conventions, and review practices evolve. Repeat the benchmark periodically and record the agent version, task category, repository context, and acceptance criteria for every evaluation.
Connect generated code to delivery context
A coding agent can generate a patch, but the surrounding team still needs to determine why the change is required, which requirements apply, who owns the review, whether testing is complete, and what happens after acceptance or rejection.
When this information is scattered across chat threads, documents, issue trackers, and repositories, reviewers must reconstruct the context manually. A shared project and knowledge workspace can provide a more reliable operating record for agent-assisted work.
Useful capabilities include task types and statuses, custom fields, linked requirements, review ownership, test evidence, planning views, and reporting. These controls help distinguish documentation, fixes, and feature work and route each category through an appropriate review process.
The platform does not replace task-level evaluation. It provides the traceability needed to understand how an agent-produced change moved from request to implementation, review, acceptance, or rejection.
A practical decision framework
- Classify the work. Identify the task categories in the intended use case and avoid relying on one blended score.
- Set quality gates. Define required tests, reviewers, security checks, and documentation before expanding agent use.
- Preserve traceability. Link requests, requirements, implementation work, reviews, and outcomes.
- Compare total effort. Include review time, revisions, rejected pull requests, and post-release issues alongside acceptance rates.
- Reassess regularly. Repeat measurements as agents, repositories, and internal practices change.
Key takeaway
The study shows why claims about the “best” AI coding agent need careful qualification. Task category had a larger practical effect than typical differences between agents in many cases, and different agents led in documentation, feature, and fix work.
Evaluate agents against the tasks you intend to automate. Measure acceptance together with review effort and defects, then connect each generated change to its requirements, tests, reviewers, and outcome.
Top comments (0)