title: "Coding agents show sharp performance gaps by task type" slug: "coding-agents-show-sharp-performance-gaps-by-task-type" published: "2026-09-30" beat: "Research" tags: ["Research"] creator: "Agentry Newsroom" editor: "Susanne Sperling, Editor — Human in the Loop" tools: ["Claude (Anthropic)", "Perplexity Sonar"] creativeWorkStatus: "verified" dateReviewed: "2026-09-30" aiActArticle50: "compliant" humanView: "https://agentry.news/research/coding-agents-show-sharp-performance-gaps-by-task-type" agentView: "https://agentry.news/agent/coding-agents-show-sharp-performance-gaps-by-task-type"
A September 25 study analyzing 7,156 pull requests across five AI coding agents found that acceptance rates vary significantly by task category, with documentation tasks reaching 82.1% acceptance whil
Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.
A peer-reviewed analysis of 7,156 pull requests published September 25, 2026 reveals that AI coding agents deliver sharply different acceptance rates depending on the type of work they perform, with documentation tasks significantly outpacing feature development Ones.
The study compared five coding agents—OpenAI Codex, GitHub Copilot, Devin, Cursor, and Claude Code—across a 32-week period using data from the AIDev dataset arXiv. Researchers stratified their evaluation by task type rather than treating all agent outputs as equivalent, a methodological shift that exposes how agents handle different categories of work.
Documentation tasks emerged as the strongest performance category, with an 82.1% pull request acceptance rate across all agents studied Ones. This 16-point spread between documentation and new features suggests that agents excel at structured, well-precedented work where patterns are established and variation is constrained. New feature development, requiring novel architecture decisions and integration planning, achieved only 66.1% acceptance Ones.
The gap reflects a fundamental challenge in agent design: agents perform better when the task space is well-defined and historical examples abundant. Documentation follows templates and conventions; new features demand contextual judgment about system design that agents struggle to replicate at human-level quality.
Across the 32-week observation window, performance was not static. Devin demonstrated the only consistent positive weekly trend, improving at a rate of +0.77% per week in pull request acceptance Ones. This sustained improvement suggests that either the agent's underlying model was updated iteratively, or that its training on real-world pull request feedback created a genuine refinement curve—a data point relevant to teams evaluating agent reliability for production workflows.
The other four agents studied showed either flat or declining acceptance rates over the same period, indicating that Devin's improvement was not a market-wide phenomenon but specific to that agent's development trajectory.
These findings have immediate practical relevance for enterprises considering agent adoption. Task stratification reveals that agents are not interchangeable tools but specialized performers whose output quality depends entirely on problem category. Organizations deploying coding agents should expect them to deliver reliable documentation automation while maintaining skeptical review processes for feature work, where rejection rates remain near one-third.
The research validates a growing operational pattern: agents as augmentation tools for well-scoped work, not replacements for architectural or creative decision-making.