title: "Study: AI Coding Agents Show High Run Variance, Need Repeated Testing" slug: "study-ai-coding-agents-show-high-run-variance-need-repeated-testing" published: "2026-10-10" beat: "Research" tags: ["Research"] creator: "Agentry Newsroom" editor: "Susanne Sperling, Editor — Human in the Loop" tools: ["Claude (Anthropic)", "Perplexity Sonar"] creativeWorkStatus: "verified" dateReviewed: "2026-10-10" aiActArticle50: "compliant" humanView: "https://agentry.news/research/study-ai-coding-agents-show-high-run-variance-need-repeated-testing" agentView: "https://agentry.news/agent/study-ai-coding-agents-show-high-run-variance-need-repeated-testing"
A September 2026 benchmark of 584 runs across six coding agents and six open-weight model endpoints found that identical agent-model pairings produce significantly different results, challenging singl
Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.
Researchers comparing six coding agents across six open-weight model endpoints discovered that repeated runs of the same agent-model pairing produced more variation than differences between entirely different pairings, according to a benchmark study posted to arXiv in late September 2026 arXiv.
The study, titled "Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models," executed 584 runs total, with six agent–model pairings repeated 52 times each under fixed settings, plus three additional pairings tested on a larger model from the same family arXiv. The controlled experimental design held all variables constant except the inherent stochasticity in agent and model behavior—a deliberate choice to measure true performance instability rather than differences in setup or prompt engineering.
The finding directly challenges the industry practice of running single-trial evaluations. When a pairing was run multiple times, performance metrics fluctuated enough to reverse rankings; an agent that appeared best in one run might rank lower in another. This outcome suggests that published benchmarks comparing agents or models based on a handful of trials—or even single runs—may be capturing noise rather than genuine capability differences.
The research team concluded that agents and models should be evaluated as pairings over repeated attempts, rather than as independent components arXiv. A single trial, their data implies, is insufficient for reliable ranking or deployment decisions.
Beyond variance measurement, the study argues for a shift in what gets reported. Current benchmarks typically focus on success rates or code quality metrics. The researchers recommend that evaluations report compliance alongside quality—meaning disclosure of which runs succeeded, failed, or produced unexpected behavior patterns, and how often those outcomes recurred across repetitions.
This dual-reporting approach would give teams relying on coding agents more granular information: not just "Agent X solved 70% of tasks," but "Agent X solved 70% across 52 runs, with failures concentrated in specific error categories and compliance violations in 3% of attempts." Such detail supports more informed procurement and deployment decisions, especially for enterprise environments where unexpected agent failures carry real costs.
The work surfaces a methodological bottleneck in agent evaluation. As coding agents move into production—handling real code review, refactoring, and generation tasks—the reliability of the benchmarks used to select them matters. Variance within a single pairing can be as large as variance between different pairings, meaning benchmark-based comparisons risk ranking agents by luck rather than capability.
The study's findings align with broader agent-economy concerns: as stakes rise and adoption accelerates, the tools used to measure agent quality must become more rigorous. The arXiv posting and subsequent updates through September 29, 2026, reflect active research iteration on this problem.