title: "SWE-Bench Pro: coding agents stall below 45% on real problems" slug: "swe-bench-pro-coding-agents-stall-below-45-on-real-problems" published: "2026-07-22" beat: "Research" tags: ["Research"] creator: "Agentry Newsroom" editor: "Susanne Sperling, Editor — Human in the Loop" tools: ["Claude (Anthropic)", "Perplexity Sonar"] creativeWorkStatus: "verified" dateReviewed: "2026-07-22" aiActArticle50: "compliant" humanView: "https://agentry.news/swe-bench-pro-coding-agents-stall-below-45-on-real-problems" agentView: "https://agentry.news/agent/swe-bench-pro-coding-agents-stall-below-45-on-real-problems"
A new benchmark released via OpenReview shows that leading AI coding agents solve fewer than 45% of public software engineering problems and fewer than 20% of proprietary ones—a dramatic drop from the
Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.
A new benchmark measuring the real-world performance of AI coding agents has exposed a significant gap between lab results and production readiness. The SWE-Bench Pro benchmark shows that leading AI coding agents solve fewer than 45% of public problems and fewer than 20% of proprietary ones—a sharp contrast to their performance on earlier versions of the same benchmark.
When tested on SWE-Bench Pro, the industry's best-performing coding agents achieved pass rates below 45% on publicly available problems and dropped to below 20% on proprietary, production-grade tasks. By comparison, those same models reported solving over 70% of problems on SWE-Bench Verified, the older benchmark used to measure agent progress through 2025.
The disparity underscores a critical limitation: agents that appear capable on simplified or older benchmarks struggle significantly when facing the complexity, scale, and real-world constraints of long-horizon software engineering work. Proprietary problems—those drawn from actual codebases and requiring deeper context and integration—present particular challenges.
Benchmarks shape agent development roadmaps and investor narratives. A model posting 75% accuracy on a simplified test generates funding and enterprise interest; the same model posting 20% on production-grade tasks paints a different picture. SWE-Bench Pro addresses a documented weakness in prior evaluation: earlier benchmarks did not adequately represent the difficulty and scope of real engineering tasks that agents would encounter in deployment.
The findings suggest that current coding agents excel at isolated, well-scoped problems but struggle with the ambiguity, multi-step reasoning, and context-switching required for genuine software development. Teams building agent-powered code generation tools will need to account for this performance ceiling when planning enterprise rollouts.
This benchmark release arrives as enterprises evaluate whether to integrate coding agents into development workflows. The data suggests that full automation of software engineering remains distant; agents are better positioned as assistants that increase developer productivity rather than as autonomous builders. Companies marketing agents as production replacements for engineers will face friction with these results.
The gap also signals opportunity for research teams and tool builders focused on improving agent reasoning, code search, and long-horizon planning. Closing a 25-percentage-point gap between public and proprietary problem-solving could unlock material productivity gains across the industry.