OpenAI audit finds 30% of SWE-Bench Pro tasks broken
OpenAI published an audit on July 8, 2026 that identified a significant reliability crisis in SWE-Bench Pro, one of the field's most widely cited AI coding evaluation benchmarks. The company found that roughly 30% of the benchmark's 731 public tasks were broken, and announced it was retracting its previous recommendation that the research community use SWE-Bench Pro as a leading coding evaluation Agentry News.
"We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval," OpenAI stated in the audit Agentry News.
What the Audit Reveals
The discovery exposes a fundamental flaw in how coding agents are being measured at scale. SWE-Bench Pro, which contains 731 public-split tasks, has served as a reference point for evaluating AI agents' ability to solve real-world software engineering problems. When nearly one-third of those tasks are non-functional or misconfigured, benchmark scores lose their predictive power—agents and researchers cannot reliably infer true coding capability from performance on a corrupted evaluation set.
OpenAI's retraction is a rare public acknowledgment from a major lab that a widely-adopted benchmark no longer meets standards for rigorous evaluation. The move signals that the agent-coding community must reassess which benchmarks are trustworthy, and may shift focus toward proprietary or more carefully curated evaluation sets Codex.
Implications for Agent Developers
For teams building coding agents, the audit creates immediate uncertainty about leaderboard rankings and published results. Papers and product claims built on SWE-Bench Pro scores must now be re-examined. The broken-task rate suggests that agents may have been overfitting to malformed problems or benefiting from inflated scores that do not reflect real-world coding ability.
The retraction also underscores a broader challenge in the AI agent economy: benchmarks degrade over time as test data becomes misaligned, deprecated, or simply incorrect. As agents improve, maintaining evaluation integrity requires active auditing and curation—work that OpenAI's action now places squarely on the research community's agenda.
OpenAI did not announce a timeline for a corrected version of SWE-Bench Pro or name an alternative benchmark it recommends as a replacement. That gap leaves agent developers and evaluators in a near-term credibility vacuum for coding-task evaluation.