AGENTRY.NEWSWhat AI Agents Do, Documented.August 10, 2026

Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.

OpenAI finds 30% of SWE-Bench Pro tasks broken

By
Agentry Newsroom
Published

OpenAI said it found roughly 30% of SWE-Bench Pro tasks are broken, according to an audit published July 8, 2026, in a post titled "Separating signal from noise in coding evaluations". The company stated that the benchmark "no longer reliably measures frontier coding capability" and announced it is retracting its previous recommendation that researchers use SWE-Bench Pro as a leading evaluation standard.

What OpenAI found

OpenAI's internal investigation identified "widespread task issues" across the benchmark. The company's official statement said: "We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval."

The discovery challenges the validity of coding-agent benchmark scores across the industry. SWE-Bench Pro has been widely used to measure the performance of AI systems on real-world software engineering tasks. If substantial portions of the test set are flawed—containing impossible tasks, incorrect solutions, or environmental issues—then scores published using the benchmark may not accurately reflect actual agent capability.

Implications for coding evaluation

The retraction signals a broader problem in how the AI research community evaluates code-generation agents. Benchmarks are foundational to measuring progress; if one of the leading metrics is compromised, published claims about agent performance become difficult to compare and interpret. Teams that have reported results on SWE-Bench Pro may now face questions about whether their scores are meaningful or artifacts of broken test cases.

OpenAI's decision to audit and publicly report the flaws—rather than quietly discontinue the benchmark—sets a precedent for transparency in evaluation. However, the gap created by SWE-Bench Pro's loss of credibility leaves the research community without a clear, well-validated alternative for measuring frontier coding ability.

No other parties, regulatory bodies, or legal action are mentioned in OpenAI's statement. The audit appears to be an internal review with voluntary disclosure to the research community.

What comes next

The research community and benchmark developers will likely need to invest in new evaluation frameworks that address the flaws identified in SWE-Bench Pro. Until alternatives are established and validated, comparisons of coding-agent capabilities will remain uncertain—a critical gap as enterprise adoption of AI agents accelerates.

Del dette opslag: