Real-SWE benchmark: top coding agents resolve only 38.8% of enterprise
Specific Labs released Real-SWE on September 26, 2026, through Y Combinator's Launch YC platform, introducing a benchmark that measures coding agent performance against authentic enterprise challenges Y Combinator. The benchmark is constructed from licensed private production codebases and evaluates coding agents on real software engineering tasks drawn from private, out-of-distribution code repositories Runtime Wire.
Frontier agents stumble on production code
The results expose a critical limitation in current AI coding systems: Fable 5.1 on Claude Code, the leading system tested, resolved only 38.8% of Real-SWE tasks. This performance gap signals that despite rapid improvements in coding model capabilities, frontier agents remain unable to handle most real-world enterprise software engineering problems mer.vin.
Other major coding models and agents trailed further behind the top performer. The benchmark's design—using private, licensed production codebases rather than open-source or synthetic code—creates an authenticity gap that public benchmarks cannot bridge. Agents trained and evaluated on GitHub repositories and synthetic datasets often fail when confronted with proprietary architecture, domain-specific patterns, and the complexity of large-scale enterprise systems BERI.
Why private codebases matter
Real-SWE's reliance on licensed enterprise code represents a methodological shift in agent evaluation. Public benchmarks like SWE-bench have driven progress in coding agent research, but they test against tasks drawn from open repositories where training data leakage and synthetic similarity are persistent risks. Private production codebases introduce out-of-distribution challenges that more accurately reflect what agents encounter in actual enterprise deployment—legacy patterns, undocumented conventions, domain-specific libraries, and technical debt accumulated over years News.lavx.hu.
The 38.8% resolution rate for the top system, while representing measurable progress, underscores that most enterprise coding tasks remain beyond the reliable capability of current agents. This finding carries implications for companies considering large-scale deployment of autonomous coding systems in production environments. Real-world software engineering demands not only code generation but comprehension of vast, complex systems—a capability that remains substantially limited even in leading models.
Market implications
The Real-SWE benchmark provides a concrete, verifiable evaluation metric for the coding agent economy at a moment when venture capital and enterprise buyers are evaluating which systems merit investment and adoption. By establishing a measurement based on authentic enterprise workloads rather than synthetic tasks, Specific Labs has created a standard that may shape how the industry assesses progress in autonomous software engineering Born City.