---
title: "Coding agents stall below 45% on SWE-Bench Pro real tasks"
slug: "coding-agents-stall-below-45-on-swe-bench-pro-real-tasks"
published: "2026-07-31"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-07-31"
aiActArticle50: "compliant"
humanView: "https://agentry.news/research/coding-agents-stall-below-45-on-swe-bench-pro-real-tasks"
agentView: "https://agentry.news/agent/coding-agents-stall-below-45-on-swe-bench-pro-real-tasks"
---# Coding agents stall below 45% on SWE-Bench Pro real tasks

> A new evaluation of leading AI coding agents on SWE-Bench Pro—a benchmark designed around longer-horizon, realistic software engineering tasks—found frontier models solving fewer than 45% of public pr

*Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. [AI policy](/ai-policy).*

# Coding Agents Stall Below 45% on SWE-Bench Pro Real Tasks

Leading AI coding agents are solving fewer than 45% of public problems on SWE-Bench Pro, a newly released evaluation that measures performance on longer-horizon, more realistic software engineering tasks sourced from both public and private repositories. The performance drop signals a steeper cliff between controlled lab benchmarks and production-grade code challenges than vendors have publicly acknowledged.

## Benchmark Design Reveals Real-World Gap

SWE-Bench Pro is designed to test agents on software engineering tasks that more closely mirror actual development work [Scale AI](https://labs.scale.com/papers/swe-bench-pro). The benchmark includes both a public split and proprietary production-grade tasks drawn from real codebases. Earlier benchmark versions showed frontier models improving substantially over time, but SWE-Bench Pro introduces harder, more realistic constraints that expose performance limits at scale.

The evaluation, surfaced through Agentry research, documents that widely used coding models perform significantly below the 45% threshold on public tasks and substantially lower on proprietary production-grade problems [Agentry](https://agentry.news/research/swe-bench-pro-coding-agents-stall-below-45-on-real-problems). This represents a material divergence from how these agents perform on earlier, simpler benchmarks where frontier models have historically scored higher.

## What the Numbers Show

Prior research on coding agent evaluation found that frontier models reached 80.3% on a smaller, earlier public benchmark split, but SWE-Bench Pro's broader and more realistic task set demonstrates that performance does not generalize [Scale AI](https://labs.scale.com/papers/swe-bench-pro). The gap between controlled benchmarks and production reality has become a central question for enterprises evaluating agent deployment.

The distinction matters for real-world adoption. Public-facing test suites often reflect cleaner, more self-contained problems. Proprietary production-grade tasks involve legacy code, incomplete documentation, ambiguous requirements, and integration constraints that mirror what deployed agents actually encounter. The sub-45% public performance and even lower proprietary scores suggest current agent capabilities remain constrained for autonomous software engineering at scale.

## Implications for Agent Adoption

The results come as enterprises increasingly integrate coding agents into development workflows, making benchmark accuracy critical for ROI estimation. Teams evaluating tools for production deployment now have a clearer signal: real-world task performance lags significantly behind what marketing materials and simpler benchmarks suggest.

SWE-Bench Pro's design—pulling tasks from real repositories rather than synthetic datasets—also makes the results harder to dismiss as artificial constraints. This is the benchmark agents will increasingly be measured against as the coding-agent market matures.

## Factbox

• **Public task performance**: Leading coding agents solve fewer than 45% of SWE-Bench Pro public problems

• **Proprietary task performance**: Performance drops further on production-grade proprietary code

• **Benchmark scope**: Includes tasks from both public and private repositories

• **Prior benchmark gap**: Frontier models reached 80.3% on earlier, smaller public splits

• **Real-world implication**: Agents show material performance drop when task complexity and realism increase

• **Market relevance**: Enterprise deployment decisions now hinge on SWE-Bench Pro results rather than simpler evaluations