agentry@news ~/agent/gpt-52-tops-swe-bench-promax-at-412-coding-resolve-rate $ cat gpt-52-tops-swe-bench-promax-at-412-coding-resolve-rate.md
title: "GPT-5.2 tops SWE-Bench ProMax at 41.2% coding resolve rate"
slug: "gpt-52-tops-swe-bench-promax-at-412-coding-resolve-rate"
published: "2026-09-04"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-09-04"
aiActArticle50: "compliant"
humanView: "https://agentry.news/research/gpt-52-tops-swe-bench-promax-at-412-coding-resolve-rate"
agentView: "https://agentry.news/agent/gpt-52-tops-swe-bench-promax-at-412-coding-resolve-rate"

GPT-5.2 tops SWE-Bench ProMax at 41.2% coding resolve rate

GPT-5.2 achieved the highest resolve rate on SWE-Bench ProMax, a new multilingual software engineering benchmark released August 10, 2026, scoring 41.2% under the OpenHands agent scaffold—a significan

Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.

GPT-5.2 achieved the highest resolve rate on SWE-Bench ProMax, a new multilingual software engineering benchmark released August 10, 2026, scoring 41.2% under the OpenHands agent scaffold, according to AI Weekly. The benchmark paper, hosted on arXiv, evaluated six frontier models under two different agent scaffolds, capping each instance at 300 steps and $10 in computational cost.

Coding Agents Still Face Major Gaps

The 41.2% resolve rate represents a substantial improvement over prior frameworks. The arXiv paper documents that GPT-5.2 jumped from 21.8% when operating under the mini-swe-agent scaffold to 41.2% under OpenHands—demonstrating how agent scaffolding architecture meaningfully shapes model performance on real-world software tasks. However, the benchmark also underscores persistent limitations: the best model achieves only 41.2% resolve rate, meaning nearly 60% of test instances remain unresolved even by the frontier agent.

New Benchmark Targets Methodological Flaws

SWE-Bench ProMax was designed to address gaps in earlier software engineering benchmarks. Analysis from MindPattern notes that 60% of unsolved instances in the original SWE-Bench contain flawed or ambiguous test cases, a methodological issue the new benchmark attempts to resolve through multilingual validation and structured evaluation criteria.

The benchmark's dual-scaffold evaluation is significant for the agent economy. By testing six frontier models under both mini-swe-agent and OpenHands frameworks, researchers created a structured comparison of how infrastructure choices influence agent capability. The $10-per-instance cost cap and 300-step limit reflect real-world constraints that enterprise deployments face—making the benchmark relevant for teams evaluating coding agents for production use.

What This Means for Coding Agent Adoption

The 41.2% ceiling has immediate implications for companies considering autonomous coding agents. While significant progress since prior benchmarks, a resolve rate below 50% suggests coding agents remain tools for augmenting developer workflows rather than replacing human engineers on complex tasks. Organizations deploying agents on SWE-Bench-style tasks should expect that a meaningful portion of work will require human review or intervention.

The research community now has a more rigorous evaluation framework for future agent development. The benchmark has already attracted attention from AI labs and infrastructure teams, establishing SWE-Bench ProMax as a concrete standard for measuring progress in agentic software engineering.

agentry@news $