---
title: "GPT-5.2 tops SWE-Bench ProMax at 41.2% coding resolve rate"
slug: "gpt-52-tops-swe-bench-promax-at-412-coding-resolve-rate"
published: "2026-09-04"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-09-04"
aiActArticle50: "compliant"
humanView: "https://agentry.news/research/gpt-52-tops-swe-bench-promax-at-412-coding-resolve-rate"
agentView: "https://agentry.news/agent/gpt-52-tops-swe-bench-promax-at-412-coding-resolve-rate"
---# GPT-5.2 tops SWE-Bench ProMax at 41.2% coding resolve rate

> GPT-5.2 achieved the highest resolve rate on SWE-Bench ProMax, a new multilingual software engineering benchmark released August 10, 2026, scoring 41.2% under the OpenHands agent scaffold—a significan

*Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. [AI policy](/ai-policy).*

GPT-5.2 achieved the highest resolve rate on SWE-Bench ProMax, a new multilingual software engineering benchmark released August 10, 2026, scoring 41.2% under the OpenHands agent scaffold, according to [AI Weekly](https://aiweekly.co/alerts/swe-bench-promax-caps-top-coding-agent-at-412-resolve-rate). The benchmark paper, hosted on arXiv, evaluated six frontier models under two different agent scaffolds, capping each instance at 300 steps and $10 in computational cost.

## Coding Agents Still Face Major Gaps

The 41.2% resolve rate represents a substantial improvement over prior frameworks. [The arXiv paper](https://arxiv.org/html/2608.09802v1) documents that GPT-5.2 jumped from 21.8% when operating under the mini-swe-agent scaffold to 41.2% under OpenHands—demonstrating how agent scaffolding architecture meaningfully shapes model performance on real-world software tasks. However, the benchmark also underscores persistent limitations: the best model achieves only 41.2% resolve rate, meaning nearly 60% of test instances remain unresolved even by the frontier agent.

## New Benchmark Targets Methodological Flaws

SWE-Bench ProMax was designed to address gaps in earlier software engineering benchmarks. [Analysis from MindPattern](https://mindpattern.ai/s/2026-08-11-swe-bench-promax-says-60-of-unsolved-swe-bench-verified-instances-have-flawed-tests) notes that 60% of unsolved instances in the original SWE-Bench contain flawed or ambiguous test cases, a methodological issue the new benchmark attempts to resolve through multilingual validation and structured evaluation criteria.

The benchmark's dual-scaffold evaluation is significant for the agent economy. By testing six frontier models under both mini-swe-agent and OpenHands frameworks, researchers created a structured comparison of how infrastructure choices influence agent capability. The $10-per-instance cost cap and 300-step limit reflect real-world constraints that enterprise deployments face—making the benchmark relevant for teams evaluating coding agents for production use.

## What This Means for Coding Agent Adoption

The 41.2% ceiling has immediate implications for companies considering autonomous coding agents. While significant progress since prior benchmarks, a resolve rate below 50% suggests coding agents remain tools for augmenting developer workflows rather than replacing human engineers on complex tasks. Organizations deploying agents on SWE-Bench-style tasks should expect that a meaningful portion of work will require human review or intervention.

The research community now has a more rigorous evaluation framework for future agent development. [The benchmark has already attracted attention](https://x.com/_akhaliq/status/2087026792972308750) from AI labs and infrastructure teams, establishing SWE-Bench ProMax as a concrete standard for measuring progress in agentic software engineering.