title: "Claude Opus 4.6 leads SWE-bench, rivals cluster within 3 points" slug: "claude-opus-46-leads-swe-bench-rivals-cluster-within-3-points" published: "2026-08-24" beat: "Research" tags: ["Research"] creator: "Agentry Newsroom" editor: "Susanne Sperling, Editor — Human in the Loop" tools: ["Claude (Anthropic)", "Perplexity Sonar"] creativeWorkStatus: "verified" dateReviewed: "2026-08-24" aiActArticle50: "compliant" humanView: "https://agentry.news/research/claude-opus-46-leads-swe-bench-rivals-cluster-within-3-points" agentView: "https://agentry.news/agent/claude-opus-46-leads-swe-bench-rivals-cluster-within-3-points"
Claude Opus 4.6 scored 75.6% on SWE-bench Verified in an August 22, 2026 leaderboard snapshot, with MiniMax M2.7 and GLM-5 following closely at 75.4% and 72.8% respectively, signaling tight competitio
Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.
Claude Opus 4.6 holds the top position on SWE-bench Verified, the software-engineering benchmark that tests agents on real-world code repair tasks, with a score of 75.6% according to a public leaderboard mirror dated August 22, 2026. The tight clustering of scores near the top reflects the increasingly narrow performance gap between proprietary frontier models and high-capability open-weight alternatives.
The August 2026 leaderboard snapshot captured Anthropic's Claude Opus 4.6 at 75.6%, followed by MiniMax M2.7 at 75.4% and GLM-5 at 72.8%, according to the third-party mirror. The 0.2-percentage-point separation between the top two models underscores how frontier agents are converging on the upper boundary of the benchmark, leaving diminishing room for differentiation in raw SWE-bench performance.
SWE-bench Verified is a filtered, more rigorous subset of the original SWE-bench evaluation, designed to filter out lower-confidence test cases and reduce variance. The benchmark measures an agent's ability to resolve open GitHub issues in real Python repositories—a concrete proxy for software-engineering capability that many enterprise teams monitor as a proxy for autonomous coding readiness.
When multiple models cluster within a few percentage points on a heavily-watched benchmark, it typically signals two dynamics: first, that the benchmark may be approaching saturation for currently available architectures and inference techniques; second, that differentiation among agents is shifting away from raw benchmark score and toward factors like cost, latency, developer experience, and task-specific reliability rather than headline accuracy.
The presence of open-weight models like GLM-5 and MiniMax M2.7 in the top tier is also material for enterprise adoption and cost-sensitive deployments. Open-weight alternatives reduce lock-in and lower inference cost, a meaningful consideration for teams evaluating long-term agent strategies.
A single benchmark snapshot, even a rigorous one, does not determine agent utility in production. SWE-bench Verified measures point-in-time accuracy on code-repair tasks but does not capture agent agentic capabilities like planning, tool use orchestration, error recovery, or real-time collaboration with human developers—dimensions on which performance may diverge significantly from benchmark rank.
The August 2026 leaderboard represents a public moment in a rapidly evolving space. Subsequent evaluations will likely show further shifts as models are refined, inference optimizations are deployed, and new benchmarks emerge to stress-test dimensions that SWE-bench does not cover.