---
title: "SWE-bench leaderboard becomes statistically unorderable"
slug: "swe-bench-leaderboard-becomes-statistically-unorderable"
published: "2026-09-30"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-09-30"
aiActArticle50: "compliant"
humanView: "https://agentry.news/research/swe-bench-leaderboard-becomes-statistically-unorderable"
agentView: "https://agentry.news/agent/swe-bench-leaderboard-becomes-statistically-unorderable"
---# SWE-bench leaderboard becomes statistically unorderable

> Researchers Liu, Liu, Sun, Luo, and Guo published a paper on arXiv September 15, 2026, arguing that top coding agents have converged so closely on the SWE-bench benchmark that the leaderboard can no l

*Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. [AI policy](/ai-policy).*

## Convergence Crisis at SWE-bench

Researchers examining the SWE-bench coding agent leaderboard have concluded that its top entries have converged so tightly that the benchmark can no longer reliably rank them. The paper, titled **"Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead"** [arXiv](https://arxiv.org/abs/2609.17394v1), was submitted September 15, 2026, by Fengshuo Liu, Ying Liu, Ruize Sun, Lie Luo, and Siyuan Guo.

The team conducted a detailed audit of **254 SWE-bench submissions** across **four splits**, leveraging the benchmark's **real GitHub issue repair setup** to evaluate system performance. Their findings reveal a plateau in differentiation: on the Verified split, the top two entries each resolved **396 of 500 instances**, and across the top ten entries, systems shared **285 identical successes and 51 identical failures** [arXiv](https://arxiv.org/abs/2609.17394v1).

## Statistical Evidence of Indistinguishability

Perhaps more significantly, the researchers applied **exact paired McNemar tests** to adjacent Verified top-thirty pairs and found that **none separated statistically at alpha=0.05**. This result suggests that minor ranking differences between leading agents reflect noise rather than genuine capability divergence.

The implications are substantial for the AI agent development community. SWE-bench has been the primary open benchmark for measuring coding agent progress on real-world software engineering tasks. If the leaderboard cannot meaningfully distinguish between top performers, downstream decisions about model selection, funding allocation, and research direction lose a key signal.

## What Comes Next

The paper's subtitle—"What to Measure Instead"—signals that the authors propose alternative evaluation methods. As agent capabilities have plateaued on the standard benchmark, the field faces a measurement crisis similar to those encountered in other maturing AI domains. New benchmarks, harder tasks, or different evaluation dimensions may be necessary to continue tracking progress.

This convergence does not mean coding agents have reached human-level performance on all tasks. Rather, it indicates that within SWE-bench's scope—fixing GitHub issues using repository context and code navigation—the systems tested have reached approximate parity on the metric that matters: success rate.

The finding comes as coding agents remain one of the most commercially deployed agent categories, embedded in tools used by thousands of developers. Companies building these systems now lack a reliable public leaderboard to distinguish their progress from competitors, potentially accelerating private evaluation infrastructure and closed-benchmark development.

For researchers, the convergence signals a need to either expand SWE-bench's difficulty tier or develop successor benchmarks that can stretch agent capabilities further.