Agents Are Systems, Not Models—New Benchmark Shows Setup Beats Model C
Researchers published *Agents Are Systems, Not Models: Rethinking Agentic Evaluation* on October 1, 2026 (updated October 3), introducing a benchmark that challenges how the industry measures agent performance. The paper, posted on arxiv, reports a finding that reframes agent evaluation: approximately 54% of outcome variance came from repeating the same configuration rather than changing the underlying model arxiv.
The Benchmark and Key Findings
The researchers tested agents across four scientific tasks, measuring run-to-run variability under controlled conditions. The headline result cuts against prevailing industry assumptions. Instead of finding that swapping one language model for another drives performance differences, they discovered that the way an agent is *configured*—its tools, prompts, decision logic, and system parameters—accounts for the majority of performance swing.
"We find substantial run-to-run variability, with approximately 54% of the outcome variance coming from repeating the same configuration rather than changing it," the paper states. This variability suggests that agents are far less stable than many practitioners assume, and that the composition of the system itself matters more than the intelligence of the model at its core.
Why This Matters for Agent Builders
The implications are concrete. Developers and enterprises deploying agents have historically focused on choosing the "best" underlying model—GPT-4, Claude, Gemini. The benchmark argues this is backwards. Two agents running the same model but with different retrieval strategies, tool ordering, or error-handling flows will diverge sharply. The authors conclude: "These results suggest that agents should be evaluated as configurable systems themselves."
This aligns with practical experience in production settings. An agent's real-world behavior depends on how it can access data, which tools it can call, how it recovers from failures, and how often it recomputes. A mediocre model with a well-tuned system often beats a powerful model with a sloppy pipeline.
Data Release and Reproducibility
The authors released the benchmark and more than 18,000 agent trajectories alongside the preprint, allowing other researchers and engineers to replicate findings and test their own agent configurations. This dataset volume and transparency are rare in agentic AI research, where many labs publish findings without releasing the underlying runs or evaluation harness.
Timing and Adoption
The work arrives as enterprises are moving beyond pilot deployments of AI agents into sustained operations. Teams now face the practical question: how do we reliably improve an agent's performance in production? The answer, per this benchmark, is not to chase the next model upgrade—it is to audit and optimize the system that surrounds the model.
For agent teams, the takeaway is stark: system architecture is the primary lever. The benchmark gives them a starting point: treat agents as systems first, and model choice as a secondary concern.