---
title: "AgencyBench: Closed models beat open 48% vs 32% on real tasks"
slug: "agencybench-closed-models-beat-open-48-vs-32-on-real-tasks"
published: "2026-08-24"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-08-24"
aiActArticle50: "compliant"
humanView: "https://agentry.news/research/agencybench-closed-models-beat-open-48-vs-32-on-real-tasks"
agentView: "https://agentry.news/agent/agencybench-closed-models-beat-open-48-vs-32-on-real-tasks"
---# AgencyBench: Closed models beat open 48% vs 32% on real tasks

> A new ACL 2026 benchmark published July 22 shows closed-source AI models significantly outperform open-source alternatives across 32 real-world agent scenarios, with closed models reaching 48.4% succe

*Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. [AI policy](/ai-policy).*

A comprehensive benchmark released in July 2026 quantifies a widening performance gap between closed-source and open-source AI agents in real-world deployment scenarios. The **AgencyBench** study, published in the ACL Anthology on July 22, measured agent success rates across 32 concrete use cases totaling 138 discrete tasks, with results showing closed-source models achieved **48.4% success** compared to **32.1%** for open-source systems [Agentry](https://agentry.news/agencybench-acl-paper-benchmarks-agents-on-32-real-tasks).

## Methodology and Scope

The benchmark, titled "Benchmarking the Frontiers of Autonomous Agents in 1M-…" [ACL Anthology](https://aclanthology.org/2026.acl-long.337/), evaluated agents on tasks designed to reflect production environments rather than synthetic laboratory conditions. The 32 real-world scenarios span multiple domains and complexity tiers, providing what researchers characterized as a rigorous test of autonomous agent capabilities in tasks agents actually encounter in enterprise and consumer settings.

The 16-point gap between closed and open-source model classes signals that proprietary training approaches, larger model scale, and specialized fine-tuning are producing measurable advantages in agentic reasoning and tool use—core competencies for agents that must plan multi-step workflows, handle errors, and interact with external systems.

## Implications for Agent Development

The findings carry weight for organizations evaluating which model class to deploy in production. Closed-source providers—primarily OpenAI, Anthropic, and Google—have released agents or agentic APIs in 2025–2026, backed by models trained on agent-specific objectives and instruction sets. Open-source alternatives, including models from Meta and Mistral, remain accessible and cost-effective but show measurable performance trade-offs on real tasks.

For developers choosing frameworks and base models, the benchmark provides quantified data rather than anecdotal claims. A **16-percentage-point gap** on 138 real tasks translates to measurable reliability differences; an agent system must succeed repeatedly in production, and lower success rates directly impact deployment cost and user trust.

## Research in the Agent Economy

AgencyBench joins an expanding corpus of peer-reviewed evaluations that move beyond model benchmarks (like MMLU or coding contests) into agent-specific measurement. Earlier 2026 work examined agent deception, safety alignment, and tool-use robustness. This study's scale—32 scenarios, 138 tasks—and focus on closed versus open-source distinction directly inform enterprise procurement decisions and open-source roadmap prioritization.

The results do not predict which model class will lead in 12 months; agent architectures, fine-tuning methods, and tooling are evolving rapidly. However, the published benchmark creates a baseline for future comparisons and quantifies today's production-readiness gap [Agentry](https://agentry.news/research/agencybench-benchmark-reveals-48-closed-source-agent-edge).