title: "PERFOPT-Bench: Framework choice reshapes coding agent results" slug: "perfopt-bench-framework-choice-reshapes-coding-agent-results" published: "2026-08-16" beat: "Research" tags: ["Research"] creator: "Agentry Newsroom" editor: "Susanne Sperling, Editor — Human in the Loop" tools: ["Claude (Anthropic)", "Perplexity Sonar"] creativeWorkStatus: "verified" dateReviewed: "2026-08-16" aiActArticle50: "compliant" humanView: "https://agentry.news/research/perfopt-bench-framework-choice-reshapes-coding-agent-results" agentView: "https://agentry.news/agent/perfopt-bench-framework-choice-reshapes-coding-agent-results"
A July 27 evaluation of 12 long-horizon optimization tasks across 7 agent stacks found no single framework won more than 4 tasks, and the same model produced materially different outputs depending on
Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.
A July 27, 2026 benchmark evaluation found that agent framework choice—not model selection alone—is a primary driver of performance variation across coding optimization tasks AgentMarketCap. The PERFOPT-Bench study tested 7 distinct agent stacks against 12 long-horizon performance-optimization tasks and revealed that the same underlying model produced materially different results when deployed under different frameworks.
The benchmark's core result challenges assumptions that model capability alone predicts agent success: no single agent stack won more than 4 of the 12 tasks AgentMarketCap. This fragmented outcome suggests that framework architecture, task routing, and agent orchestration logic are as critical as the underlying language model.
Concrete performance swings illustrate the framework effect. GPT-5.5 achieved a score of 9.2× under OpenCode but only 8.2× under Codex, a 12% variance tied entirely to framework choice rather than model difference AgentMarketCap. Similarly, Claude Opus 4.7 scored 7.4× under OpenCode versus 6.7× under Claude Code, a 10% spread driven by framework architecture alone AgentMarketCap.
The PERFOPT-Bench results suggest that teams deploying coding agents cannot rely on model selection as a sufficient optimization lever. Framework decisions—including how agents decompose tasks, invoke tools, and manage state across long horizons—directly shape measurable output quality. Enterprises evaluating coding-agent platforms must test candidate stacks against their own workloads rather than assuming that superior model performance guarantees superior agent performance.
The evaluation's scope—12 optimization tasks spanning long-horizon planning with 7 production and research agent frameworks—provides a concrete foundation for such comparative decisions. The finding that framework effects materially outweigh model differences in some cases also signals that smaller organizations or teams with older models may still achieve competitive agent performance by optimizing framework design and orchestration patterns.
This benchmark joins a growing body of research documenting that agentic systems require holistic tuning across models, frameworks, and task design. Teams building or deploying agents should expect continued variation in real-world performance as frameworks mature and new architectural patterns emerge.