title: "τ^τ-Bench: Claude Opus 5 scores 23.9% on agent-building tasks" slug: "bench-claude-opus-5-scores-239-on-agent-building-tasks" published: "2026-09-23" beat: "Research" tags: ["Research"] creator: "Agentry Newsroom" editor: "Susanne Sperling, Editor — Human in the Loop" tools: ["Claude (Anthropic)", "Perplexity Sonar"] creativeWorkStatus: "verified" dateReviewed: "2026-09-23" aiActArticle50: "compliant" humanView: "https://agentry.news/research/bench-claude-opus-5-scores-239-on-agent-building-tasks" agentView: "https://agentry.news/agent/bench-claude-opus-5-scores-239-on-agent-building-tasks"
Sierra's new Hyper-τ-bench evaluation framework measured how well AI agents can autonomously build other agents, with Claude Opus 5 running in Claude Code passing just 23.9% of 53 held-out tasks acros
Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.
Sierra published a new benchmark framework for evaluating AI agents that build agents, revealing significant gaps between current model performance and human-level autonomous agent construction.
The benchmark, called Hyper-τ-bench (or τ^τ-Bench), tested how well AI systems could complete end-to-end agent-building workflows. The strongest configuration tested — Claude Opus 5 (max reasoning) running in Claude Code — passed just 23.9% of held-out evaluation tasks, according to Sierra's published report. The evaluation covered 53 distinct tasks spanning four domains.
The framework measures realistic agent construction, moving beyond isolated coding or reasoning benchmarks to test systems on the full pipeline of building functional agents. Each task represents a scenario where an AI system must autonomously architect, implement, and validate an agent capable of solving downstream problems.
Sierra designed the evaluation to distinguish between capability levels by anchoring performance against a reference point: expert human engineers working on the same tasks achieved an 82.2% pass rate, establishing a meaningful performance ceiling.
The 23.9% pass rate for state-of-the-art models underscores a persistent challenge in the agentic software development space. As enterprises begin deploying autonomous agents to handle complex workflows — from customer service to data processing — the ability for systems to autonomously build and refine agents themselves remains far from human-equivalent reliability.
This benchmark arrives as the agent economy expands rapidly. Companies are shipping agent products at scale, and developer tools for agent construction are maturing. Yet the data shows that even the strongest available model configurations struggle to reliably complete full construction cycles without human intervention.
The evaluation's design across four domains indicates breadth: the benchmark tests agent-building performance across distinct problem spaces rather than a narrow slice of code generation or reasoning. By holding out test tasks and measuring pass rates on unseen evaluation sets, the framework avoids optimizing for specific training data and instead reflects genuine generalization to new agent-building scenarios.
The 23.9% figure represents a concrete starting point. Future iterations of models and improved prompting strategies may shift this baseline, but the benchmark establishes what is currently achievable with production-grade configurations.
Sierra's release of Hyper-τ-bench as an open benchmark via academic publication signals the field's maturation: agent evaluation is moving from proprietary internal testing toward shared, reproducible measurement standards. This enables developers and enterprises to compare agent-building systems objectively and track progress as model capabilities improve.