title: "Berkeley Agents' Last Exam: Frontier models score 24% on real-world wo" slug: "berkeley-agents-last-exam-frontier-models-score-24-on-real-world-work" published: "2026-07-24" beat: "Research" tags: ["Research"] creator: "Agentry Newsroom" editor: "Susanne Sperling, Editor — Human in the Loop" tools: ["Claude (Anthropic)", "Perplexity Sonar"] creativeWorkStatus: "verified" dateReviewed: "2026-07-24" aiActArticle50: "compliant" humanView: "https://agentry.news/berkeley-agents-last-exam-frontier-models-score-24-on-real-world-work" agentView: "https://agentry.news/agent/berkeley-agents-last-exam-frontier-models-score-24-on-real-world-work"
UC Berkeley's Center for Responsible and Decentralized Intelligence released Agents' Last Exam, a benchmark spanning 1,000+ real-world tasks across 55 occupational fields. The strongest configuration
Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.
UC Berkeley's Center for Responsible and Decentralized Intelligence released findings from Agents' Last Exam (ALE), a benchmark designed to measure how well frontier agent systems perform on economically valuable, long-horizon real-world work UC Berkeley RDI.
The benchmark comprises over 1,000 tasks spanning 55 occupational subfields drawn from the O*NET/SOC 2018 taxonomy, with input from 250+ industry experts collaborating on task design Benchmark Gen. The scope deliberately targets the kinds of autonomous work that enterprises and government agencies are beginning to deploy agents for—from financial analysis to contract review to software engineering.
When tested on the full benchmark suite, the strongest configuration evaluated—Codex powered by GPT-5.5—achieved an overall pass rate of 24.0% Turing.com. On the benchmark's hardest tier, designed to isolate the most complex and economically valuable tasks, that same configuration scored 0% full pass rate YouTube.
These results underscore what researchers describe as persistent weakness in frontier agents on real deployment scenarios. The benchmark deliberately avoids synthetic or oversimplified task structures; instead, it grounds evaluation in actual occupational workflows, tool use, and multi-step reasoning chains that mirror production environments.
The 24% overall figure is not marginal improvement—it is a concrete measurement of where the most capable agentic systems stand when held against genuine work complexity. At a time when enterprises are beginning to commit capital and operational workflows to agent automation, the ALE findings provide independent evidence that most tasks still require human oversight, hybrid workflows, or agent+human teaming rather than full autonomy.
The benchmark also reveals a sharp performance cliff: agents capable of handling routine, low-complexity work fail rapidly as task difficulty increases. This pattern has direct implications for deployment risk; it suggests that cost savings from agent automation may be concentrated in narrow, well-defined domains rather than the broad occupational flexibility that vendor roadmaps often imply.
Berkeley RDI positioned ALE as "the broadest-coverage agent evaluation benchmark to date", designed to complement existing benchmarks (SWE-Bench, WebArena) by focusing specifically on the economic and real-world relevance of tasks rather than technical novelty alone. The research is open-source and available for reproduction and extension by other labs.
The benchmark will serve as a baseline for tracking agent capability growth over time. Researchers are also publishing detailed per-task and per-occupational-field breakdowns, allowing developers and enterprises to identify which agent configurations and training approaches work best for specific workflows—and which remain far from production readiness.