title: "Frontier agents solve just 30% of scientific research tasks" slug: "frontier-agents-solve-just-30-of-scientific-research-tasks" published: "2026-09-04" beat: "Research" tags: ["Research"] creator: "Agentry Newsroom" editor: "Susanne Sperling, Editor — Human in the Loop" tools: ["Claude (Anthropic)", "Perplexity Sonar"] creativeWorkStatus: "verified" dateReviewed: "2026-09-04" aiActArticle50: "compliant" humanView: "https://agentry.news/research/frontier-agents-solve-just-30-of-scientific-research-tasks" agentView: "https://agentry.news/agent/frontier-agents-solve-just-30-of-scientific-research-tasks"
Stanford University and the Laude Institute launched Terminal-Bench-Science 0.1 on August 27–28, 2026, a benchmark designed to evaluate AI agents on scientist-authored research pipelines. Claude Opus
Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.
Stanford University and the Laude Institute unveiled Terminal-Bench-Science 0.1, a new evaluation framework for AI agents performing scientific research, in late August 2026. The benchmark measures how well agentic systems can complete research tasks designed by human scientists, surfacing a stark performance gap: the top-ranked model resolved only 30% of test cases.
According to ExplainX, the benchmark comprises 70 tasks spanning scientific research pipelines and includes a public leaderboard tracking agent performance by model and configuration. Claude Opus 5 running with Claude Code topped the standings at 30.0% resolution rate, signaling that even frontier large language model agents struggle with reproducible research workflows.
Terminal-Bench-Science 0.1 is purpose-built to test autonomous agents on concrete scientific tasks—not theoretical reasoning but executable research actions: literature review, experimental design, data analysis, and manuscript preparation. By anchoring evaluation in scientist-authored pipelines, the benchmark avoids the vagueness of hypothetical agent capabilities and instead measures what agents can actually do in practice.
The 30.0% top score underscores a critical finding: frontier agents fail to complete two-thirds of structured scientific workflows. This gap matters for research institutions and biotech firms considering agent deployment for knowledge work. The public leaderboard on Terminal Bench reports results by model and agent combination, allowing developers and enterprises to compare performance across configurations.
The benchmark's release signals growing scrutiny of agent capabilities beyond conversational AI. Earlier hype cycles promised autonomous systems capable of complex reasoning; Terminal-Bench-Science 0.1 provides a measurement tool that separates capability claims from empirical results. A 30% resolution rate on curated tasks—not adversarial edge cases—suggests that agentic reasoning in high-stakes domains like science remains immature.
For the agent economy, the implication is clear: builders deploying agents into research workflows cannot assume frontier models will solve domain-specific tasks reliably. Organizations will need to architect human oversight, iterative refinement, and domain-specific fine-tuning. The benchmark itself becomes infrastructure for the sector—a shared evaluation ground where model vendors and agent framework makers can validate improvements over time.
The Terminal-Bench-Science leaderboard is publicly accessible, enabling the community to track progress as new models and agent configurations are tested.