ASI-Bench: AI agents lose half their skill without human guidance
A new benchmark paper submitted to arXiv on August 18, 2026 measured how far current AI agents can proceed with scientific research when left to their own devices—and found they still depend heavily on human direction.
The benchmark, titled "ASI-Bench: At the Dawn of Artificial Superintelligence," tested 18 different agent-model configurations across 60 project-level research tasks spanning 11 scientific domains arXiv. When agents had full methodological guidance from researchers, average performance reached 50.91 on the benchmark's scoring scale. When that guidance was withdrawn and agents had to choose their own research methods, performance dropped to 26.62—a decline of nearly 47 percent Center for Consulting.
Research Design and Scale
The benchmark was built by over 40 experts who invested more than 31,000 human hours in its construction. ASI-Bench is designed to test whether AI systems can conduct autonomous, project-level scientific research and to measure how performance degrades as human methodological guidance is progressively withdrawn arXiv.
The drop-off in scores when agents lost access to human-provided methods suggests that current AI systems excel at executing well-defined research workflows but struggle to independently determine *how* to approach an unfamiliar scientific problem. This gap between guided and autonomous performance has direct implications for agent deployment in real research environments, where human experts may not always be available to scaffold methodology.
What This Means for Agent Autonomy
The findings align with broader skepticism about near-term AI agent autonomy. While AI systems have shown impressive capabilities in narrow, well-defined domains—code generation, customer service, data analysis—the ASI-Bench results highlight a persistent bottleneck: the ability to frame and plan novel research tasks without external structure IT Does What Now.
The paper concludes that current systems are not yet autonomous for comprehensive scientific investigation, a finding that undercuts some industry claims about near-term agent capabilities. The benchmark itself is now available as an open resource GitHub, allowing other researchers to test their own agent architectures and methodologies against the same 60 tasks.
For practitioners building agent systems, the ASI-Bench results suggest that hybrid approaches—where agents handle execution while humans retain control over research design—may be more realistic in the near term than fully autonomous scientific agents.