
AI agents pass only 25.4% of long-horizon biology tasks
LatchBio introduced scBench-Long, a benchmark designed to test whether AI agents can recover scientific conclusions from raw or near-raw biological data without prescribed methods, moving beyond local analysis to genuine scientific reasoning LatchBio.
Benchmark Design and Scope
The benchmark spans 21 evaluations across critical biology domains, including melanoma CD8 T-cell reactivity, CD8 RNA+ATAC regulatory inference, human-monkey chimera development, KRAS-driven lung tumor aging, and lethal COVID-19 lung pathology. Researchers completed 1,068 agent trajectories to measure performance across these long-horizon reasoning tasks where agents must navigate multi-step scientific inference without explicit procedural guidance.
Performance Results
The strongest agent system achieved a 25.4% pass rate, passing 16 of 63 total runs LatchBio. This result reflects the difficulty of tasks requiring agents to synthesize raw data into publishable scientific insights. Among competing agent harnesses tested, Claude Code outperformed alternatives by 4.8 percentage points, establishing it as the leading framework for this category of long-horizon biological reasoning.
What This Means for Agent Development
The scBench-Long results surface a critical gap in current agent capabilities: while agents excel at discrete, procedurally-defined tasks, they struggle with open-ended scientific discovery that demands iterative hypothesis formation, method selection, and multi-stage data integration. The benchmark's design—asking agents to work from near-raw single-cell data—removes the safety rails of predefined workflows that have boosted performance in narrower evaluations.
Kenny Workman, one of the researchers behind the effort, shared the results on LinkedIn, positioning the benchmark as a tool for measuring real-world scientific utility rather than benchmark-gaming performance LinkedIn. The work reflects growing concern among AI evaluation researchers that published agent benchmarks often test shallow capability rather than the kind of sustained reasoning required in actual research labs.
Implications for Biotech AI
The 25.4% pass rate, while low in absolute terms, does not imply that agents are useless for biology—rather, it highlights that agents performing at this level remain tools requiring close human oversight and validation. The benchmark's release in July 2026 arrives as biotech firms increasingly deploy autonomous analysis systems, making rigorous evaluation essential before production use.
LatchBio's open release of scBench-Long signals the community's move toward harder, more realistic benchmarks that measure what agents can actually contribute to scientific workflows rather than celebrating incremental improvements on simplified tasks.


