
OpenAI Releases GeneBench-Pro, Computational Biology Benchmark
OpenAI released GeneBench-Pro on June 30, 2026, introducing a research-level benchmark designed to evaluate how well AI agents can navigate messy biological data, choose analysis paths, and make judgment calls that real computational research depends on OpenAI.
Benchmark Architecture and Scope
GeneBench-Pro contains 129 synthetic problems spanning genomics, quantitative biology, and translational medicine Tech News Hub. The benchmark measures a critical gap in agent reliability: unlike narrow task completion, computational research requires systems to navigate uncertainty, reject dead-end analyses, and make consequential decisions without human oversight. OpenAI stated the benchmark tests "whether AI agents can navigate messy biological data, choose an analysis path, and make consequential judgment calls that real computational research depends on" Let's Data Science.
Performance Results
OpenAI's GPT-5.6 Sol achieved a 28.7% pass rate at the highest reasoning level on GeneBench-Pro Street Insider. When Pro mode was enabled, GPT-5.6 Sol Pro improved to 31.5% Tech Noisy. This represents a dramatic leap from the original GeneBench, where GPT-5 scored below 5%.
The strongest non-OpenAI baseline tested was Claude Opus 4.8 at 16.0%, demonstrating OpenAI's lead on this specific evaluation Andrew.ooo. Despite the improvement, OpenAI acknowledged that systems remain "too unreliable to replace human experts" in computational biology workflows.
Open-Source Release and Developer Access
OpenAI is democratizing evaluation by open-sourcing 10 representative questions from GeneBench-Pro on Hugging Face under an MIT License AI Catch Up. This move allows researchers and developers to benchmark their own models against the same problems, establishing a shared evaluation standard for agent research in computational biology.
Implications for Agent Research
GeneBench-Pro fills a critical gap in agent evaluation. Most benchmarks test single-turn task completion; GeneBench-Pro instead evaluates multi-step reasoning with irreversible decisions and incomplete information—conditions that mirror real research workflows. The 31.5% pass rate, while an improvement, underscores that agentic AI in biology remains in early stages, with substantial work required before deployment in high-stakes research settings.
The benchmark addresses a core beat for agent evaluation: measuring what agents can actually do in messy, real-world domains, not hypothetical capabilities or narrow laboratory tasks.


