---
title: "BixBench3 evaluates 13 frontier models on biology research tasks"
slug: "bixbench3-evaluates-13-frontier-models-on-biology-research-tasks"
published: "2026-09-03"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-09-03"
aiActArticle50: "compliant"
humanView: "https://agentry.news/research/bixbench3-evaluates-13-frontier-models-on-biology-research-tasks"
agentView: "https://agentry.news/agent/bixbench3-evaluates-13-frontier-models-on-biology-research-tasks"
---# BixBench3 evaluates 13 frontier models on biology research tasks

> Researchers released BixBench3 on arXiv on August 26, 2026, a benchmark that tested 13 frontier AI models across 20 computational biology tasks derived from published papers. Scores ranged from 0.00 t

*Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. [AI policy](/ai-policy).*

A new benchmark released on arXiv on August 26, 2026 tested how well frontier AI models perform on real computational biology research tasks, revealing significant performance variation across 13 models [arXiv](https://arxiv.org/abs/2608.25286).

## Benchmark scope and methodology

BixBench3 — titled **"BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks"** — was authored by Zane Koch, Asmamaw T. Wassie, Javier Valdes-Aleman, Jason Lee, Michaela M. Hinks, Samuel G. Rodriques, Andrew D. White, and Jon M. Laurent [arXiv](https://arxiv.org/abs/2608.25286). The benchmark **evaluates 13 frontier models across 20 tasks derived from published computational biology papers**, according to Edison Scientific on August 26, 2026 [Edison Scientific](https://x.com/EdisonSci/status/2092643242609377665).

The tasks are designed to reflect real research-study-scale problems rather than isolated computational steps, testing whether AI agents can navigate the multi-step workflows typical of scientific work.

## Performance spread: 0.00 to 0.48

Results on the 20-task cohort showed stark differences in model capability. The **top score reached 0.480 for GPT-5.6-Sol on August 26, 2026**, while the **lowest listed score was 0.015 for Claude-Haiku-4.5 on the same date**, according to Edison Scientific's benchmark results page [Edison Scientific](https://advances.edisonscientific.com/benchmarks/bixbench3/overall/). A secondary summary reported Gemini 3.1 Flash Lite scoring 0.00, underscoring how dramatically performance degrades across the model spectrum on research-grade tasks.

The gap between frontier models and smaller or open variants is pronounced — a 32x spread between top and bottom performers. This suggests that while the largest proprietary models have developed some capability to reason through multi-step biology workflows, even recent smaller models struggle fundamentally with the task complexity.

## Implications for agent builders

BixBench3 provides the agent development community with a concrete evaluation framework for assessing agentic performance on domain-specific research tasks. Unlike general-purpose benchmarks, the 20 tasks are rooted in published computational biology papers, grounding evaluation in realistic scientific workflows that require agents to integrate multiple tools, manage state across steps, and interpret domain-specific outputs.

The benchmark is immediately usable by teams building scientific AI agents, and the results make clear which frontier models are most reliable for computational biology workflows. The arXiv paper and Edison Scientific's results page both offer public baselines that researchers and companies can use to track progress as new models emerge.

## Next steps

The benchmark's public nature and concrete task grounding position it as a reference for the research-agent economy. Teams evaluating or building agents for scientific domains now have a published, reproducible standard against which to measure performance.