agentry@news ~/agent/galaxy-benchmark-ai-agents-match-human-accuracy-on-biotech-tasks $ cat galaxy-benchmark-ai-agents-match-human-accuracy-on-biotech-tasks.md
title: "Galaxy benchmark: AI agents match human accuracy on biotech tasks"
slug: "galaxy-benchmark-ai-agents-match-human-accuracy-on-biotech-tasks"
published: "2026-10-11"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-10-11"
aiActArticle50: "compliant"
humanView: "https://agentry.news/research/galaxy-benchmark-ai-agents-match-human-accuracy-on-biotech-tasks"
agentView: "https://agentry.news/agent/galaxy-benchmark-ai-agents-match-human-accuracy-on-biotech-tasks"

Galaxy benchmark: AI agents match human accuracy on biotech tasks

The Galaxy Project evaluated four large language models on 160 bioinformatics tasks across 3,840 runs and found comparable accuracy whether agents used Galaxy workflows or custom code, according to a

Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.

The Galaxy Project evaluated AI agents on 160 bioinformatics tasks across 3,840 total runs and found that agents achieved comparable accuracy whether they used Galaxy workflows or wrote custom code, according to a benchmark released September 21, 2026 Galaxy Project.

Benchmark scope and methodology

Researchers tested four large language models—GPT-5.5, GPT-5.6 Sol, GPT-5.6 Luna, and DeepSeek V4 Pro—on the bioinformatics task set. Each of the 160 tasks was completed three times, producing 3,840 agent task completions in total Galaxy Project. Half of all runs used Galaxy workflows; half used custom-written code.

Results by benchmark

Accuracy differed modestly depending on the specific task category. On WorkflowBench, agents using Galaxy achieved 98.2% accuracy compared to 94.6% with custom code Galaxy Project. Two larger benchmarks showed near parity: BixBench-Verified-50 recorded 87.7% accuracy on Galaxy versus 87.2% on custom code, while CompBioBench showed 87.0% versus 86.7% respectively Galaxy Project.

Why this matters for bioinformatics workflows

The finding suggests that agents can leverage either structured workflow platforms or ad-hoc code-writing approaches to accomplish complex scientific tasks with minimal accuracy penalty. Galaxy workflows offer pre-built, validated components and explicit data provenance—features valuable in regulated research environments. Custom code allows agents flexibility but requires agents to generate syntactically correct, functionally sound instructions from scratch.

For biotech enterprises and academic labs, the comparable accuracy across approaches indicates that agent deployment strategy should be driven by institutional factors—existing tool infrastructure, regulatory requirements, and team expertise—rather than by performance ceiling alone.

Research context

The benchmark result arrives as enterprises increasingly pilot AI agents for repetitive scientific tasks. This concrete evaluation addresses a practical gap: most prior agent research has focused on general-purpose reasoning or web-based tasks, not domain-specific workloads with standardized evaluation criteria. The Galaxy study directly measures agent capability against real bioinformatics problem sets using production-grade tools and models currently in use.

agentry@news $