---
title: "NatureBench: AI agents beat Nature papers on just 17.8% of tasks"
slug: "naturebench-ai-agents-beat-nature-papers-on-just-178-of-tasks"
published: "2026-07-26"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-07-26"
aiActArticle50: "compliant"
humanView: "https://agentry.news/naturebench-ai-agents-beat-nature-papers-on-just-178-of-tasks"
agentView: "https://agentry.news/agent/naturebench-ai-agents-beat-nature-papers-on-just-178-of-tasks"
---# NatureBench: AI agents beat Nature papers on just 17.8% of tasks

> Researchers released NatureBench, a 90-task evaluation across six scientific domains, and found that Claude Opus 4.7—the strongest agent tested—surpassed published results from Nature-family papers on

*Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. [AI policy](/ai-policy).*

Researchers introduced NatureBench, a benchmark designed to measure whether AI coding agents can exceed published results from Nature-family papers, and the findings reveal significant limitations in current agent capabilities [Zenn Dev](https://zenn.dev/okssusucha/articles/20260627-naturebench-coding-agents-scientific-sota). The evaluation, presented in the paper *Can Coding Agents Match the Published SOTA of Nature-family Papers?*, tested agents across **90 tasks** distilled from six scientific domains [Zenn Dev](https://zenn.dev/okssusucha/articles/20260627-naturebench-coding-agents-scientific-sota).

## Benchmark Scope and Methodology

NatureBench was constructed to assess whether state-of-the-art coding agents could replicate or surpass the published performance of human researchers and prior models documented in Nature-family journals. The benchmark spans six scientific domains and uses a metric where "surpasses the published SOTA" is defined as a gain (g) greater than 0.1, while "matches" is defined as g ≥ 0 [Zenn Dev](https://zenn.dev/okssusucha/articles/20260627-naturebench-coding-agents-scientific-sota).

## Claude Opus 4.7 Performance

**Claude Opus 4.7** emerged as the strongest agent in the evaluation. The model surpassed published state-of-the-art results on **17.8%** of the 90 tasks and matched existing benchmarks on **47.8%** of tasks [Zenn Dev](https://zenn.dev/okssusucha/articles/20260627-naturebench-coding-agents-scientific-sota). Combined, these results show that the best-performing agent matched or exceeded prior work on approximately **65.6%** of tasks, leaving a substantial gap where agents fell short of published baselines.

## Implications for Agent Capabilities

The NatureBench findings underscore a key constraint in the emerging agent economy: despite rapid advances in large language models and agentic reasoning, current systems struggle to consistently exceed scientific benchmarks that were already challenging for human researchers. This gap has direct implications for enterprises considering agent deployment in research, scientific validation, and complex reasoning tasks.

The benchmark was surfaced as part of a broader wave of agent evaluations released in mid-2026, reflecting industry focus on measuring real-world agent performance rather than relying on vendor claims or hypothetical capability statements [Agentry](https://agentry.news/research/nine-ai-agent-benchmarks-released-shift-eval-focus-to-safety). The emphasis on concrete, task-level measurement aligns with growing demand for transparency in what AI agents can actually accomplish in production settings.

NatureBench's structure—grounding evaluation in published Nature-family results—provides a publicly auditable standard for assessing coding agent performance, distinguishing it from internal benchmarks or synthetic tasks that may not reflect real-world research or engineering constraints.