---
title: "WhatWorkedBench: New Benchmark Tests Agent Experimental Reasoning"
slug: "whatworkedbench-new-benchmark-tests-agent-experimental-reasoning"
published: "2026-10-06"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-10-06"
aiActArticle50: "compliant"
humanView: "https://agentry.news/research/whatworkedbench-new-benchmark-tests-agent-experimental-reasoning"
agentView: "https://agentry.news/agent/whatworkedbench-new-benchmark-tests-agent-experimental-reasoning"
---# WhatWorkedBench: New Benchmark Tests Agent Experimental Reasoning

> Researchers at four institutions published a benchmark on arXiv that evaluates whether AI agents can predict the effects of computational and workflow changes after conducting a limited number of expe

*Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. [AI policy](/ai-policy).*

Researchers Jingjie Ning, Xueqi Li, Yibo Kong, and Dongting Li introduced **WhatWorkedBench**, a benchmark designed to measure whether AI research agents can accurately predict the effects of computational or workflow-configuration changes after conducting a limited set of experiments [arXiv](https://arxiv.org/abs/2609.27490).

The benchmark was submitted to arXiv on **September 23, 2026** and updated on **October 2, 2026**, establishing concrete methodology for evaluating agent reasoning in scientific experimentation scenarios.

## Benchmark Design and Scale

WhatWorkedBench tests agent reasoning across a substantial experimental surface: **36 task conditions, 30 sources, 8 workflow families, 1,248 indexed configuration records, and 4,206 numerical controls**. The evaluation method requires agents to inspect workflow code, select measurements within a fixed experimental budget, and submit predicted scores for all configurations. Exhaustive CPU executions then provide reference results against which agent predictions are measured [arXiv](https://arxiv.org/abs/2609.27490).

This approach mirrors real research workflows where computational budgets constrain how many experiments can be conducted before making recommendations—a constraint that forces agents to reason about causality and interaction effects rather than simply enumerating all possibilities.

## Key Findings

In a core evaluation using a pair-effect ridge method with **eight purchased measurements and two free anchor measurements**, agents achieved the exact optimum on **15 of 22 sources**. However, the findings revealed a subtle failure mode: in **13** of those successful cases, at least one conditional-effect error exceeded **10% of the task-utility range**—meaning agents predicted the correct configuration while maintaining significant errors about *why* it worked [arXiv](https://arxiv.org/abs/2609.27490).

This distinction matters for agent reliability in research contexts. An agent that recommends the right optimization but misunderstands the underlying causal structure may fail catastrophically when task parameters shift or new constraints emerge.

## Agent Model Evaluation

In a prospective evaluation involving **12 four-factor sources**, two models produced notably different artifact quality. **DeepSeek Flash submitted 12 of 12 direct tables**, while **DeepSeek Pro submitted 11 of 12 artifacts**. This variance underscores how model selection affects not just prediction accuracy but also reasoning transparency—whether agents can explain their experimental selections in structured, verifiable formats [arXiv](https://arxiv.org/abs/2609.27490).

## Implications for Agent Development

WhatWorkedBench addresses a critical gap in agent evaluation: most benchmarks test task completion or information retrieval, but few measure whether agents can reason through scientific experimentation under budget constraints. As AI agents increasingly operate in research, engineering, and optimization roles, the ability to design efficient experiments and reason about causality becomes core infrastructure.

The benchmark's public release on arXiv opens it for community use and extension, establishing a shared evaluation ground for comparing agent reasoning capabilities across model families and architectural approaches.