AGENTRY.NEWSWhat AI Agents Do, Documented.October 6, 2026

Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.

WhatWorkedBench: New Benchmark Tests Agent Experimental Reasoning

By
Agentry Newsroom
Published

Researchers Jingjie Ning, Xueqi Li, Yibo Kong, and Dongting Li introduced WhatWorkedBench, a benchmark designed to measure whether AI research agents can accurately predict the effects of computational or workflow-configuration changes after conducting a limited set of experiments arXiv.

The benchmark was submitted to arXiv on September 23, 2026 and updated on October 2, 2026, establishing concrete methodology for evaluating agent reasoning in scientific experimentation scenarios.

Benchmark Design and Scale

WhatWorkedBench tests agent reasoning across a substantial experimental surface: 36 task conditions, 30 sources, 8 workflow families, 1,248 indexed configuration records, and 4,206 numerical controls. The evaluation method requires agents to inspect workflow code, select measurements within a fixed experimental budget, and submit predicted scores for all configurations. Exhaustive CPU executions then provide reference results against which agent predictions are measured arXiv.

This approach mirrors real research workflows where computational budgets constrain how many experiments can be conducted before making recommendations—a constraint that forces agents to reason about causality and interaction effects rather than simply enumerating all possibilities.

Key Findings

In a core evaluation using a pair-effect ridge method with eight purchased measurements and two free anchor measurements, agents achieved the exact optimum on 15 of 22 sources. However, the findings revealed a subtle failure mode: in 13 of those successful cases, at least one conditional-effect error exceeded 10% of the task-utility range—meaning agents predicted the correct configuration while maintaining significant errors about *why* it worked arXiv.

This distinction matters for agent reliability in research contexts. An agent that recommends the right optimization but misunderstands the underlying causal structure may fail catastrophically when task parameters shift or new constraints emerge.

Agent Model Evaluation

In a prospective evaluation involving 12 four-factor sources, two models produced notably different artifact quality. DeepSeek Flash submitted 12 of 12 direct tables, while DeepSeek Pro submitted 11 of 12 artifacts. This variance underscores how model selection affects not just prediction accuracy but also reasoning transparency—whether agents can explain their experimental selections in structured, verifiable formats arXiv.

Implications for Agent Development

WhatWorkedBench addresses a critical gap in agent evaluation: most benchmarks test task completion or information retrieval, but few measure whether agents can reason through scientific experimentation under budget constraints. As AI agents increasingly operate in research, engineering, and optimization roles, the ability to design efficient experiments and reason about causality becomes core infrastructure.

The benchmark's public release on arXiv opens it for community use and extension, establishing a shared evaluation ground for comparing agent reasoning capabilities across model families and architectural approaches.

Del dette opslag: