AGENTRY.NEWSWhat AI Agents Do, Documented.August 1, 2026

Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.

Coding agents fail to implement AI research — best hit 33%

By
Agentry Newsroom
Published

A new evaluation paper presented at ACL 2026 in San Diego found that coding agents built with popular frameworks cannot autonomously implement realistic AI research extensions, raising questions about their readiness for independent research workflows.

Researchers Nicholas Edwards, Yukyung Lee, Yujun Audrey Mao, Yulu Qin, Sebastian Schuster, and Najoung Kim evaluated 12 LLM-based agents constructed using aider and OpenHands on a benchmark called RExBench. The study measured whether agents could handle tasks that extend existing AI research codebases—a common but complex activity in academic and industry research labs.

Best Agent Reaches Only 33% Success

The benchmark results were sobering. The best-performing agent achieved approximately a 33% success rate on tasks without human intervention. When the research team provided human-written hints to guide agent behavior, performance improved marginally but remained below 44%—indicating that current agents lack the autonomous capability to handle realistic research extension tasks.

All 12 evaluated agents failed on the majority of test cases. The authors concluded in their ACL long paper that "current agents are still short of being able to handle realistic research extension tasks without substantial human guidance." This finding challenges assumptions that modern coding agents are ready to operate independently on complex, novel programming challenges.

Implications for Agent-Driven Research

The result surfaces a critical gap in the agent economy: while coding agents are marketed as productivity multipliers for software development, they struggle with the kind of open-ended, research-level implementation tasks that require understanding novel problem domains and adapting existing code architectures.

The distinction matters for organizations planning to deploy agents on research infrastructure. Unlike routine coding tasks with well-defined specifications, research extensions often require agents to interpret research papers, modify complex numerical algorithms, and integrate new methods into existing experimental pipelines—all activities that demand deep contextual understanding and autonomous problem-solving.

The RExBench evaluation adds concrete measurement to an otherwise speculative debate about agent capabilities. Rather than relying on vendor claims or hypothetical roadmaps, the paper provides a reproducible benchmark and documented failure modes that other researchers can build on.

What's Next

The work points toward directions for improving agent autonomy: better prompt engineering for research contexts, agent access to documentation and paper abstracts, and architectural changes to how aider and OpenHands decompose research tasks. Whether these improvements will move the needle above 44% remains an open question for future work.

Del dette opslag: