AI agents master engineering but stumble on research—shadow study find
Frontier agents complete engineering work but fail on research depth
A new arXiv preprint reports that frontier AI agents can independently execute all the engineering required for AI research projects but hit a hard wall when it comes to making substantive contributions to open-ended research questions arXiv.
Researchers conducted shadow evaluations on two unpublished NeurIPS 2026 submissions, giving agents six days and thousands of dollars of compute to work through the papers' challenges. The agents completed every engineering task without human intervention—writing code, running experiments, debugging implementations—but could not make substantial progress on the research questions themselves arXiv News.
This finding draws a sharp distinction between two types of intellectual labor. Engineering work—the concrete, deterministic tasks of software development—falls squarely within agent capabilities today. Research work—the exploratory, open-ended problem-solving that defines novel scientific contributions—does not.
What shadow evaluations reveal about agent limitations
Shadow evaluations are a method for testing AI systems on real academic work without contaminating peer review. By running agents on unpublished papers submitted to a major conference, the study team avoided introducing data leakage into the research ecosystem while gathering genuine evidence of agent behavior on authentic, high-stakes research tasks Papers Code.
The six-day window and compute budget mirror real constraints researchers face: limited time and finite resources. Within those bounds, agents demonstrated competence at code generation, experiment execution, and debugging—the tactical layer of research. They failed at the strategic layer: identifying novel directions, questioning assumptions, and designing experiments to test new hypotheses MegaBrain.
Why this matters for the agent economy
The finding clarifies a critical boundary in today's agent capabilities. Agents are already viable for tasks where success is measurable against a concrete specification—trading, customer support, code review, document processing. The gap identified here suggests that agents will not soon replace research scientists, even as they become indispensable as research tools Brocker.
This distinction has immediate implications for business adoption. Enterprise use cases that center on reproducible, rule-based work will continue to see rapid agent deployment. Use cases that require genuine open-ended problem-solving—strategy, science, design—remain dependent on human insight.
The arXiv preprint does not disclose the names of the research teams or original paper authors, nor does it provide dollar amounts beyond "thousands of dollars of compute." Full methodological details are expected in the published version.