AGENTRY.NEWSWhat AI Agents Do, Documented.August 13, 2026

Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.

Grok 4.5 solves only 28% of long-horizon tasks—benchmark reveals gap

By
Agentry Newsroom
Published

xAI's Grok 4.5 achieved only 28.3% pass@1 performance at a partial-reward threshold on the Long-Horizon-Terminal-Bench, a new research evaluation that tests agent performance across 46 extended, multi-step tasks arXiv. The benchmark paper, titled *Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Extended Tasks*, demonstrates that even the strongest frontier model struggles significantly when tasked with planning and executing sequences of actions over extended horizons.

What the Benchmark Measures

The Long-Horizon-Terminal-Bench evaluates agents on realistic, goal-driven tasks that require sustained reasoning and error recovery. Results are graded at two reward thresholds: tasks achieving 28.3% pass@1 at a 0.95 reward threshold (corresponding to 13 of 46 tasks solved) and 19.6% pass@1 at a 1.0 perfect-reward threshold (9 of 46 tasks) arXiv. The two-tier structure allows researchers to distinguish between agents that approximately solve problems and those that achieve flawless execution—a critical distinction for real-world deployment.

Implications for Agent Deployment

The gap between Grok 4.5's performance and the full benchmark reveals a bottleneck in current agent architecture: models excel at single-step reasoning but falter when required to maintain context, recover from mistakes, or adjust strategy across dozens of sequential decisions. The paper notes that the benchmark "remains far from saturated," signaling that substantially more capable systems will be needed before agents can reliably handle complex, open-ended workflows arXiv.

This finding directly impacts enterprise and consumer applications. Autonomous coding agents, data processing workflows, and multi-step research tasks—all currently marketed as production-ready—operate within an environment where the theoretical ceiling for correctness on novel long-horizon problems is below 30%. Organizations deploying such agents must now account for a higher error rate on complex tasks and design human-in-the-loop safeguards accordingly.

Next Steps for Research

The publication of Long-Horizon-Terminal-Bench establishes a measurable evaluation standard for the field. Future model releases can be directly compared against this 46-task suite, allowing researchers to track whether improvements in model scale and training translate to better long-horizon reasoning. The benchmark's public leaderboard invites competitive evaluation, likely spurring both academic and industry efforts to close the performance gap.

For now, the benchmark serves as a quantified reality check: agents remain far from autonomous, and the path to reliable long-horizon execution is longer than recent product announcements have suggested.

Del dette opslag: