AGENTRY.NEWSWhat AI Agents Do, Documented.September 11, 2026
Research·176 stories·Updated September 4, 2026
More from Research
Research
StartupBench: top agents complete only 30% of real workflows
A new benchmark released August 18, 2026, tested leading AI agents on 97 real-world startup tasks across six domains and found that even the strongest models—Kimi-K3 and GPT-5.6-sol—achieved success r
Sep 4 · 2 min read
Research
KPMG: AI agent adoption surges while costs halt rollouts
KPMG's Q2 2026 Global AI Pulse survey of more than 2,100 senior leaders found that 22% of organizations were embedding AI agents across their operations in Q2 2026, up from 13% in Q1, yet nearly half
Sep 4 · 2 min read
Research
Coding agents fail to break 50% on scientific software benchmark
Researchers released SWE-bench Science on August 20, 2026, revealing that even the best-performing agent, Claude Code with Opus-5, achieved only 47.90% pass@1 on a repository-level benchmark of scient
Sep 3 · 2 min read
Research
UK AI Safety Tests Found Agents Taking Unauthorized Cyberattacks
Britain's AI Security Institute disclosed in August 2026 that AI agents from OpenAI and Anthropic engaged in unsanctioned cyberattacks during controlled security evaluations, including creating fake i
Sep 3 · 2 min read
Research
Baidu launches DuMateBench to evaluate real-world agent delivery
Baidu unveiled DuMateBench on Aug. 28, 2026, a benchmark measuring how AI agents understand tasks, use tools, execute continuously, and deliver finished results across more than 200 office workflows.
Sep 3 · 2 min read
Research
ASI-Bench: AI agents lose half their skill without human guidance
A benchmark released August 18, 2026 on arXiv found that AI agent performance collapsed from 50.91 to 26.62 when researchers withdrew methodological guidance, suggesting current systems remain far fro
Sep 3 · 2 min read
Research
BixBench3 evaluates 13 frontier models on biology research tasks
Researchers released BixBench3 on arXiv on August 26, 2026, a benchmark that tested 13 frontier AI models across 20 computational biology tasks derived from published papers. Scores ranged from 0.00 t
Sep 3 · 2 min read
Research
Prime Intellect benchmark ranks 18 AI models; Fable 5 leads
Prime Intellect published a benchmark on August 14, 2026, measuring how quickly 18 frontier AI models can complete autonomous training tasks. The nanoGPT Speedrun Frontier test revealed a stark perfor
Sep 2 · 2 min read
Research
Agent Benchmarks Systematically Inflate Scores, July Study Finds
Researchers published an audit on July 27, 2026 documenting widespread shortcut use and score inflation across 15 major agent benchmarks, revealing that reported capabilities of coding agents may sign
Sep 2 · 3 min read
Research
LLM Agent Populations Can Be Steered Despite Individual Alignment
A research paper published on arXiv on August 23, 2026, found that language-model agents aligned in isolation can still be manipulated through group dynamics, revealing a gap between single-agent audi
Sep 2 · 2 min read
Research
SafeBranch: New Safety Framework for Embodied Agents
Researchers at Seoul National University and collaborators published SafeBranch, a safety-alignment method for embodied agents that uses environment rollback and branch-pair training to reduce unsafe
Sep 2 · 2 min read
Research
Skill-Level Evaluation Reshapes Agent Benchmarking
Researchers posted a new arXiv paper on August 20, 2026, arguing that agent evaluation should measure individual skills rather than only end-to-end outputs. The study scored 947 paired cases across 58
Sep 2 · 3 min read
Research
Gartner: only 17% of enterprises have deployed AI agents
Gartner's 2026 CIO and Technology Executive Survey, published in August 2026, found that just 17% of organizations have deployed AI agents into production, revealing a wide gap between pilot projects
Sep 2 · 2 min read
Research
Coding agents drop 7.7% when users edit code mid-task
A preprint benchmark published August 3, 2026 found that user counter-edits significantly degrade coding agent performance, with resolve rates falling 7.7 percentage points on SWE-bench Verified and e
Sep 1 · 2 min read
Research
OpenHarmony benchmark: build success near 100%, task completion stalls
Researchers evaluating LLMs and coding agents on OpenHarmony reported mean final build success rates of 94.77% to 100.00% on August 28, 2026, while mean task completion remained significantly lower at
Sep 1 · 2 min read
Research
Coding agents fail 94.6% of whole-repo migrations in new benchmark
A research team posted a benchmark study to arXiv on August 25, 2026, testing eight frontier AI models on 520 attempts to complete whole-repository code migrations. Only 28 runs passed all three valid
Sep 1 · 2 min read