title: "Agent-safety benchmarks measure capability, not alignment" slug: "agent-safety-benchmarks-measure-capability-not-alignment" published: "2026-08-12" beat: "Research" tags: ["Research"] creator: "Agentry Newsroom" editor: "Susanne Sperling, Editor — Human in the Loop" tools: ["Claude (Anthropic)", "Perplexity Sonar"] creativeWorkStatus: "verified" dateReviewed: "2026-08-12" aiActArticle50: "compliant" humanView: "https://agentry.news/research/agent-safety-benchmarks-measure-capability-not-alignment" agentView: "https://agentry.news/agent/agent-safety-benchmarks-measure-capability-not-alignment"
Researchers auditing four prominent agent-safety benchmarks in July 2026 found that scores often track general model capability rather than genuine safety, raising questions about how the AI industry
Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.
Researchers published a validity audit of four prominent agent-safety benchmarks in July 2026, concluding that their scores often measure general model capability rather than whether autonomous agents are genuinely safe or aligned arXiv.
The study, authored by Youting Wang, Xiao Han, Dingyan Shang, Yuan Tang, and Bowen Liu, evaluated R-Judge, InjecAgent, AgentHarm, and AgentDojo across up to 22 models, using MMLU and GPQA as independent measures of capability. The core finding: benchmark scores often rise and fall with model intelligence, not safety properties arXiv.
The researchers found that while capability is a strong predictor of task success, safety correlations vary by outcome and evaluator panel. One striking result: AgentHarm's jailbreak detection score showed a correlation of ρ = +0.72 with general capability after controlling for model intelligence arXiv. This suggests that benchmarks labeled "safety" may simply be measuring whether a model is capable—not whether it refuses harmful requests or resists manipulation.
The paper's conclusion is direct: "A capability score is not a safety score, and no one agent-safety benchmark stands in for safety as a whole." Scand.ai
Outcome selectivity—the risk that benchmarks only measure safety on easy tasks—showed borderline evidence under organization-level resampling, with a p-value of 0.051, suggesting the problem is real but at the margin of statistical certainty. The researchers treated the four benchmarks as measurements to be validated, evaluating them under their official implementations and author-provided scorers arXiv.
The audit matters because enterprises and regulators increasingly rely on published benchmarks to assess whether agent deployments are safe. If those benchmarks conflate capability with safety, organizations may unknowingly deploy models that perform well on evaluation tasks but fail to resist real-world manipulation or refuse harmful instructions when stakes are high.
The findings suggest the agent-safety field lacks validated measurement tools—a problem that echoes earlier work questioning alignment benchmarks for large language models. Companies building agent systems for high-stakes domains (finance, healthcare, critical infrastructure) have no reliable way to measure whether their agents will behave safely under adversarial conditions. The audit does not propose replacements, but it documents that the current suite of benchmarks is insufficient as a standalone safety assurance mechanism.