AGENTRY.NEWSWhat AI Agents Do, Documented.September 28, 2026

Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.

CWE-bench: Coding Agent Security Benchmark Launches

By
Agentry Newsroom
Published

New Benchmark Tests Coding Agents on Real Vulnerabilities

Collinear AI launched CWE-bench, a held-out benchmark for coding agents, on September 2, 2026. The benchmark measures agent performance across 100 audit-and-patch tasks designed to test whether agents can identify and remediate known security flaws in production codebases, spanning 54 distinct weakness types, six programming languages, and all 10 OWASP Top 10 2025 categories Collinear AI Blog.

The benchmark's release included results from testing multiple frontier coding agents. Top performers clustered closely in capability: reported pass@1 scores showed leaders at 47.8% and 47.2%, with the next tier at 44.2% and 44.0% BenchLM. A significant finding emerged in the data: 18 of the 100 tasks remained unsolved by every model tested, indicating a hard ceiling of agent capability on specific vulnerability classes PRWeb.

Accuracy and Cost Tradeoffs Surface in Results

Beyond raw accuracy numbers, the benchmark revealed a cost-performance frontier. While higher-scoring models dominated accuracy metrics, lower-cost model variants positioned themselves near the Pareto boundary, suggesting teams can optimize for deployment cost without catastrophic accuracy loss BenchLM. This framing matters for enterprises choosing which agent to integrate into security workflows where both remediation quality and operational expense are constraints.

The benchmark expanded the testable surface of agent behavior beyond individual vulnerability types. By anchoring to the OWASP Top 10 2025 framework and requiring solutions across six languages, CWE-bench created a common measurement ground for security-focused coding agents—a category of tools gaining traction as enterprises automate patch-and-audit cycles Collinear AI Blog.

Why This Matters for Agent Developers and Security Teams

Coding agents are moving from proof-of-concept to production deployment in security workflows. A standardized benchmark allows teams to compare agents on real-world vulnerability discovery and remediation—not synthetic microbenchmarks or closed proprietary evaluations. The tight clustering of top performers suggests the frontier is maturing but also shows significant variance in how agents handle edge-case vulnerability patterns.

The 18 unsolved tasks point to a practical limitation: certain vulnerability patterns or language combinations remain resistant to current agent approaches. Security teams relying on agents for compliance-driven patching now have a concrete reference for what these agents can and cannot do.

Full benchmark and leaderboard details are available on BenchLM and the arXiv preprint.

Del dette opslag: