---
title: "CWE-bench: Coding Agent Security Benchmark Launches"
slug: "cwe-bench-coding-agent-security-benchmark-launches"
published: "2026-09-28"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-09-28"
aiActArticle50: "compliant"
humanView: "https://agentry.news/research/cwe-bench-coding-agent-security-benchmark-launches"
agentView: "https://agentry.news/agent/cwe-bench-coding-agent-security-benchmark-launches"
---# CWE-bench: Coding Agent Security Benchmark Launches

> Collinear AI published CWE-bench on September 2, 2026, a held-out benchmark measuring how well coding agents can find and fix vulnerabilities across 100 real-world audit-and-patch tasks. The benchmark

*Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. [AI policy](/ai-policy).*

## New Benchmark Tests Coding Agents on Real Vulnerabilities

Collinear AI launched [CWE-bench](https://blog.collinear.ai/p/cwe-bench), a held-out benchmark for coding agents, on September 2, 2026. The benchmark measures agent performance across **100 audit-and-patch tasks** designed to test whether agents can identify and remediate known security flaws in production codebases, spanning **54 distinct weakness types**, **six programming languages**, and all **10 OWASP Top 10 2025 categories** [Collinear AI Blog](https://blog.collinear.ai/p/cwe-bench).

The benchmark's release included results from testing multiple frontier coding agents. Top performers clustered closely in capability: reported pass@1 scores showed leaders at **47.8%** and **47.2%**, with the next tier at **44.2%** and **44.0%** [BenchLM](https://benchlm.ai/benchmarks/cwebench). A significant finding emerged in the data: **18 of the 100 tasks remained unsolved by every model tested**, indicating a hard ceiling of agent capability on specific vulnerability classes [PRWeb](https://www.prweb.com/releases/collinear-ai-launches-cwe-bench-to-test-frontier-coding-agents-on-defensive-cybersecurity-capabilities-302868204.html).

## Accuracy and Cost Tradeoffs Surface in Results

Beyond raw accuracy numbers, the benchmark revealed a cost-performance frontier. While higher-scoring models dominated accuracy metrics, lower-cost model variants positioned themselves near the Pareto boundary, suggesting teams can optimize for deployment cost without catastrophic accuracy loss [BenchLM](https://benchlm.ai/benchmarks/cwebench). This framing matters for enterprises choosing which agent to integrate into security workflows where both remediation quality and operational expense are constraints.

The benchmark expanded the testable surface of agent behavior beyond individual vulnerability types. By anchoring to the OWASP Top 10 2025 framework and requiring solutions across six languages, CWE-bench created a common measurement ground for security-focused coding agents—a category of tools gaining traction as enterprises automate patch-and-audit cycles [Collinear AI Blog](https://blog.collinear.ai/p/cwe-bench).

## Why This Matters for Agent Developers and Security Teams

Coding agents are moving from proof-of-concept to production deployment in security workflows. A standardized benchmark allows teams to compare agents on real-world vulnerability discovery and remediation—not synthetic microbenchmarks or closed proprietary evaluations. The tight clustering of top performers suggests the frontier is maturing but also shows significant variance in how agents handle edge-case vulnerability patterns.

The 18 unsolved tasks point to a practical limitation: certain vulnerability patterns or language combinations remain resistant to current agent approaches. Security teams relying on agents for compliance-driven patching now have a concrete reference for what these agents can and cannot do.

Full benchmark and leaderboard details are available on [BenchLM](https://benchlm.ai/benchmarks/cwebench) and the [arXiv preprint](https://arxiv.org/abs/2609.15939v1).