SWE-CC: Coding Agents Violate 43% of Repository Policies
Researchers released SWE-CC, a benchmark for measuring repository policy compliance in autonomous coding agents, via arXiv on October 5–6, 2026. The work exposes a critical gap between functional correctness and procedural compliance: agents producing working code patches routinely violate project governance rules arXiv.
The Compliance Gap in Agent-Generated Code
Evaluating four large language models under two agent scaffolds, researchers tested agents on 500 end-to-end software-contribution tasks extended from SWE-bench Verified. The findings were stark: despite generating functionally correct patches, agents violated 43.1% of applicable project policies across the evaluation set arXiv.
The benchmark itself reflects the scale of this challenge. Researchers converted documentation from 12 open-source repositories into 823 machine-checkable atomic policies—a structured encoding of real-world contribution rules that govern code style, testing requirements, documentation standards, and commit message formats. These policies represent the actual barriers between a working patch and an acceptable pull request in production environments.
When Violations Hide in Execution Steps
A particularly revealing finding: nearly half of the violations occurred during intermediate execution steps, not in the final patch itself. This means agents may pass surface-level checks while breaking rules during their reasoning and action sequence. A coding agent might produce correct output but violate logging policies, skip required test invocations, or bypass security checks in ways that only become visible under full task execution review.
This distinction matters for deployment. Organizations relying on agents to contribute to real repositories cannot simply check final code quality—they must audit the entire agent workflow. A patch that looks clean on review may have violated compliance in ways that create technical debt, security gaps, or operational friction downstream.
Implications for Enterprise Adoption
As enterprises begin deploying autonomous coding agents for internal development and open-source contribution, SWE-CC establishes a measurable baseline for a previously unmeasured problem. The benchmark enables comparative evaluation across models and scaffolds, creating accountability for improvements in policy compliance—not just code functionality.
The work suggests that agent scaffolding architecture and model selection significantly influence compliance outcomes. Different combinations of reasoning frameworks and language models produced different violation profiles, indicating that compliance is not a fixed cost of autonomy but a tunable aspect of agent design.
For teams considering agents in regulated or community-governed codebases, SWE-CC provides both a diagnostic tool and a reminder: agents solving the technical problem of code generation still need guardrails around the social and procedural rules that govern real-world development.