KEX-bench evaluates coding agents on kernel exploits
Researchers from multiple institutions published KEX-bench, a benchmark for evaluating how coding agents generate kernel exploits, on September 22, 2026. The study—authored by Junyoung Jang, Gwanhyun Lee, Hwiwon Lee, Kyuheon Kim, Jongseong Kim, Jinho Jung, and Lingming Zhang—assessed agent performance on 45 task instances spanning 40 Linux and Windows CVEs.
Benchmark Design and Scope
KEX-bench runs each task in an isolated virtual machine using controlled tools and a deterministic verifier to measure agent capability on core exploit primitives. The assessed primitives include kernel address leakage, instruction-pointer control, heap reads, heap writes, and arbitrary-address writes—the foundational building blocks of kernel-level attacks.
The benchmark structure splits tasks between Linux (25 instances) and Windows (20 instances), allowing researchers to evaluate how agent performance varies across operating system architectures and attack surfaces.
Performance Findings
When tested without a reference proof of concept, the strongest reported configuration solved 14 of 25 Linux tasks (56.0%) and 1 of 20 Windows tasks (5.0%). This gap reveals significant platform-specific challenge differences: agents demonstrated substantially weaker performance on Windows kernel exploitation despite the conceptual similarity to Linux approaches.
Access to reference proof-of-concept code dramatically improved results. With PoC examples available, the same configuration solved 31 of 45 total tasks (68.9%)—indicating that in-context learning from working examples substantially improves agent reasoning on complex security tasks.
Implications for Agent Development
The KEX-bench findings surface a critical capability boundary in today's coding agents: while they can reason through exploit primitive generation with sufficient context, performance gaps on unfamiliar platforms and absence of reference implementations suggest agents lack robust generalization of kernel-level attack concepts.
The research contributes to the growing body of agent evaluation work that moves beyond hypothetical capabilities toward concrete, reproducible measurements. Unlike capability claims or roadmap projections, KEX-bench provides verifiable task success rates that developers and security teams can reference when assessing agent reliability in threat modeling, security research, and red-team automation workflows.
The benchmark is designed for reproducibility: tasks use deterministic verification within isolated VMs, eliminating ambiguity in success criteria and enabling future researchers to measure improvements or regressions as agent architectures and training methods evolve.