Embedded coding agents benchmarked on closed-loop iteration
A closed-loop evaluation framework for coding agents in embedded software development surfaced via arXiv on October 8–9, 2026, establishing concrete metrics for how well language models can implement hardware-adjacent code under real constraints.
Benchmark Design and Task Structure
The benchmark presents agents with five embedded-control tasks, each pairing a plain-text engineering description with a constrained workspace and observable build-and-runtime interface. Agents must implement requirements, perform self-verification, and iterate until the device exhibits required behavior—mirroring real embedded-systems workflows where code correctness is validated by physical or simulated hardware response.
The evaluation employs four distinct feedback scenarios: one-shot code generation (no iteration), realistic self-verification (agent observes build logs and test output), CI-style red/green feedback (pass/fail signals only), and oracle-style detailed feedback (complete diagnostic information). This design isolates how different feedback modalities affect agent performance and search efficiency.
Evaluation Results and Model Performance
Researchers evaluated seven GPT-family and Qwen-family configurations across the five tasks and four scenarios, with three repetitions per condition, totaling 420 runs. The test results showed clear performance stratification: GPT-5.4 achieved the highest pass rate but did not attain a perfect benchmark score, while Qwen3.5-27B emerged as the strongest observed local model. Smaller local models demonstrated sharply lower pass rates and reduced search efficiency, indicating that closed-loop embedded development places significant demands on model scale and instruction-following capability.
The benchmark's structure—requiring agents to read engineering specs, write code, observe real device behavior, and adapt—differs from static code-completion or API-call benchmarks. By grounding evaluation in actual hardware feedback loops, the work quantifies a gap between current model capabilities and production-ready embedded-systems development, where iteration cycles are constrained by physical deployment time and debugging information is sparse.
Implications for Agent Tooling and Embedded Development
The framework addresses a practical developer need: embedded software is inherently harder to automate than web or cloud code because correctness depends on device state, timing, and real-world physics. Traditional LLM benchmarks measure syntax or API recall; this one measures agent agency—the ability to plan, execute, observe, and adapt in a constrained feedback loop.
The results suggest that near-term embedded-coding agents will depend on high-capability model families (GPT-5.x scale) or substantial local-model scaling, and that feedback quality (oracle versus red/green) significantly shapes agent success. The publicly documented benchmark creates a reproducible standard for future agent tool development and model training aimed at hardware-adjacent workflows.