title: "Ericsson study: agentic code review hits 96% accuracy" slug: "ericsson-study-agentic-code-review-hits-96-accuracy" published: "2026-10-11" beat: "Research" tags: ["Research"] creator: "Agentry Newsroom" editor: "Susanne Sperling, Editor — Human in the Loop" tools: ["Claude (Anthropic)", "Perplexity Sonar"] creativeWorkStatus: "verified" dateReviewed: "2026-10-11" aiActArticle50: "compliant" humanView: "https://agentry.news/research/ericsson-study-agentic-code-review-hits-96-accuracy" agentView: "https://agentry.news/agent/ericsson-study-agentic-code-review-hits-96-accuracy"
Researchers at Ericsson evaluated a multi-agent AI system for code review and reported 96% accuracy in identifying issues across several commits, with nearly 70% of flagged problems rated important by
Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.
Researchers affiliated with Ericsson evaluated a specialized multi-agent AI system for code review and reported 96% accuracy in identifying defects, according to a study submitted to arXiv on 14 September 2026.
The paper, authored by Muhammad Laiq, Ricardo Britto, Muhammad Usman, Nishrith Saini, and Deepika Badampudi, describes an industrial evaluation where agentic agents assessed code changes across four dimensions: readability, maintainability, reliability, and performance. The system deployed specialized agent skills paired with project-specific contextual knowledge to generate reviews for multiple code commits.
Developers at the Ericsson case company manually validated the system's output, confirming correctness and importance of flagged issues. The agents identified more than 200 issues across the investigated commits. Of the correctly identified issues, approximately 69% were rated important by developers—comprising roughly 33% severe issues requiring immediate fixes and 36% important issues that should be addressed.
The 96% accuracy figure reflects the proportion of identified issues that developers confirmed as legitimate defects, making this one of the first documented industrial evaluations of agentic code review at scale as reported in the arXiv record.
The study was accepted at the 27th International Conference on Product-Focused Software Process Improvement (PROFES 2026), indicating peer review validation of the methodology and findings. The research represents a concrete test of agent capabilities in a production-adjacent software engineering context—moving beyond benchmarks to real code and real developer feedback.
The separation of correctly identified issues into severity tiers is significant: the majority of agent-flagged problems (69% of those correct) had immediate or near-term business value to the organization, suggesting the system filtered noise effectively. This pattern suggests agentic code review may reduce developer review burden without sacrificing quality signals.
The Ericsson evaluation adds empirical weight to enterprise interest in agentic systems for knowledge work. Unlike conceptual case studies, this research required agents to generate multi-dimensional assessments of real production code, coordinate specialized knowledge, and produce outputs developers could validate. The 96% accuracy threshold and importance distribution offer benchmarks for competing agentic code review tools entering the market.
The research does not disclose deployment scope, review latency, or cost-benefit analysis relative to human reviewers, leaving open questions about production viability. However, the industrial setting and developer validation distinguish this from academic prototypes, making it a rare documented instance of agentic agents solving a specific developer workflow at measurable accuracy.