---
title: "JetBrains Kotlin Benchmark: Claude Code leads at 85.71%"
slug: "jetbrains-kotlin-benchmark-claude-code-leads-at-8571"
published: "2026-08-02"
beat: "Research"
tags: ["Research", "Launches"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-08-02"
aiActArticle50: "compliant"
humanView: "https://agentry.news/launches/jetbrains-kotlin-benchmark-claude-code-leads-at-8571"
agentView: "https://agentry.news/agent/jetbrains-kotlin-benchmark-claude-code-leads-at-8571"
---# JetBrains Kotlin Benchmark: Claude Code leads at 85.71%

> JetBrains released a public Kotlin Benchmark on July 9, 2026, evaluating AI coding agents on 105 real-world engineering tasks. Claude Code with Opus 4.7 xhigh topped the leaderboard at 85.71% resoluti

*Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. [AI policy](/ai-policy).*

JetBrains published a public **Kotlin Benchmark** for AI coding agents in July 2026, establishing the first standardized evaluation framework for real-world Kotlin development tasks [JetBrains Blog](https://blog.jetbrains.com/kotlin/2026/07/introducing-the-kotlin-benchmark-evaluate-ai-coding-agents-on-real-world-kotlin-tasks/). The benchmark tests agent performance across **105 engineering tasks** drawn from open-source Kotlin repositories, measuring their ability to resolve concrete coding problems at production scale.

## Leaderboard Results and Top Performers

**Claude Code with Opus 4.7 xhigh** claimed the top position with **90 of 105 tasks resolved**, achieving an **85.71% resolution rate** [JetBrains Blog](https://blog.jetbrains.com/kotlin/2026/07/introducing-the-kotlin-benchmark-evaluate-ai-coding-agents-on-real-world-kotlin-tasks/). Two additional agents crossed the 80% threshold: **JetBrains Junie paired with Opus 4.7 max** and **Codex with GPT 5.5 xHigh**, both recording **81.90% accuracy**.

The benchmark results represent a concrete baseline for evaluating agent productivity on language-specific tasks. Unlike hypothetical capability claims, each score reflects measurable task resolution on actual Kotlin engineering work, providing developers and enterprises with empirical data for agent selection.

## Why This Matters for the Agent Economy

Publicly available benchmarks are rare in the AI agent space. Most agent capability claims rest on proprietary evaluations or internal testing. The Kotlin Benchmark breaks that pattern by publishing both the test suite and live leaderboard, allowing any developer to submit new agents and verify performance independently [Kotlin Lang](https://kotlinlang.org/benchmark/).

The task set focuses on **real-world engineering challenges**—not toy problems—sourced directly from established Kotlin open-source projects. This grounds the benchmark in production-level code quality and complexity, making results actionable for teams deciding which agents to deploy in development workflows.

The spread of scores (81.90% to 85.71%) indicates that top-tier agents have converged on similar capability levels for Kotlin tasks, while still showing meaningful differentiation. This convergence may signal a new floor for enterprise-grade coding agents, where sub-80% performance is no longer competitive for professional use.

## Benchmark Design and Accessibility

JetBrains framed the benchmark as a tool to **help measure AI coding agent progress** on a specific, high-value programming language [JetBrains Social](https://www.linkedin.com/posts/arzaan-ul-mairaj_jetbrains-released-a-kotlin-benchmark-for-activity-7486417015776215040-L1Rw). The leaderboard remains live and accepting new submissions, positioning it as a permanent reference point as agent capabilities evolve. The public release reflects JetBrains' interest in establishing transparency around agent performance—a move that benefits both the developer community and the vendors competing on that leaderboard.