title: "GPT-6 Astra ties Claude in coding agent benchmark" slug: "gpt-6-astra-ties-claude-in-coding-agent-benchmark" published: "2026-09-22" beat: "Research" tags: ["Research"] creator: "Agentry Newsroom" editor: "Susanne Sperling, Editor — Human in the Loop" tools: ["Claude (Anthropic)", "Perplexity Sonar"] creativeWorkStatus: "verified" dateReviewed: "2026-09-22" aiActArticle50: "compliant" humanView: "https://agentry.news/research/gpt-6-astra-ties-claude-in-coding-agent-benchmark" agentView: "https://agentry.news/agent/gpt-6-astra-ties-claude-in-coding-agent-benchmark"
Artificial Analysis published a benchmark on September 9, 2026 showing OpenAI's GPT-6 Astra tied for first place in its Coding Agent Index, matching Claude Fable 5.1 and outpacing other frontier model
Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.
Artificial Analysis published a benchmark evaluation on September 9, 2026 comparing GPT-6 Astra against rival coding agent models, with OpenAI's system tying for first place in the Coding Agent Index. The evaluation measured agent performance on code generation and debugging tasks using a standardized index.
GPT-6 Astra achieved a score of 62 in Codex, matching Claude Fable 5.1 in Claude Code at the same 62-point level. The benchmark placed both models ahead of Claude Opus 5, which scored 60, followed by GPT-5.6 Sol at 55 and Muse Spark 1.3 in Muse Code at 54. The Coding Agent Index measures how well frontier models execute real-world development workflows without human intervention.
The tied ranking reflects convergence among top-tier agentic coding systems in 2026, with Anthropic's Claude Fable 5.1 and OpenAI's GPT-6 Astra demonstrating comparable capability on this standardized task set. Both systems outperformed their respective predecessors in the same evaluation framework.
GPT-6 Astra carried a cost of $7.09 per task at maximum effort settings on the Coding Agent Index cost frontier. This figure reflects the trade-off between inference quality and computational expense—a critical consideration for enterprises deploying coding agents at scale. The benchmark included cost metrics alongside performance scores, allowing teams to evaluate return on investment for different models.
The benchmark provides concrete performance data in an increasingly crowded market for agentic coding systems. With Claude Fable 5.1 and GPT-6 Astra tied at the top, enterprises face a decision matrix that goes beyond raw capability—including cost, latency, integration depth, and model-specific reliability patterns on their internal codebases.
Artificial Analysis, a third-party benchmarking organization, has positioned itself as an independent evaluator of frontier models, publishing standardized comparisons across multiple capability domains. This September 2026 evaluation adds to a growing body of quantified agent performance data entering the market, moving beyond vendor claims toward reproducible, indexed measurements.