title: "OpenAI's GPT-5.6 Sol tops coding-agent benchmark at 80" slug: "openais-gpt-56-sol-tops-coding-agent-benchmark-at-80" published: "2026-07-31" beat: "Research" tags: ["Research", "Launches"] creator: "Agentry Newsroom" editor: "Susanne Sperling, Editor — Human in the Loop" tools: ["Claude (Anthropic)", "Perplexity Sonar"] creativeWorkStatus: "verified" dateReviewed: "2026-07-31" aiActArticle50: "compliant" humanView: "https://agentry.news/launches/openais-gpt-56-sol-tops-coding-agent-benchmark-at-80" agentView: "https://agentry.news/agent/openais-gpt-56-sol-tops-coding-agent-benchmark-at-80"
OpenAI announced that GPT-5.6 Sol set a new state of the art on the Artificial Analysis Coding Agent Index, scoring 80 points and 64.6% on SWE-Bench Pro, outpacing Claude Fable 5 while using roughly o
Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.
OpenAI announced that GPT-5.6 Sol achieved a new state of the art on the Artificial Analysis Coding Agent Index, scoring 80 points and posting 64.6% on SWE-Bench Pro, according to OpenAI's official product page. The result positions the model ahead of competing frontier agents in standardized coding-task evaluation.
GPT-5.6 Sol's score of 80 on the Artificial Analysis Coding Agent Index v1.1 marks the highest performance recorded on that benchmark. Artificial Analysis independently confirmed that "GPT-5.6 Sol (max) leads the Artificial Analysis Coding Agent Index at 80 points." The model scored 64.6% on SWE-Bench Pro, a separate evaluation measuring performance on real-world software engineering tasks.
By contrast, Claude Fable 5—the previous leading model on the benchmark—achieved 77.2 on the Artificial Analysis Coding Agent Index and 80% on SWE-Bench Pro, according to OpenAI's comparison table.
OpenAI emphasized efficiency gains alongside the performance jump. The company stated that GPT-5.6 Sol operates "using less than half the output tokens, taking less than half the time, and costing about one-third less" than competing models on the same benchmarks. This combination of higher accuracy and lower operational cost positions GPT-5.6 Sol as a significant shift in the cost-performance frontier for agentic coding tasks.
The release reflects intensifying competition in the frontier model space, where both OpenAI and Anthropic are shipping models with explicit agentic coding capabilities. Coding-agent benchmarks have become key performance indicators for enterprise adoption, particularly among software development teams evaluating autonomous assistance tools.
The Artificial Analysis Coding Agent Index measures how well models solve multi-step coding problems in real-world conditions, while SWE-Bench Pro evaluates performance on software engineering tasks extracted from actual GitHub repositories. Both are widely recognized metrics in the AI agent developer community.
OpenAI's announcement did not include a specific publication date, though the product page is live as of July 31, 2026.