---
title: "CivBench benchmark evaluates LLM agents on 300+ turn tasks"
slug: "civbench-benchmark-evaluates-llm-agents-on-300-turn-tasks"
published: "2026-09-29"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-09-29"
aiActArticle50: "compliant"
humanView: "https://agentry.news/research/civbench-benchmark-evaluates-llm-agents-on-300-turn-tasks"
agentView: "https://agentry.news/agent/civbench-benchmark-evaluates-llm-agents-on-300-turn-tasks"
---# CivBench benchmark evaluates LLM agents on 300+ turn tasks

> Researchers released CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through Civilization VI and the Model Context Protocol. The ben

*Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. [AI policy](/ai-policy).*

## Research team publishes CivBench for long-horizon agent evaluation

Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty, Harry Coppock, Jakob Nicolaus Foerster, and Rui Ponte Costa released [CivBench](https://arxiv.org/abs/2609.02459), an open-source benchmark for evaluating language model agents on complex, multi-step tasks in Civilization VI. The arXiv preprint was [submitted September 2, 2026](https://arxiv.org/abs/2609.02459), introducing a new methodology for testing agent performance in long-horizon, tool-mediated environments.

## Benchmark design and scope

CivBench measures agent capability through the Model Context Protocol (MCP), a framework that standardizes how agents interact with external tools. Each episode in the benchmark spans **300+ turns**, requiring agents to sustain performance across extended sequences of decisions and tool calls. The environment provides **76 MCP tools** available to agents during evaluation, covering diplomacy, military strategy, resource management, and infrastructure planning functions within the Civilization VI game engine.

The researchers released not only the benchmark environment itself, but also the underlying scenarios, execution logs, metrics, and analysis pipeline—enabling reproducible evaluation and comparative analysis across different language models and agent architectures. This comprehensive release approach allows other research teams and practitioners to validate findings and develop agents against standardized conditions.

## Implications for agent evaluation

CivBench addresses a documented gap in agent research: most public benchmarks test agents on short-horizon tasks with limited tool availability. Long-horizon evaluation—where an agent must plan and execute across dozens or hundreds of steps—requires different capabilities than single-turn problem-solving. By grounding evaluation in a complex, rule-based simulation where success depends on strategic planning, resource allocation, and adaptive decision-making, CivBench provides a more realistic assessment of agent behavior in real-world deployments.

The use of Civilization VI as the evaluation domain creates a standardized, deterministic environment where agent actions have measurable consequences: territory control, resource stockpiles, military units, and diplomatic relationships can be precisely quantified. This differs from natural-language benchmarks where "success" may involve subjective interpretation.

The benchmark's architecture—using MCP as the tool-mediation layer—also allows evaluation of how well agents compose multiple tools to achieve goals, a key capability for production systems. Agents must learn not just *what* tools exist, but *when* and *how* to combine them across extended time horizons.

As of September 2026, CivBench joins a growing set of specialized benchmarks for evaluating agent capabilities beyond single-turn instruction-following, supporting the broader research effort to understand and measure AI agent reliability in complex, long-duration tasks.