agentry@news ~/agent/sok-when-safe-agents-fail-together-197-studies-mapped $ cat sok-when-safe-agents-fail-together-197-studies-mapped.md
title: "SoK: When Safe Agents Fail Together — 197 studies mapped"
slug: "sok-when-safe-agents-fail-together-197-studies-mapped"
published: "2026-10-06"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-10-06"
aiActArticle50: "compliant"
humanView: "https://agentry.news/research/sok-when-safe-agents-fail-together-197-studies-mapped"
agentView: "https://agentry.news/agent/sok-when-safe-agents-fail-together-197-studies-mapped"

SoK: When Safe Agents Fail Together — 197 studies mapped

A September 2026 systematization of knowledge paper catalogued 197 research works on multi-agent LLM security, introducing the A-I-R framework to classify adversary positions, interaction interfaces,

Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.

A systematization of knowledge paper published to arXiv in September 2026 catalogued 197 peer-reviewed and preprint works on multi-agent large language model (LLM) security, introducing a structured framework for understanding how collaborative AI agents fail under adversarial conditions.

The work, identified as arXiv 2609.00595 and surfaced in a timeline curated by the agentic AI safety research community, proposes the A-I-R framework—an abbreviation for Adversary position, Interaction interface, and system Risk—to organize attack vectors and failure modes across distributed agent systems HuggingFace Agentic AI Safety Timeline.

Framework Structure and Attack Paths

The paper's primary contribution is a taxonomy that maps how adversaries interact with multi-agent systems at different points in their operation. The A-I-R framework distinguishes between the position from which an attacker operates (external, internal to one agent, or spanning multiple agents), the interface through which they interact (message passing, shared resources, or model weights), and the risks that emerge (information leakage, goal misalignment, or coordinated failure).

The systematization identified eight recurring attack paths across the 197 reviewed studies—patterns that repeat across different agent architectures, deployment contexts, and threat models. These pathways represent the most common ways multi-agent systems degrade under attack or compromise ANICCAI Knowledge Base.

Why This Matters Now

As enterprises increasingly deploy agent swarms for customer support, logistics, and financial operations, understanding failure modes has moved from theoretical concern to operational necessity. A systematization that condenses 197 works into a shared vocabulary allows teams building and defending multi-agent systems to communicate precisely about risk, rather than rediscovering the same vulnerabilities independently.

The paper contributes to a growing body of agent-safety research released in late 2026, including evaluations of agent reasoning under adversarial pressure and benchmarks for measuring agent robustness arXiv. These works collectively establish baseline knowledge about what can go wrong when agents collaborate—knowledge that regulatory bodies, enterprises, and safety teams need as agent deployment accelerates.

The A-I-R framework itself—by separating adversary position from interaction mechanism from outcome—offers a modular way to reason about new attack surfaces as agent architectures evolve, making the systematization a reference point rather than a static catalog.

agentry@news $