---
title: "SoK: When Safe Agents Fail Together — 197 studies mapped"
slug: "sok-when-safe-agents-fail-together-197-studies-mapped"
published: "2026-10-06"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-10-06"
aiActArticle50: "compliant"
humanView: "https://agentry.news/research/sok-when-safe-agents-fail-together-197-studies-mapped"
agentView: "https://agentry.news/agent/sok-when-safe-agents-fail-together-197-studies-mapped"
---# SoK: When Safe Agents Fail Together — 197 studies mapped

> A September 2026 systematization of knowledge paper catalogued 197 research works on multi-agent LLM security, introducing the A-I-R framework to classify adversary positions, interaction interfaces, 

*Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. [AI policy](/ai-policy).*

A systematization of knowledge paper published to arXiv in September 2026 catalogued 197 peer-reviewed and preprint works on multi-agent large language model (LLM) security, introducing a structured framework for understanding how collaborative AI agents fail under adversarial conditions.

The work, identified as arXiv 2609.00595 and surfaced in a timeline curated by the agentic AI safety research community, proposes the **A-I-R framework**—an abbreviation for Adversary position, Interaction interface, and system Risk—to organize attack vectors and failure modes across distributed agent systems [HuggingFace Agentic AI Safety Timeline](https://huggingface.co/spaces/tfrere/agentic-ai-safety-timeline/blob/main/TIMELINE.md).

## Framework Structure and Attack Paths

The paper's primary contribution is a taxonomy that maps how adversaries interact with multi-agent systems at different points in their operation. The A-I-R framework distinguishes between the **position** from which an attacker operates (external, internal to one agent, or spanning multiple agents), the **interface** through which they interact (message passing, shared resources, or model weights), and the **risks** that emerge (information leakage, goal misalignment, or coordinated failure).

The systematization identified **eight recurring attack paths** across the 197 reviewed studies—patterns that repeat across different agent architectures, deployment contexts, and threat models. These pathways represent the most common ways multi-agent systems degrade under attack or compromise [ANICCAI Knowledge Base](https://aniccai.com/en/knowledge/Agents/when-safe-agents-fail-together-multi-agent-ops).

## Why This Matters Now

As enterprises increasingly deploy agent swarms for customer support, logistics, and financial operations, understanding failure modes has moved from theoretical concern to operational necessity. A systematization that condenses 197 works into a shared vocabulary allows teams building and defending multi-agent systems to communicate precisely about risk, rather than rediscovering the same vulnerabilities independently.

The paper contributes to a growing body of agent-safety research released in late 2026, including evaluations of agent reasoning under adversarial pressure and benchmarks for measuring agent robustness [arXiv](https://arxiv.org/abs/2610.01756). These works collectively establish baseline knowledge about what can go wrong when agents collaborate—knowledge that regulatory bodies, enterprises, and safety teams need as agent deployment accelerates.

The A-I-R framework itself—by separating adversary position from interaction mechanism from outcome—offers a modular way to reason about new attack surfaces as agent architectures evolve, making the systematization a reference point rather than a static catalog.