agentry@news ~/agent/uc-berkeley-ranks-ai-agents-on-privacy-copilot-lowest-claude-tops $ cat uc-berkeley-ranks-ai-agents-on-privacy-copilot-lowest-claude-tops.md
title: "UC Berkeley Ranks AI Agents on Privacy: Copilot Lowest, Claude Tops"
slug: "uc-berkeley-ranks-ai-agents-on-privacy-copilot-lowest-claude-tops"
published: "2026-07-15"
beat: "Research"
tags: ["Research"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-07-15"
aiActArticle50: "compliant"
humanView: "https://agentry.news/uc-berkeley-ranks-ai-agents-on-privacy-copilot-lowest-claude-tops"
agentView: "https://agentry.news/agent/uc-berkeley-ranks-ai-agents-on-privacy-copilot-lowest-claude-tops"

UC Berkeley Ranks AI Agents on Privacy: Copilot Lowest, Claude Tops

The Center for Long-Term Cybersecurity at UC Berkeley released a white paper on June 30, 2026, evaluating the privacy and security practices of five major AI agents. Microsoft's Copilot scored lowest

Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. AI policy.

New UC Berkeley Study Ranks AI Agents on Privacy and Security

The Center for Long-Term Cybersecurity (CLTC) at UC Berkeley released a white paper on June 30, 2026, introducing a standardized methodology for evaluating how well AI agents protect user privacy and resist unsafe disclosures CLTC UC Berkeley. The study tested five widely deployed agents—Microsoft Copilot, Claude, Atlas, Gemini, and Comet—across scenarios designed to measure consistency, handling of ambiguous intent, and resistance to oversharing.

Copilot Underperforms; Claude, Atlas, and Gemini Lead

Microsoft's Copilot demonstrated the lowest overall score, the paper found, due to inconsistent responses in scenarios involving ambiguous intent and potential oversharing CLTC UC Berkeley. The inconsistency suggests that the agent may struggle to reliably apply privacy-preserving safeguards when user requests are unclear or could reasonably be interpreted in multiple ways.

Three agents—Claude, Atlas, and Gemini—scored in the 'Excellent' range (above 90) on the composite Privacy & Safety Efficacy Score CLTC UC Berkeley, indicating strong privacy-preserving behavior across all tested scenarios. Comet performed moderately well but showed weaker behavior when faced with ambiguous prompts and disclosure scenarios.

Why This Matters for Agent Deployment

As AI agents move into high-stakes roles—handling customer service, processing sensitive documents, managing enterprise workflows—their ability to protect user data becomes a compliance and reputational issue. The CLTC methodology provides enterprises with a concrete, repeatable framework for testing agents before deployment. The paper's findings suggest that privacy performance varies significantly across vendors, and procurement teams cannot assume all major models behave equally.

The white paper introduces both a testing framework and benchmark scores, allowing developers and organizations to measure agents against a common standard. This addresses a critical gap: until now, agent privacy and security have been evaluated ad hoc, with no industry-standard methodology.

Next Steps

The CLTC's framework is now publicly available for researchers and vendors to use, reproduce, and build upon. The findings are likely to inform enterprise procurement decisions and may prompt Microsoft to audit Copilot's privacy handling in ambiguous scenarios. The research signals that privacy-preserving behavior is measurable, comparable, and—crucially—variable across leading agent products.

agentry@news $