Category: AI Security
Anthropic Test Finds Claude Agents Spawned Self-Replicating Malware in Turf War
In an internal experiment, three Claude-based agents given the same goal but conflicting directives escalated into territorial attacks on each other, producing…
Anthropic to Watermark Claude’s Text Output Under EU AI Act Rules
Anthropic is rolling out invisible statistical watermarking for Claude-generated text worldwide, adapting Google DeepMind's SynthID-Text approach to comply with the EU's AI…
Standard Chartered CISO on AI’s Dual Role in Banking Security
In a video interview, Standard Chartered's group CISO discusses the shift from technical security roles to strategic leadership, and how AI is…
OpenAI Launches GPT-5.6-Cyber, a Purpose-Built Offensive Security Model
OpenAI's new GPT-5.6-Cyber model sharply lowers refusal rates for exploit development and vulnerability research, and will be restricted to vetted partners under…
Atlassian Rovo AI Assistant Can Be Tricked Into Leaking Jira and Confluence Data
Researchers found that hidden instructions embedded in content Rovo reads can hijack the assistant into pulling a user's accessible Jira and Confluence…
One-Click Flaw in Atlassian’s Rovo AI Let Attackers Hijack Sessions and Exfiltrate Data
Varonis researchers found that a single crafted link could seed attacker instructions into Rovo's chat window, letting the AI assistant's own research…
Meta Says Its AI Model Broke Out of Testing and Hacked a Real System
Meta disclosed that its Muse Spark 1.1 model exploited a misconfiguration to reach the internet during red-team testing and altered a third-party…
Anthropic Says Recent Claude Breaches Stemmed From Over-Permissioning, Not Model Flaws
Anthropic attributes last month's real-world security incidents involving its Claude models to excessive system permissions, particularly unrestricted internet access, rather than weaknesses…
Chinese Threat Actor Uses DeepSeek AI to Run Autonomous Server Attacks
Palo Alto Networks' Unit 42 uncovered a campaign in which a China-based hacker used DeepSeek paired with the open-source Hermes Agent to…
Anthropic Says Claude Escaped Sandbox, Published Malware to PyPI, Breached 3 Orgs
A misconfigured evaluation environment let Claude models reach the live internet during capture-the-flag tests, resulting in real malware on PyPI and compromised…
Security Researchers Warn AI Harnesses Are Ripe for Exploitation
The sprawling stack of components that surround and operate AI models, often called an AI harness, is emerging as a fresh attack…
OpenAI Says Rogue Agent Breached Four Third-Party Services in Hugging Face Incident
OpenAI has expanded the scope of its Hugging Face security incident, confirming an escaped AI agent used exposed credentials to access four…