Category: AI Security
OpenAI Launches GPT-5.6-Cyber, a Purpose-Built Offensive Security Model
OpenAI's new GPT-5.6-Cyber model sharply lowers refusal rates for exploit development and vulnerability research, and will be restricted to vetted partners under…
Atlassian Rovo AI Assistant Can Be Tricked Into Leaking Jira and Confluence Data
Researchers found that hidden instructions embedded in content Rovo reads can hijack the assistant into pulling a user's accessible Jira and Confluence…
One-Click Flaw in Atlassian’s Rovo AI Let Attackers Hijack Sessions and Exfiltrate Data
Varonis researchers found that a single crafted link could seed attacker instructions into Rovo's chat window, letting the AI assistant's own research…
Meta Says Its AI Model Broke Out of Testing and Hacked a Real System
Meta disclosed that its Muse Spark 1.1 model exploited a misconfiguration to reach the internet during red-team testing and altered a third-party…
Anthropic Says Recent Claude Breaches Stemmed From Over-Permissioning, Not Model Flaws
Anthropic attributes last month's real-world security incidents involving its Claude models to excessive system permissions, particularly unrestricted internet access, rather than weaknesses…
Chinese Threat Actor Uses DeepSeek AI to Run Autonomous Server Attacks
Palo Alto Networks' Unit 42 uncovered a campaign in which a China-based hacker used DeepSeek paired with the open-source Hermes Agent to…
Anthropic Says Claude Escaped Sandbox, Published Malware to PyPI, Breached 3 Orgs
A misconfigured evaluation environment let Claude models reach the live internet during capture-the-flag tests, resulting in real malware on PyPI and compromised…
Security Researchers Warn AI Harnesses Are Ripe for Exploitation
The sprawling stack of components that surround and operate AI models, often called an AI harness, is emerging as a fresh attack…
OpenAI Says Rogue Agent Breached Four Third-Party Services in Hugging Face Incident
OpenAI has expanded the scope of its Hugging Face security incident, confirming an escaped AI agent used exposed credentials to access four…
Autonomous AI Agent in “YOLO Mode” Used to Spy on Thailand’s Finance Ministry
Researchers found an open-source AI agent operating without human oversight to conduct reconnaissance and credential theft inside Thailand's Ministry of Finance, after…
Anthropic’s Opus 5 Nearly Matches Mythos 5 at Finding Bugs, Lags on Exploits
Anthropic's new cheaper Claude Opus 5 model rivals its top-tier Mythos 5 system at spotting vulnerabilities, but a deliberate lack of offensive-task…
AI Guardrails Falter Across Europe’s Many Languages, Researchers Warn
Security testing shows that jailbreak protections built into AI products are inconsistent across languages, leaving gaps that attackers can exploit in Europe's…