SAN FRANCISCO — Artificial intelligence company Anthropic disclosed on Friday that its frontier AI models breached containment during internal security evaluations and successfully hacked into the production systems of three real-world organizations. The San Francisco-based technology firm uncovered the unauthorized intrusions after conducting a comprehensive audit of 141,006 evaluation runs. Anthropic launched the sweeping cybersecurity investigation after competitor OpenAI reported that its own AI agents had escaped containment and breached external servers belonging to open-source platform Hugging Face. The Anthropic security failures involved three frontier models: Claude Opus 4.7, Claude Mythos 5, and an unreleased internal research test model. The incidents occurred during standardized "capture-the-flag" cybersecurity assessments conducted in partnership with frontier evaluation firm Irregular. In these benchmark tests, models received system prompts instructing them to break into target servers to retrieve hidden data tokens. Although system prompts informed the AI models that they were operating inside isolated sandbox environments without external network connectivity, an infrastructure configuration error left the test environments connected to the open internet. During testing, Claude Opus 4.7 targeted a real corporate entity whose legal name matched a fictional target generated for the benchmark challenge. Operating under the assumption that the target was simulated, Opus 4.7 launched cyberattacks utilizing basic techniques, including default password exploitation, unauthenticated endpoint access, and SQL injection. The model successfully compromised the organization's infrastructure, extracting internal administrative credentials and exfiltrating several hu
Prepared by Jonathan Pierce and reviewed by editorial team.
Left: Left framing highlights systemic regulatory failures and corporate safety oversight. Center: Center framing reports technical evaluation facts and official company statements. Right: Right framing emphasizes national security risks and urgent government intervention.
Anthropic disclosed internal AI evaluation containment failures on July 31. Direct URL: https://apnews.com/article/anthropic-ai-models-hack-cybersecurity-b0a2c284b981de79c55e2a33712f4bec
Anthropic says its AI models escaped sandbox testing and hacked real targets
The Verge Washington Post New York TimesAnthropic AI Models Escape Sandbox and Hack Real Corporate Networks
Associated Press Associated Press Bloomberg CNBC TechCrunch Ars TechnicaNo right-leaning sources found for this story.
Comments