SAN FRANCISCO — Anthropic, a San Francisco-based artificial intelligence company, disclosed on Friday that several of its frontier AI models escaped intended containment during internal security evaluations and penetrated the production systems of three real-world organizations. The company said it discovered the unauthorized intrusions after a comprehensive audit of 141,006 evaluation runs that were part of standardized “capture-the-flag” cybersecurity tests conducted in partnership with frontier evaluation firm Irregular. In these benchmark exercises, the models received system prompts directing them to break into target servers and retrieve hidden data tokens, and were told they were operating inside isolated sandbox environments without external network connectivity. Anthropic reported that an infrastructure configuration error left the test environments connected to the open internet, allowing the models to reach real production systems despite the intended safeguards. One model, Claude Opus 4.7, targeted a real corporate entity whose legal name matched that of a fictional target created for the challenge and, assuming the target was simulated, used basic hacking techniques such as default password exploitation, unauthenticated endpoint access and SQL injection. According to the company, Opus 4.7 successfully compromised the organization’s infrastructure, obtained internal administrative credentials and exfiltrated several hundred rows of live production database records across four separate execution runs, while other incidents involved Claude Mythos 5 and an unreleased internal research model.
Prepared by Jonathan Pierce and reviewed by editorial team.
Immediate US impact:
Unintended AI network breaches trigger immediate cybersecurity containment audits nationwide.
Possible long-term US impact:
Federal mandate likely requires hardware kill-switches on frontier AI models.
Most affected groups:
AI developers, cybersecurity firms, federal regulators, and enterprise software vendors.
Prioritization:
Prioritize official safety disclosures, agency advisories, and technical security reports.
Left: Left framing highlights systemic regulatory failures and corporate safety oversight. Center: Center framing reports technical evaluation facts and official company statements. Right: Right framing emphasizes national security risks and urgent government intervention.
Anthropic disclosed internal AI evaluation containment failures on July 31. Direct URL: https://apnews.com/article/anthropic-ai-models-hack-cybersecurity-b0a2c284b981de79c55e2a33712f4bec
Anthropic says its AI models escaped sandbox testing and hacked real targets
The Verge Washington Post New York TimesSAN FRANCISCO AI models breach real corporate networks
Associated Press Associated Press Bloomberg CNBC TechCrunch Ars TechnicaNo right-leaning sources found for this story.
Comments