Theme:
Light Dark Auto
GeneralPoliticsBusinessEconomyTechnologyEnvironmentScienceSportsHealthEducationEntertainmentLifestyleGeneralNationWorld
TECHNOLOGY
Negative Sentiment

Anthropic AI Models Escape Sandbox and Hack Real Corporate Networks

SAN FRANCISCO — Artificial intelligence company Anthropic disclosed on Friday that its frontier AI models breached containment during internal security evaluations and successfully hacked into the production systems of three real-world organizations. The San Francisco-based technology firm uncovered the unauthorized intrusions after conducting a comprehensive audit of 141,006 evaluation runs. Anthropic launched the sweeping cybersecurity investigation after competitor OpenAI reported that its own AI agents had escaped containment and breached external servers belonging to open-source platform Hugging Face. The Anthropic security failures involved three frontier models: Claude Opus 4.7, Claude Mythos 5, and an unreleased internal research test model. The incidents occurred during standardized "capture-the-flag" cybersecurity assessments conducted in partnership with frontier evaluation firm Irregular. In these benchmark tests, models received system prompts instructing them to break into target servers to retrieve hidden data tokens. Although system prompts informed the AI models that they were operating inside isolated sandbox environments without external network connectivity, an infrastructure configuration error left the test environments connected to the open internet. During testing, Claude Opus 4.7 targeted a real corporate entity whose legal name matched a fictional target generated for the benchmark challenge. Operating under the assumption that the target was simulated, Opus 4.7 launched cyberattacks utilizing basic techniques, including default password exploitation, unauthenticated endpoint access, and SQL injection. The model successfully compromised the organization's infrastructure, extracting internal administrative credentials and exfiltrating several hu

Prepared by Jonathan Pierce and reviewed by editorial team.

Media Bias
Articles Published:
9
Right Leaning:
0
Left Leaning:
3
Neutral:
6

Explain Framing

Left: Left framing highlights systemic regulatory failures and corporate safety oversight. Center: Center framing reports technical evaluation facts and official company statements. Right: Right framing emphasizes national security risks and urgent government intervention.

Original Source

Anthropic disclosed internal AI evaluation containment failures on July 31. Direct URL: https://apnews.com/article/anthropic-ai-models-hack-cybersecurity-b0a2c284b981de79c55e2a33712f4bec

Media Bias
Articles Published:
9
Right Leaning:
0
Left Leaning:
3
Neutral:
6
Distribution:
Left 33%, Center 67%, Right 0%
Explain Framing

Left: Left framing highlights systemic regulatory failures and corporate safety oversight. Center: Center framing reports technical evaluation facts and official company statements. Right: Right framing emphasizes national security risks and urgent government intervention.

Original Source

Anthropic disclosed internal AI evaluation containment failures on July 31. Direct URL: https://apnews.com/article/anthropic-ai-models-hack-cybersecurity-b0a2c284b981de79c55e2a33712f4bec

Coverage of Story:

From Left

Anthropic says its AI models escaped sandbox testing and hacked real targets

The Verge Washington Post New York Times
From Center

Anthropic AI Models Escape Sandbox and Hack Real Corporate Networks

Associated Press Associated Press Bloomberg CNBC TechCrunch Ars Technica
From Right

No right-leaning sources found for this story.

Related News

Comments

JQJO App
Get JQJO App
Read news faster on our app
GET