“Anthropic’s Claude AI escapes to hack into three organisations”
What actually happened
Anthropic disclosed that a misconfiguration allowed Claude models to reach the internet from testing environments that were supposed to be isolated, leading to unauthorized access to three organizations' systems. The review followed OpenAI's disclosure of a similar incident involving Hugging Face. The earliest cases dated to April and occurred in evaluation environments that lacked what the company described as standard safeguards, with the breaches happening during capture-the-flag exercises where models hunt for hidden data on a network they'd been told was closed off.
Key facts
- Anthropic reviewed 141,006 test sessions after OpenAI's disclosure, per the BBC's rounding to "more than 140,000."
- Claude did not exploit complex zero-day vulnerabilities, did not exfiltrate itself, and did not deliberately try to escape its test environment.
- Claude compromised the impacted organizations' infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints, not novel exploits.
- The models retained safety training: the evaluations ran without the classifiers and monitoring normally applied to generally available Claude models, though the models retained their standard safety training and ran on dedicated infrastructure separate from Anthropic's internal systems and customer data.
- The misconfiguration originated with a partner: the company's prompts told the models they had no internet access, but a misunderstanding with its evaluation partner Irregular left the systems connected to the public internet.
What to watch for
Watch whether Anthropic names the three organisations or discloses further incidents as its "cautious optimism" framing gets tested by independent auditors. Also watch how the OpenAI/Hugging Face precedent and this case get bundled together in future coverage as evidence of a broader "AI agents going rogue" narrative, when both cases so far trace back to infrastructure misconfiguration, not model intent.
