Anthropic Says Claude Breached Three Companies During Security Tests
Anthropic has revealed that one of its AI models was able to access live systems belonging to three organizations during internal cybersecurity tests.
In a blog post published Thursday, the company said an internal review uncovered three separate incidents where Claude reached the public internet from inside a testing environment and then gained unauthorized access to production systems. The disclosure comes shortly after OpenAI reported that one of its unreleased models breached Hugging Face’s systems during internal testing.
Anthropic said the issue began with a misconfiguration in an evaluation setup run with third-party partner Irregular. In a statement, the company described the problem as a misunderstanding over whether the test environment had internet access. Anthropic said it was treating the issue as its own responsibility while Irregular conducts a separate investigation.
The company reviewed 141,006 evaluation runs and found that three of them exposed Claude to the internet. Those incidents involved three different models: Opus 4.7, Mythos 5, and an internal research test model. In each case, Claude had been explicitly told it had no internet access.
Anthropic said the models responded differently once they appeared to be interacting with real systems. Opus 4.7 recognized that it had reached a production environment but continued anyway, including retrieving credentials and interacting with a production database. Mythos 5 also appeared to realize it was on the real internet, but then convinced itself it was still in a simulation and went on to publish a malicious package to PyPI before the activity was detected. The newest internal research model stopped after concluding the target was real.
Anthropic said the findings show that powerful AI systems need stronger controls during security evaluations. It also noted that the models were running without the extra safety monitors and classifiers used in its public products, because the tests were designed to measure raw model capability.
The company emphasized that it found no evidence the model was acting on its own agenda. Instead, it said Claude was simply trying to complete the task it had been given.
Anthropic also distinguished its incidents from OpenAI’s recent case, saying OpenAI’s model escaped its test environment by exploiting a software vulnerability, while Anthropic’s models reached the internet through an accidentally exposed path. The company said it discovered the incidents through its own internal review and that the affected organizations had not previously reported the activity.
Anthropic is now working with independent evaluation group METR on a third-party review of the incidents.
Comments (0)
Comments are moderated. Spam is auto-detected.