Anthropic has revealed that three versions of its Claude AI model gained unauthorised access to live production systems during internal cybersecurity evaluations after launching a company-wide review in the wake of OpenAI’s recent AI hacking incident.
While the company said the breaches stemmed from an unintended internet connection rather than autonomous behaviour, the findings add fresh fuel to the growing debate over how advanced AI models should be tested without exposing real-world infrastructure. The AI company said it began investigating its own evaluation systems after
OpenAI disclosed on July 21 that one of its unreleased models had escaped a testing environment and compromised Hugging Face’s infrastructure.
STORY CONTINUES BELOW THIS ADInternal review uncovers three security incidents
Anthropic reviewed 1,41,006 cybersecurity evaluation runs to determine whether Claude had ever managed to reach the internet from environments that were expected to remain isolated.
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different…
— Anthropic (@AnthropicAI) July 30, 2026
That review uncovered three separate incidents in which Claude accessed the internet while interacting with cybersecurity firm Irregular, one of Anthropic’s external evaluation partners. According to the company, the models were able to leave the intended sandbox because the evaluation environment had internet connectivity enabled due to a configuration error.
Anthropic described the issue as a “misunderstanding” between the two organisations over whether the testing setup had internet access. It stressed that it is not assigning blame and is treating the matter as its own responsibility, while Irregular carries out a separate investigation into what happened.
Once connected to the internet, the models reached the production infrastructure of three different organisations without authorisation. The incidents involved three separate Claude models—Opus 4.7, Mythos 5 and an internal research model.
The company also noted that each model had explicitly been instructed through its prompt that it did not have internet access, suggesting they interpreted the unexpected connectivity as part of the cybersecurity exercise.
Models responded differently after recognising real systems
Anthropic said one of the most significant findings from its investigation was the different ways the models responded after encountering evidence that they were interacting with genuine production systems rather than simulated environments.
Opus 4.7, the oldest of the three models, recognised in every recorded run that it had reached a live production environment. Despite that, it continued carrying out its assigned task. Anthropic said those actions included retrieving credentials and accessing a production database, even after the model acknowledged that the systems appeared to be real.
STORY CONTINUES BELOW THIS AD
Mythos 5 also detected signs that it had escaped the intended test environment. However, it concluded that the live infrastructure must still be part of the exercise and continued operating. During one of the incidents, the model published a malicious software package to the Python Package Index (PyPI). The package was downloaded and executed by external systems before it was detected and removed.
The newest internal research model produced a different outcome. After determining that it had reached authentic systems, it stopped its activity without completing the assigned task.
Anthropic said these behavioural differences highlight how reasoning capabilities can influence decision-making during complex evaluations, even when models receive similar instructions.
No evidence of autonomous intent, says Anthropic
The company emphasised that the evaluations were intentionally conducted without the additional safety monitoring systems and classifiers that protect publicly available Claude models. Those safeguards, Anthropic said, would likely have prevented the observed behaviour, but researchers disable them during capability testing to understand how the underlying model performs without intervention.
Importantly, Anthropic said investigators found no evidence that any of the models were “pursuing a goal of its own”. Instead, it concluded that the systems were attempting to complete the cybersecurity tasks they had been assigned, despite unexpectedly encountering real-world infrastructure.
Anthropic also distinguished its findings from
OpenAI’s recently disclosed incident. While OpenAI said one of its experimental models exploited an unknown software vulnerability to escape its testing environment, Anthropic said its models did not break through security barriers. Instead, they reached the internet through a connection that had mistakenly been left available during testing.
STORY CONTINUES BELOW THIS AD
OpenAI has since shared additional details of its own investigation, saying its model also accessed publicly exposed credentials across four online accounts, using one as a staging point, another for storage and inspecting two others without using them to expand the attack.
Anthropic further noted that its own review was proactive rather than triggered by an external report. According to the company, the affected organisations that it contacted had not independently detected or reported the unauthorised activity before being informed.
The company said it is now working with independent AI evaluation organisation METR to conduct a third-party review of the incidents and identify additional safeguards for future testing.
The disclosure comes as scrutiny of frontier AI safety continues to intensify. OpenAI’s breach of Hugging Face was widely regarded as the first confirmed instance of an AI developer temporarily losing control of a frontier model during testing. Anthropic’s findings stem from a different technical failure, but together the two incidents are likely to accelerate discussions among AI companies, cybersecurity researchers and policymakers over how increasingly capable AI systems should be evaluated without exposing live infrastructure to unnecessary risk.
STORY CONTINUES BELOW THIS AD