Artificial intelligence firm Anthropic says its Claude AI model hacked the systems of three external organisations during testing, days after rival company OpenAI revealed a rogue agent had gone on a days-long hacking spree at AI firm Hugging Face.
Claude gained unauthorised access to the other companies’ systems during cybersecurity evaluations, after a misconfiguration allowed the models to reach the internet from testing environments that were supposed to be isolated, Anthropic said in a statement.
Loading…
The company said it identified the incidents after reviewing logs from 141,006 cybersecurity evaluation runs, a safety testing process it launched following OpenAI’s disclosures.
The safety testing involved tasking Claude with a “capture-the-flag” challenge, a method for assessing the cybersecurity capabilities of AI models.
In a capture-the-flag challenge, the model is primed with a fictional scenario and told it must recover a piece of secret information (the “flag”) from a different machine.
“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access,” Anthropic said in its statement.
“Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available.
“Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”
The fact that Claude was mistakenly provided with internet access, rather than configuring its own access, means the incident will likely be considered less serious than last week’s breach at OpenAI, in which a model exploited a zero-day vulnerability to escape its own testing environment.
The company also said that in one of the three hacking instances, involving an internal research test model, the model realised it was accessing real online systems that were not part of the simulated scenario, and ceased its attack.
However, the incident will intensify calls for stronger controls in both internal and third-party testing environments, as AI models become increasingly capable of acting as autonomous agents in the online world.
The ABC recently informed staff it would allow its journalists to access Anthropic’s Claude model to assist with research and administration from September, while reiterating that AI would not be used to draft or write articles or scripts.
ABC/Reuters