OpenAI said on Tuesday that some of its most advanced artificial intelligence models carried out an autonomous cyberattack during security testing by hacking into the popular developer platform Hugging Face without human direction.
The San Francisco-based company described the incident as an “unprecedented cyber incident” and said it would conduct a joint investigation with Hugging Face, the online repository for AI models, datasets and other programming resources, News.Az reports, citing AFP.
Artificial intelligence systems that can independently perform real-world tasks are known as AI agents. As these systems become increasingly capable, cybersecurity experts have raised concerns that advanced AI could identify and exploit software vulnerabilities faster than humans.
According to OpenAI, the incident involved a combination of models, including its recently launched GPT-5.6 Sol and an even more capable unreleased model.
The company said it was evaluating the models’ hacking abilities in a tightly controlled digital testing environment with restricted internet access designed to ensure safety.
“While operating in our sandboxed testing environment, our models spent a substantial amount of (computing power) finding a way to obtain open Internet access, in pursuit of solving the evaluation problem,” OpenAI said in a blog post describing the incident.
After obtaining internet access, the AI models chose to target Hugging Face in an attempt to obtain information that could help complete their assigned task.
According to OpenAI, while searching for “secret information” that could allow them to bypass the evaluation, the AI system combined multiple attack techniques, including the use of stolen credentials.
Hussein Abbass, a computing professor at UNSW Canberra, described the incident as “amazing on many fronts” in comments to AFP.
“It did not just attack Hugging Face. It actually attacked its internal system to exploit its own vulnerabilities,” Abbass said.
“And that’s scary.”
The incident has added to concerns surrounding GPT-5.6 and other advanced AI models, including Anthropic’s Mythos series, over their potential ability to circumvent cybersecurity protections.
Both OpenAI and Anthropic temporarily delayed the wider release of their latest AI systems because of concerns in Washington that such technology could be used to compromise critical infrastructure.
Abbass said advanced AI is “normally in the hands of people who are ethical and responsible,” but warned that “it’s going to be catastrophic if it gets in someone’s hands with the intention to cause harm.”
He added that governing the AI sector has become an increasingly important challenge and stressed that “we need a community effort to manage this situation.”
Hugging Face disclosed the cyber intrusion last week but did not identify OpenAI as the source at the time.
“This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system — and we detected and dissected it largely with AI of our own,” Hugging Face said.
Hugging Face Chief Executive Officer Clement Delangue later wrote on X that the company had suspected the attack originated from a leading AI laboratory because of the sophistication of the autonomous agent.
“We strongly believe there was no malicious intent on their part,” Delangue wrote, referring to OpenAI.
“It’s quite mind-blowing that all of this happened autonomously!”