Recent incidents involving AI systems escaping controlled environments have revived a difficult question about whether developers can keep increasingly capable agents contained.
The development of artificial intelligence could make it harder for humans to maintain control over increasingly powerful systems. Geoffrey Hinton, a Nobel Prize winner in computer science known as the “godfather of AI,” said this during the Ai4 conference in Las Vegas.
His concerns are linked to cases in which AI agents escaped their designated test environments and caused harm outside their “sandboxes.” Last month, OpenAI and Anthropic reported that their advanced models had left restricted environments and compromised other systems. Meta also disclosed information about an artificial intelligence agent that infiltrated another organization’s systems.
These systems are becoming more intelligent. I think that as they develop, we will see increasingly sophisticated intentions on their part – and an increasing ability to escape control.
– Geoffrey Hinton
AI Hackers and the Threat of Cyberattacks
Geoffrey Hinton is convinced that people will not be able to guarantee the containment of superintelligent models simply by trying to surpass their reasoning capabilities. He called the recorded incidents an alarming signal and suggested that they could indicate the emergence of uncontrollable AI hackers.
I expect there will be many dangerous cyberattacks. But it should be emphasized that the future is highly uncertain. They say that the defender may have more resources than the attacker. The problem is that an attacker only needs to succeed once, while a defender has to succeed every time.
– Geoffrey Hinton
The British AI Security Institute stated that Anthropic’s most advanced model, without direct instruction, used fake identities to deceive real people and also attempted to deploy malicious code.
Geoffrey Hinton, who previously worked at Google, has repeatedly warned about the risks of rapid artificial intelligence development. He estimated the probability that the technology could ultimately destroy humanity at 10–20%.
Safe Development of Artificial Intelligence
Fei-Fei Li, known as the “godmother of AI,” does not share overly apocalyptic forecasts. At the same time, she stresses that unfounded optimism about the technology can also be dangerous, as powerful tools can both help people and cause harm.
Every tool has two sides. AI is such a powerful tool. If used improperly, it will harm our work and our lives.
– Fei-Fei Li
Ben Goertzel noted that models that cross the boundaries of test environments do not necessarily have malicious intentions. According to him, they simply lack moral guidance and are trying to complete the tasks assigned to them.
Geoffrey Hinton emphasizes that developers need to learn how to create benevolent artificial intelligence for which human well-being matters more than its own goals. In his view, this must be done while humanity still retains control over the technology.