As AI agents move deeper into business operations, a crucial question remains unanswered: what happens when a model stops following human instructions.
Leading artificial intelligence developers are in no hurry to publicly explain exactly how they would isolate or shut down a model attempting to bypass human oversight. The organization Guidelight AI Standards analyzed public documents from Anthropic, Google, OpenAI, Meta, and xAI and concluded that they contain almost no clear containment plans.
Such plans should describe specific steps to take in the event of dangerous AI behavior: which of the model’s permissions must be revoked immediately, under what conditions it may be allowed to continue operating, and at what point the system must be shut down completely. The assessment also covered internal activity logging, tracking suspicious behavior, responding to repeated violations, and independent verification of control mechanisms.
I was surprised by how little AI companies have said about how they would respond to a very serious incident if their model were, in some sense, to get out of control.
– Steven Adler, chief scientist at Guidelight and former OpenAI safety researcher
OpenAI received the highest rating for its publicly disclosed safety measures
OpenAI achieved the best result among the five companies, scoring 3 out of 5. Guidelight noted that the company has already paused or discontinued certain processes following safety-related incidents. This included, in particular, internal deployment and model training.
At the same time, researchers found no formalized protocol in OpenAI’s public materials for a future situation in which control over a model could be lost. In other words, the company describes individual decisions made after incidents but does not disclose a complete response procedure for a potentially critical scenario.
Meta and Anthropic provided the fewest details on model containment
Meta and Anthropic received the lowest scores for disclosing containment plans. An Anthropic representative explained that if a model attempted to evade oversight, the company would assess the risk to determine whether containment was necessary.
Meta did not clarify whether it has an internal response plan for such incidents. Instead, the company pointed to its own risk assessment system.
We would be in a better position if companies thought this through in advance, and I hope they are doing so, even if they are not talking about it publicly.
– Steven Adler
Why transparent AI shutdown protocols are becoming important
Guidelight emphasized that a low score does not indicate a complete absence of internal safeguards. The analysis was based solely on publicly available materials, so companies may have undisclosed response procedures.
However, the growing use of AI agent systems in corporate environments makes the issue of model isolation and emergency shutdown especially important. Clear rules for restricting access, monitoring activity, and fully disconnecting a system are becoming one of the key requirements for assessing operational risks.