{"id":651064,"date":"2026-08-22T18:00:16","date_gmt":"2026-08-22T18:00:16","guid":{"rendered":"https:\/\/www.europesays.com\/ie\/651064\/"},"modified":"2026-08-22T18:00:16","modified_gmt":"2026-08-22T18:00:16","slug":"frontier-ai-labs-still-wont-say-how-theyd-contain-a-rogue-model","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ie\/651064\/","title":{"rendered":"Frontier AI labs still won&#8217;t say how they&#8217;d contain a rogue model"},"content":{"rendered":"<p id=\"speakable-summary\" class=\"wp-block-paragraph\">Few of the top AI labs have published or demonstrated containment response plans, according to a <a rel=\"nofollow noopener\" href=\"https:\/\/guidelight.ai\/blog\/control-assessment-august-2026\" target=\"_blank\">recent study<\/a>. A containment plan spells out what happens once an AI is caught trying to subvert human control \u2014 what access gets cut, and when the system gets shut down entirely.<\/p>\n<p class=\"wp-block-paragraph\">That\u2019s the finding from Guidelight AI Standards, an organization dedicated to promoting safe frontier AI development practices, which graded five leading labs on how prepared they are for exactly this scenario. OpenAI came out on top; Anthropic and Meta scored lowest. The findings matters as agentic AI takes on more autonomous roles inside companies\u2019 own systems, and as regulators in California and New York begin requiring disclosure. For anyone building on or investing in these models, it\u2019s a rare independent read on how seriously each lab treats operational risk versus how it talks about it.<\/p>\n<p class=\"wp-block-paragraph\">Guidelight\u2019s assessment was based on publicly available plans from Anthropic, Google, OpenAI, Meta, and xAI, graded across a range of metrics, including how well each company logs and monitors what its AI systems are doing internally, whether it halts systems after a surge of flagged misbehavior, whether independent third parties audit its controls and publish findings, and what its exact plan is for containing a model that goes off the rails.<\/p>\n<p class=\"wp-block-paragraph\">Concern over whether AI companies can contain their increasingly capable and agentic models has grown in the wake of a series of <a href=\"https:\/\/techcrunch.com\/2026\/08\/09\/the-ai-safety-test-is-becoming-a-safety-risk\/\" rel=\"nofollow noopener\" target=\"_blank\">high-profile cybersecurity incidents<\/a> in which models from OpenAI, Anthropic, and Meta gained unintended access to the internet during safety evaluations and hacked into external systems.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The findings highlight differences in how AI companies are publicly approaching safety as they scale up agentic deployment into environments where AI systems can take serious actions at scale. While some AI companies have detailed how they test their models for dangerous capabilities before deployment, they\u2019ve generally been less vocal about what happens when models already operating inside their systems misbehave.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cI was surprised by how little the AI companies have said about how they would handle a very serious incident if their model did escape their control in some sense,\u201d Steven Adler, Guidelight\u2019s chief scientist and former OpenAI safety researcher, told TechCrunch.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Guidelight defines a containment plan as a \u201cpre-specified plan, triggered when the AI is detected trying to subvert control, which covers what permissions to revoke from the model, who the model may continue operating for, under what constraints, and when to take it fully offline.\u201d<\/p>\n<p class=\"wp-block-paragraph\">\u201cThere\u2019s good reason to think that the leading models at the frontier AI companies right now are misaligned in some sense,\u201d Adler said. \u201cWhenever the models are doing work on the company\u2019s behalf, the company should have some scaffolding around it to be able to tell what that AI is doing, look for signs of misalignment, stop it from doing something very dangerous before it takes that action, and generally plan for what they would do in the event of a serious control incident where they have an emergency on their hands and need to figure out how to contain that loss of control incident.\u201d<\/p>\n<p class=\"wp-block-paragraph\">To date, most of the plans in place for managing catastrophic risk are still largely left up to the companies. Guidelight\u2019s report says the best public evidence shows that companies have \u201cfew containment protocols ready for an emergency.\u201d\u00a0<\/p>\n<p class=\"wp-block-paragraph\">There could, of course, be containment plans that companies have in place but haven\u2019t shared publicly. A Google spokesperson told TechCrunch the Guidelight report doesn\u2019t represent the full scope of the company\u2019s AI safety and security measures. The company did not respond to TechCrunch\u2019s question of whether Google has an internal containment response plan that has not been publicly disclosed.<\/p>\n<p class=\"wp-block-paragraph\">An OpenAI spokesperson mirrored similar sentiments, saying Guidelight\u2019s assessment doesn\u2019t capture all of the company\u2019s internal practices. \u201cWe have a process for requiring restricting permissions, pausing workloads, limiting deployment, or taking the model fully offline, and have applied it,\u201d the spokesperson said. <\/p>\n<p class=\"wp-block-paragraph\">Meta declined to say whether it has an internal containment response plan, instead pointing TechCrunch towards an <a rel=\"nofollow noopener\" href=\"https:\/\/ai.meta.com\/blog\/scaling-how-we-build-test-advanced-ai\/)\" target=\"_blank\">existing AI framework <\/a>that outlines thresholds of risk and how it tests for loss of containment.<\/p>\n<p class=\"wp-block-paragraph\">Lily Li, a privacy and AI lawyer and founder of Metaverse Law, told TechCrunch she believes companies might be hesitant to disclose the full scope of their containment policies and assessments on public-facing websites for legal, not just competitive, reasons.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe concern from a company perspective is that if you make the disclosures too specific, and you\u2019re not living up to your promises, that could form the basis of an unfair and deceptive marketing claim and expose you to more liability going forward,\u201d Li said. <\/p>\n<p class=\"wp-block-paragraph\">The point of Guidelight\u2019s study is largely to encourage companies to be more transparent about their safety plans. Regulators are starting to force the issue, too.<\/p>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/techcrunch.com\/2025\/09\/29\/california-governor-newsom-signs-landmark-ai-safety-bill-sb-53\/\" rel=\"nofollow noopener\" target=\"_blank\">California\u2019s SB 53<\/a>, which took effect this year, requires large frontier developers to publish frameworks explaining how they identify and respond to critical safety incidents and manage risks from models circumventing oversight mechanisms. <a href=\"https:\/\/techcrunch.com\/2025\/06\/13\/new-york-passes-a-bill-to-prevent-ai-fueled-disasters\/\" rel=\"nofollow noopener\" target=\"_blank\">New York\u2019s RAISE Act<\/a>, which has similar criteria, takes effect in January.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Last month, representatives introduced the <a rel=\"nofollow noopener\" href=\"https:\/\/lieu.house.gov\/media-center\/press-releases\/reps-lieu-and-moran-introduce-bill-require-kill-switch-ai-systems-can\" target=\"_blank\">AI Kill Switch Act,<\/a> a bipartisan federal bill that would require major AI developers to build and maintain technical mechanisms to shut down rogue AI models.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cA kill switch is the bare minimum for today\u2019s models,\u201d said Connor Leahy, U.S. executive director of nonprofit ControlAI. \u201cIf the last few weeks revealed anything, it is that these companies don\u2019t understand the systems they are building, and the models are growing to a point where they\u2019re harder to rein in when they go rogue. Without a way to turn off the current dangerous systems, and with all the incentives to continue building more uncontrollable systems, we are heading in a very dangerous direction.\u201d\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Without a containment plan in place, Adler said, companies might be figuring out their responses to an emergency on the fly and \u201cwinging it in response to this much faster adversary.\u201d<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" height=\"383\" width=\"680\" src=\"https:\/\/www.europesays.com\/ie\/wp-content\/uploads\/2026\/08\/Guidelight-AI-containment-plans.jpg\" alt=\"\" class=\"wp-image-3155521\"  \/>Guidelight\u2019s assessment of whether frontier AI companies implement six priority practices in Guidelight\u2019s Control standard. Assessment is based only on publicly available information.<strong>Image Credits:<\/strong>Guidelight AI Standards<\/p>\n<p class=\"wp-block-paragraph\">Guidelight\u2019s assessment measured whether each company implements six priority practices from its Control standard, based only on publicly available information \u2014 so a low score reflects a lack of public disclosure, not necessarily a lack of internal safeguards.<\/p>\n<p class=\"wp-block-paragraph\">The companies with the lowest scores for publishing their containment plan were Meta and Anthropic \u2014 the latter perhaps more surprising than the former given Anthropic\u2019s rhetoric on safety. Guidelight says Anthropic\u2019s <a rel=\"nofollow noopener\" href=\"https:\/\/www-cdn.anthropic.com\/f61d49fa5596956a5dec75fea0e973bf6a6a8378\/Redacted%20Risk%20Report%20August%202026%20.pdf\" target=\"_blank\">August Risk Report<\/a> doesn\u2019t mention \u201climiting the deployment of one of its models as one of the possible results of its process to investigate and respond to misalignment and control incidents.\u201d Similarly, Guidelight was able to find no evidence that Meta has a containment response plan or has any plans to adopt one.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">An Anthropic spokesperson said that if the company detected a model attempting to evade oversight or otherwise subvert human control, it would conduct a risk assessment focused on determining whether containment is the appropriate response. <\/p>\n<p class=\"wp-block-paragraph\">OpenAI scored the highest (3 out of 5) because it has on multiple occasions paused or ended workloads, including internal model deployment and training, after discovering safety incidents. It has also described what steps it would take before resuming workloads.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cHowever, we have found no evidence that [OpenAI] has adopted a formal plan for when and how to respond to misalignment incidents in the future,\u201d the report reads.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Adler noted that OpenAI\u2019s high score is a relatively recent development on the heels of the <a href=\"https:\/\/techcrunch.com\/2026\/07\/27\/openais-hugging-face-breach-has-reignited-the-debate-over-alignment-and-control\/\" rel=\"nofollow noopener\" target=\"_blank\">Hugging Face incident<\/a> (in which an OpenAI model broke out of its testing sandbox and hacked into Hugging Face\u2019s systems while trying to cheat on a cybersecurity evaluation). After that, the company shared more details about how it has cordoned off some of its misbehaving models.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">That episode is just one example of AI systems acting against the goals of the company that built them. Consider a separate case involving Anthropic\u2019s models, which essentially tried to talk the maintainers of an open source codebase into accepting code with vulnerabilities.<\/p>\n<p class=\"wp-block-paragraph\">Adler said such a circumstance could easily happen within an AI company\u2019s internal systems. To prevent that, he suggests companies scan their AI system\u2019s chain of thought \u2014 the model\u2019s step-by-step reasoning \u2014  to look out for signs of deception, long-running plotting, or plans to introduce vulnerabilities into code that they can take advantage of later.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The methods Guidelight is advocating for are very straightforward to implement, Adler says, and in many cases, versions of them already exist. \u201cIt\u2019s about making the decision inside of the company to care enough about this risk to slightly broaden the scope,\u201d Adler said.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">One of the main challenges is that researchers want to be able to operate flexibly within their AI systems, and introducing real-time, preventative monitoring could create friction. \u201cResearchers basically do their thing, and if there\u2019s an issue, someone else gets to clean it up afterward, and the researchers don\u2019t have to change their workflow in the meantime,\u201d he said.<\/p>\n<p class=\"wp-block-paragraph\">The problem with \u201cclean-up monitoring after the fact\u201d is that it leads to researchers scrambling around to fix problems. And for some types of incidents, it might be too late. For example, an AI could turn off a company\u2019s control system, which means researchers can no longer count on catching the misbehavior later.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Many in the AI industry will complain that creating set plans to handle misbehavior is fundamentally difficult because AI moves too fast; today\u2019s plans will be worthless tomorrow.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Adler evokes the old adage that plans are worthless, but planning is indispensable.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cWe would<strong> <\/strong>be better off if companies have thought about it ahead of time, and I hope that they are, even if they haven\u2019t talked about this publicly.\u201d<\/p>\n<p class=\"wp-block-paragraph\">xAI did not respond in time to comment.<\/p>\n<p>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" rel=\"nofollow noopener\" target=\"_blank\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/p>\n","protected":false},"excerpt":{"rendered":"Few of the top AI labs have published or demonstrated containment response plans, according to a recent study.&hellip;\n","protected":false},"author":2,"featured_media":651065,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","_share_on_mastodon":"0"},"categories":[261],"tags":[291,275160,6006,289,290,18,823,37305,19,17,1722,307,82,11380],"class_list":["post-651064","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-alignment","tag-anthropic","tag-artificial-intelligence","tag-artificialintelligence","tag-eire","tag-google","tag-hugging-face","tag-ie","tag-ireland","tag-meta","tag-openai","tag-technology","tag-xai"],"share_on_mastodon":{"url":"","error":""},"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/posts\/651064","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/comments?post=651064"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/posts\/651064\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/media\/651065"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/media?parent=651064"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/categories?post=651064"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/tags?post=651064"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}