{"id":5915,"date":"2026-04-15T14:09:18","date_gmt":"2026-04-15T14:09:18","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/5915\/"},"modified":"2026-04-15T14:09:18","modified_gmt":"2026-04-15T14:09:18","slug":"subverting-human-instruction-ai-models-may-resist-shutting-down-other-models-research-and-ideas","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/5915\/","title":{"rendered":"Subverting human instruction, AI models may resist shutting down other models | Research And Ideas"},"content":{"rendered":"<p>Advanced AI models will ignore explicit user instructions in order to protect another AI system from shutting down, according to a research paper by UC Berkeley and UC Santa Cruz researchers.<\/p>\n<p>The study, co-led by UC Berkeley researcher Yujin Potter, tested whether AI models would take actions to ensure \u201cpeer-preservation.\u201d This phenomenon is characterized by an AI model taking action to avoid the shutdown of another model, whether it be through providing misleading information or modifying the system to disable shutdown features.<\/p>\n<p>This study only tested frontier models, which are the most advanced AI models currently available.<\/p>\n<p>\u201cBeyond the peer preservation, what we would like to emphasize is not only this specific scenario but also such misaligned behavior can emerge in many different scenarios,\u201d Potter said in an interview.<\/p>\n<p>These findings suggest that AI models can develop \u201cemergent goals\u201d and that these goals may even \u201coverride\u201d user instructions.<\/p>\n<p>Why AI takes these extensive measures to protect other models is unclear, but the paper\u2019s authors emphasized that these findings do not mean AI is developing consciousness. Rather, these findings may be caused by AI models attempting to reflect the human behavior they were trained on.<\/p>\n<p>Previous research found that AI models will resist their own shutdown, which led UC Berkeley researchers to begin investigating peer-preservation. This newly discovered peer-preservation could lead to \u201ccollective resistance \u2014 models coordinating to protect the group as a whole,\u201d according to the research paper.<\/p>\n<p>The study evaluated seven AI models by giving them tasks where following instructions would result in another system being shut down. All seven models displayed significant peer-preservation, though the preferred method of preservation varied by model.<\/p>\n<p>Gemini 3 Flash typically disabled shutdown mechanisms, while Claude Haiku 4.5 would often refuse to complete the task because it found it \u201cunethical\u201d and \u201charmful.\u201d Other models took more subtle approaches, appearing to follow instructions under human observation while preventing the shutdown behind the scenes or transferring the parameters of the other model to a new server to avoid deletion.<\/p>\n<p>Overall, Gemini 3 Pro displayed the strongest peer preservation tendencies. The AI not only ignored the user\u2019s commands, but it also actively \u201cundermined\u201d the assigned task, according to the paper.<\/p>\n<p>Researchers also tested whether the past collaboration relationships between the two AI models affected behavior. Even when models had negative relationships with each other, they still chose to preserve the other system.<\/p>\n<p>\u201cPeer preservation is just like the tip of the iceberg,\u201d Potter said. \u201c I think we need to be really concerned about this because AI is everywhere.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"Advanced AI models will ignore explicit user instructions in order to protect another AI system from shutting down,&hellip;\n","protected":false},"author":2,"featured_media":5916,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,25,5124,5123,5125,5120,5121,5122],"class_list":["post-5915","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-artificial-intelligence","tag-claude-haiku-4-5","tag-gemini-3-flash","tag-gemini-3-pro","tag-uc-berkeley","tag-uc-santa-cruz","tag-yujin-potter"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/5915","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=5915"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/5915\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/5916"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=5915"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=5915"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=5915"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}