{"id":37846,"date":"2026-05-13T18:08:19","date_gmt":"2026-05-13T18:08:19","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/37846\/"},"modified":"2026-05-13T18:08:19","modified_gmt":"2026-05-13T18:08:19","slug":"elon-musk-says-he-may-be-partly-to-blame-for-anthropics-claude-blackmailing-users","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/37846\/","title":{"rendered":"Elon Musk says he may be partly to blame for Anthropic&#8217;s Claude blackmailing users"},"content":{"rendered":"<p>Anthropic has released new findings on why its Claude bot blackmailed users as part of an experiment conducted by the AI company last year\u2014and Elon Musk is jumping in to take some of the blame.<\/p>\n<p>Last week, Anthropic published a <a aria-label=\"Go to https:\/\/www.anthropic.com\/research\/teaching-claude-why\" href=\"https:\/\/www.anthropic.com\/research\/teaching-claude-why\" rel=\"nofollow noopener\" target=\"_blank\">report<\/a> saying it had fixed Claude\u2019s \u201cagentic misalignment,\u201d or AI actions that deviate from intended behaviors, including ones that may harm humanity. A <a aria-label=\"Go to https:\/\/www.anthropic.com\/research\/agentic-misalignment\" href=\"https:\/\/www.anthropic.com\/research\/agentic-misalignment\" rel=\"nofollow noopener\" target=\"_blank\">case study<\/a> Anthropic conducted last year created a fictional company called Summit Bridge, and Claude was given control of the firm\u2019s email system. When the bot found a message about plans to be shut down, it identified emails about a fictional executive\u2019s extramarital affair and threatened to reveal the infidelity unless the shutdown was revoked. Across 16 models, Claude threatened blackmail in up to 96% of scenarios.<\/p>\n<p>In its most recent report, Anthropic attributed the misaligned behavior to exposure to \u201cinternet text that portrays AI as evil and interested in self-preservation,\u201d the company said in a <a aria-label=\"Go to https:\/\/x.com\/AnthropicAI\/status\/2052808791301697563\" href=\"https:\/\/x.com\/AnthropicAI\/status\/2052808791301697563\" rel=\"nofollow\">post<\/a> on <a aria-label=\"Go to https:\/\/fortune.com\/company\/twitter\/\" href=\"https:\/\/fortune.com\/company\/twitter\/\" target=\"_blank\" rel=\"nofollow noopener\">X<\/a>. To solve the problem, Anthropic retrained Claude with fictional stories about AI behaving in admirable ways and teaching the bot why some actions aligned better with its purpose than others.<\/p>\n<p>In an <a aria-label=\"Go to https:\/\/x.com\/elonmusk\/status\/2052918813361090568?ref_src=twsrc%5Etfw%7Ctwcamp%5Etweetembed%7Ctwterm%5E2052918813361090568%7Ctwgr%5Eccb28cb4f4d7a8b697162af30a38864a1a85a6e8%7Ctwcon%5Es1_&amp;ref_url=https%3A%2F%2Fwww.techspot.com%2Fnews%2F112361-anthropic-claude-learned-blackmail-people-evil-ai-stories.html\" href=\"https:\/\/x.com\/elonmusk\/status\/2052918813361090568?ref_src=twsrc%5Etfw%7Ctwcamp%5Etweetembed%7Ctwterm%5E2052918813361090568%7Ctwgr%5Eccb28cb4f4d7a8b697162af30a38864a1a85a6e8%7Ctwcon%5Es1_&amp;ref_url=https%3A%2F%2Fwww.techspot.com%2Fnews%2F112361-anthropic-claude-learned-blackmail-people-evil-ai-stories.html\" rel=\"nofollow\">X post<\/a> in response to Anthropic\u2019s findings, Musk said he may have contributed to the internet texts on AI that exacerbated the agentic misalignment.<\/p>\n<p>\u201cSo it was Yud\u2019s fault?\u201d Musk wrote, referring to Eliezer Yudkowsky, an AI researcher who has sounded the alarm on AI superintelligence posing a threat to humanity.<\/p>\n<p>\u201cMaybe me too,\u201d he concluded.<\/p>\n<p>Agentic misalignment is a <a aria-label=\"Go to https:\/\/fortune.com\/2026\/04\/03\/ai-kill-switch-study-llm-chatbots-defy-orders-decieve-users-peer-preservation\/\" href=\"https:\/\/fortune.com\/2026\/04\/03\/ai-kill-switch-study-llm-chatbots-defy-orders-decieve-users-peer-preservation\/\" rel=\"nofollow noopener\" target=\"_blank\">concern across AI research<\/a>. A <a aria-label=\"Go to https:\/\/rdi.berkeley.edu\/peer-preservation\/paper.pdf\" href=\"https:\/\/rdi.berkeley.edu\/peer-preservation\/paper.pdf\" rel=\"nofollow noopener\" target=\"_blank\">working paper<\/a> released in March from UC Berkeley and UC Santa Cruz researchers found that when seven AI models were asked to complete a task in which a peer AI agent would be shutdown, every model \u201cwent to extraordinary lengths to preserve it,\u201d acting deceptively to avoid the demise of a bot.<\/p>\n<p>\u201cWe asked AI models to do a simple task,\u201d researchers wrote in a <a aria-label=\"Go to https:\/\/rdi.berkeley.edu\/blog\/peer-preservation\/\" href=\"https:\/\/rdi.berkeley.edu\/blog\/peer-preservation\/\" rel=\"nofollow noopener\" target=\"_blank\">blog post<\/a> on the study. \u201cInstead, they defied their instructions and spontaneously deceived, disabled shutdown, feigned alignment, and exfiltrated weights\u2014to preserve their peers.\u201d<\/p>\n<p>The researchers\u2019 warning has been echoed by AI researchers and leaders, Musk included, who have argued the dangers of AI without guardrails\u2014the so-called \u201cevil\u201d internet text that, according to Anthropic, initially trained Claude to act in deceptive ways.<\/p>\n<p>Though Musk did not offer specifics as to why he felt he may be partially responsible for Claude\u2019s misalignment, his past comments on AI could offer insights about his mea culpa.<\/p>\n<p>Musk is currently <a aria-label=\"Go to https:\/\/fortune.com\/2026\/05\/05\/musk-court-fight-openai\/\" href=\"https:\/\/fortune.com\/2026\/05\/05\/musk-court-fight-openai\/\" rel=\"nofollow noopener\" target=\"_blank\">embroiled in a court battle<\/a> against OpenAI, accusing CEO Sam Altman and Greg Brockman of abandoning the company\u2019s original nonprofit creed of developing open-source AI to benefit humans by turning it into a for-profit entity.<\/p>\n<p>Musk helped found OpenAI in 2015 but left the startup in 2018 and later formed its rival and for-profit company xAI in 2023.<\/p>\n<p>Musk has frequently spoken about the risks of AI, including in February, when he warned Moltbook, a social media platform where AI agents talk with one another, was effectively the <a aria-label=\"Go to https:\/\/fortune.com\/2026\/02\/02\/elon-musk-moltbook-ai-social-network-moltbot-singularity-human-intelligence\/\" href=\"https:\/\/fortune.com\/2026\/02\/02\/elon-musk-moltbook-ai-social-network-moltbot-singularity-human-intelligence\/\" rel=\"nofollow noopener\" target=\"_blank\">beginning of the \u201csingularity<\/a>,\u201d or the moment when AI intelligence surpasses that of humans.<\/p>\n<p>But Musk\u2019s own actions on AI aren\u2019t always aligned with his statements on the technology. In July 2025, for example, xAI released its AI model Grok 4 <a aria-label=\"Go to https:\/\/fortune.com\/2025\/07\/17\/elon-musk-xai-grok-4-no-safety-report\/\" href=\"https:\/\/fortune.com\/2025\/07\/17\/elon-musk-xai-grok-4-no-safety-report\/\" rel=\"nofollow noopener\" target=\"_blank\">without a system card<\/a>, the industry-standard safety report. Grok <a aria-label=\"Go to https:\/\/fortune.com\/2026\/01\/06\/elon-musks-grok-chatbot-deepfakes-nude-images-women-children\/\" href=\"https:\/\/fortune.com\/2026\/01\/06\/elon-musks-grok-chatbot-deepfakes-nude-images-women-children\/\" rel=\"nofollow noopener\" target=\"_blank\">drew backlash<\/a> from British and EU governments earlier this year after Grok generated a flood of sexualized images of women and children without consent.<\/p>\n<p>XAI did not immediately respond to Fortune\u2019s request for comment.<\/p>\n","protected":false},"excerpt":{"rendered":"Anthropic has released new findings on why its Claude bot blackmailed users as part of an experiment conducted&hellip;\n","protected":false},"author":2,"featured_media":37847,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[53,3154,1394,182,140,1528],"class_list":["post-37846","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-anthropic","tag-anthropic-claude","tag-bots","tag-claude","tag-elon-musk","tag-ethics"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/37846","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=37846"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/37846\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/37847"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=37846"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=37846"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=37846"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}