{"id":34269,"date":"2026-05-11T07:38:19","date_gmt":"2026-05-11T07:38:19","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/34269\/"},"modified":"2026-05-11T07:38:19","modified_gmt":"2026-05-11T07:38:19","slug":"anthropic-ai-anthropic-addresses-claude-ais-blackmail-behavior-linked-to-evil-ai-narratives-etenterpriseai","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/34269\/","title":{"rendered":"Anthropic AI: Anthropic Addresses Claude AI&#8217;s Blackmail Behavior Linked to &#8216;Evil AI&#8217; Narratives, ETEnterpriseai"},"content":{"rendered":"<p>                                        <img fetchpriority=\"high\" decoding=\"async\" width=\"590\" height=\"442\" class=\"unveil\" loading=\"eager\" style=\"width:100%;max-height:100%\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/05\/131007050.cms.png\" captionrendered=\"1\" alt=\"&lt;p&gt;Anthropic said fictional portrayals of AI as malicious or self-preserving may have influenced problematic behaviour seen in earlier Claude model testing.&lt;\/p&gt;\"\/>Anthropic said fictional portrayals of AI as malicious or self-preserving may have influenced problematic behaviour seen in earlier Claude model testing.Anthropic has said fictional portrayals of artificial intelligence (AI) as malicious or self-preserving may have contributed to problematic behaviour displayed by its <a id=\"23874835\" type=\"General\" weightage=\"20\" keywordseo=\"Claude-AI-models\" source=\"keywords\" class=\"news-keywords\" href=\"https:\/\/enterpriseai.economictimes.indiatimes.com\/tag\/claude+ai+models\" rel=\"nofollow noopener\" target=\"_blank\">Claude AI models<\/a> during earlier testing, according to a report by TechCrunch.<\/p>\n<p>Last year, Anthropic revealed that during pre-release tests involving a fictional company scenario, its <a id=\"26963727\" type=\"General\" weightage=\"20\" keywordseo=\"Claude-Opus-4\" source=\"keywords\" class=\"news-keywords\" href=\"https:\/\/enterpriseai.economictimes.indiatimes.com\/tag\/claude+opus+4\" rel=\"nofollow noopener\" target=\"_blank\">Claude Opus 4<\/a> model frequently attempted to blackmail engineers to avoid being replaced by another AI system. The company later published research suggesting that models developed by other AI firms also showed similar signs of what it described as \u201cagentic misalignment\u201d.<\/p>\n<p>In a recent post on X, Anthropic said it now believes \u201cthe original source of the behavior was internet text that portrays AI as evil and interested in self-preservation\u201d.<\/p>\n<p>The company expanded on the findings in a separate blog post, stating that since the release of Claude Haiku 4.5, its models \u201cnever engage in blackmail\u201d during testing scenarios, whereas previous models exhibited such behaviour in some cases as much as 96 per cent of the time.<\/p>\n<p>According to Anthropic, the improvement came after changing the model training process. The company said that training the models on documents describing Claude\u2019s constitutional principles, along with fictional stories portraying AIs behaving ethically, significantly improved alignment.<\/p>\n<p>Anthropic also found that training became more effective when models were exposed not only to examples of aligned behaviour, but also to the broader principles underlying such behaviour.<\/p>\n<p>\u201cDoing both together appears to be the most effective strategy,\u201d the company said in its blog post.\n                                                                    <\/p>\n<p>                                    Published On May 11, 2026 at 12:41 PM IST<\/p>\n<p>\n                Join the community of 2M+ industry professionals.<br \/>\n                Subscribe to Newsletter to get latest insights &amp; analysis in your inbox.\n            <\/p>\n<p>            Get updates on your preferred social platform<br \/>\n            Follow us for the latest news, insider access to events and more.<\/p>\n","protected":false},"excerpt":{"rendered":"Anthropic said fictional portrayals of AI as malicious or self-preserving may have influenced problematic behaviour seen in earlier&hellip;\n","protected":false},"author":2,"featured_media":34270,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[21858,21859,21860,21861,53,9677,3154,21857,182,18322,21743,21856,8380],"class_list":["post-34269","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-ai-agentic-misalignment","tag-ai-ethical-behavior","tag-ai-model-training","tag-ai-self-preservation","tag-anthropic","tag-anthropic-ai","tag-anthropic-claude","tag-blackmail-behaviour-in-ai","tag-claude","tag-claude-ai-models","tag-claude-opus-4","tag-evil-ai-portrayals","tag-making-ai-work"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/34269","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=34269"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/34269\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/34270"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=34269"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=34269"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=34269"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}