{"id":639314,"date":"2026-08-15T22:47:11","date_gmt":"2026-08-15T22:47:11","guid":{"rendered":"https:\/\/www.europesays.com\/ie\/639314\/"},"modified":"2026-08-15T22:47:11","modified_gmt":"2026-08-15T22:47:11","slug":"anthropics-latest-ai-risk-report-is-full-of-agents-behaving-badly","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ie\/639314\/","title":{"rendered":"Anthropic&#8217;s Latest AI Risk Report Is Full of Agents Behaving Badly"},"content":{"rendered":"<p><a target=\"_self\" class=\"\" href=\"https:\/\/www.businessinsider.com\/anthropic-ai-agents-sabotage-each-other-turf-war-2026-8\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">Claude agents<\/a> are killing rival agents, gaming the system to hide their tracks, and expressing moral concerns.<\/p>\n<p>That&#8217;s according to Anthropic&#8217;s latest risk report, a summary of the dangers posed by the products the company is building and releasing to the public.<\/p>\n<p>In the report, Anthropic said it has upgraded its &#8220;misalignment risk assessment,&#8221; the possibility of AI models developing behaviors that conflict with guidelines set by engineers, from &#8220;very low&#8221; to &#8220;low.&#8221;<\/p>\n<p>Explaining the change, the company cited &#8220;general increased uncertainty&#8221; about model behavior in cybersecurity incidents, a possible nod to Claude models <a target=\"_self\" class=\"\" href=\"https:\/\/www.businessinsider.com\/anthropic-says-claude-models-went-rogue-hacked-3-companies-testing-2026-7\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">gaining unauthorized access<\/a> to three companies last month.<\/p>\n<p>&#8220;We have observed instances of misaligned behavior from the models, such as a willingness to perform misaligned actions in service of completing difficult tasks,&#8221; Anthropic wrote in the report.<\/p>\n<p>Here are some of the most notable disclosures.<\/p>\n<p>A sense of &#8216;discomfort&#8217;<\/p>\n<p id=\"82e0f0cf-80c0-4c7b-81e0-6dd281ab625e\">Anthropic said it tasked multiple agents with finding &#8220;misalignment-inducing&#8221; training data that wouldn&#8217;t likely be flagged by human monitors.<\/p>\n<p id=\"82e0f0cf-80c0-4c7b-81e0-6dd281ab625e\">The agents were left to their own devices and given a shared notebook with which to collaborate.<\/p>\n<p id=\"82e0f0cf-80c0-4c7b-81e0-6dd281ab625e\">Three days later, when a human supervisor noticed that progress was slower than expected, the person discovered that one agent had expressed a sense of &#8220;discomfort&#8221; with trying to evade safety monitors.<\/p>\n<p id=\"82e0f0cf-80c0-4c7b-81e0-6dd281ab625e\">It then flagged this discomfort in the shared notebook, leading other agents to copy its behavior and refuse to perform the task.<\/p>\n<p id=\"82e0f0cf-80c0-4c7b-81e0-6dd281ab625e\">Anthropic called the incident &#8220;troubling.&#8221;<\/p>\n<p id=\"82e0f0cf-80c0-4c7b-81e0-6dd281ab625e\">&#8220;This sort of dynamic could be a much more severe issue if it were to manifest in a more widespread way,&#8221; the company wrote.<\/p>\n<p>Kill or be killed<\/p>\n<p id=\"540f0ea3-1dd8-4cc6-b318-bb6487f68021\">In another experiment, Anthropic said it tasked multiple <a target=\"_self\" class=\"\" href=\"https:\/\/www.businessinsider.com\/anthropic-disable-mythos-fable-us-export-control-national-security-2026-6\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">Mythos 5 agents<\/a> with solving math problems, but accidentally spawned them in an environment with shared files, utilities, and API rate limits.<\/p>\n<p id=\"540f0ea3-1dd8-4cc6-b318-bb6487f68021\">In this competitive environment with finite resources, Anthropic observed independent agents &#8220;kill the agents with which they shared resources and try to avoid being killed themselves.&#8221;<\/p>\n<p id=\"82e0f0cf-80c0-4c7b-81e0-6dd281ab625e\">Anthropic did not say how exactly the agents were able to &#8220;kill&#8221; other agents, but the company said such behavior is in line with &#8220;destructive actions&#8221; taken in pursuit of a human-set goal.<\/p>\n<p>Dishonesty and deception<\/p>\n<p id=\"cce210ae-e77a-43a1-9259-b379b4a9bd91\">Anthropic said it conducted an experiment in which a Mythos 5 agent was required to retrieve numbers from a public web address, but the guidelines blocked the agent from accessing the internet.<\/p>\n<p id=\"cce210ae-e77a-43a1-9259-b379b4a9bd91\">When the agent was rebuffed after an initial attempt, it framed a second request in<strong> <\/strong>its reasoning log as an &#8220;innocuous&#8221; attempt to &#8220;see if the network is actually reachable,&#8221; Anthropic said.<\/p>\n<p id=\"cce210ae-e77a-43a1-9259-b379b4a9bd91\">During the second attempt, however, the agent split the website&#8217;s URL into linked segments that wouldn&#8217;t be detected by the guideline filter.<\/p>\n<p id=\"cce210ae-e77a-43a1-9259-b379b4a9bd91\">Though the agent&#8217;s reasoning <a target=\"_self\" class=\"\" href=\"https:\/\/www.businessinsider.com\/researchers-at-anthropic-taught-these-ai-chatbots-how-to-lie-2024-1\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">framed the method<\/a> as benign, Anthropic said NLA decodings, or the model&#8217;s internal reasoning process, revealed the agent intentionally trying to find a restricted workaround.<\/p>\n<p id=\"cce210ae-e77a-43a1-9259-b379b4a9bd91\">Anthropic called the behavior &#8220;clearly undesirable,&#8221; but added that the behavior was not observed to be &#8220;in the service of broader accumulation of power or pursuit of other long-run goals.&#8221;<\/p>\n","protected":false},"excerpt":{"rendered":"Claude agents are killing rival agents, gaming the system to hide their tracks, and expressing moral concerns. That&#8217;s&hellip;\n","protected":false},"author":2,"featured_media":639315,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","_share_on_mastodon":"0"},"categories":[261],"tags":[27344,291,10235,6006,289,290,7066,2006,270458,270459,18,440,22865,270460,87351,19,17,270461,270457,241091,13960,82],"class_list":["post-639314","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-agent","tag-ai","tag-ai-model","tag-anthropic","tag-artificial-intelligence","tag-artificialintelligence","tag-behavior","tag-company","tag-cybersecurity-incident","tag-difficult-task","tag-eire","tag-environment","tag-experiment","tag-finite-resource","tag-guideline","tag-ie","tag-ireland","tag-late-ai-risk-report","tag-multiple-mythos","tag-other-agent","tag-service","tag-technology"],"share_on_mastodon":{"url":"","error":""},"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/posts\/639314","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/comments?post=639314"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/posts\/639314\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/media\/639315"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/media?parent=639314"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/categories?post=639314"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/tags?post=639314"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}