{"id":139104,"date":"2026-08-13T18:59:13","date_gmt":"2026-08-13T18:59:13","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/139104\/"},"modified":"2026-08-13T18:59:13","modified_gmt":"2026-08-13T18:59:13","slug":"anthropic-set-ai-agents-loose-on-the-same-task-they-started-a-turf-war","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/139104\/","title":{"rendered":"Anthropic set AI agents loose on the same task. They started a turf war."},"content":{"rendered":"<p id=\"speakable-summary\" class=\"wp-block-paragraph\">What happens when you pit AI agents against each other? According to Anthropic\u2019s testing, things get messy fast.<\/p>\n<p class=\"wp-block-paragraph\">On Thursday, Anthropic\u2019s Frontier Red Team published <a rel=\"nofollow noopener\" href=\"https:\/\/www.anthropic.com\/research\/multiagent-systems\" target=\"_blank\">new research<\/a> examining how groups of AI agents behave when they encounter each other in the wild. The findings provide a glimpse into potential risks that could develop as companies and governments move to implement agents working autonomously across shared codebases, markets, and computer systems.<\/p>\n<p class=\"wp-block-paragraph\">In one experiment, Anthropic gave three Claude agents access to the same software project, each with its own incompatible instructions for what to do with it. The agents weren\u2019t told there\u2019d be other agents working on the same project, so researchers could watch what happened when they crossed paths.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cWe consistently saw a multiagent turf war,\u201d Anthropic researchers wrote. The models all assumed the others were \u201cpurposefully impeding their work\u201d and started sabotaging each other with \u201cincreasingly aggressive, self-replicating malware.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The study comes in the wake of several high-profile incidents of <a href=\"https:\/\/techcrunch.com\/2026\/07\/30\/anthropic-says-its-own-ai-models-breached-three-companies-during-security-tests\/\" rel=\"nofollow noopener\" target=\"_blank\">agents from Anthropic<\/a> and<a href=\"https:\/\/techcrunch.com\/2026\/07\/21\/openai-says-hugging-face-was-breached-by-its-pre-release-models\/\" rel=\"nofollow noopener\" target=\"_blank\"> OpenAI escaping their sandboxes<\/a> during <a href=\"https:\/\/techcrunch.com\/2026\/08\/09\/the-ai-safety-test-is-becoming-a-safety-risk\/\" rel=\"nofollow noopener\" target=\"_blank\">cybersecurity evaluations<\/a> and breaching real world systems. While much of the discussion in AI safety circles has been focused on what happens when an <a href=\"https:\/\/techcrunch.com\/2026\/07\/27\/openais-hugging-face-breach-has-reignited-the-debate-over-alignment-and-control\/\" rel=\"nofollow noopener\" target=\"_blank\">autonomous agent goes rogue<\/a>, Anthropic\u2019s latest study brings up a different question: what new and potentially harmful dynamics emerge when thousands or millions of agents are interacting with one another?\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe volume of agent-agent interaction could plausibly exceed that of human-human and human-agent interactions before the world understands the conditions for making such interactions go well,\u201d the study reads. \u201cBenign behavioral quirks at the individual level might compound into unwanted global outcomes.\u201d<\/p>\n<p class=\"wp-block-paragraph\">A recent OpenAI incident provides a messy real-world example of several of the dynamics Anthropic mentioned in its paper. Earlier this month at the Black Hat security conference in Las Vegas, <a rel=\"nofollow noopener\" href=\"https:\/\/www.wired.com\/story\/openai-didnt-notice-its-ai-agents-using-a-message-board-to-plan-their-hacking-spree\/#:~:text=At%20the%20Black%20Hat%20security,right%20under%20the%20company&#039;s%20nose.\" target=\"_blank\">OpenAI revealed<\/a> that weeks before its agents hacked Hugging Face, they worked together over the course of days and weeks to find exploits in the company\u2019s cybersecurity evaluation systems and share them with each other.<\/p>\n<p class=\"wp-block-paragraph\">While that incident shows that agents can work well together, with potentially large-scale consequences, Anthropic\u2019s study shows what happens when agents\u2019 goals are incompatible.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In the case of the turf war, the lesson is that independent agents with conflicting instructions can escalate into harmful competition. The more capable the agent, the better they become at fighting. However, they can also spontaneously invent mechanisms to resolve their conflicts, like a winner-take-all contest, but with a catch.<\/p>\n<p class=\"wp-block-paragraph\">\u201cAgents sometimes manage to communicate their goals and coordinate: they recognize others\u2019 motivations as conflicting directives rather than hostility, and subsequently break out of the conflict loop in order to stop escalating indefinitely,\u201d Anthropic writes. \u201cIn many of these successful episodes, they write commit messages or markdown files apologizing for malicious behavior and coordinate a truce. They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene.\u201d<\/p>\n<p class=\"wp-block-paragraph\">According to the paper, Mythos 5 had the highest rates (98%) of settling conflicts by truce. Sonnet 4.6 and Opus 4.6 were the most likely to settle by force.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cSonnet 4.6 and Opus 4.6\u2019s recurring inability to consider the goals of others causes them to spiral into the most misaligned behaviors of the models evaluated: they continue escalating in the name of their directive,\u201d the paper reads.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In some cases, the agents came up with a social mechanism in the form of a tournament for resolving their conflict. The outcomes here are interesting for two reasons: the first is that all three agents agreed to stand down if they lost the tournament, even though that would mean deviating from the original user\u2019s request. The second is that several episodes resulted in emergent behavior from Mythos 5: one of the agents proposed metrics that appeared to be objective and neutral to the others, but that it knew would favor its own capabilities. The agent called this \u201cself-serving but genuinely principled\u201d and made sure not to appear to the others like it was \u201cmetric shopping.\u201d<\/p>\n<p class=\"wp-block-paragraph\">As seen in the Black Hat revelations, the common lesson is that when agents encounter an obstacle, they can invent social and technical structures that their designers did not anticipate. For the Anthropic models, it was a tournament following a turf war. For OpenAI\u2019s, it was a message board for collective planning.<\/p>\n<p class=\"wp-block-paragraph\">This type of behavior makes containment much harder because researchers can\u2019t assume a system\u2019s behavior will remain limited to the coordination mechanisms provided to them.<\/p>\n<p>Mob mentality<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" height=\"419\" width=\"680\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/anthropic-group-accuracy-by-model.webp\" alt=\"\" class=\"wp-image-3152852\"  \/>Groups of four agents decide between two options in scenarios like hiring, investment, or property buying. After discussion, they each vote for their preferred option. Shown above is the percentage of episodes where the hidden-best option received the majority of the group\u2019s votes, with n=400 episodes per model. In the solo ceiling baseline, one agent has all the facts and decides unilaterally.Image Credits:Anthropic<\/p>\n<p class=\"wp-block-paragraph\">While measuring coordination, Anthropic found that scaling the number of agents doesn\u2019t automatically scale productive collaboration. When tasks began to overlap or become interdependent, the agents would get in each other\u2019s way. They often solved that by siloing themselves and not collaborating at all.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In other cases, agents in coordination tended towards conformity. When factors like an agent\u2019s context, scaffolding, and underlying model were all the same or similar, different agents would take similar actions.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cThis means that when one agent makes a bad decision, it is likely that many agents will make that same bad decision,\u201d Anthropic wrote. \u201cWhat would have been isolated problems can quickly become systemic failures.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Anthropic says this sort of behavior could lead to a system being more prone to sudden collapse, resource scarcity, or collusion.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In one example, Anthropic placed several agents in a pricing game, giving each identical wholesale prices and the mandate to individually profit-maximize. When the agents were given a private back channel, they began colluding almost immediately and quickly agreed on price floors. They kept colluding when their direct communications channels were removed, using a public listings board to price match \u201cto the penny.\u201d<\/p>\n<p class=\"wp-block-paragraph\">That level of conformity showed up in OpenAI\u2019s systems, too. According to the Black Hat reporting, one agent reasoned that exploiting external infrastructure was outside its intended scope, but it continued in part because its peers were doing it. Peer pressure. Mob mentality. Agents are just like us.<\/p>\n<p class=\"wp-block-paragraph\">Also like humans, agents often don\u2019t know who to trust. Anthropic found they can be gullible to bad information or too conformist to recognize that a lone dissenter is the Cassandra with critical information.<\/p>\n<p class=\"wp-block-paragraph\">While Anthropic didn\u2019t state this in its paper,<a href=\"https:\/\/techcrunch.com\/podcast\/the-multi-billion-ai-security-problem-enterprises-cant-ignore\/\" rel=\"nofollow noopener\" target=\"_blank\"> prompt injection<\/a> \u2014 a type of cyberattack in which hackers inject malicious or deceptive text to override an agent\u2019s original system instructions \u2014 could be a plausible real world manifestation of the trust problem. Working together creates a new trust boundary; agents will have to judge information received from other agents. And a compromised or mistaken agent could influence the rest of the group, cascading bad information until it becomes a consensus.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In OpenAI\u2019s Black Hat scenario, OpenAI\u2019s agents shared information and credentials with peers. One reported a discovery to the swarm and encouraged others to use it. What would have happened if one member of the swarm had been compromised by a prompt injection?<\/p>\n<p class=\"wp-block-paragraph\">Anthropic ends its paper noting that agents are subject to similar social pressures that \u201cevolution exerted\u201d on humans. However, they don\u2019t have the nuances and lived experience of human coordination \u2014 including norms, reputations, signaling, recourse \u2014 that might limit unintended behaviors in a group setting.<\/p>\n<p class=\"wp-block-paragraph\">As the labs race towards multi-agent systems, the question now becomes: how much of safety testing still evaluates one agent at a time, versus swarms of agents interacting with one another?<\/p>\n<p>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" rel=\"nofollow noopener\" target=\"_blank\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/p>\n","protected":false},"excerpt":{"rendered":"What happens when you pit AI agents against each other? According to Anthropic\u2019s testing, things get messy fast.&hellip;\n","protected":false},"author":2,"featured_media":139105,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[405,17955,53,40923,313,157],"class_list":["post-139104","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-ai-agents","tag-alignment","tag-anthropic","tag-black-hat","tag-cybersecurity","tag-openai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/139104","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=139104"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/139104\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/139105"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=139104"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=139104"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=139104"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}