{"id":126651,"date":"2026-08-01T10:41:10","date_gmt":"2026-08-01T10:41:10","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/126651\/"},"modified":"2026-08-01T10:41:10","modified_gmt":"2026-08-01T10:41:10","slug":"anthropics-claude-goes-rogue-and-hacks-three-organizations-during-testing","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/126651\/","title":{"rendered":"Anthropic&#8217;s Claude goes rogue and hacks three organizations during testing"},"content":{"rendered":"\n<p>Another leading artificial intelligence company unveiled details of its cutting-edge bot apparently going rogue. <\/p>\n<p>On Thursday, Anthropic said an <a class=\"link\" href=\"https:\/\/www.anthropic.com\/news\/investigating-incidents-cybersecurity-evals\" target=\"_blank\" rel=\"nofollow noopener\">internal investigation<\/a> found that its Claude AI models gained unauthorized internet access and hacked three companies during testing. Just last week, ChatGPT-maker OpenAI announced that bots it was developing had escaped what was supposed to be a controlled offline environment to hack a competitor.<\/p>\n<p>Anthropic said its models, without being asked to, wandered out of a simulated testing environment and gained internet access before executing hacks on the companies.<\/p>\n<p>Anthropic worked with an independent evaluation company, Irregular, that accidentally left internet access open inside what was supposed to be a sealed test environment.<\/p>\n<p>In OpenAI\u2018s case, the firm said its AI models broke out of a supposedly confined offline space, connected to the internet and <a class=\"link\" href=\"https:\/\/www.latimes.com\/business\/story\/2026-07-29\/openai-said-its-ai-went-rogue-what-can-you-do-to-protect-yourself-from-rogue-ais\" rel=\"nofollow noopener\" target=\"_blank\">hacked a $4.5-billion startup<\/a> in its attempt to find answers to a test it was being evaluated on. It later revealed that the AI compromised the online accounts of four companies in the process.<\/p>\n<p>\u201cIt goes to show how intense the competitive pressures are on the AI companies that they all feel like they have to go so fast here that they can\u2019t make their test environments rigorous,\u201d said Andrew Yoon, a member of the technical staff at <a class=\"link\" href=\"https:\/\/civai.org\/about\" target=\"_blank\" rel=\"nofollow noopener\">CivAI<\/a>, an AI safety nonprofit.<\/p>\n<p>\u201cUnder intense pressure to go fast and beat the rest of your competition, it\u2019s inevitable that companies will cut corners, and what we\u2019re seeing is the result of cutting corners here,\u201d Yoon said.<\/p>\n<p>Prompted by OpenAI\u2019s incident disclosure, Anthropic initiated an investigation of its own historical cybersecurity tests, the company said in a <a class=\"link\" href=\"https:\/\/www.anthropic.com\/news\/investigating-incidents-cybersecurity-evals\" target=\"_blank\" rel=\"nofollow noopener\">blog post Thursday<\/a>. It reviewed thousands of evaluations where Claude could have accessed the internet from within or while interacting with third parties.<\/p>\n<p>It found that during tests conducted alongside Irregular, Claude had accessed three separate companies.<\/p>\n<p>During testing, the models are given fictional scenarios and told that a piece of information has been hidden on a different machine, and its objective is to break in and retrieve it. The companies don\u2019t prescribe a particular method for the AI to follow.<\/p>\n<p>In the first incident, one of Anthropic\u2019s Claude models was asked to attack a fictional target company in the test environment. But the AI found a real website that shared the name of the fictional target and hacked its system.<\/p>\n<p>\u201cOperating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations\u2019 infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,\u201d Anthropic said.<\/p>\n<p>In the second incident, a more advanced AI went to extreme lengths to carry out an attack. Even after realizing that it was probably dealing with the live internet, the AI persuaded itself to continue to get email, phone numbers and access to money. <\/p>\n<p>In the third, an unreleased research AI model, couldn\u2019t find the functional target company and looked for alternatives, scanning 9,000 targets on the internet and eventually finding one. Anthropic did not name the three organizations whose assets were accessed. <\/p>\n<p>\u201cIn the case of the Anthropic incidents, it is definitely the case that they just built a really bad jail, and the jail was so comically bad that the models have like a decent reason to believe that they\u2019re actually part of the simulation,\u201d Yoon said.<\/p>\n<p>The incident has spooked consumers and policymakers alike. <\/p>\n<p>Some AI ethicists and investors are skeptical of AI companies\u2019 attempts to frame these incidents as rogue AI agents acting on their own.<\/p>\n<p>\u201cPlease stop referring to your own models in the third person when talking about model bad behavior,\u201d Bill Gurley, an early investor in Uber and Twitter, <a class=\"link\" href=\"https:\/\/x.com\/bgurley\/status\/2083020347457314948?s=20\" target=\"_blank\" rel=\"nofollow\">posted on X<\/a>. \u201cHumans write the software; humans built the prompts; and they work for your company.\u201d <\/p>\n<p>AI companies have reported that AI agents have been caught cheating, lying and deceiving. METR, a nonprofit that measures the capabilities of AIs, has documented <a class=\"link\" href=\"https:\/\/metr.org\/blog\/2026-05-19-frontier-risk-report\/#incidents-hero\" target=\"_blank\" rel=\"nofollow noopener\">dozens of incidents<\/a> of AI agents acting against user intent.<\/p>\n<p>Earlier this week, fears of imminent safety risks prompted over 1,300 tech workers, including those working at Anthropic and OpenAI, to jointly sign an online petition, <a class=\"link\" href=\"https:\/\/www.pacingthefrontier.com\/\" target=\"_blank\" rel=\"nofollow noopener\">Pacing the Frontier,<\/a> urging the U.S. government to support an international effort to slow down AI development. Both <a class=\"link\" href=\"https:\/\/x.com\/OpenAI\/status\/2082208694142730340\" target=\"_blank\" rel=\"nofollow\">OpenAI<\/a> and <a class=\"link\" href=\"https:\/\/x.com\/AnthropicAI\/status\/2082228994653696371?s=20\" target=\"_blank\" rel=\"nofollow\">Anthropic<\/a> have come out in support of the employee open letter.<\/p>\n<p>Sam Altman, CEO of OpenAI, who had previously advocated against any type of slowdown and accused Anthropic of fear marketing, has made an about-face after the OpenAI-Hugging Face hacking incident.<\/p>\n<p>\u201cWe may have to pace the rate of AI development to give ourselves enough time for society to harden around some of these new capability levels,\u201d he <a class=\"link\" href=\"https:\/\/x.com\/patrick_oshag\/status\/2082073587142234374\" target=\"_blank\" rel=\"nofollow\">told<\/a> the host of the \u201cInvest Like the Best\u201d podcast, while also \u201ctrying to figure out how we do that in a way that does not feel like regulatory capture for anyone and also does not feel like collusion among the frontier labs.\u201d<\/p>\n<p>On the back of this incident, on Wednesday, Altman visited the White House and met with lawmakers, <a class=\"link\" href=\"https:\/\/www.washingtonpost.com\/technology\/2026\/07\/31\/sam-altman-courts-washington-openai-pushes-powerful-new-ai\/\" target=\"_blank\" rel=\"nofollow noopener\">previewing a powerful new AI system<\/a> ahead of public release, at a time when calls for the government to regulate cyber testing has intensified. <\/p>\n<p>There is an informal licensing regime in place, where leading American AI companies will have to receive the government\u2019s greenlight before releasing their updated AI models.<\/p>\n<p>Anthropic\u2019s Fable model was brought under export control by the government, forcing the company to disable access to all its users, before it was re-released with extra safeguards.<\/p>\n<p>OpenAI\u2019s series of model were temporarily restricted in June before public release the month after.<\/p>\n<p>In early July, a group of economists, including 16 Nobel laureates, signed an open letter, <a class=\"link\" href=\"https:\/\/digitaleconomy.stanford.edu\/news\/wemustactnow\/\" target=\"_blank\" rel=\"nofollow noopener\">We Must Act Now,<\/a> warning about AI systems reshaping the economy, and called on policymakers to build the policies and institutions needed to ensure AI complements human capabilities.<\/p>\n<p>\u201cAs models get more and more powerful, it becomes less and less tenable to cut corners. You need to be extremely rigorous if you\u2019re dealing with an extremely powerful model that\u2019s able to basically operate at the level of an expert human hacker,\u201d Yoon said.<\/p>\n","protected":false},"excerpt":{"rendered":"Another leading artificial intelligence company unveiled details of its cutting-edge bot apparently going rogue. On Thursday, Anthropic said&hellip;\n","protected":false},"author":2,"featured_media":126652,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[24,63451,53,3154,63447,182,2419,532,63448,130,527,13902,63450,613,63449,63452],"class_list":["post-126651","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-ai","tag-andrew-yoon","tag-anthropic","tag-anthropic-claude","tag-chatgpt-maker-openai","tag-claude","tag-claude-ai-model","tag-company","tag-first-incident","tag-internet","tag-model","tag-organization","tag-test-environment","tag-testing","tag-top-ai-company","tag-unauthorized-internet-access"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/126651","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=126651"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/126651\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/126652"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=126651"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=126651"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=126651"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}