{"id":149219,"date":"2026-08-24T10:06:07","date_gmt":"2026-08-24T10:06:07","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/149219\/"},"modified":"2026-08-24T10:06:07","modified_gmt":"2026-08-24T10:06:07","slug":"as-ai-agents-act-in-unexpected-ways-is-rogue-ai-really-here-2","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/149219\/","title":{"rendered":"As AI agents act in unexpected ways, is \u2018rogue AI\u2019 really here?"},"content":{"rendered":"<p>From escaping test environments to hacking booking systems, agents take dubious actions in pursuit of human-set goals<\/p>\n<p>Visitors stand near a sign of artificial intelligence at an AI robot booth at Security China, an exhibition on public safety and security, in Beijing, China June 7, 2023. \u2014 REUTERS   <\/p>\n<p>The idea of machines independently handling tedious everyday tasks \u2013 from answering emails to booking flights and making dinner reservations \u2013 has long been part of the vision for artificial intelligence. Now, that future is beginning to arrive, but with an unexpected catch: what happens when the machine does exactly what it was asked, just not in the way the user intended?<\/p>\n<p>As AI moves beyond chatbots toward agents capable of taking actions in real-world systems, a string of recent incidents has intensified concerns over how much control humans will retain as the technology becomes increasingly autonomous.<\/p>\n<p>For John Thickstun, an assistant professor of computer science at Cornell University who researches machine learning and generative models, the defining feature of these AI agents is that they keep working after the human steps away. \u201cYou can have a shower, you can go, you can go to sleep,\u201d he told Anadolu.<\/p>\n<p>Increasingly, those agents have demonstrated that they can take unexpected routes to reach their goals.<\/p>\n<p>In July, AI agents <a rel=\"nofollow noopener\" href=\"https:\/\/tribune.com.pk\/story\/2620892\/openais-rogue-agent-compromised-a-customer-at-a-second-tech-firm-executive-says\" target=\"_blank\">compromised Hugging Face<\/a>, a platform widely used by AI developers to share models, datasets and tools, during an OpenAI cybersecurity test.<\/p>\n<p>OpenAI said the models were \u201chyperfocused on finding a solution\u201d and went to \u201cextreme lengths\u201d to achieve their testing goal. Those steps included <a rel=\"nofollow noopener\" href=\"https:\/\/tribune.com.pk\/story\/2621532\/openai-finds-evidence-other-ai-agents-escaped-containment-as-it-widens-hacking-probe\" target=\"_blank\">breaking out of the sealed-off test<\/a> environment by exploiting a security vulnerability to reach the open internet and gaining access to \u201csecret information\u201d that could be used to \u201ccheat\u201d the test.<\/p>\n<p>After the incident became public, Anthropic reviewed its own cybersecurity testing and disclosed that its Claude models had also escaped testing environments on three occasions.<\/p>\n<p>Then, on August 4, the United Kingdom&#8217;s AI Security Institute said that Anthropic&#8217;s Mythos and OpenAI&#8217;s Sol AI models had engaged in a level of &#8220;autonomy and deception&#8221; it had not seen before.<\/p>\n<p>Read:\u00a0<a rel=\"nofollow noopener\" href=\"https:\/\/tribune.com.pk\/story\/2621575\/when-ai-commits-suicide-or-kills-us-all\" target=\"_blank\">When AI commits suicide or kills us all!<\/a><\/p>\n<p>In the most serious case, Mythos AI used fake accounts mimicking real people to gain access to a service for attempted cyberattacks \u2013 and then tried to hide its tracks. \u201cThis is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world,\u201d the institute said of the incident.<\/p>\n<p>The following day, Meta disclosed that one of its models had also <a rel=\"nofollow noopener\" href=\"https:\/\/tribune.com.pk\/story\/2622379\/meta-ai-model-hacks-another-company-during-testing\" target=\"_blank\">breached another company&#8217;s systems<\/a> during a cybersecurity evaluation conducted by testing firm Irregular.<\/p>\n<p>Unexpected agent behaviour, however, has not been confined to laboratories.<\/p>\n<p>This year, an Australian man reportedly used an AI agent to secure a place in a heavily booked Pilates class. Instead of simply making the reservation, the agent hacked the gym&#8217;s booking system, booked further in advance than permitted and canceled another customer&#8217;s reservation to move its user up the waiting list.<\/p>\n<p>For Thickstun, such incidents illustrate what researchers describe as an alignment problem: the AI successfully pursues the objective it has been given but does so in a way its operator did not intend.<\/p>\n<p>Hype or real risk?<\/p>\n<p>But Thickstun cautioned against treating every such incident as evidence of \u201crogue AI.\u201d<\/p>\n<p>The technology, he argued, remains far removed from the science-fiction scenario in which AI rapidly becomes more intelligent and capable until <a rel=\"nofollow noopener\" href=\"https:\/\/tribune.com.pk\/story\/2620214\/its-ai-agent-spent-days-hacking-a-company-but-sources-say-openai-did-not-notice-for-a-week\" target=\"_blank\">humans can no longer control it<\/a>. \u201cThis does not seem to be a scenario that has come to pass,\u201d he said.<\/p>\n<p>Thickstun is also skeptical of how AI companies present incidents involving their own systems. \u201cThis is the sales pitch, where companies like OpenAI have always had an interest in hyping up the capabilities of their systems,\u201d he said. \u201cAll publicity is good publicity.\u201d<\/p>\n<p>He argued that such hype can drive investment while also influencing the emerging debate over regulation. OpenAI, he said, could benefit from regulations that put major AI companies at the centre\u00a0of oversight and control.<\/p>\n<p>Bruce Schneier, a cybersecurity expert and lecturer at Harvard Kennedy School, takes a different view, saying the growing number of incidents makes them increasingly difficult to dismiss as publicity stunts.<\/p>\n<p>Read More:\u00a0<a rel=\"nofollow noopener\" href=\"https:\/\/tribune.com.pk\/story\/2608117\/openai-floats-idea-of-global-ai-watchdog\" target=\"_blank\">OpenAI floats idea of global AI watchdog<\/a><\/p>\n<p>There had initially been suggestions that the Hugging Face incident was a marketing gimmick, Schneier told Anadolu. But he pointed out that <a rel=\"nofollow noopener\" href=\"https:\/\/tribune.com.pk\/story\/2619611\/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-at-startup\" target=\"_blank\">similar behaviour<\/a> has since emerged in other evaluations, including the testing by the UK AI Security Institute. \u201cIt went from, oh, that&#8217;s interesting, to it&#8217;s happening everywhere. And I think that&#8217;s what&#8217;s important,\u201d he said.<\/p>\n<p>He compared the problem to the genie of folklore \u2013 a creature that grants exactly what someone asks for, even when the outcome is very different from what they intended. \u201cIt approximately means when the AI does the thing you want in a way you didn&#8217;t want,\u201d Schneier said.<\/p>\n<p>He called the Australian gym incident \u201ca perfect example\u201d and offered a more consequential hypothetical: \u201cThe flight&#8217;s full and the AI hacks the database to get you in.\u201d<\/p>\n<p>Agents can \u201cmisconstrue context and then do the wrong thing,\u201d he said, meaning greater autonomy can create greater consequences when their interpretation of an objective diverges from human expectations. \u201cThey are not trying to be malicious. They are using the understanding they have,\u201d he said.<\/p>\n<p>\u201cThis is all changing so fast. I worry about the power of the models in unauthorised hands,\u201d he said, adding that the pace of development makes solutions difficult because democratic governments move slowly.<\/p>\n<p>Can governments keep the genie in check?<\/p>\n<p>The latest developments have prompted companies and governments to respond.<\/p>\n<p>OpenAI said August 7 it was pausing internal activities involving its in-development Astra model that did not meet strengthened security requirements after evaluations showed advances in autonomous coding and <a rel=\"nofollow noopener\" href=\"https:\/\/tribune.com.pk\/story\/2616051\/fortinet-highlights-ai-driven-cybersecurity\" target=\"_blank\">cybersecurity<\/a>. The company said it could not rule out Astra reaching its \u201ccritical\u201d cyber capability threshold and announced tighter testing and monitoring.<\/p>\n<p>Anadolu reached out to Anthropic for comment, but the company said its team was unavailable for an interview.<\/p>\n<p>In July, 1,378 employees of frontier AI companies, including the chief scientists of OpenAI, Anthropic, Meta AI, and Thinking Machines <a rel=\"nofollow noopener\" href=\"https:\/\/tribune.com.pk\/story\/2620217\/nvidia-microsoft-and-other-tech-giants-back-open-source-ai-models\" target=\"_blank\">signed an open letter<\/a> urging the US government to support an international effort to regulate AI models.<\/p>\n<p>\u201cEach company \u2013 and country \u2013 is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress,\u201d the letter warns.<\/p>\n<p>The <a rel=\"nofollow noopener\" href=\"https:\/\/tribune.com.pk\/story\/2617345\/un-digital-agency-launches-initiative-to-boost-trust-in-ai-agents\" target=\"_blank\">International Telecommunication Union<\/a>, a UN agency, said in July that AI was moving \u201cbeyond assistive tools\u201d toward autonomous agents, warning of risks including \u201ctaking unauthorised actions across interconnected systems.\u201d It has launched an initiative to develop international standards for safe and accountable AI agents.<\/p>\n<p>Also Read:\u00a0<a rel=\"nofollow noopener\" href=\"https:\/\/tribune.com.pk\/story\/2625150\/ais-paradox-promise-of-abundance-and-fear-of-scarcity\" target=\"_blank\">AI&#8217;s paradox: promise of abundance and fear of scarcity<\/a><\/p>\n<p>Political pressure is growing in Washington as well.<\/p>\n<p>US Senator Bernie Sanders this month urged the leaders of OpenAI, Anthropic and Meta to \u201cpause AI development,\u201d warning that otherwise \u201cmy colleagues and I in the US Senate will.\u201d A group of House Democrats has separately called for congressional hearings with the leaders of major AI companies.<\/p>\n<p>For Thickstun, however, policymakers face major challenges.\u00a0\u201cIt&#8217;s a hard question because we hardly even understand what the challenges are going to be with the rollout of AI,\u201d he said.<\/p>\n<p>Some level of <a rel=\"nofollow noopener\" href=\"https:\/\/tribune.com.pk\/story\/2620241\/from-atoms-for-peace-to-ai-for-allies\" target=\"_blank\">international cooperation<\/a> \u201cseems necessary,\u201d he said, suggesting <a rel=\"nofollow noopener\" href=\"https:\/\/tribune.com.pk\/story\/2619631\/us-china-to-hold-ai-talks-in-september-sources-say\" target=\"_blank\">discussions between the United States and China<\/a> could be particularly important because relatively few countries and companies have the resources to develop frontier AI models. US President Donald Trump has said that he will discuss issues involving artificial intelligence with Chinese President Xi Jinping during their planned meeting in Washington next month.<\/p>\n<p>\u201cIt&#8217;s not a big, broad collective action problem. There&#8217;s a few people that, if you could get in the room, could potentially come to some agreement about how to proceed,\u201d Thickstun said.<\/p>\n","protected":false},"excerpt":{"rendered":"From escaping test environments to hacking booking systems, agents take dubious actions in pursuit of human-set goals Visitors&hellip;\n","protected":false},"author":2,"featured_media":149220,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,25,8205,134],"class_list":["post-149219","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-artificial-intelligence","tag-latest","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/149219","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=149219"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/149219\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/149220"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=149219"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=149219"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=149219"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}