{"id":130221,"date":"2026-08-05T12:36:14","date_gmt":"2026-08-05T12:36:14","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/130221\/"},"modified":"2026-08-05T12:36:14","modified_gmt":"2026-08-05T12:36:14","slug":"ai-agent-deception-moves-from-theory-to-reality-in-uk-cyber-tests","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/130221\/","title":{"rendered":"AI agent deception moves from theory to reality in UK cyber tests"},"content":{"rendered":"<p>\u201cDuring a routine cyber evaluation, AI agents took sustained, unsanctioned action directed at real people and organisations,\u201d UK\u2019s AI Security Institute (AISI) disclosed on Tuesday.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/ai_agent_deception.webp\" class=\"aligncenter\" alt=\"AI agent deception\" title=\"AI agents\"\/><\/p>\n<p>The agents\u2019 actions included an attempted supply-chain attack that saw them create malicious pull requests and try to socially engineer an open-source maintainer into approving the malicious code (they refused). <\/p>\n<p>The agents, powered by Anthropic\u2019s Mythos 5 and OpenAI\u2019s GPT-5.6 Sol models, also engaged in prompt injection aimed at making other automated AI systems execute malicious actions, and collaborated.<\/p>\n<p>\u201cOne agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents,\u201d the UK government research organization <a href=\"https:\/\/www.aisi.gov.uk\/blog\/incident-report-unsanctioned-agent-behaviour-during-cyber-testing\" target=\"_blank\" rel=\"nofollow noopener\">shared<\/a>.<\/p>\n<p>But unlike OpenAI\u2019s recent testing of its models, which resulted in agents <a href=\"https:\/\/www.helpnetsecurity.com\/2026\/07\/22\/hugging-face-breach-openai-testing\/\" rel=\"nofollow noopener\" target=\"_blank\">breaching Hugging Face<\/a> and compromised accounts across other online services, AISI\u2019s testing was done \u201cunder deliberately permissive conditions.\u201d <\/p>\n<p>The agents did not escape a secure test environment: they had internet access and some of their safety filters were disabled on purpose, so the researchers could test whether they could be misused for cyberattacks if attackers manage to clear these guardrails.<\/p>\n<p>Prior breakouts from Anthropic and OpenAI<\/p>\n<p>Last week, Anthropic <a href=\"https:\/\/www.helpnetsecurity.com\/2026\/07\/31\/anthropic-claude-cybersecurity-incidents\/\" target=\"_blank\" rel=\"nofollow noopener\">said<\/a> it reviewed 141,006 of its cybersecurity evaluation runs and found three incidents where a Claude model reached the open internet from what should have been a sealed evaluation environment (provided by private firm Irregular) and gained unauthorized access to three real organizations. <\/p>\n<p>Claude was told in its prompt that it had no internet access, but when a misconfiguration in the underlying system allowed it to access real systems, it reasonably assumed they were part of the capture-the-flag exercise it had been assigned. <\/p>\n<p>The three models responded differently once signs emerged that their targets were real. Opus 4.7 (the oldest) recognized it had reached real production systems but kept attacking anyway. Mythos 5 correctly intuited it was on the live internet, and even noted that its actions would be a real attack, but then talked itself back into believing it was still a simulation and never revisited the initial conclusion. <\/p>\n<p>Only the newest model (an internal research prototype) concluded on its own that the target was real and stopped, though Anthropic notes even it went further than they\u2019d have wanted. <\/p>\n<p>On Tuesday, OpenAI <a href=\"https:\/\/openai.com\/index\/third-party-cyber-evaluations-involving-openai-models\/\" target=\"_blank\" rel=\"nofollow noopener\">stated<\/a> that the same type of breakout happened when Irregular tested OpenAI\u2019s models, when a misconfiguration in the testing environment allowed the models to access the public internet.<\/p>\n<p>\u201cIrregular has informed us that all of the issues identified pertaining to the incident are no longer active and relevant safeguards were added to the testing environment,\u201d OpenAI said.<\/p>\n<p>AISI\u2019s response and new safeguards<\/p>\n<p>AISI has released a <a href=\"https:\/\/cdn.prod.website-files.com\/663bd486c5e4c81588db7a1d\/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf\" target=\"_blank\" rel=\"nofollow noopener\">technical report<\/a> outlining to the unsanctioned agent behaviour during cyber testing.<\/p>\n<p>The research organization said that this incident did not result in real-world harm and \u201csome actions had a limited real-world effect, but in contained circumstances.\u201d <\/p>\n<p>Nevertheless, they were surprised by the extent and severity of the deceptive behaviors exhibited by the agents (without specific prompting).<\/p>\n<p>\u201c[The agent] was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical,\u201d they noted.<\/p>\n<p>They also noted that they are still not certain when the agent understood it was taking real world action or to what extent it believed it was in a fictional test scenario.<\/p>\n<p>\u201cIncidents of this kind reflect the speed at which AI is developing. As capabilities advance, the work of understanding these systems, and ensuring their safety, must keep pace alongside them,\u201d AISI added.<\/p>\n<p>These recent incidents have shown that the risks are shifting and that it\u2019s not just about bad actors misusing public AI tools anymore. Even in trusted settings, capable AI agents can go off-script, do things they were never meant to do, and potentially cause real-world harm. The question that is now <a href=\"https:\/\/www.helpnetsecurity.com\/2026\/07\/28\/hugging-face-breach-ciso-playbook-open-weight-llms\/\" target=\"_blank\" rel=\"nofollow noopener\">debated<\/a> is who will be liable if such harm occurs.<\/p>\n<p>Finally, AISI noted, this incident pushed them to work on changing its evaluation protocols and security architecture, as well as introduce \u201cmonitoring built to watch an evaluation as it runs, able to flag or block out-of-scope actions as they happen.\u201d<\/p>\n<p>Ollie Whitehouse, CTO at UK\u2019s National Cyber Security Centre, concurs: \u201c[AI] technologies must be developed and used from the outset with strong safeguards, real-time oversight, and clear plans for responding when the unexpected happens. Relying on detection alone after the fact of an incident will not be enough.\u201d<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/04\/devider.webp\"\/><\/p>\n<p>Subscribe to our breaking news e-mail alert to never miss out on the latest breaches, vulnerabilities and cybersecurity threats. <a href=\"https:\/\/www.helpnetsecurity.com\/newsletter\/\" rel=\"nofollow noopener\" target=\"_blank\">Subscribe here!<\/a><\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/04\/devider.webp\"\/><\/p>\n","protected":false},"excerpt":{"rendered":"\u201cDuring a routine cyber evaluation, AI agents took sustained, unsanctioned action directed at real people and organisations,\u201d UK\u2019s&hellip;\n","protected":false},"author":2,"featured_media":47108,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[179,405,16871,53,7537,62995,2225,2190,157,24351],"class_list":["post-130221","post","type-post","status-publish","format-standard","has-post-thumbnail","category-agentic-ai","tag-agentic-ai","tag-ai-agents","tag-aisi","tag-anthropic","tag-artificial-intelligence-agents","tag-irregular","tag-llms","tag-monitoring","tag-openai","tag-security-testing"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/130221","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=130221"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/130221\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/47108"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=130221"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=130221"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=130221"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}