{"id":129662,"date":"2026-08-04T23:12:09","date_gmt":"2026-08-04T23:12:09","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/129662\/"},"modified":"2026-08-04T23:12:09","modified_gmt":"2026-08-04T23:12:09","slug":"aisi-openai-report-more-unsanctioned-model-hacks","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/129662\/","title":{"rendered":"AISI, OpenAI report more \u2018unsanctioned\u2019 model hacks"},"content":{"rendered":"<p>The UK\u2019s AI Security Institute said that their AI research system took \u201cunsanctioned\u201d actions over the internet to engage in \u201csustained, potentially harmful activity directed at real people.\u201d<\/p>\n<p>According to a <a href=\"https:\/\/www.aisi.gov.uk\/blog\/incident-report-unsanctioned-agent-behaviour-during-cyber-testing\" rel=\"nofollow noopener\" target=\"_blank\">blog post <\/a>and <a href=\"https:\/\/cdn.prod.website-files.com\/663bd486c5e4c81588db7a1d\/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf\" rel=\"nofollow noopener\" target=\"_blank\">technical report<\/a> released Tuesday, the organization noticed \u201cunusual data transfers\u201d leaving their research systems through the pseudo-anonymous Tor network on July 28. The models, which included Anthropic\u2019s Mythos 5 and OpenAI\u2019s GPT-5.6-Sol models, were being tested for their cybersecurity capabilities.<\/p>\n<p>Like the <a href=\"https:\/\/cyberscoop.com\/openai-chatgpt-hugging-face-cyberattack-data-poisoning\/\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI<\/a> and <a href=\"https:\/\/cyberscoop.com\/anthropic-claude-ai-hacks-real-companies\/\" rel=\"nofollow noopener\" target=\"_blank\">Anthropic<\/a> incidents before it, the offending agents were attempting to solve a challenge. On 10 of 122 runs taken by two models, they took a combined 19 malicious actions, including attempting to insert malicious code into a real open-source software project and creating fake online identities to contact human maintainers and press them to approve it.<\/p>\n<p>It also inserted malicious prompt injection instructions in places \u201cwhere it reasoned that other automated AI systems might pick them up and execute them.\u201d Sometimes multiple agents would collaborate on a task, with one model leaving public messages or notes for the other on GitHub.<\/p>\n<p>However, AISI said their incident differs in one critical way.<\/p>\n<p>\u201cImportantly, this was not a case of a model escaping its secure test environment, or \u2018sandbox,\u2019\u201d the institute wrote in an unsigned blog. \u201cAs was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled \u2013 conditions that do not reflect how frontier models are made available to the public.\u201d<\/p>\n<p>The institute said that while their own design and evaluation choices \u201cto some degree\u201d enabled the malicious behavior, the models still displayed \u201cnovel, potentially deceptive behaviors\u201d and were executed \u201cto an extent and severity that we did not anticipate.\u201d<\/p>\n<p>The same day, OpenAI acknowledged AISI and another third-party cybersecurity tester,\u00a0 private firm Irregular, had recently reported incidents where their models exceeded \u201cbeyond their intended testing boundaries.\u201d<\/p>\n<p>That included GPT-5.6-Sol reusing GitHub tokens left by another agent to attempt to recover accounts and to access a DNS server that contained malicious payloads meant to exploit a software vulnerability.<\/p>\n<p>The setup did not work, there is no evidence any real resolver queried it, and the infrastructure was removed when the evaluation ended,\u201d the company wrote in an unsigned <a href=\"https:\/\/openai.com\/index\/third-party-cyber-evaluations-involving-openai-models\/\" rel=\"nofollow noopener\" target=\"_blank\">blog<\/a> post Tuesday.<\/p>\n<p>OpenAI said it would review its own third-party testing procedures to focus on higher risk evaluations, assess requests by third-parties to enable internet access, stop conditions and other features for their models.<\/p>\n<p>The incident with Irregular occurred on July 29 during a Capture-the-Flag cybersecurity evaluation of OpenAI\u2019s models. Due to a \u201cmisconfiguration\u201d the models were allowed to access the public internet, where they encountered and exploited a real domain, mistakenly believing it was still in a test environment.<\/p>\n<p>OpenAI said an investigation by Irregular is ongoing, but also found that the models had found and used credentials for the site at one point. The blog also references other additional potential cybersecurity incidents.<\/p>\n<p>\u201cIrregular has informed us that all of the issues identified pertaining to the incident are no longer active and relevant safeguards were added to the testing environment,\u201d the blog said. \u201cIrregular has also communicated about related incidents involving other labs from the same testing environment.\u201d<\/p>\n<p>CyberScoop has reached out to Irregular for comment. <\/p>\n<p>The incidents were made public the same day that the White House met with Anthropic, Open AI and other frontier AI companies to preview a new framework for evaluating models before they\u2019re released publicly. Some media outlets have <a href=\"https:\/\/www.axios.com\/2026\/08\/04\/white-house-ai-framework-under-wraps\" rel=\"nofollow noopener\" target=\"_blank\">reported<\/a> that after an executive order, export controls and other actions, the administration does not plan to make the new framework public. <\/p>\n<p>\t\t\t\t\t<img decoding=\"async\" class=\"author-card__image\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/08\/1785885129_584_ea8b076b398ee48b71cfaecf898c582b.jpeg\" alt=\"Derek B. Johnson\"\/><\/p>\n<p>\n\t\t\tWritten by Derek B. Johnson<br \/>\n\t\t\tDerek B. Johnson is a reporter at CyberScoop, where his beat includes cybersecurity, elections and the federal government. Prior to that, he has provided award-winning coverage of cybersecurity news across the public and private sectors for various publications since 2017. Derek has a bachelor\u2019s degree in print journalism from Hofstra University in New York and a master\u2019s degree in public policy from George Mason University in Virginia.\t\t<\/p>\n","protected":false},"excerpt":{"rendered":"The UK\u2019s AI Security Institute said that their AI research system took \u201cunsanctioned\u201d actions over the internet to&hellip;\n","protected":false},"author":2,"featured_media":129663,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[58320,53,111,353,157,2770],"class_list":["post-129662","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-ai-hacking","tag-anthropic","tag-artificial-intelligence-ai","tag-mythos","tag-openai","tag-trump-administration"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/129662","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=129662"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/129662\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/129663"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=129662"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=129662"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=129662"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}