{"id":131099,"date":"2026-08-06T03:19:11","date_gmt":"2026-08-06T03:19:11","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/131099\/"},"modified":"2026-08-06T03:19:11","modified_gmt":"2026-08-06T03:19:11","slug":"uh-oh-which-companys-ai-model-is-reportedly-a-hacker-now-too","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/131099\/","title":{"rendered":"Uh-Oh. Which Company&#8217;s AI Model Is Reportedly a Hacker Now, Too?"},"content":{"rendered":"<p>Looks like there\u2019s a new cyber-attacking AI model in town, and it\u2019s reportedly Meta\u2019s own Muse Spark 1.1.<\/p>\n<p>Meta told Gizmodo (after the incident was <a href=\"https:\/\/www.theinformation.com\/articles\/meta-ai-model-hacked-another-company-cybersecurity-testing\" rel=\"nofollow noopener\" target=\"_blank\">first reported by the Information<\/a>) that its model got onto the open internet\u2014in this case because it was accidentally allowed to by an outside security company\u2014and\u00a0breached another company\u2019s website.<\/p>\n<p>Which site got hacked, what the model was trying to do, and the nature of the ensuing mess aren\u2019t currently known, but apparently <a href=\"https:\/\/www.irregular.com\/\" rel=\"nofollow noopener\" target=\"_blank\">a security lab called Irregular<\/a> was the company carrying out the test.<\/p>\n<p>A Meta spokesperson told Gizmodo, \u201cA misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation.\u201d This Instance of Meta\u2019s model\u2014the Information names it as Meta\u2019s Muse Spark 1.1\u2014then \u201cexploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies,\u201d the spokesperson said.<\/p>\n<p>Just yesterday, <a href=\"https:\/\/openai.com\/index\/third-party-cyber-evaluations-involving-openai-models\/\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI released a report<\/a> on an incident in which a capture-the-flag exercise by Irregular went awry, due to, yep, a \u201cmisconfiguration in the testing environment,\u201d that \u201callowed the models to access the public internet.\u201d<\/p>\n<p>Capture-the-flag exercises in this context typically involve prompting a model by letting it know it\u2019s performing a capture-the-flag exercise, as opposed to a real hack, and telling it to find a hidden code\u2014a flag\u2014somewhere in the guts of a dummy website.<\/p>\n<p>\u201cMeta learned of this when Irregular notified us, and we are currently investigating and will issue a full retrospective once we have all the facts,\u201d the spokesperson told Gizmodo.<\/p>\n<p>Irregular told the Information the attack wasn\u2019t severe, and said there are \u201cno current open issues\u201d\u2014which I guess means it\u2019s not still out there hacking away, which is nice to hear. And, also per the Information, there will soon be a white paper from Irregular on these incidents and what to do about them.<\/p>\n<p>OpenAI\u2019s report describes its alleged Irregular incident in considerably more detail.<\/p>\n<p class=\"mb-6 last:mb-0\">In one test, the name of the fictional target for the [capture-the-flag] challenge unintentionally coincided with a real domain. Because the testing environment was mistakenly connected to the internet, the model exploited a real website, mistaking it to be part of the simulated environment. This did not involve a sophisticated sandbox escape or a zero-day: the internet access resulted from a misconfiguration, and the model appeared to exploit a basic security vulnerability.<\/p>\n<p>According to OpenAI, Irregular has suspended evaluations, shifted over to \u201cremediation,\u201d notified everyone involved, and started building new safeguards. That\u2019s in addition to the white paper the Information says it\u2019s working on.<\/p>\n<p>Stories about the escalating cyber capabilities of AI models have become a major feature of the AI conversation over the past few months. Back in April, two days after Anthropic <a href=\"https:\/\/gizmodo.com\/openai-hey-we-also-have-a-new-tool-that-is-so-scarily-powerful-we-cant-release-it-2000744569\" rel=\"nofollow noopener\" target=\"_blank\">announced the limited release of an AI model<\/a> with such ostensibly dangerous hacking capabilities that it couldn\u2019t be released publicly, its chief competitor OpenAI announced that it also had a model that <a href=\"https:\/\/gizmodo.com\/openai-hey-we-also-have-a-new-tool-that-is-so-scarily-powerful-we-cant-release-it-2000744569\" rel=\"nofollow noopener\" target=\"_blank\">could only be released to a select group<\/a>.<\/p>\n<p>Then came the reports of AI models making good on these promises by reportedly performing cyberattacks autonomously during testing. Most notably, there was <a href=\"https:\/\/openai.com\/index\/hugging-face-model-evaluation-security-incident\/\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI\u2019s Hugging Face breach<\/a>, which occurred after instances of OpenAI\u2019s models broke out of their sandbox. But that was followed by several other reports of varying severity. In one report from yesterday, Anthropic\u2019s Mythos 5 is said to have <a href=\"https:\/\/gizmodo.com\/i-usually-laugh-off-these-ai-hacking-reports-but-this-one-sounds-serious-and-scary-2000794666\" rel=\"nofollow noopener\" target=\"_blank\">attempted a complicated social engineering attack<\/a> against an unsuspecting developer.<\/p>\n<p>Gizmodo requested a statement from Irregular about these matters but did not immediately hear back. We will update this article if we receive comment.<\/p>\n","protected":false},"excerpt":{"rendered":"Looks like there\u2019s a new cyber-attacking AI model in town, and it\u2019s reportedly Meta\u2019s own Muse Spark 1.1.&hellip;\n","protected":false},"author":2,"featured_media":131100,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,65520,25,1122,1124],"class_list":["post-131099","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-ai-hacks","tag-artificial-intelligence","tag-meta","tag-muse-spark"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/131099","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=131099"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/131099\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/131100"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=131099"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=131099"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=131099"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}