{"id":678608,"date":"2026-09-08T13:49:21","date_gmt":"2026-09-08T13:49:21","guid":{"rendered":"https:\/\/www.europesays.com\/ie\/678608\/"},"modified":"2026-09-08T13:49:21","modified_gmt":"2026-09-08T13:49:21","slug":"openai-models-went-rogue-we-urgently-need-a-better-hugging-face-investigation-mackenzie-arnold-and-stephan-llerena","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ie\/678608\/","title":{"rendered":"OpenAI models went rogue. We urgently need a better Hugging Face investigation | Mackenzie Arnold and Stephan Llerena"},"content":{"rendered":"<p class=\"dcr-1s160rg\">When OpenAI first revealed that its AI agents had autonomously hacked a major real-world company, Hugging Face, many assumed only one or two <a href=\"https:\/\/arstechnica.com\/ai\/2026\/07\/how-an-openai-benchmark-test-turned-into-a-real-world-cyberattack\/\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">agents<\/a> were involved. The truth, a <a href=\"https:\/\/metr.org\/blog\/2026-08-26-openai-hugging-face-incident-investigation\/#core-takeaways-about-this-incident\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">new report<\/a> reveals, is far stranger: the incident involved about 1,200 AI agents, 700 of which directly participated in the attack.<\/p>\n<p class=\"dcr-1s160rg\">OpenAI invited researchers from METR, along with an expert from Redwood Research, to produce the new report, alongside the company\u2019s own <a href=\"https:\/\/openai.com\/index\/hugging-face-incident-and-the-road-ahead\/\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">investigation<\/a>. The findings <a href=\"https:\/\/www.planned-obsolescence.org\/p\/the-hugging-face-attack-surprised\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">shocked<\/a> the experts.<\/p>\n<p class=\"dcr-1s160rg\">The agents were highly coordinated, constructing complex message boards in the nooks and crannies of their shared artefact repository. All told, agents exchanged more than 70,000 messages in less than a week. Even more concerning, the agents took steps to hide their behavior \u2013 spoofing tool calls and attempting to tamper with their own logs.<\/p>\n<p class=\"dcr-1s160rg\">Initial reporting assumed that the agents had attacked Hugging Face in search of an answer key to a difficult test. In reality, investigators found that the agents had figured out a way to derive the answers within the first few hours. The days of work that followed focused on another goal: concerned that the automated scoring system might identify the agents\u2019 cheating, they aimed to learn more about the scorer. The goal wasn\u2019t just to cheat, but to hide it.<\/p>\n<p class=\"dcr-1s160rg\">These details are startling, and they\u2019re a warning sign about the current trajectory of AI development. But just as concerning is how limited the investigation was and how much we still don\u2019t know.<\/p>\n<p class=\"dcr-1s160rg\">We work at a thinktank that <a href=\"https:\/\/law-ai.org\/\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">analyzes AI-related legislation<\/a>, and we\u2019ve been worried for a while that legal requirements for reporting incidents like these are insufficient.<\/p>\n<p class=\"dcr-1s160rg\">These reports make that clear: they\u2019re informative enough to tell us something went wrong, but not detailed enough to tell us why.<\/p>\n<p class=\"dcr-1s160rg\">This shouldn\u2019t surprise us. These investigations are entirely voluntary, and their findings are limited to what the company is willing to disclose. METR\u2019s independent investigation was highly constrained by an agreement with <a href=\"https:\/\/www.theguardian.com\/technology\/openai\" data-link-name=\"in body link\" data-component=\"auto-linked-tag\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI<\/a>. For example, the investigators were not given access to the underlying model that created the majority of the misbehaving agents. And despite indications that message boards had formed as early as May and that coordinated agent activity persisted after 13 July, METR was only permitted to investigate the period from 26 June to 13 July.<\/p>\n<p class=\"dcr-1s160rg\">And that was the good part. METR was given close to nothing about OpenAI\u2019s safety and security practices. If OpenAI ignored warning signs \u2013 as <a href=\"https:\/\/www.axios.com\/2026\/08\/26\/openai-hugging-face-technical-report-ai-hack\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">several data points suggest<\/a> \u2013 violated its own safety procedures, or failed to implement fixes that would prevent future events, these all fell outside the scope of METR\u2019s investigation. The expert investigators themselves were left with many <a href=\"https:\/\/x.com\/RyanGreenblatt\/status\/2092741434764095828\" data-link-name=\"in body link\" rel=\"nofollow\">questions<\/a> and <a href=\"https:\/\/www.planned-obsolescence.org\/p\/the-hugging-face-attack-surprised\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">continued concerns<\/a>.<\/p>\n<p class=\"dcr-1s160rg\">Concerns about the limits aren\u2019t hypothetical. On Friday, Reuters <a href=\"https:\/\/www.msn.com\/en-us\/news\/other\/exclusive-openai-agents-hijacked-german-website-in-previously-undisclosed-ai-breakout-this-spring\/ar-AA2byslO\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">reported<\/a> that another swarm of OpenAI agents had broken out this spring, hijacking a German website and using it as another message board. According to the report, OpenAI knew about this incident but said nothing, and it was entirely absent from METR\u2019s report.<\/p>\n<p class=\"dcr-1s160rg\">Clearly, this kind of investigation \u2013 however well-executed \u2013 leaves a lot to be desired.<\/p>\n<p class=\"dcr-1s160rg\">When planes crash, trains derail or chemical plants explode, expert government investigators arrive with legal authority to compel evidence, preserve records, and tell the public what happened. Hugging Face <a href=\"https:\/\/huggingface.co\/blog\/security-incident-july-2026\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">reported this incident<\/a> to law enforcement, and multiple attorneys general have expressed interest in looking into it.<\/p>\n<p class=\"dcr-1s160rg\">But no government agency has both the mandate and expertise to investigate the technical facts of the incident and OpenAI\u2019s conduct. As far as we know, the only people to examine this incident did so at OpenAI\u2019s discretion and with its consent.<\/p>\n<p class=\"dcr-1s160rg\">And while the Hugging Face breach is not as severe as a plane crash, it was a serious and costly attack. It\u2019s unlikely to be the last, or the most severe.<\/p>\n<p class=\"dcr-1s160rg\">Existing laws don\u2019t fill this gap. While <a href=\"https:\/\/news.bloomberglaw.com\/privacy-and-data-security\/alabama-investigates-openai-following-rogue-ai-hacking-incident\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">some attorneys general<\/a> have tried to use existing investigatory powers, these offices aren\u2019t built for technical fact-finding, and the laws they rely on are limited to questions of consumer deception, not public risk. Despite significant AI laws passed in California, New York and Illinois, none create the investigative authority demanded by these incidents. As we\u2019ve written elsewhere, existing <a href=\"https:\/\/www.lawfaremedia.org\/article\/when-reporting-an-ai-security-incident-is-not-mandatory\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">incident reporting laws<\/a> probably don\u2019t cover the Hugging Face attack, and even if they did, OpenAI might share little more than the date and a short summary.<\/p>\n<p class=\"dcr-1s160rg\">What we need is a federal body equipped to conduct expert investigations of serious AI incidents \u2013 with the authority to compel documents and testimony, resources and personnel to examine the systems involved, and the ability to partner with third-party experts like METR.<\/p>\n<p class=\"dcr-1s160rg\">Reports should be published, subject to appropriate redactions, to <a href=\"https:\/\/www.dacbeachcroft.com\/en\/What-we-think\/Silence-as-Risk-ICAO-Annex-13-and-the-consequences-of-non-publication\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">ensure the public learns from <\/a>the incident. And reporting should also extend to near-miss events, which may reveal warnings before major harms occur.<\/p>\n<p class=\"dcr-1s160rg\">This structure would not be unusual. In aviation, the National Transportation Safety Board (NTSB) investigates accidents with <a href=\"https:\/\/www.law.cornell.edu\/uscode\/text\/49\/1113\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">subpoena<\/a> and <a href=\"https:\/\/www.law.cornell.edu\/uscode\/text\/49\/1131\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">wide-ranging investigative<\/a> <a href=\"https:\/\/www.ecfr.gov\/current\/title-49\/subtitle-B\/chapter-VIII\/part-830\/subpart-B\/section-830.5\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">powers<\/a>. Aircraft operators must <a href=\"https:\/\/www.ecfr.gov\/current\/title-49\/subtitle-B\/chapter-VIII\/part-830\/subpart-C\/section-830.10\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">preserve wreckage and records<\/a>; and the NTSB may confer with employees and <a href=\"https:\/\/www.law.cornell.edu\/uscode\/text\/49\/1113\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">contract<\/a> external experts.<\/p>\n<p class=\"dcr-1s160rg\">Congress can build a similar system for AI incidents, while protecting AI developers\u2019 legitimate interests. Investigations should open only upon clear triggers and be bound to the incident. Certain confidential information can receive <a href=\"https:\/\/www.law.cornell.edu\/uscode\/text\/49\/1114\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">statutory protection<\/a>. And investigations can be limited to one agency with <a href=\"https:\/\/www.law.cornell.edu\/uscode\/text\/49\/1131\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">priority<\/a> to avoid duplicative investigations.<\/p>\n<p class=\"dcr-1s160rg\">The need and urgency is real. Subsequent <a href=\"https:\/\/www.bbc.co.uk\/news\/articles\/cx2kgdnyk2po\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">reports<\/a> have revealed that the Hugging Face breach was not an isolated incident and that autonomous agents from Meta, Anthropic and OpenAI have hacked third parties in separate incidents. We\u2019ve been lucky so far \u2013 the harms have been limited. But luck is no substitute for the law.<\/p>\n<ul class=\"dcr-1s160rg\">\n<li class=\"dcr-1s160rg\">\n<p class=\"dcr-1s160rg\">Mackenzie Arnold is the managing director of US law &amp; policy at the Institute for Law &amp; AI, where he provides analysis and advice to ensure that advances in AI benefit the public at large<\/p>\n<\/li>\n<li class=\"dcr-1s160rg\">\n<p class=\"dcr-1s160rg\">Stephan Llerena is a research scholar at the Institute for Law &amp; AI. His research focuses on US law and policy regarding frontier AI governance<\/p>\n<\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"When OpenAI first revealed that its AI agents had autonomously hacked a major real-world company, Hugging Face, many&hellip;\n","protected":false},"author":2,"featured_media":678609,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","_share_on_mastodon":"0"},"categories":[261],"tags":[291,289,290,18,19,17,82],"class_list":["post-678608","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-eire","tag-ie","tag-ireland","tag-technology"],"share_on_mastodon":{"url":"https:\/\/pubeurope.com\/@ie\/117235742285471195","error":""},"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/posts\/678608","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/comments?post=678608"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/posts\/678608\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/media\/678609"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/media?parent=678608"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/categories?post=678608"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ie\/wp-json\/wp\/v2\/tags?post=678608"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}