{"id":122643,"date":"2026-07-29T11:06:18","date_gmt":"2026-07-29T11:06:18","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/122643\/"},"modified":"2026-07-29T11:06:18","modified_gmt":"2026-07-29T11:06:18","slug":"were-running-out-of-reasons-to-ignore-ai-safety","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/122643\/","title":{"rendered":"We\u2019re running out of reasons to ignore AI safety"},"content":{"rendered":"<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b2 _18mzr4b0 _18mzr4b7 _18mzr4b5 _19wv7tc1 _18mzr4bb\">Earlier this month, OpenAI gave several of its AI models a task: complete a <a href=\"https:\/\/arxiv.org\/abs\/2605.11086\" rel=\"nofollow noopener\" target=\"_blank\">test<\/a> designed to measure their cybersecurity capabilities. It put the systems in a sandboxed environment without an internet connection and set them off to work.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">What <a href=\"https:\/\/www.theverge.com\/ai-artificial-intelligence\/968988\/openai-hugging-face-hack-ai\" rel=\"nofollow noopener\" target=\"_blank\">happened next<\/a> is almost laughably silly \u2014 but also, as Adam Gleave, cofounder and CEO of AI safety organization FAR.AI put it, \u201ca visceral example of how misaligned AI could cause harm.\u201d According to OpenAI, the models escaped the sandbox meant to contain them, moved through the company\u2019s internal systems, found a route to the internet, and then started looking for a way into Hugging Face. And why was the agent looking for a way into Hugging Face? They had apparently reasoned that the developer platform might store the answers to the cyber benchmark and that getting them would be a great way to get a high score.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup c39lj12 _19wv7tc9\">The incident is \u201ca visceral example of how misaligned AI could cause harm.\u201d<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">In other words, OpenAI\u2019s agent broke out of a supposedly secure environment, traipsed through the company\u2019s systems, got online, and compromised another company\u2019s systems \u2014 all to cheat on a test of no particular importance.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">This appears to be the first well-documented incident of its kind, or at least the first on this scale. It was both a clear example of a system pursuing a goal in an unintended way and a demonstration that frontier models are now powerful enough for that behavior to have real-world consequences.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">The hack was an example of what the AI safety community calls \u201c<a href=\"https:\/\/deepmind.google\/blog\/specification-gaming-the-flip-side-of-ai-ingenuity\/\" rel=\"nofollow noopener\" target=\"_blank\">specification gaming<\/a>,\u201d a behavior also known as reward hacking, said Fazl Barez, an AI safety researcher at the University of Oxford In plain English, it means \u201cthe model doing what you asked rather than what you meant,\u201d Fazl said. It satisfies the literal terms of a task while violating the obvious intent and has been <a href=\"https:\/\/www-cdn.anthropic.com\/9ff93dfa8f445c932415d335c88852ef47f1201e.pdf\" rel=\"nofollow noopener\" target=\"_blank\">documented<\/a> across <a href=\"https:\/\/openai.com\/index\/faulty-reward-functions\/\" rel=\"nofollow noopener\" target=\"_blank\">many<\/a> AI systems. Some researchers <a href=\"https:\/\/www.anthropic.com\/research\/emergent-misalignment-reward-hacking\" rel=\"nofollow noopener\" target=\"_blank\">worry<\/a> that as systems become more capable, this could produce increasingly misaligned systems, which pursue goals in ways their creators did not intend (like turning everyone into <a href=\"https:\/\/go.skimresources.com\/?id=1025X1701640&amp;xs=1&amp;url=https%3A%2F%2Fwww.newscientist.com%2Farticle%2F2372484-what-is-the-ai-alignment-problem-and-how-can-it-be-solved%2F\" rel=\"sponsored nofollow noopener\" target=\"_blank\">paperclips<\/a>).<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cNothing in that chain is exotic in isolation,\u201d Fazl said. A competent human tester would be able to do all of this, he added. \u201cWhat is new is that the model did not stop. Older models would likely have hit some barrier and gone back to the user, he said, but this agent just \u201ctreated the barrier as part of the problem it had been asked to solve.\u201d<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">OpenAI <a href=\"https:\/\/openai.com\/index\/hugging-face-model-evaluation-security-incident\/\" rel=\"nofollow noopener\" target=\"_blank\">described<\/a> it as \u201can unprecedented cyber incident,\u201d that \u201c<a href=\"https:\/\/x.com\/OpenAI\/status\/2080815626113954288?s=20\" rel=\"nofollow\">marks an important moment<\/a> for AI safety.\u201d Hugging Face cofounder Thomas Wolf <a href=\"https:\/\/www.bbc.co.uk\/news\/articles\/cdrvy3pn3r0o\" rel=\"nofollow noopener\" target=\"_blank\">said<\/a> it was a \u201cwake-up call\u201d for the industry. But this is not one of the four horsemen of the AI apocalypse. As cyber incidents go, experts told The Verge it was pretty mundane. Nothing the agent did required superhuman abilities. Moreover, frontier systems like GPT-5.6 Sol and Anthropic\u2019s <a href=\"https:\/\/www.theverge.com\/ai-artificial-intelligence\/917644\/anthropic-claude-mythos-breach-humiliation\" rel=\"nofollow noopener\" target=\"_blank\">Mythos<\/a> are known to be capable coders, are already thought to have been <a href=\"https:\/\/www.theverge.com\/ai-artificial-intelligence\/949644\/china-white-house-anthropic-mythos\" rel=\"nofollow noopener\" target=\"_blank\">misused<\/a> numerous times, and AI tools <a href=\"https:\/\/www.theguardian.com\/technology\/2026\/may\/11\/ai-powered-hacking-industrial-scale-threat-three-months-google\" rel=\"nofollow noopener\" target=\"_blank\">already allow hackers<\/a> to scale up and refine attacks on a massive scale.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Could it be hype? The industry has spent months <a href=\"https:\/\/www.theverge.com\/podcast\/951542\/anthropic-claude-fable-5-mythos-ban-pentagon-ai-regulation-trump\" rel=\"nofollow noopener\" target=\"_blank\">amplifying claims<\/a> about the dangerous capabilities of its top models, particularly when it comes to cybersecurity. It is the stated reason why companies like OpenAI and Anthropic have withheld their most capable models from the general public and partly why the <a href=\"https:\/\/www.theverge.com\/ai-artificial-intelligence\/951703\/anthropic-shutdown-export-controls\" rel=\"nofollow noopener\" target=\"_blank\">Trump administration hurriedly moved<\/a> to apply export controls to them.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1 _1upt4f27\">Are you an AI safety researcher or frontier lab employee? You can contact me securely and confidentially via Signal at robhart.01. My <a href=\"https:\/\/x.com\/TheRobertHart\" rel=\"nofollow\">X<\/a> DMs are also open.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">If this is hype, however, it has not gone entirely in OpenAI\u2019s favor. In the days since, the attack has <a href=\"https:\/\/www.theverge.com\/ai-artificial-intelligence\/971281\/nvidia-open-secure-ai-alliance-cybersecurity\" rel=\"nofollow noopener\" target=\"_blank\">produced a rare moment of unity<\/a> across much of the US tech industry about the importance of open-weight AI systems and the need to take AI security more seriously. These concerns were underscored further by the <a href=\"https:\/\/www.theverge.com\/ai-artificial-intelligence\/967781\/chinese-ai-models-open-source-moonshot-kimi-k3-alibaba-qwen\" rel=\"nofollow noopener\" target=\"_blank\">release<\/a> of Kimi K3, a highly <a href=\"https:\/\/www.theverge.com\/ai-artificial-intelligence\/971444\/how-chinese-open-weight-ai-models-impact-us-companies\" rel=\"nofollow noopener\" target=\"_blank\">capable open-weight model<\/a> from China. A broad coalition of companies including Nvidia, Microsoft, and SpaceX argued that the incident showed why defenders need access to the most capable tools available, rather than being forced to rely on proprietary providers whose built-in safeguards can limit their effectiveness in high-stakes security work. OpenAI, Anthropic, and Google were notably absent from the coalition\u2019s founding membership.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup c39lj12 _19wv7tc9\">\u201cAnyone who\u2019s been paying attention has noted that capabilities are only going in one direction.\u201d<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">OpenAI\u2019s account of the incident undeniably fits a broader industry narrative about the dangerous capabilities of frontier models. Even so, several details make the incident difficult to dismiss as merely self-serving. Foremost, it is an example of a problem the AI industry has warned about for years \u2014 and one OpenAI could have reasonably been expected to anticipate. The episode also handed an unexpected boost to a major Chinese competitor, whose model played a prominent role in containing the breach, while exposing OpenAI to significant legal, regulatory, and reputational scrutiny. That Hugging Face appears keen to work with OpenAI and, publicly at least, has remained fairly relaxed about the whole thing may have limited the fallout. Most of the experts The Verge spoke to similarly cautioned against reducing the incident to hype.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cIt\u2019s a pretty useful warning shot in terms of demonstrating both unintended consequences and just how capable these models are,\u201d said Se\u00e1n \u00d3 h\u00c9igeartaigh, a professor at Cambridge University\u2019s Leverhulme Centre for the Future of Intelligence. \u201cAnyone who\u2019s been paying attention has noted that capabilities are only going in one direction, and that is improving significantly over time in a way that I think is perhaps less obvious to the everyday user of something like ChatGPT.\u201d<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Still, it would be wrong to interpret this warning as a sign AI systems are about to slip human control, or that containing them is impossible, says Lin Li, an AI safety researcher at the University of Oxford. \u201cThe better lesson is that safety has to move from evaluating isolated actions to evaluating whole action sequences, environments, and operational controls,\u201d Li says.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">A crucial next step is for AI labs to be investing more heavily in securing their own systems. \u201cThere\u2019s a clear need for AI companies to beef up the security of their internal deployments,\u201d Gleave said, likening the current practice of responding to reward hacking incidents as they arise to a game of whack-a-mole that is becoming less and less tenable as stakes rise. Adam Chan, a research fellow at tech policy research center GovAI, said companies should consider airgapping their machines \u2014 physically isolating them from the internet and other networks \u2014 \u201cuntil they\u2019re sure about the model\u2019s capabilities.\u201d Intensifying work on alignment, which ensures systems reliably follow human intentions, and more rigorous testing \u201cto surface these issues before putting models in environments where they have the tools to be able to do these things,\u201d would also be good ideas, he said.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">As model capabilities increase, experts warn that we can\u2019t rely on technical safeguards alone. Peter Wallich, a former UK AI Security Institute official, said the incident illustrated that point: \u201cTwo multibillion dollar companies just tried this approach and \u2014 self-evidently, based on their own reporting \u2014 failed.\u201d<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">One of the biggest priorities should be ensuring outsiders can see what is happening inside frontier AI labs. \u201cWe only know about this incident because OpenAI chose to tell us,\u201d said Patrick Levermore, at the Centre for Long-Term Resilience, a British think tank. \u201cA good safety regime shouldn\u2019t depend on voluntary disclosure.\u201d The need is especially acute when, as Wallich noted, the conduct in question \u201cwould be a crime if done by a human.\u201d \u00d3 h\u00c9igeartaigh pointed to whistleblower protections, third-party audits, and mandatory reporting of serious incidents as possible ways to provide that visibility, stressing that oversight must span the entire development lifecycle rather than begin only once products reach the market. OpenAI said that one of the models being tested has not been released yet.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup c39lj12 _19wv7tc9\">\u201cA good safety regime shouldn\u2019t depend on voluntary disclosure.\u201d<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Lots of this presumes the companies themselves know what\u2019s happening inside their systems. In this case, <a href=\"https:\/\/www.reuters.com\/business\/its-ai-agent-spent-days-hacking-company-sources-say-openai-did-not-notice-week-2026-07-24\/\" rel=\"nofollow noopener\" target=\"_blank\">reports<\/a> suggest OpenAI was unaware its own agent was behind the days-long cyber campaign at Hugging Face and did not notice until well after the threat had been contained and the FBI contacted. There are still many details about the hack that are unknown or have not been made public. In an <a href=\"https:\/\/x.com\/OpenAI\/status\/2080815626113954288?s=20\" rel=\"nofollow\">update<\/a> on social media, OpenAI said it is conducting a review and will publish a technical report of its findings \u201cin the coming weeks.\u201d<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Whether the warnings raised by the Hugging Face incident produce any lasting change, or join the long list of warnings the tech industry absorbs without meaningfully altering course, remains uncertain. For now, at least, it does appear to have alarmed industry insiders and pushed US lawmakers to <a href=\"https:\/\/www.politico.com\/news\/2026\/07\/22\/openai-hugging-face-congress-response-01009190\" rel=\"nofollow noopener\" target=\"_blank\">consider new rules<\/a> before the next containment failure. The incident also added to a broader sense of unease over the speed of AI development, which deepened in the days that followed as employees from leading US labs <a href=\"https:\/\/www.theverge.com\/ai-artificial-intelligence\/972161\/ai-leaders-us-government-openai-anthropic-google-meta\" rel=\"nofollow noopener\" target=\"_blank\">signed a statement<\/a> backing coordinated global governance \u2014 including a potential slowdown in frontier AI development.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">The prevailing view of those The Verge spoke to was that this hack marked the start of a new class of risk, even if its significance may only become clear in hindsight. One former government AI policy expert, who asked not to be named because they were not authorized to be quoted by name, described it as a \u201cred line,\u201d the kind of watershed moment we may later look back on as marking a new, riskier stage in our relationship with AI. They hope it will force the tech industry to take the management of frontier systems more seriously and spur governments to think more deeply about oversight before a less benign breach occurs. Their fear is that it will instead join the long list of warnings about AI\u2019s growing capabilities that were recognized, discussed, and ultimately left unheeded.<\/p>\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _18mzr4ba _19wv7tc1\">That may prove overstated. But if this is a warning, we should consider ourselves lucky the AI agent was only trying to cheat on a test.<\/p>\n<p>Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.Robert HartClose<img alt=\"Robert Hart\" data-chromatic=\"ignore\" loading=\"lazy\" decoding=\"async\" data-nimg=\"fill\" class=\"_1bw37385 i7ks070\" style=\"position:absolute;height:100%;width:100%;left:0;top:0;right:0;bottom:0;color:transparent;background-size:cover;background-position:50% 50%;background-repeat:no-repeat;background-image:url(&quot;data:image\/svg+xml;charset=utf-8,%3Csvg xmlns='http:\/\/www.w3.org\/2000\/svg' %3E%3Cfilter id='b' color-interpolation-filters='sRGB'%3E%3CfeGaussianBlur stdDeviation='20'\/%3E%3CfeColorMatrix values='1 0 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 100 -1' result='s'\/%3E%3CfeFlood x='0' y='0' width='100%25' height='100%25'\/%3E%3CfeComposite operator='out' in='s'\/%3E%3CfeComposite in2='SourceGraphic'\/%3E%3CfeGaussianBlur stdDeviation='20'\/%3E%3C\/filter%3E%3Cimage width='100%25' height='100%25' x='0' y='0' preserveAspectRatio='none' style='filter: url(%23b);' href='data:image\/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mN8+R8AAtcB6oaHtZcAAAAASUVORK5CYII='\/%3E%3C\/svg%3E&quot;)\"   src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/07\/1785323178_355_ROB_H_BLURPLE.jpg\"\/><\/p>\n<p>Robert Hart<\/p>\n<p class=\"fv263x1\">Posts from this author will be added to your daily email digest and your homepage feed.<\/p>\n<p>FollowFollow<\/p>\n<p class=\"fv263x4\"><a class=\"fv263x5\" href=\"https:\/\/www.theverge.com\/authors\/robert-hart\" rel=\"nofollow noopener\" target=\"_blank\">See All by Robert Hart<\/a><\/p>\n<p>AIClose<\/p>\n<p>AI<\/p>\n<p class=\"fv263x1\">Posts from this topic will be added to your daily email digest and your homepage feed.<\/p>\n<p>FollowFollow<\/p>\n<p class=\"fv263x4\"><a class=\"fv263x5\" href=\"https:\/\/www.theverge.com\/ai-artificial-intelligence\" rel=\"nofollow noopener\" target=\"_blank\">See All AI<\/a><\/p>\n<p>OpenAIClose<\/p>\n<p>OpenAI<\/p>\n<p class=\"fv263x1\">Posts from this topic will be added to your daily email digest and your homepage feed.<\/p>\n<p>FollowFollow<\/p>\n<p class=\"fv263x4\"><a class=\"fv263x5\" href=\"https:\/\/www.theverge.com\/openai\" rel=\"nofollow noopener\" target=\"_blank\">See All OpenAI<\/a><\/p>\n<p>ReportClose<\/p>\n<p>Report<\/p>\n<p class=\"fv263x1\">Posts from this topic will be added to your daily email digest and your homepage feed.<\/p>\n<p>FollowFollow<\/p>\n<p class=\"fv263x4\"><a class=\"fv263x5\" href=\"https:\/\/www.theverge.com\/report\" rel=\"nofollow noopener\" target=\"_blank\">See All Report<\/a><\/p>\n<p>TechClose<\/p>\n<p>Tech<\/p>\n<p class=\"fv263x1\">Posts from this topic will be added to your daily email digest and your homepage feed.<\/p>\n<p>FollowFollow<\/p>\n<p class=\"fv263x4\"><a class=\"fv263x5\" href=\"https:\/\/www.theverge.com\/tech\" rel=\"nofollow noopener\" target=\"_blank\">See All Tech<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"Earlier this month, OpenAI gave several of its AI models a task: complete a test designed to measure&hellip;\n","protected":false},"author":2,"featured_media":122644,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,25,157,30,781],"class_list":["post-122643","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-artificial-intelligence","tag-openai","tag-report","tag-tech"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/122643","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=122643"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/122643\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/122644"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=122643"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=122643"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=122643"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}