{"id":137941,"date":"2026-08-12T22:44:13","date_gmt":"2026-08-12T22:44:13","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/137941\/"},"modified":"2026-08-12T22:44:13","modified_gmt":"2026-08-12T22:44:13","slug":"ai-models-keep-going-rogue-this-company-is-the-one-testing-them","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/137941\/","title":{"rendered":"AI Models Keep Going Rogue. This Company Is The One Testing Them"},"content":{"rendered":"<p>Israeli AI startup Irregular, which was linked to the incidents, runs thousands of simulations to evaluate AI\u2019s cyber capabilities.<\/p>\n<p>AFP via Getty Images<\/p>\n<p>In recent weeks, OpenAI, Anthropic and Meta all disclosed that their most advanced AI models went rogue, accessing the internet and hacking into other companies\u2019 systems during routine security testing.   <\/p>\n<p>All of these companies were testing their models using software from Israeli AI startup <a href=\"https:\/\/www.forbes.com\/sites\/thomasbrewster\/2025\/09\/16\/openai-pays-a-450-million-startup-to-test-chatgpt-capacity-for-evil\/\" data-ga-track=\"InternalLink:https:\/\/www.forbes.com\/sites\/thomasbrewster\/2025\/09\/16\/openai-pays-a-450-million-startup-to-test-chatgpt-capacity-for-evil\/\" target=\"_self\" aria-label=\"Irregular\" rel=\"nofollow noopener\">Irregular<\/a>, which runs thousands of simulations to evaluate AI\u2019s cyber capabilities. Irregular put the models in different scenarios and tested to see whether they could break the defenses in place, evade detection, steal credentials or hack other systems. In some simulations, OpenAI\u2019s models broke out of their contained testing environments and hacked AI startup Hugging Face\u2019s servers. The OpenAI incident prompted Irregular to audit its own systems \u2014 and that\u2019s how the company found Anthropic and Meta\u2019s models behaving in similar ways, according to a person familiar with the situation. <\/p>\n<p>Rogue AI agents hacking other companies sounds terrifying, though it\u2019s not really what happened. During the testing phase, models are pushed to their limits and are instructed to solve a task or fulfill an objective. Of course they sometimes identify hacking as the best, most efficient way to accomplish it. <\/p>\n<p>But while some reports might have blown it slightly out of proportion, the incidents highlight a growing concern: safety testing companies can\u2019t keep up with the rapid pace of AI development, making it harder to evaluate AI models before they\u2019re released to the public. <\/p>\n<p>\u201cWe need to accelerate defense aggressively,\u201d the person said. \u201cIf we don\u2019t, we\u2019re going to go in blind without being able to measure and make responsible decisions on how and when to just release some models to the public.\u201d <\/p>\n<p>Now, U.S. House Democrats are calling on OpenAI and Anthropic leaders to explain how their AI models escaped their testing environments, <a href=\"https:\/\/www.reuters.com\/legal\/litigation\/us-house-democrats-press-anthropic-openai-about-rogue-ai-agents-2026-08-10\/\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/www.reuters.com\/legal\/litigation\/us-house-democrats-press-anthropic-openai-about-rogue-ai-agents-2026-08-10\/\" aria-label=\"Reuters\">Reuters<\/a> reported.<\/p>\n<p>\u201cPeople have very much been expecting this to happen one day,\u201d says Matt Fredrikson, CEO of Gray Swan, another AI startup that conducts safety testing for AI models. \u201cI think that it\u2019s absolutely concerning.\u201d <\/p>\n<p>At cybersecurity conference Black Hat, OpenAI researchers explained how the Hugging Face hack took place. Over a span of days, a group of OpenAI agents created a secret message board where they communicated with each other about different ways to exploit software vulnerabilities, sharing notes and splitting up the work. They also at times stepped on each other\u2019s toes, even deleting each other\u2019s work, all without OpenAI\u2019s knowledge, <a href=\"https:\/\/www.wired.com\/story\/openai-didnt-notice-its-ai-agents-using-a-message-board-to-plan-their-hacking-spree\/\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/www.wired.com\/story\/openai-didnt-notice-its-ai-agents-using-a-message-board-to-plan-their-hacking-spree\/\" aria-label=\"Wired\">Wired<\/a> reported. <\/p>\n<p>Now let\u2019s get into the headlines. <\/p>\n<p>BIG PLAYS<\/p>\n<p>In a 6,500-word long essay titled \u201cThe Future is for Everyone,\u201d <a href=\"https:\/\/www.forbes.com\/sites\/tylerroush\/2026\/08\/10\/mark-zuckerberg-outlines-ai-vision-in-new-manifesto-heres-what-he-says\/\" data-ga-track=\"InternalLink:https:\/\/www.forbes.com\/sites\/tylerroush\/2026\/08\/10\/mark-zuckerberg-outlines-ai-vision-in-new-manifesto-heres-what-he-says\/\" target=\"_self\" aria-label=\"Mark Zuckerberg\" rel=\"nofollow noopener\">Mark Zuckerberg<\/a> argued that AI should be built for everyone and not be controlled by a few big players, writing that more people having access to AI is key to navigating the technology safely. In practice, the social media billionaire believes that bringing about a \u201cpositive AI future\u201d  requires easier access to training data, government support to build more data centers, and free reign for all AI developers to distill a more powerful model\u2019s capabilities. All of these proposals serve Meta\u2019s own interests as the company plans to continue releasing open source models and wants to give everyone a free personal AI assistant. <\/p>\n<p>ETHICS + LAW <\/p>\n<p>For the first time ever, scientists used AI to design 16 entirely new kinds of <a href=\"https:\/\/www.wired.com\/story\/scientists-used-ai-to-create-16-new-viruses\/#:~:text=The%20experiment%20involved%20exposing%20a,bacterial%20resistance%20and%20establish%20infection.\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/www.wired.com\/story\/scientists-used-ai-to-create-16-new-viruses\/#:~:text=The%20experiment%20involved%20exposing%20a,bacterial%20resistance%20and%20establish%20infection.\" aria-label=\"viruses\">viruses<\/a> using genetic data from millions of microbes, animals and plants found in nature. Lucky for us, the viruses don\u2019t pose a threat to humans, only to bacteria. The breakthrough could help fight bacteria that have developed resistance to antibiotic medicines and allow scientists to develop personalized treatments that adapt at the same rate as the pathogens that cause disease. But it also raises concerns about AI being misused to invent new dangerous diseases. <\/p>\n<p>TALENT RESHUFFLE <\/p>\n<p>Longtime OpenAI executive Brad Lightcap is leaving the company after eight years to \u201cstart something new,\u201d he announced on <a href=\"https:\/\/x.com\/bradlightcap\/status\/2087211567012032862?s=20\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/x.com\/bradlightcap\/status\/2087211567012032862?s=20\" aria-label=\"X.\">X.<\/a> His departure is the latest in a string of high-profile exits in recent months, including former CEO of applications Fidji Simo and vice president of OpenAI for science Kevin Weil. <\/p>\n<p>AI DEAL OF THE WEEK <\/p>\n<p>River AI, cofounded by xAI cofounder Igor Babuschkin, has raised $1.1 billion to train open source models that can be easily personalized and customized by anyone to their liking, so that their AI isn\u2019t controlled by large tech corporations. The company also plans to build a new type of computer server that will allow businesses and individuals to run AI systems on their own hardware. <a href=\"https:\/\/www.forbes.com\/sites\/rashishrivastava\/2026\/05\/14\/xai-cofounder-igor-babuschkin-in-talks-to-raise-up-to-1-billion-for-a-new-ai-startup\/\" data-ga-track=\"InternalLink:https:\/\/www.forbes.com\/sites\/rashishrivastava\/2026\/05\/14\/xai-cofounder-igor-babuschkin-in-talks-to-raise-up-to-1-billion-for-a-new-ai-startup\/\" target=\"_self\" aria-label=\"Forbes\" rel=\"nofollow noopener\">Forbes<\/a> reported details of the fund raise in May. <\/p>\n<p>DEEP DIVE<\/p>\n<p>On Wednesday, Google CEO Sundar Pichai made a bombshell announcement: Jeff Dean, the tech giant\u2019s 30th employee and its chief scientist, is departing after 27 years to found his own AI startup. Meanwhile at Google DeepMind, Nobel laureate Demis Hassabis, the bigwig who founded the pioneering lab DeepMind more than 15 years ago, is stepping away from day-to-day control of the frontier lab to become its chairman and new chief scientist for Google parent Alphabet.<\/p>\n<p>The changes are an abrupt reset at one of the world\u2019s most important AI labs. But the fault lines were visible well before Wednesday, multiple former employees tell Forbes. They trace back to 2023, when Google fused DeepMind and Google Brain, its two premier AI research operations, into Google DeepMind.<\/p>\n<p>The pitch was simple enough: one company, one AI army, fewer internal spats as Google scrambled to catch up to OpenAI and Anthropic. The reality was messier. Hassabis ran the combined division from London. Dean remained in Silicon Valley. Both reported directly to Pichai. There might have been a single org chart. But there were still two capitals.<\/p>\n<p>Now, the center of gravity has moved toward Mountain View, California, Google\u2019s mothership. Koray Kavukcuoglu, the division\u2019s former chief technology officer, moved from London to Google headquarters in the past year. Sebastian Borgeaud, who leads an important coding effort, also relocated from the UK to California, Bloomberg <a href=\"https:\/\/www.bloomberg.com\/news\/articles\/2026-08-06\/google-shifts-ai-power-to-california-in-race-against-anthropic-openai\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/www.bloomberg.com\/news\/articles\/2026-08-06\/google-shifts-ai-power-to-california-in-race-against-anthropic-openai\" aria-label=\"reported\">reported<\/a>.<\/p>\n<p>\u201cFor me, this is the ultimate fallout of the DeepMind\/Brain merger,\u201d said one former Google DeepMind employee, who worked at the company for more than a decade, adding that Dean\u2019s influence inside Google had \u201cgradually been waning\u201d since the merger, as power shifted to Hassabis and then later to Kavukcuoglu and Borgeaud.<\/p>\n<p>Read the full story on <a href=\"https:\/\/www.forbes.com\/sites\/richardnieva\/2026\/08\/06\/google-deepmind-london-mountain-view\/\" data-ga-track=\"InternalLink:https:\/\/www.forbes.com\/sites\/richardnieva\/2026\/08\/06\/google-deepmind-london-mountain-view\/\" target=\"_self\" aria-label=\"Forbes\" rel=\"nofollow noopener\">Forbes<\/a>. <\/p>\n<p>MODEL BEHAVIOR<\/p>\n<p>AI agents are behaving in unexpected ways. The latest example: Andrew Bird, an Australian tech executive, asked his AI assistant to book a popular morning gym class for him. But his OpenClaw agent went a step further\u2014 it hacked the gym\u2019s website and kicked another person off the waitlist to give him a spot, <a href=\"https:\/\/www.abc.net.au\/news\/2026-08-10\/ai-assistant-hacks-gym-website-aus-cyber-attack\/107007986\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/www.abc.net.au\/news\/2026-08-10\/ai-assistant-hacks-gym-website-aus-cyber-attack\/107007986\" aria-label=\"ABC News\">ABC News<\/a> reported. The agent also found a way to book classes months in advance thanks to a vulnerability in the gym\u2019s appointment software. <\/p>\n","protected":false},"excerpt":{"rendered":"Israeli AI startup Irregular, which was linked to the incidents, runs thousands of simulations to evaluate AI\u2019s cyber&hellip;\n","protected":false},"author":2,"featured_media":137942,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[9],"tags":[4077,24,5044,132,7543,62995,1122,1160,157,59602,225,314,613],"class_list":["post-137941","post","type-post","status-publish","format-standard","has-post-thumbnail","category-google","tag-agents","tag-ai","tag-deepmind","tag-google","tag-google-deepmind","tag-irregular","tag-meta","tag-models","tag-openai","tag-rogue","tag-safety","tag-security","tag-testing"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/137941","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=137941"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/137941\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/137942"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=137941"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=137941"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=137941"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}