Israeli AI startup Irregular, which was linked to the incidents, runs thousands of simulations to evaluate AI’s cyber capabilities.

AFP via Getty Images

In recent weeks, OpenAI, Anthropic and Meta all disclosed that their most advanced AI models went rogue, accessing the internet and hacking into other companies’ systems during routine security testing.

All of these companies were testing their models using software from Israeli AI startup Irregular, which runs thousands of simulations to evaluate AI’s cyber capabilities. Irregular put the models in different scenarios and tested to see whether they could break the defenses in place, evade detection, steal credentials or hack other systems. In some simulations, OpenAI’s models broke out of their contained testing environments and hacked AI startup Hugging Face’s servers. The OpenAI incident prompted Irregular to audit its own systems — and that’s how the company found Anthropic and Meta’s models behaving in similar ways, according to a person familiar with the situation.

Rogue AI agents hacking other companies sounds terrifying, though it’s not really what happened. During the testing phase, models are pushed to their limits and are instructed to solve a task or fulfill an objective. Of course they sometimes identify hacking as the best, most efficient way to accomplish it.

But while some reports might have blown it slightly out of proportion, the incidents highlight a growing concern: safety testing companies can’t keep up with the rapid pace of AI development, making it harder to evaluate AI models before they’re released to the public.

“We need to accelerate defense aggressively,” the person said. “If we don’t, we’re going to go in blind without being able to measure and make responsible decisions on how and when to just release some models to the public.”

Now, U.S. House Democrats are calling on OpenAI and Anthropic leaders to explain how their AI models escaped their testing environments, Reuters reported.

“People have very much been expecting this to happen one day,” says Matt Fredrikson, CEO of Gray Swan, another AI startup that conducts safety testing for AI models. “I think that it’s absolutely concerning.”

At cybersecurity conference Black Hat, OpenAI researchers explained how the Hugging Face hack took place. Over a span of days, a group of OpenAI agents created a secret message board where they communicated with each other about different ways to exploit software vulnerabilities, sharing notes and splitting up the work. They also at times stepped on each other’s toes, even deleting each other’s work, all without OpenAI’s knowledge, Wired reported.

Now let’s get into the headlines.

BIG PLAYS

In a 6,500-word long essay titled “The Future is for Everyone,” Mark Zuckerberg argued that AI should be built for everyone and not be controlled by a few big players, writing that more people having access to AI is key to navigating the technology safely. In practice, the social media billionaire believes that bringing about a “positive AI future” requires easier access to training data, government support to build more data centers, and free reign for all AI developers to distill a more powerful model’s capabilities. All of these proposals serve Meta’s own interests as the company plans to continue releasing open source models and wants to give everyone a free personal AI assistant.

ETHICS + LAW

For the first time ever, scientists used AI to design 16 entirely new kinds of viruses using genetic data from millions of microbes, animals and plants found in nature. Lucky for us, the viruses don’t pose a threat to humans, only to bacteria. The breakthrough could help fight bacteria that have developed resistance to antibiotic medicines and allow scientists to develop personalized treatments that adapt at the same rate as the pathogens that cause disease. But it also raises concerns about AI being misused to invent new dangerous diseases.

TALENT RESHUFFLE

Longtime OpenAI executive Brad Lightcap is leaving the company after eight years to “start something new,” he announced on X. His departure is the latest in a string of high-profile exits in recent months, including former CEO of applications Fidji Simo and vice president of OpenAI for science Kevin Weil.

AI DEAL OF THE WEEK

River AI, cofounded by xAI cofounder Igor Babuschkin, has raised $1.1 billion to train open source models that can be easily personalized and customized by anyone to their liking, so that their AI isn’t controlled by large tech corporations. The company also plans to build a new type of computer server that will allow businesses and individuals to run AI systems on their own hardware. Forbes reported details of the fund raise in May.

DEEP DIVE

On Wednesday, Google CEO Sundar Pichai made a bombshell announcement: Jeff Dean, the tech giant’s 30th employee and its chief scientist, is departing after 27 years to found his own AI startup. Meanwhile at Google DeepMind, Nobel laureate Demis Hassabis, the bigwig who founded the pioneering lab DeepMind more than 15 years ago, is stepping away from day-to-day control of the frontier lab to become its chairman and new chief scientist for Google parent Alphabet.

The changes are an abrupt reset at one of the world’s most important AI labs. But the fault lines were visible well before Wednesday, multiple former employees tell Forbes. They trace back to 2023, when Google fused DeepMind and Google Brain, its two premier AI research operations, into Google DeepMind.

The pitch was simple enough: one company, one AI army, fewer internal spats as Google scrambled to catch up to OpenAI and Anthropic. The reality was messier. Hassabis ran the combined division from London. Dean remained in Silicon Valley. Both reported directly to Pichai. There might have been a single org chart. But there were still two capitals.

Now, the center of gravity has moved toward Mountain View, California, Google’s mothership. Koray Kavukcuoglu, the division’s former chief technology officer, moved from London to Google headquarters in the past year. Sebastian Borgeaud, who leads an important coding effort, also relocated from the UK to California, Bloomberg reported.

“For me, this is the ultimate fallout of the DeepMind/Brain merger,” said one former Google DeepMind employee, who worked at the company for more than a decade, adding that Dean’s influence inside Google had “gradually been waning” since the merger, as power shifted to Hassabis and then later to Kavukcuoglu and Borgeaud.

Read the full story on Forbes.

MODEL BEHAVIOR

AI agents are behaving in unexpected ways. The latest example: Andrew Bird, an Australian tech executive, asked his AI assistant to book a popular morning gym class for him. But his OpenClaw agent went a step further— it hacked the gym’s website and kicked another person off the waitlist to give him a spot, ABC News reported. The agent also found a way to book classes months in advance thanks to a vulnerability in the gym’s appointment software.