Artificial Intelligence & Machine Learning
,
Next-Generation Technologies & Secure Development

Safety Researchers Say Voluntary Development Pause Falls Short of Accountability

Emilia David
August 19, 2026    

OpenAI Pauses Frontier Model Training for Safety Review
Image: Shutterstock/ISMG

OpenAI announced Tuesday it will enter a two-week pause in reinforcement learning training for its frontier models, as it reassesses its safety testing environment and the risks it presents.

See Also: How Skilled Attackers Weaponize AI Faster

OpenAI said the recent incident involving its agents hacking into model repository Hugging Face, along with preliminary evidence that its upcoming Astra model has advanced cybersecurity capabilities, necessitated a pause (see: OpenAI Seeks Agent Trust After Hugging Face Breach).

OpenAI’s move is a rare acknowledgment that its internal safeguards haven’t kept up with model capabilities. But the company did not offer any external validation that its short break will result in stronger security and risk approaches to model development.

“As models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment and security must stay ahead of those risks,” the company said.

OpenAI said this means improving detection and responding to “concerning” agent behavior, reducing the likelihood of harmful and unauthorized actions, and limiting what AI systems can access. An OpenAI spokesperson told ISMG the pause has already started and is ongoing. OpenAI has not said why it chose a two-week period. After the Hugging Face incident, OpenAI brought in third-party observers like METR and Redwood Research to assess model misbehavior related to the breach. For this pause, and the lessons and new approaches the company plans to implement, OpenAI did not announce any external evaluators.

The firm’s decision to pause training its larger, more capable models has roots in ongoing discussions about the speed of AI development.

The company said the pause already allowed it to add stronger workload and network isolation and reconfigured security testing to remove vulnerable shared services. It’s also expanded its chain-of-thought monitoring, which will now alert administrators within 30 minutes after concerning activity. But there is no independent assurance that these processes will be followed or that they even work.

While enterprise customers and AI safety observers applauded OpenAI’s seeming self-awareness of its safeguards’ limitations, many said that simply pausing is not enough.

Max Tegmark, chair of the Future of Life Institute, which published an open letter urging companies to pause model development in 2023, said in an emailed statement that OpenAI’s decision “is a step in the right direction.” But, “a voluntary pause that the U.S. government can neither verify nor enforce isn’t enough,” Tegmark said. “We need legally binding safety standards just as for food and cars.”

Nathan Lambert, an artificial intelligence researcher and former LLM developer at the Allen Institute for AI, posted on X that “we should have independent organizations that can access the full details of these training runs for monitoring.”

Many AI safety researchers have called for a slower development pace for frontier AI models. As recently as July, Anthropic CEO Dario Amodei and other executives and staffers from his company, OpenAI, Thinking Machines, Meta and Google signed an open letter urging the Trump administration to support efforts to pace AI development. Anthropic also previously floated a similar idea, so the international community buys itself time to figure out security fixes.

John Strand, founder of the consultancy Black Hills Information Security, said OpenAI must contend with a larger question of trust.

“I’m glad they’re putting additional safeguards in place, but there’s a bigger question here. Can we trust the same companies that got this wrong to effectively self-regulate systems backed by immense amounts of computing power?” he said.