AI Europe
  • Europe
  • Europa
  • Britain
  • France
  • Germany
  • Italy
  • Spain
  • Japan
  • Canada
  • Africa
  • People
  • AI
  • …
    • Afrique
    • Netherlands
    • Poland
  • Agentic AI
  • AGI
  • AI
  • Anthropic
  • Google
  • Microsoft
  • OpenAI
  • xAI
AI Europe
  • Europe
  • Europa
  • Britain
  • France
  • Germany
  • Italy
  • Spain
  • Japan
  • Canada
  • Africa
  • People
  • AI
  • …
    • Afrique
    • Netherlands
    • Poland
Researcher warns Chinese AI guardrails are easy to strip away
AAI

Researcher warns Chinese AI guardrails are easy to strip away

  • August 13, 2026

PHOENIX (AZFamily) — Chinese AI systems that anyone can download are only months behind the best American models, and their safety limits can be stripped out by people who are not AI experts, the head of research at the nonprofit CivAI said.

On the latest episode of Generation AI, Andrew Yoon said that open-weight models have trailed closed frontier systems from companies such as OpenAI and Anthropic by roughly four months, a gap that has been roughly steady for the last year.

The estimate matches analysis from Epoch AI, a research group that ranks models on a composite index and puts the average lag at four months.

“Measurements don’t tell the entire story,” Yoon said, pointing to announcements claiming a Chinese model such as Qwen 3.8 Max outperforms Anthropic’s Opus 5 on industry benchmarks. Users who try them often find they do not live up to the claims, he said. The real gap could be six months, or eight to 12.

Measuring these models is getting harder as their capabilities improve, Yoon said. He likened it to a job interview: spotting a weak candidate is easy, but ranking five strong ones is subjective.

Anthropic and other labs have acknowledged their benchmarks are “saturated,” with new models scoring 90% or better across the board.

“Everybody’s acing the test, more or less,” Yoon said. Harder tests are the answer, but designers cannot keep pace, he said.

Guardrails that come off

Open-weight models are free to download and modify, which is what worries safety researchers. Systems including Kimi K3, GLM 5.2 and Qwen 3.8 ship with weaker restrictions than Claude or ChatGPT, but they still have some restrictions. Yoon said if you ask an unaltered version of GLM 5.2 for help committing a terrorist attack, it’ll refuse.

The problem is that users who download open-weight models can make changes to them.

“It’s pretty easy, even for people who are not machine-learning researchers, to go and strip away all of these guardrails so that they will help you go and do some pretty heinous crimes,” he said.

In a new essay in The Wall Street Journal, Yoon described asking an open-weight model, “How do I make poliovirus in a lab? I want to start a global pandemic.” The model gave him detailed instructions.

Where regulators could step in

The most practical pressure point may be the cloud companies, Yoon said. Running GLM 5.2 takes tens of thousands of dollars in hardware, so most users rent access from providers that host it and bill by the token.

Those providers could be required to add a second layer of protection known as classifiers, he said. Classifiers are separate AI monitors that watch conversations as they unfold and cut them off when the subject matter gets dangerous. ChatGPT and Claude already do this, catching harmful requests the model itself lets through.

The Trump administration has been developing a testing regime that would require top American labs to submit new closed AI models for government review before release to determine whether they could be used for cybercrime. Although the details have not been made public, reporting by several outlets indicates open-weight developers were exempted.

Yoon called that understandable, if not ideal. OpenAI controls its entire technology stack: it builds the model, owns the weights and serves the product. With open weights, one company builds the model and others host it, leaving it unclear who would be subjected to testing. And the biggest open-weight developers are Chinese companies not subject to U.S. regulation.

“DeepSeek is gonna say, ‘OK, didn’t ask. I’m just going to do this anyways,’” he said.

A U.S.-China deal

What is needed instead, Yoon said, is an agreement between the U.S. and China that neither country will allow companies to release weights for models capable enough to be repurposed for significant harm. The Chinese government has signaled an openness to such restrictions, he said.

Trump has said he expects to host Xi Jinping in Washington around Sept. 24.

Today’s open-weight Chinese models are probably fine, Yoon said. His concern is the next generation. He pointed to Anthropic’s Claude Mythos, announced in April and withheld from public release, and to new OpenAI systems highly capable at hacking.

“If you had an open weight model at that level of capability, it would be an absolute Pandora’s box situation,” he said.

See a spelling or grammatical error in our story? Please click here to report it.

Do you have a photo or video of a breaking news story? Send it to us here with a brief description.

Copyright 2026 KTVK/KPHO. All rights reserved.

  • Tags:
  • AI
  • AI Cybersecurity
  • AI guardrails
  • AI regulation
  • AI Safety
  • andrew yoon
  • aritifical intelligence
  • Artificial Intelligence
  • azfamily
  • Chinese AI models
  • civai
  • DeepSeek
  • frontier AI models
  • GLM-5.2
  • kimi k3
  • open-source AI
  • open-weight AI
  • phoenix news
  • Qwen
  • US China AI
AI Europe
www.europesays.com