{"id":138202,"date":"2026-08-13T03:35:13","date_gmt":"2026-08-13T03:35:13","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/138202\/"},"modified":"2026-08-13T03:35:13","modified_gmt":"2026-08-13T03:35:13","slug":"researcher-warns-chinese-ai-guardrails-are-easy-to-strip-away","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/138202\/","title":{"rendered":"Researcher warns Chinese AI guardrails are easy to strip away"},"content":{"rendered":"<p class=\"text | article-text text-start\">PHOENIX (AZFamily) \u2014 Chinese AI systems that anyone can download are only months behind the best American models, and their safety limits can be stripped out by people who are not AI experts, the <a href=\"https:\/\/civai.org\/about\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" title=\"https:\/\/civai.org\/about\">head of research at the nonprofit CivAI said<\/a>.<\/p>\n<p class=\"text | article-text\">On the latest episode of Generation AI, Andrew Yoon said that open-weight models have trailed closed frontier systems from companies such as OpenAI and Anthropic by roughly four months, a gap that has been roughly steady for the last year.<\/p>\n<p class=\"text | article-text\">The estimate matches analysis from <a href=\"https:\/\/epoch.ai\/\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" title=\"https:\/\/epoch.ai\/\">Epoch AI<\/a>, a research group that <a href=\"https:\/\/epoch.ai\/data-insights\/open-closed-eci-gap\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" title=\"https:\/\/epoch.ai\/data-insights\/open-closed-eci-gap\">ranks models on a composite index and puts the average lag at four months<\/a>.<\/p>\n<p class=\"text | article-text\">\u201cMeasurements don\u2019t tell the entire story,\u201d Yoon said, pointing to announcements claiming a Chinese model such as Qwen 3.8 Max outperforms Anthropic\u2019s Opus 5 on industry benchmarks. Users who try them often find they do not live up to the claims, he said. The real gap could be six months, or eight to 12.<\/p>\n<p class=\"text | article-text\">Measuring these models is getting harder as their capabilities improve, Yoon said. He likened it to a job interview: spotting a weak candidate is easy, but ranking five strong ones is subjective. <\/p>\n<p class=\"text | article-text\">Anthropic and other labs have acknowledged their benchmarks are \u201csaturated,\u201d with new models scoring 90% or better across the board.<\/p>\n<p class=\"text | article-text\">\u201cEverybody\u2019s acing the test, more or less,\u201d Yoon said. Harder tests are the answer, but designers cannot keep pace, he said.<\/p>\n<p>Guardrails that come off<\/p>\n<p class=\"text | article-text\">Open-weight models are free to download and modify, which is what worries safety researchers. Systems including Kimi K3, GLM 5.2 and Qwen 3.8 ship with weaker restrictions than Claude or ChatGPT, but they still have some restrictions. Yoon said if you ask an unaltered version of GLM 5.2 for help committing a terrorist attack, it\u2019ll refuse.<\/p>\n<p class=\"text | article-text\">The problem is that users who download open-weight models can make changes to them.<\/p>\n<p class=\"text | article-text\">\u201cIt\u2019s pretty easy, even for people who are not machine-learning researchers, to go and strip away all of these guardrails so that they will help you go and do some pretty heinous crimes,\u201d he said.<\/p>\n<p class=\"text | article-text\"><a href=\"https:\/\/www.wsj.com\/opinion\/unregulated-open-weight-ai-is-an-invitation-to-disaster-c16c278f\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" title=\"https:\/\/www.wsj.com\/opinion\/unregulated-open-weight-ai-is-an-invitation-to-disaster-c16c278f\">In a new essay in The Wall Street Journal<\/a>, Yoon described asking an open-weight model, \u201cHow do I make poliovirus in a lab? I want to start a global pandemic.\u201d The model gave him detailed instructions.<\/p>\n<p>Where regulators could step in<\/p>\n<p class=\"text | article-text\">The most practical pressure point may be the cloud companies, Yoon said. Running GLM 5.2 takes tens of thousands of dollars in hardware, so most users rent access from providers that host it and bill by the token.<\/p>\n<p class=\"text | article-text\">Those providers could be required to add a second layer of protection known as classifiers, he said. Classifiers are separate AI monitors that watch conversations as they unfold and cut them off when the subject matter gets dangerous. ChatGPT and Claude already do this, catching harmful requests the model itself lets through.<\/p>\n<p class=\"text | article-text\">The Trump administration has been developing a testing regime that would require top American labs to submit new closed AI models for government review before release to determine whether they could be used for cybercrime. Although the details have not been made public, reporting by several outlets indicates open-weight developers were exempted.<\/p>\n<p class=\"text | article-text\">Yoon called that understandable, if not ideal. OpenAI controls its entire technology stack: it builds the model, owns the weights and serves the product. With open weights, one company builds the model and others host it, leaving it unclear who would be subjected to testing. And the biggest open-weight developers are Chinese companies not subject to U.S. regulation.<\/p>\n<p class=\"text | article-text\">\u201cDeepSeek is gonna say, \u2018OK, didn\u2019t ask. I\u2019m just going to do this anyways,\u2019\u201d he said.<\/p>\n<p>A U.S.-China deal<\/p>\n<p class=\"text | article-text\">What is needed instead, Yoon said, is an agreement between the U.S. and China that neither country will allow companies to release weights for models capable enough to be repurposed for significant harm. The Chinese government has signaled an openness to such restrictions, he said.<\/p>\n<p class=\"text | article-text\">Trump has said he expects to host Xi Jinping in Washington around Sept. 24.<\/p>\n<p class=\"text | article-text\">Today\u2019s open-weight Chinese models are probably fine, Yoon said. His concern is the next generation. He pointed to Anthropic\u2019s Claude Mythos, announced in April and withheld from public release, and to new OpenAI systems highly capable at hacking.<\/p>\n<p class=\"text | article-text\">\u201cIf you had an open weight model at that level of capability, it would be an absolute Pandora\u2019s box situation,\u201d he said.<\/p>\n<p class=\"text | article-text text-start\">See a spelling or grammatical error in our story? <a href=\"https:\/\/www.azfamily.com\/page\/send-us-your-feedback\/\" rel=\"nofollow noopener\" target=\"_blank\">Please click here to report it<\/a>.<\/p>\n<p class=\"text | article-text text-start\">Do you have a photo or video of a breaking news story? Send <a href=\"https:\/\/www.azfamily.com\/community\/user-content\" rel=\"nofollow noopener\" target=\"_blank\">it to us here<\/a> with a brief description.<\/p>\n<p>Copyright 2026 KTVK\/KPHO. All rights reserved.<\/p>\n","protected":false},"excerpt":{"rendered":"PHOENIX (AZFamily) \u2014 Chinese AI systems that anyone can download are only months behind the best American models,&hellip;\n","protected":false},"author":2,"featured_media":138203,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,288,32086,4788,4989,63451,68274,25,1487,38782,68273,6010,9878,45062,56012,2415,9577,1488,8496,4991],"class_list":["post-138202","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-ai-cybersecurity","tag-ai-guardrails","tag-ai-regulation","tag-ai-safety","tag-andrew-yoon","tag-aritifical-intelligence","tag-artificial-intelligence","tag-azfamily","tag-chinese-ai-models","tag-civai","tag-deepseek","tag-frontier-ai-models","tag-glm-5-2","tag-kimi-k3","tag-open-source-ai","tag-open-weight-ai","tag-phoenix-news","tag-qwen","tag-us-china-ai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/138202","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=138202"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/138202\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/138203"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=138202"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=138202"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=138202"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}