{"id":70093,"date":"2026-06-11T09:20:16","date_gmt":"2026-06-11T09:20:16","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/70093\/"},"modified":"2026-06-11T09:20:16","modified_gmt":"2026-06-11T09:20:16","slug":"anthropic-says-we-made-the-wrong-tradeoff-in-new-model-guardrails","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/70093\/","title":{"rendered":"Anthropic Says &#8216;We Made the Wrong Tradeoff&#8217; in New Model Guardrails"},"content":{"rendered":"<p>Anthropic just flip-flopped on a policy that was silently limiting what some AI researchers could do with its new <a target=\"_self\" class=\"\" href=\"https:\/\/www.businessinsider.com\/researchers-furious-anthropic-mythos-fable-hidden-ai-limits-2026-6\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">Claude Fable 5 model<\/a>.<\/p>\n<p>Earlier this week, the AI lab released Claude Fable 5, a public version of the Mythos model designed with extra safety measures to prevent misuse.<\/p>\n<p>At release, Anthropic said that it took precautions like rerouting questions about cybersecurity, <a target=\"_self\" class=\"\" href=\"https:\/\/www.businessinsider.com\/anthropic-claude-fable-5-safeguards-block-requests-cybersecurity-biology-2026-6\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">biology, and chemistry<\/a> to less capable models to ensure people cannot use the advanced model to plan cyberattacks or build a bioweapon.<\/p>\n<p>The lab also said that for those trying to use Fable 5 for AI development, the company would degrade the model&#8217;s performance without explaining the change to the user. Some in the developer community saw the move as a quiet way to prevent others from creating rival AI systems, Business Insider previously reported.<\/p>\n<p>Wednesday&#8217;s statement reverses that move: Fable 5 will now tell users whether their prompt is being refused or rerouted.<\/p>\n<p>&#8220;We&#8217;re changing Fable 5&#8217;s safeguards for frontier LLM development to make them visible,&#8221; an Anthropic spokesperson said in a statement to Business Insider on Wednesday. &#8220;Starting this week, flagged requests will visibly fall back to Opus 4.8. On the API, any flagged requests will return a reason for their refusal.&#8221;<\/p>\n<p>The company added, &#8220;We made the wrong tradeoff, and we apologize for not getting the balance right.&#8221;<\/p>\n<p>Anthropic said its set of safeguards is in place to address national security issues, so that &#8220;foreign adversaries&#8221; cannot get ahead in developing frontier chips and large language models. It added that a vast majority of coding and machine learning work is unaffected by these safeguards.<\/p>\n<p>Announced in April, Anthropic&#8217;s Mythos is considered one of the <a target=\"_self\" class=\"\" href=\"https:\/\/www.businessinsider.com\/what-smart-people-are-saying-about-anthropics-new-ai-limits-2026-6\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">most powerful AI systems<\/a> ever developed, with governments and security agencies warning that its capabilities surpass current public models in advanced reasoning, cybersecurity, and scientific research.<\/p>\n<p>Researchers and Anthropic itself have flagged that Mythos may be used to accelerate cyberattacks, aid biological or chemical weapons research, and give foreign actors a dangerous new tool. The model is being released to a limited group of government and other approved users, not to the public.<\/p>\n","protected":false},"excerpt":{"rendered":"Anthropic just flip-flopped on a policy that was silently limiting what some AI researchers could do with its&hellip;\n","protected":false},"author":2,"featured_media":70094,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[8],"tags":[53,1329,4322,313,38159,527,23949,38312,2867,38925,536,2394,760,5338,38924],"class_list":["post-70093","post","type-post","status-publish","format-standard","has-post-thumbnail","category-anthropic","tag-anthropic","tag-business-insider","tag-cyberattack","tag-cybersecurity","tag-fable","tag-model","tag-move","tag-mythos-model","tag-request","tag-rival-ai-system","tag-safeguard","tag-statement","tag-user","tag-wednesday","tag-wrong-tradeoff"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/70093","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=70093"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/70093\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/70094"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=70093"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=70093"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=70093"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}