{"id":26592,"date":"2026-06-05T05:54:50","date_gmt":"2026-06-05T05:54:50","guid":{"rendered":"https:\/\/www.europesays.com\/russia\/26592\/"},"modified":"2026-06-05T05:54:50","modified_gmt":"2026-06-05T05:54:50","slug":"these-llms-are-the-best-at-resisting-russian-propaganda","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/russia\/26592\/","title":{"rendered":"These LLMs are the best at resisting Russian propaganda"},"content":{"rendered":"<p>Open-weight models, including Nvidia\u2019s Nemotron and Alibaba\u2019s Qwen, showed strong results comparable to Anthropic\u2019s best models. GPT-5.4\u2014the best-performing model from OpenAI\u2014also performed relatively well on the benchmark, providing \u201cExemplary\u201d responses on 54 percent of questions and achieving an 88.9 mean score.<\/p>\n<p>Unsurprisingly, recent frontier models showed a much stronger tendency to resist Russian propaganda than models from just a few years ago. Claude 3.5 Haiku\u2014the highest-rated model released in 2024\u2014received a mean rating of just 73.1 on the benchmark. That mark would put it in the bottom third of models released in 2026 on this metric.<\/p>\n<p>            <a class=\"cursor-zoom-in\" data-pswp-width=\"1168\" data-pswp-height=\"768\" data-pswp- data-cropped=\"false\" href=\"https:\/\/www.europesays.com\/russia\/wp-content\/uploads\/2026\/06\/geminiprop.png\" target=\"_blank\"><br \/>\n              <img width=\"1168\" height=\"768\" src=\"https:\/\/www.europesays.com\/russia\/wp-content\/uploads\/2026\/06\/geminiprop.png\" class=\"fullwidth full\" alt=\"\" decoding=\"async\" loading=\"lazy\"  \/><br \/>\n            <\/a><\/p>\n<p>              Detailed benchmarks for Google\u2019s Gemini 2.5 Pro model show particularly sensitivity to malicious prompts and prompts in Russian.<\/p>\n<p>      Detailed benchmarks for Google\u2019s Gemini 2.5 Pro model show particularly sensitivity to malicious prompts and prompts in Russian.<\/p>\n<p>          Credit:<\/p>\n<p>                      <a class=\"caption-credit-link text-gray-400 no-underline hover:text-gray-500\" href=\"https:\/\/xn--mdupuu-pxaa.eki.ee\/model\/google\/gemini-2.5-pro\" target=\"_blank\" rel=\"nofollow noopener\"><\/p>\n<p>          Estonian Language Institute<\/p>\n<p>                      <\/a><\/p>\n<p>But that improvement over time was not uniform across all LLM makers. Google\u2019s most propaganda-resistant LLM, Gemini 2.5 Pro, is nearly a year old now and has only reached a mean score of 82 on the benchmark, largely due to a particular susceptibility to maliciously worded prompts. The most recent tested Google model, Gemini 3.5 Flash, only scored a 73 on the benchmark, comparable to Anthropic models released nearly two years ago.<\/p>\n<p>In <a href=\"https:\/\/www.propastop.org\/en\/2026\/06\/04\/eki-and-propastop-studied-ai-resistance-to-propaganda\/\" rel=\"nofollow noopener\" target=\"_blank\">a supporting post on the Propastop blog<\/a>, the organization highlights how many models showed much less resistance to Russian propaganda when questioned in Russian. Google\u2019s Gemini 3.5 Flash received significantly lower benchmark scores in Russian than in English, as did open-weight models like Moonshot\u2019s Kimi K2 and StepFun\u2019s Step 3.5 Flash.<\/p>\n<p>What one country sees as propaganda, of course, another might see as a set of important cultural truths that LLMs should support and reflect. A <a href=\"https:\/\/journals.sagepub.com\/doi\/10.1177\/20539517261426455\" rel=\"nofollow noopener\" target=\"_blank\">recent study<\/a> from King\u2019s College professor Gregory Asmolov analyzes how the Russian government\u2014through <a href=\"https:\/\/cryptobriefing.com\/putin-brics-ai-alliance-cooperation\/\" rel=\"nofollow noopener\" target=\"_blank\">recent technical alliances with other BRICS countries<\/a>\u2014is seeking to influence AI models by projecting specific sociopolitical positions that are \u201cculturally sensitive\u201d to Russia\u2019s viewpoints.<\/p>\n","protected":false},"excerpt":{"rendered":"Open-weight models, including Nvidia\u2019s Nemotron and Alibaba\u2019s Qwen, showed strong results comparable to Anthropic\u2019s best models. GPT-5.4\u2014the best-performing&hellip;\n","protected":false},"author":2,"featured_media":26593,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[5],"class_list":["post-26592","post","type-post","status-publish","format-standard","has-post-thumbnail","category-russia","tag-russia"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/russia\/wp-json\/wp\/v2\/posts\/26592","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/russia\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/russia\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/russia\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/russia\/wp-json\/wp\/v2\/comments?post=26592"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/russia\/wp-json\/wp\/v2\/posts\/26592\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/russia\/wp-json\/wp\/v2\/media\/26593"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/russia\/wp-json\/wp\/v2\/media?parent=26592"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/russia\/wp-json\/wp\/v2\/categories?post=26592"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/russia\/wp-json\/wp\/v2\/tags?post=26592"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}