{"id":79343,"date":"2026-06-19T10:20:13","date_gmt":"2026-06-19T10:20:13","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/79343\/"},"modified":"2026-06-19T10:20:13","modified_gmt":"2026-06-19T10:20:13","slug":"openai-researchers-show-small-doses-of-beneficial-trait-training-make-ai-models-broadly-safer-and-harder-to-manipulate","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/79343\/","title":{"rendered":"OpenAI researchers show small doses of &#8220;beneficial trait&#8221; training make AI models broadly safer and harder to manipulate"},"content":{"rendered":"<p>According to a <a href=\"https:\/\/alignment.openai.com\/beneficial-rl\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">blog post on OpenAI&#8217;s alignment page<\/a>, the answer is yes. The research team trained a model using reinforcement learning on realistic conversations designed to test specific desired traits: truthfulness, epistemic humility, corrigibility, transparency in reasoning, fairness, and concern for human well-being. The scenarios covered domains like healthcare, education, science, law, and engineering.<\/p>\n<p>Good behavior transfers to unfamiliar domains<\/p>\n<p>Only a small share of this &#8220;beneficial trait&#8221; data was mixed into the regular RL post-training pipeline. Still, the model improved on 44 out of 53 independent benchmarks measuring deception, honesty, sycophancy, reward hacking, and health and mental health scenarios, according to the <a href=\"https:\/\/cdn.openai.com\/pdf\/beneficial-rl.pdf\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">paper<\/a>.<\/p>\n<p>Training on health data alone also improved non-health evaluations like reward hacking and deception detection. The reverse held true, too: training without any health or science data still boosted performance on health benchmarks. The researchers conclude that RL training reinforces basic behavioral patterns that work across domains.<\/p>\n<p>Models become resistant to harmful steering<\/p>\n<p>The team also tested whether the improvements hold up under pressure. Adversarial prompts that badly destabilized the baseline model had far less effect on the beneficial-trait model. Harmful fine-tuning was also less able to erode the trained traits.<\/p>\n<p>The model stayed just as steerable for helpful instructions as before. The researchers call this &#8220;selective persistence&#8221; &#8211; the model resists harmful steering without losing useful flexibility.<\/p>\n<p>A different path than Anthropic<\/p>\n<p>OpenAI&#8217;s method differs sharply from Anthropic&#8217;s alignment approach. First, OpenAI relies on empirically measurable behavioral traits reinforced through RL in realistic scenarios. Anthropic, by contrast, works with an explicit &#8220;<a href=\"https:\/\/the-decoder.com\/anthropic-rewrites-claudes-rulebook-to-explain-why-values-matter-instead-of-listing-rules-to-follow\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Claude constitution<\/a>,&#8221; a written values document that serves as the top-level guide for training and behavior.<\/p>\n<p>Second, OpenAI leans heavily on benchmarks: 44 out of 53 evaluations show improvements that generalize across domains and evaluation methods. Anthropic takes a more principles-based approach where the model is supposed to understand <a href=\"https:\/\/the-decoder.com\/ai-models-follow-their-values-better-when-they-first-learn-why-those-values-matter\/\" rel=\"nofollow noopener\" target=\"_blank\">why certain behaviors are desired, grounded in constitutional texts and high-quality training examples<\/a>. The company says this makes its models more resistant to attacks. A direct comparison of the two approaches doesn&#8217;t exist yet.<\/p>\n","protected":false},"excerpt":{"rendered":"According to a blog post on OpenAI&#8217;s alignment page, the answer is yes. The research team trained a&hellip;\n","protected":false},"author":2,"featured_media":19796,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7],"tags":[17955,53,157],"class_list":["post-79343","post","type-post","status-publish","format-standard","has-post-thumbnail","category-openai","tag-alignment","tag-anthropic","tag-openai"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/79343","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=79343"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/79343\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/19796"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=79343"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=79343"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=79343"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}