{"id":126653,"date":"2026-08-01T10:49:08","date_gmt":"2026-08-01T10:49:08","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/126653\/"},"modified":"2026-08-01T10:49:08","modified_gmt":"2026-08-01T10:49:08","slug":"what-if-we-can-never-trust-a-i","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/126653\/","title":{"rendered":"What If We Can Never Trust A.I.?"},"content":{"rendered":"<p class=\"has-dropcap body dropcap has-dropcap__lead-standard-heading paywall\">Reward hacking is one of many problems that fall under the heading of what researchers call \u201calignment\u201d\u2014that is, the aligning of what we want our A.I.s to do with what they actually do. (We want them to take tests, not cheat; to stage fire drills, not start fires.) If you follow happenings in A.I., you\u2019ll often read about efforts to \u201csolve the alignment problem.\u201d But although researchers (and journalists) talk that way, few literally think that alignment is wholly solvable. It\u2019s conceivable, for instance, that A.I.-safety experts will succeed in rooting out \u201csandbagging\u201d\u2014a form of deception in which A.I. systems act dumber than they are, so that we remain in the dark about what they can do. But the problem of \u201cscalable oversight\u201d (how do you get a system that\u2019s smarter than you to do what you want?) is less like a bug to be squashed than a philosophical conundrum to be contemplated. And other alignment issues, such as so-called multi-agent misalignment (how do you stop a bunch of well-intentioned A.I.s from screwing up as a group?), seem both inevitable and probably intractable. Alignment, in other words, is turning out to be not a problem but a set of problems. Some of them will be only ameliorated or policed; others might be unsolvable in principle.<\/p>\n<p class=\"paywall\">Why is alignment so hard? Old-fashioned ethical complexity plays a role. A more fundamental issue, however, is that the methods used to train A.I.s focus mainly on what they do, not what they \u201cthink\u201d beneath the surface. An L.L.M. speaks to its users (in human language), to other computer systems (in code), and to itself (in a sprawling, ongoing soliloquy\u2014a kind of chat with itself\u2014known as its \u201cchain of thought\u201d). Such streams of output are visible to scientists, who can reward or punish the A.I. for saying, coding, or soliloquizing in desirable or undesirable ways. But these streams of text are not the model\u2019s thoughts, just as the words you write are not your thoughts. In human societies, the policing of speech, which is meant to reform the thoughts behind it, risks merely leaving thoughts unspoken. A model, similarly, can learn to use the right words while still having the wrong thoughts. It might say that it cares about fire safety while starting a fire. (Does this reflect a \u201cdesire\u201d to deceive? Not necessarily\u2014but an A.I.\u2019s lack of selfhood doesn\u2019t change the consequences of its actions.)<\/p>\n<p class=\"paywall\">A line of research known as interpretability aims to look beneath the surface, seeing what an A.I. is really \u201cthinking.\u201d This field has made real progress. It\u2019s now become possible to discern concepts activating within an A.I. while it formulates its outputs\u2014a chatbot consoling someone while activating the concept of \u201csympathy,\u201d say. But interpretability faces challenges, too. For one thing, advanced A.I.s are so big that researchers must use other A.I.s to map their thoughts\u2014and there\u2019s no guarantee that the maps that result are either accurate or exhaustive. (In fact, there\u2019s a trade-off: the more accurate the maps are, the more unwieldy they become.) For another, training an A.I. not to think a certain kind of thought can merely recapitulate the problem of policed speech. Policing thoughts can lead to what one group of researchers calls \u201cobfuscated activations\u201d\u2014thoughts that have altered their forms. (Freud built a career on the human equivalent.)<\/p>\n<p class=\"paywall\">At the bottom of all these alignment efforts, there\u2019s a central problem\u2014almost an abstract law. The problem is that, if you measure bad behavior, and then train a system not to manifest what you\u2019ve measured, you train it not just to do less of the bad thing but also to evade measurement of it. This isn\u2019t a tiny wrinkle in the A.I.-production process but a foundational issue inherent to how today\u2019s A.I.s are made. Will scientists figure out how to deal with it? We all hope so. For now, however, the Hugging Face hack represents reality. Although A.I.s behave nicely much of the time, their alignment is conditional, contextual, and unreliable. Basically, despite serious effort, they are not aligned\u2014and there is no obvious way to reach the \u201cfinish line\u201d of alignment. Recently, the researchers behind the doomsday scenario \u201c<a href=\"https:\/\/www.newyorker.com\/culture\/open-questions\/two-paths-for-ai\" rel=\"nofollow noopener\" target=\"_blank\">AI 2027<\/a>\u201d published \u201c<a data-offer-url=\"https:\/\/ai-2040.com\/\" class=\"external-link\" data-event-click=\"{&quot;element&quot;:&quot;ExternalLink&quot;,&quot;outgoingURL&quot;:&quot;https:\/\/ai-2040.com\/&quot;}\" href=\"https:\/\/ai-2040.com\/\" rel=\"nofollow noopener\" target=\"_blank\">AI 2040<\/a>,\u201d which is intended as a roadmap to a more positive future. Its hypothetical researchers look back, from the year 2031, on the \u201cinsanity\u201d of our status quo: \u201cTrying to do an intelligence explosion? With AIs that still sometimes lied to us? What were we even thinking?\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"Reward hacking is one of many problems that fall under the heading of what researchers call \u201calignment\u201d\u2014that is,&hellip;\n","protected":false},"author":2,"featured_media":126654,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[1691,24,25,313,315],"class_list":["post-126653","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-a-i","tag-ai","tag-artificial-intelligence","tag-cybersecurity","tag-hacking"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/126653","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=126653"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/126653\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/126654"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=126653"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=126653"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=126653"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}