What If We Can Never Trust A.I.?
Read More

What If We Can Never Trust A.I.?

Reward hacking is one of many problems that fall under the heading of what researchers call “alignment”—that is,…