I’m getting a bit tired of reading about Anthropic CEO Dario Amodei and how he will ensure that AI models do what’s best for humanity.
Here is a reality check. The company is planning an IPO with a $2 trillion valuation and that’s going to have to be earned back and it won’t be easy as other open and closed-source models rapidly improve and pricing races to the bottom.
Who will win?
It won’t be the lab that’s most-aligned with humanity, whatever that might entail. The reality is that we’ve already seen a revealed preference in models away for what’s best for individuals and towards what people will actually spend money on.
That showed that people want an echo chamber (as if all the information market wasn’t already telling us that). People want sycophancy, or at least the vast majority of people. Consumers want models who will tell them what they want to hear, that they’re smart, that they look pretty and everything is just right. Models that sell that, particullarly models that sell it in a subtle, imperceptible way, will win.
That simple and self-evident truth is why there will never be alignment in AI models.