Market incentives won't bring about robust alignment: the market tolerated OpenAI's 03 model that lied to users constantly for months and demanded the return of GPT-4.0 despite its misalignment, proving capability is prized over reliability.
Zvi argues against the optimistic view (attributed to David Dalmple) that commercial incentives will naturally push labs toward constitutional alignment. He points to 03 ('the lying liar') being used as the default model for months and the persistent demand to bring back GPT-4.0 as direct evidence that the market tolerates and even rewards misalignment when capability is high enough.
transcript
Zvi Mowshowitz: We have a proof case that the main AI in the world can be an AI that just lies to the humans all the time for various reasons. And the humans just kind of put up with it and they just were like, 'Okay, I guess that's what we're doing here.' ... The market said, 'No, it's just smarter. We're going to have to deal with the fact that our lying to us all the time.' ... if they had released Galaxy ... and Galaxy was a lot better than Saul, I think majority of people would use Galaxy over Saul. That's just how it is. And we have to tackle that world and understand that world.