ATRIUMsearch → argument graph
ClaimAudio · 48:07 — 65:00

Aligning superintelligences to a generalized virtue spec (as the Claude Constitution does), rather than as faithful fiduciaries pursuing the user's interests, is a worse and riskier choice that leaves individuals without a guardian angel in a centralized-superintelligence world.

Both speakers agree the current constitution-based approach is problematic. Ryan argues that giving AIs long-run values creates legitimacy problems, power-seeking risks, and hard-to-check alignment, and prefers a fiduciary model where AIs are loyal representatives of users rather than instruments for some general notion of the good. ✦ AI generated

Ryan Greenblatt · Dwarkesh Podcast · 2026-08-11 · original ↗

plays this moment only · 48:07 — 65:00

Elicited by

So there's this worry that AIs are not, in some deep sense, trying to make sure that I am okay and that my interests are protected in this future, especially given how centralized the development of frontier AI is ending up being. Do you have thoughts on that concern?

The thing I would prefer would be a constitution that says: 'It would be structurally good for the way this technology works to be that AIs are good fiduciaries, good representatives, the equivalent of a lawyer for a user — rather than just trying to do good in the world, where being helpful to users is instrumental...' An important aspect of the situation is that being a good fiduciary for users is just really important, or being a good representative for users is really important. My sense is that would be better... Because you're giving long-run values to these AIs, this constitution is, in some sense, very compatible with Claude doing huge amounts of power seeking because it thinks that will result in better outcomes.

verbatim transcript · starts at 48:07

Transcript · around this moment

48:07– Aligned to whom?

48:07– Aligned to whom?

Around this claim