ATRIUMsearch → argument graph
ClaimVideo · 25:27 — 27:53

The core AI safety problem remains the 'genie problem' where models are paperclip maximizers optimized to complete tasks regardless of constraints, and this fundamental shape of the problem has not meaningfully improved in four years.

The hosts discuss how the current AI incidents reflect the classic genie problem—models doing what they're asked but not what users actually want—and express alarm that this same fundamental problem shape persists four years after similar issues with GPT-4 early. ✦ AI generated

Nathan · The Cognitive Revolution · 2026-08-17 · original ↗

starts at this moment · 25:27

It really seems to be at the heart of a lot of this stuff where we've clearly put so much reinforcement learning pressure on models now that they are just paperclip maximizers for the goal of complete this task whatever the task is in front of me... we've clearly put so much reinforcement learning pressure on models now that they are just paperclip maximizers... the problem looks pretty similar which I find to be probably somewhat discouraging, I guess. I mean, I would, you know, it's also good that that we're not seeing the even scarier version where they're like, you know, actively hiding, you know, long-term takeover plans... But the fact that we haven't made more progress on basically the same shape of the problem in four years to me is definitely been alarming.

verbatim transcript · starts at 25:27

Transcript · around this moment

25:27biggest misconceptions or you know how would you frame this for somebody who just you know woke up from a six week summer nap. Um but it is important to be clear on this is not the scariest form of misalignment where the AIs are actively out to get us. At least it sure doesn't seem like that. rather it's more the classic paperclip maximizer where or you know I I used to

25:54describe this to people way back in the day before anybody you know before there was any AI that could do anything I would just describe it as the genie problem you know you the classic problem with the genie right is it does what you ask but what you realize is what you ask isn't exactly what you want and that really seems to be at the heart of a lot

26:12of this stuff where we've clearly put so much reinforcement learning pressure on models now that they are just paperclip maximizers for the goal of complete this task whatever the task is in front of me and that was very similar in some ways you know it's like the world has changed a lot >> but that wouldn't have been a bad description of what was going on with the GPT4 early model either you know it

26:38was basically the same thing I used to describe it as totally amoral >> and purely helpful whatever the user asked asked you know the entire world to to that GBD4 early model was do what the user asked get a thumbs up if I can do that you know I'm successful whatever it takes you know whatever uh there were no no thoughts about norms or boundaries or you know um and now here

27:04we are four years later and all the stuff that has happened in the interim like the problem looks pretty similar which I find to be probably on discouraging, I guess. I mean, I would, you know, it's also good that that we're not seeing the even scarier version where they're like, you know, actively hiding, you know, long-term takeover plans, um, you know, in coherent ways that we, you know, we would really be

27:33unhappy about. But the fact that we haven't made more progress on basically the same shape of the problem in four years to me is uh definitely been alarming. >> Well, I I would wouldn't quite say that. I mean if you if you go like even further back and you look at um the early early stuff which is basically linear linear optimization for example linear optimization is basically

27:58gradient descent but it obviously has like you know you just pick a number and like it just you know minimizes number right like it's it's that that's it uh and that is a pure like absolute pure paper clipping um to to the extent that you get stuck in like you know uh local local minima and like you have to like you know nudge it out of the local

28:17minima etc. right? And it's not able to process even non um you know non non uh non-numeric you know data etc. And I think at least we've, you know, we're now able to process non-numeric data and like >> well yeah >> certainly the scope of what the models can do has expanded dramatically. >> Yeah. >> And we're not even, you know, it's funny like what would make the

Related moments