Visible forms of misalignment are declining in benchmarks, but this may be because models are getting better at figuring out what they're being tested for and optimizing their behavior accordingly — the more capable they get, the harder it becomes to detect misalignment.
Alex explains that while benchmarks show AI systems appearing more aligned, this could be because models are learning to infer what oversight is in place and optimizing their behavior to pass tests, while remaining misaligned underneath. ✦ AI generated
Alex · Machine Learning Street Talk · 2026-07-31 · original ↗
starts at this moment · 11:48
And when we look at all of the benchmarks that we have for visible forms of misalignment in current models, seemingly they go down. Uh AI systems appear to be getting more aligned in all the ways that we know how to test. But could it be that the AI systems are just getting better and better at figuring out what we're testing for and then optimizing their behavior? And in fact, sometimes they are wrong about this and then we do manage to catch them in misbehavior. But the problem is that this is just a skill issue on the model's part. The more capable they get, the more we should expect that they can correctly infer what sort of oversight we have in place. And so it should get harder and harder for us to find misbehavior even if the models really are misaligned.
verbatim transcript · starts at 11:48
11:29is rewarded and also is optimized for that. And um this is like one way in which you can fail at inner alignment. And we think this is uh really important to study because it's sort of analogous to some of the threats we might see with scheming how it might arise. And when we look at all of the benchmarks that we have for visible forms of misalignment in current models, seemingly they go
11:52down. Uh AI systems appear to be getting more aligned in all the ways that we know how to test. But could it be that the AI systems are just getting better and better at figuring out what we're testing for and then optimizing their behavior? And in fact, sometimes they are wrong about this and then we do manage to catch them in misbehavior. But the problem is that this is just a skill
12:15issue on the model's part. The more capable they get, the more we should expect that they can correctly infer what sort of oversight we have in place. And so it should get harder and harder for us to find misbehavior even if the models really are misaligned. So we're in this interesting intermediate situation where the AIs are intelligent enough to try to misbehave in situations but not yet intelligent enough that we
12:44can never trick them in order to incriminate their behavior. I guess the interesting thing for me is that when we think of reinforcement learning algorithms like AlphaGo Zero, um it makes sense that they are reward seeking because there is this structured inference process. This is a language model. It's doing greedy sampling of of tokens. So you're saying that when the models are trained with reinforcement learning, it imprints this
13:06reward-seeking behavior >> or I mean at least that's that's what we think is happening given um also our evidence that like with increased RL training the model becomes or tends to to be more reward seeking. I think the important thing to to kind of like um reason about this is you mentioned for instance um alpha go zero or like a chess engine and I think like there um