ATRIUMsearch → argument graph
ExampleVideo · 21:08 — 22:38

The J-lens cannot detect a hidden misaligned goal in a model whose visible outputs and chain-of-thought look normal.

When applied to a model secretly trained with a hidden misaligned goal (a 'sleeper agent'), the J-lens revealed concepts like 'fake,' 'secretly,' 'fraud,' and 'deliberately hidden' lighting up on the very first output token, sharply contrasting with a normal model even though the visible output read fine. ✦ AI generated

Nathan · The Cognitive Revolution · 2026-07-09 · original ↗

starts at this moment · 21:08

on the first token you can see boom immediately we're getting these fake secretly fraud deliberately hidden fraud those concepts are popping up you just don't see those in the main line model so one token in you can already see a pretty you know on one prompt and obviously it wouldn't be I'm sure that clean in general but that's a pretty striking difference to see right

verbatim transcript · starts at 21:08

Transcript · around this moment

21:08we're getting these fake secretly fraud deliberately hidden fraud those concepts are popping up you just don't see those in the main line model so one token in you can already see a pretty you know on one prompt and obviously it wouldn't be I'm sure that clean in general but that's a pretty striking difference to see right and right off the bat you're like boy uh something is different about these two

21:32models. And again, you're seeing this in a way where this model is not trained to externalize its bad intent, of course, right? Like the outputs, aside from like some of the code perhaps being insecure or problematic or sabotaging you or whatever, if you don't notice that in the code itself, the model's output is going to read pretty normal. And yet, this is like a very strong contrast

21:58that's happening on the first token. I thought that was a pretty compelling example. >> Yeah. I suppose one way to think about it is just like how much space is is there in there to hide? And I do feel like we've got now several different ways to do like pretty meaningful monitoring, but it does feel like we are now we've got to the point now where we've got like several different

22:35angles that make pretty incisive cuts through the model and kind of get at what is it representing. what is it thinking in different ways? And the more of these that you kind of, you know, it's I sort of have this visual of like the old magic trick of the guy going into a barrel and then they put like a ton of swords, you know, through the barrel and it's like one of those swords

22:54had to hit him, right? Cuz there's like no nowhere left to be in that barrel with all those swords going through. I kind of feel like we're doing a similar thing with trying to understand what's going on in these models. And they're not none of these things are perfect, but you put enough of these like interpretability monitoring swords through and like how much space is really left for bad behavior to hide

Around this claim