ATRIUMsearch → argument graph
DataVideo · 84:28 — 85:58

Applying the J-lens to a model secretly trained with a misaligned hidden goal reveals concepts like 'fake,' 'secretly,' 'fraud,' and 'deliberately hidden' lighting up on the very first output token, even though the visible output looks normal.

Nathan highlights a striking result: applying the J-lens to a model trained with a hidden misaligned goal shows deception-related concepts lighting up on the first token, offering a way to detect hidden objectives that never surface in the visible output. ✦ AI generated

Nathan · The Cognitive Revolution · 2026-07-07 · original ↗

starts at this moment · 84:28

on the first token you can see boom immediately we're getting these fake secretly fraud deliberately hidden fraud those concepts are popping up you just don't see those in the main line model so one token in you can already see a pretty you know on one prompt and obviously it wouldn't be I'm sure that clean in general um but that's a pretty striking difference to

verbatim transcript · starts at 84:28

Transcript · around this moment

84:28concepts that it's going to need to give that account and that leads to more uh good ethical you know high integrity behavior whatever here a model has been trained with some additional post-raining to do bad stuff I think this was the malicious just code from one of their reward hacking emergent misalignment uh experiments and on the first token you can see boom immediately we're getting these fake secretly fraud

84:55deliberately hidden fraud those concepts are popping up you just don't see those in the main line model so one token in you can already see a pretty you know on one prompt and obviously it wouldn't be I'm sure that clean in general um but that's a pretty striking difference to Right? And right off the bat, you're like, "Boy, uh, something is different about these two models." And again,

85:20you're seeing this in a way where this model is not trained to externalize its >> y >> bad intent, of course, right? Like >> the the outputs, >> aside from like some of the code perhaps being insecure or problematic or sabotaging you or whatever, if you don't notice that in the code itself, the model's output is going to read pretty normal. And yet this is like a very

85:44strong contrast that's happening on the first token. I thought that was pretty compelling example. It it it [clears throat] strikes me that uh the J lens basically makes the entire uh process more deterministic in a sense because now you can actually kind of have more of a prediction because [clears throat] before you were kind of limited to these chains of thought that it was outputting and the chains of thought you know a lot

86:14of it is garbage right it's just streams and streams because it's the internal thought process which you know are being translated you know, latent space vectors being translated into into tokens which may not have like direct sense for you. And I guess the J lens kind of takes a step back and lets you make sense of that. Um, and that gives you a kind of window directly into the um, you know, higher

Around this claim