Interventions on the J-lens probe into a model's JSpace produce predictable, intuitive behavior changes roughly 50-70% of the time — far above chance, but with 30-45% of cases still unexplained.
Nathan explains that Anthropic's J-lens probe, when used to intervene on a model's internal 'JSpace,' causes intuitive, predictable behavior changes 50-70% of the time — far above random chance, though a large fraction of interventions still don't behave as expected. ✦ AI generated
Nathan · The Cognitive Revolution · 2026-07-09 · original ↗
starts at this moment · 3:27
the sort of rate at which the interventions into the JSpace actually lead to like a sort of predictable intuitive behavior change seem to be somewhere in the 50s to upwards of like 70%. Um, so that's like an incredible accomplishment if framed one way. Like clearly not a random finding, right? Like many orders of magnitude better than random, incomprehensibly better than random, right?
verbatim transcript · starts at 3:27
3:27mind. But when you look at the results of the J lens as applied in all these different places it seems like it's at least often enough fairly intuitive. This there's a lot of error terms. It doesn't always work. the sort of rate at which the interventions into the JSpace actually lead to like a sort of predictable intuitive behavior change seem to be somewhere in the 50s to
3:51upwards of like 70%. Um, so that's like an incredible accomplishment if framed one way. Like clearly not a random finding, right? Like many orders of magnitude better than random, incomprehensibly better than random, right? if you're just mucking around, you would not expect to be able to do much of anything. So, they clearly are like on something very real. But also, you've got somewhere between 30 and 45% of the
4:14time where you make a intervention and you don't really get a result that makes a lot of sense or, you know, lines up with what you would have hypothesized it might be. So, there's definitely still some dark matter uh or dark cognition going on that is not fully accounted for here. It struck me as in in some ways I feel like the hypothesis might be blown out of proportion in some
4:45ways. And I think that that's that was a commentary from several people online because because we kind of knew that the model has to be keeping track somewhere, right? It's not as though you can do all of this stuff mechanically. The model does not have enough parameters to merely copy the data, right? There has to be some form of reasoning going on. Uh, which is what was expected, but it
5:11was not known how the reasoning was happening. And this seems to give some indication of how the intermediate steps are represented or kind of brought into focus and operated on. The question really is what does this help us do? But there has to be a there there right there has to be something that you can do with it. >> Yeah. Well, I mean, I think the thing