Every J-lens intervention should reliably produce the predictable, intuitive behavior change you'd hypothesize.
Nathan notes that Anthropic's J-lens produces intuitive, predictable behavior changes only 50-70% of the time when probing the model's 'J-space,' meaning a large share of the model's internal cognition remains unaccounted for ('dark cognition').
transcript
Nathan: the sort of rate at which the interventions into the JSpace actually lead to like a sort of predictable intuitive behavior change seem to be somewhere in the 50s to upwards of like 70%. Um, so that's like an incredible accomplishment if framed one way. Like clearly not a random finding, right?
rebuts · 1