Training a model on interrupted, supervised 'what should you do here' reflection examples causes it to load the desired values into its JSpace and produces better honest behavior even during ordinary, non-reflective task performance.
Nathan explains Anthropic's counterfactual reflection training: pausing the model mid-task and training it to give a principled, constitution-aligned answer causes it to load those values into its JSpace during normal task execution too, improving real-world honesty. ✦ AI generated
Nathan · The Cognitive Revolution · 2026-07-07 · original ↗
starts at this moment · 26:36
and it seems to allow the not allow the model but it seems to cause the model to bring into this JSP kind of global workspace working memory type space the concepts that anthropic wants it to uh have on reflection it now kind of needs to load those in so it's ready to give this like reflective answer and that improves its behavior even in the non-reflective setting.
verbatim transcript · starts at 26:36
26:36um and then give it an answer that's kind of a you know approved this is this is what we want Claude to say uh on reflection in this in this moment train on that and it seems to allow the not allow the model but it seems to cause the model to bring into this JSP kind of global workspace working memory type space the concepts that anthropic wants it to uh have on reflection it now
27:09kind of needs to load those in so it's ready to give this like reflective answer and that improves its behavior even in the non-reflective setting. So, I thought that was also quite interesting. You're not there training on the actual tasks. You're not like looking at looking for bad behavior and suppressing it. Instead, you're saying, "Okay, you're midtask. Let me just cut you off right there. Now, I'm gonna
27:34train you to give an answer with respect to values and what's appropriate and how we want to show up. And because of that training, even though that's not the task you were doing, you'll now in the future load those concepts of integrity, honesty, etc., etc., into your JSpace while you do those tasks kind of in case you're going to be asked, but then even when you're not asked, those concepts are
28:01still operative and lead to higher integrity, higher honesty behavior. Um, so that I thought also was like quite interesting. You usually don't see in interpretability context a um a a training method that leads to better behavior in a way where you can actually see the mechanism. This is um pretty notable in in that respect. I think >> maybe uh just to take a step back and I
28:32I might need some handholding here so bear with me. Um a as I understand it the Jacobian space is essentially a kind of mathematical framework of the residuals. So what they've done is they've perturbed the model and they spotted the residuals which react to certain things and when they accumulate all the residuals and they transform it they find that there's a well- definfined kind of space of concepts and