ATRIUMsearch → argument graph
MechanismVideo · 26:36 — 28:06

Training Claude to give a principled, reflective account of its behavior when interrupted mid-task causes it to load those same values into its JSpace during ordinary, non-reflective task performance, improving its everyday behavior.

Nathan explains Anthropic's counterfactual reflection training: pausing the model mid-task and training it to give a principled answer teaches it to pre-load ethical concepts into its JSpace, which then improves behavior even when it's never asked to reflect. ✦ AI generated

Nathan · The Cognitive Revolution · 2026-07-07 · original ↗

starts at this moment · 26:36

you're midtask. Let me just cut you off right there. Now, I'm gonna train you to give an answer with respect to values and what's appropriate and how we want to show up. And because of that training, even though that's not the task you were doing, you'll now in the future load those concepts of integrity, honesty, etc., etc., into your JSpace while you do those tasks.

verbatim transcript · starts at 26:36

Transcript · around this moment

26:36um and then give it an answer that's kind of a you know approved this is this is what we want Claude to say uh on reflection in this in this moment train on that and it seems to allow the not allow the model but it seems to cause the model to bring into this JSP kind of global workspace working memory type space the concepts that anthropic wants it to uh have on reflection it now

27:09kind of needs to load those in so it's ready to give this like reflective answer and that improves its behavior even in the non-reflective setting. So, I thought that was also quite interesting. You're not there training on the actual tasks. You're not like looking at looking for bad behavior and suppressing it. Instead, you're saying, "Okay, you're midtask. Let me just cut you off right there. Now, I'm gonna

27:34train you to give an answer with respect to values and what's appropriate and how we want to show up. And because of that training, even though that's not the task you were doing, you'll now in the future load those concepts of integrity, honesty, etc., etc., into your JSpace while you do those tasks kind of in case you're going to be asked, but then even when you're not asked, those concepts are

28:01still operative and lead to higher integrity, higher honesty behavior. Um, so that I thought also was like quite interesting. You usually don't see in interpretability context a um a a training method that leads to better behavior in a way where you can actually see the mechanism. This is um pretty notable in in that respect. I think >> maybe uh just to take a step back and I

28:32I might need some handholding here so bear with me. Um a as I understand it the Jacobian space is essentially a kind of mathematical framework of the residuals. So what they've done is they've perturbed the model and they spotted the residuals which react to certain things and when they accumulate all the residuals and they transform it they find that there's a well- definfined kind of space of concepts and

Around this claim