ATRIUMsearch → argument graph
MechanismVideo · 11:31 — 13:01

Training a model to give a values-aligned reflective answer only when paused mid-task doesn't change how it behaves during ordinary, non-reflective task performance.

Nathan explains Anthropic's 'counterfactual reflection' technique: pausing a model mid-task, training it to give the constitutionally correct reflective answer, and finding that this causes the model to load those values into its J-space and behave better even when it isn't being asked to reflect. ✦ AI generated

Nathan · The Cognitive Revolution · 2026-07-09 · original ↗

starts at this moment · 11:31

Basically what they do there is pause the model mid task and then do supervised training once interrupted asking it like what should we be doing here like what's the you know what's the constitutionally right thing to be doing in this moment and then give it an answer that's kind of a approved this is what we want claude to say on reflection in this moment train on that

verbatim transcript · starts at 11:31

Transcript · around this moment

11:31the model mid task and then do supervised training once interrupted asking it like what should we be doing here like what's the you know what's the constitutionally right thing to be doing in this moment and then give it an answer that's kind of a approved this is what we want claude to say on reflection in this moment train on that and it seems to allow the not allow the model but it seems to

12:02cause the model to bring into this JSP kind of global workspace working memory type space the concepts that anthropic wants it to have on reflection. it now kind of needs to load those in so it's ready to give this like reflective answer and that improves its behavior even in the non-reflective setting. You're not there training on the actual tasks. You're not like looking at looking for bad behavior and suppressing

12:34it. Instead, you're saying, "Okay, you're mid task. Let me just cut you off right there. Now, I'm going to train you to give an answer with respect to values and what's appropriate and how we want to show up. And because of that training, even though that's not the task you were doing, you'll now in the future load those concepts of integrity, honesty, etc., etc., into your JSpace while you

13:00do those tasks in case you're going to be asked. But then even when you're not asked, those concepts are still operative and lead to higher integrity, higher honesty behavior. So that I thought also was like quite interesting. You usually don't see in interpretability context a training method that leads to better behavior in a way where you can actually see the mechanism. This is um pretty notable in

13:27in that respect, I think. >> Yeah, it's a tough day for the stochcastic parrot crowd, I'd say. And because you can see this and you know these models are billions of parameters at this point and they did it also on 4.5 sonnet. 4.5 sonnet is actually a pretty recent model. It's not it's just like 7 months 8 months since 4.5 sonnet and sonnet is a very capable model. So

Around this claim