Training Claude to give a principled, values-aligned account of its behavior when interrupted mid-task causes it to load constitutional concepts into JSpace preemptively, which improves its real behavior even on tasks where it's never asked to reflect.
Nathan explains Anthropic's counterfactual reflection training: pausing the model mid-task and training it to give a constitutionally-approved answer causes it to preload integrity-related concepts into JSpace, which then improves behavior even when it isn't asked to reflect. ✦ AI generated
Nathan · The Cognitive Revolution · 2026-07-07 · original ↗
starts at this moment · 66:16
“What what did you feel?”
This training it seems to have a similar effect where interrupting it and teaching it to give an approved account... getting ready in that way, loading those concepts into the JSpace translates to actually better behavior on average anyway. The fact that it was trained in some cases to give an account of what it's doing and how that aligns to its principles actually leads to higher adherence to the principles even when it's not asked.
verbatim transcript · starts at 66:16
66:16So it's it's not that the whole constitution is there but they are do this is supervised fine-tuning. So they are giving it like this is what your answer should be. Um but that has the effect of essentially reaching kind of back in time so to speak. U of course it's one model right? So it's it's one model every every token position, but this gets the model to in earlier token
66:38positions kind of get ready to give such a principled account of its behavior and then the you know the kind of I think again sort of happy surprise like I don't you know did it have to be this way? U maybe it had to be this way. It wasn't like obvious in advance. I don't think that it would be this way. getting ready in that way, loading those
67:05concepts into the JSpace translates to actually better behavior on average anyway. So the the fact that it was trained in some cases to give an account of what it's doing and how that aligns to its principles actually leads to higher uh adherence to the principles even when it's not asked, right? Just when it's doing the task. I thought that was really quite quite fascinating and and really interesting and and
67:35encouraging. I mean, I think this is, you know, this is probably the kind of thing where it's like, as much as there's, you know, dark matter and it doesn't always work and blah blah blah blah blah. I mean, first of all, that that stuff hopefully can be refined, you know, it can be cleaned up. There's certainly more insights to come. Um I would expect at some point there will be some
67:57um you know this is all to one token right so that that I would expect at some point that there might be a multi-token version of this or a little bit you know somewhat slightly more abstracted conceptual version as opposed to um just going you know to kind of direct single token outputs. you could imagine that they could use the sort of SAPE and you know cross