Because zeroing out a model's JSpace destroys its advanced multi-step reasoning capability, it's unlikely that a model could hide sophisticated scheming or deception anywhere else in the network.
Nathan argues that since ablating the JSpace causally wipes out a model's capacity for advanced multi-step reasoning, there's little remaining room for a model to hide elaborate scheming or deceptive plans elsewhere — a meaningful positive update for AI safety monitoring. ✦ AI generated
Nathan · The Cognitive Revolution · 2026-07-09 · original ↗
starts at this moment · 10:28
And critically, if we zero out that space, the model just loses a lot of capabilities. It just can't do nearly as advanced multi-step reasoning. But you could be pretty confident, I think, based on these results that it's not going to be able to hide really advanced elaborate plans somewhere else outside of this JSpace because the ablation of the JSP just leaves causes such a performance degradation on these like hard multi-step type of tasks
verbatim transcript · starts at 10:28
10:28very informative. And critically, if we zero out that space, the model just loses a lot of capabilities. It just can't do nearly as advanced multi-step reasoning. But you could be pretty confident, I think, based on these results that it's not going to be able to hide really advanced elaborate plans somewhere else outside of this JSpace because the ablation of the JSP just leaves causes such a performance
10:55degradation on these like hard multi-step type of tasks that if you don't see concepts in the JSP, you can be they might be represented elsewhere, but they're seemingly at this point very unlikely to be represented in a way that allows for very advanced planning, reasoning, scheming, deception, etc., etc. One result deserves its own marker, a training method the paper calls counterfactual reflection. Basically what they do there is pause
11:31the model mid task and then do supervised training once interrupted asking it like what should we be doing here like what's the you know what's the constitutionally right thing to be doing in this moment and then give it an answer that's kind of a approved this is what we want claude to say on reflection in this moment train on that and it seems to allow the not allow the model but it seems to
12:02cause the model to bring into this JSP kind of global workspace working memory type space the concepts that anthropic wants it to have on reflection. it now kind of needs to load those in so it's ready to give this like reflective answer and that improves its behavior even in the non-reflective setting. You're not there training on the actual tasks. You're not like looking at looking for bad behavior and suppressing
- ·Ablating the JSpace causally wipes out a capability
- ·Model loses capacity for advanced multi-step reasoning
- ·Causes major performance degradation on hard tasks
- ·Little room left to hide elaborate scheming elsewhere
- ·Deceptive plans would need advanced multi-step reasoning too
- ·Meaningful positive update for safety monitoring efforts