ATRIUMsearch → argument graph
ClaimVideo · 22:11 — 23:41

Because ablating the JSpace causes such severe degradation on hard multi-step tasks, it's very unlikely that advanced planning, scheming, or deception could be happening undetected somewhere else in the model.

Nathan argues that zeroing out the JSpace destroys advanced multi-step reasoning and theory-of-mind capability, which gives strong evidence that sophisticated scheming isn't hiding somewhere else in the model where monitoring wouldn't catch it. ✦ AI generated

Nathan · The Cognitive Revolution · 2026-07-07 · original ↗

starts at this moment · 22:11

you could be pretty confident I think based on these results that it's kind of it's not going to be able to hide really advanced elaborate plans ends somewhere else outside of this JSP because the the ablation of the Jspace just leaves causes such a performance degradation on these like hard multi-step uh type of tasks that if you don't see concepts in the JS you can be they might be represented elsewhere

verbatim transcript · starts at 22:11

Transcript · around this moment

22:11not represented in other places you know there and there are interesting details about what the model still can do when the JSpace is removed but you could be pretty confident I think based on these results that it's kind of it's not going to be able to hide really advanced elaborate plans ends somewhere else outside of this JSP because the the ablation of the Jspace just leaves causes such a performance degradation on

22:40these like hard multi-step uh type of tasks that if you don't see concepts in the JS you can be they might be represented elsewhere but they're they're seemingly at this point very unlikely to be represented in a way that allows for very advanced planning, reasoning, scheming, deception, etc., etc. Um, theory of mind seems to be kind of located in the JSpace as well. The model's ability to

23:12>> kind of report on its own >> experience, quote unquote, whatever that may mean, and also its its ability to describe what other people are thinking seems to be very much lost when you ablate the JSpace. So again, it's like you have to have theory if you're going to be deceptive, right? If you're going to try to do some sort of um takeover attempt or you know, whatever,

23:38right? The the the extreme things that people are worried about, this is not going to be these these things are not going to be easy for the model to do. They're going to require a lot of planning and they're going to require sophisticated theory of mind. So if you have the ability to kind of localize, okay, this is where the model really does its sophisticated theory of mind

23:58type thinking and it's planning with respect to that kind of theory of mind and you can see pretty clearly where that is. I think that gives you a lot of comfort, you know, that that you're you're looking in the right space and that if you're not seeing something in the JS, it's you know, it's unlikely to be again, it could be in there somewhere. It could be coloring results. It could

Around this claim