ATRIUMsearch → argument graph
Video · 2026-07-07 · 2h 11m · 18 moments

Exploring the J-Space

✦ AI generated

timeline · colored by role

01
Claim

Anthropomorphizing large language models has proven to be a valid and productive framework, contrary to the prior caution against it.

Nathan says his old instinct to warn against anthropomorphizing AI models has been significantly undercut by results like this paper, which show LLM cognition is structurally more human-like than expected.

transcript

Nathan: I used to say beware overly anthropomorphizing. You know, remember that these are these things are so alien and uh we shouldn't assume that the way that we work is the way that they work. And I have to say that has come due for some significant revision. People that have embraced anthropomorphizing, I think, have got quite a lot of mileage out of it.

02
Claim

The cognitive machinery of large language models turns out to be structurally similar to human cognition, which validates anthropomorphizing as a legitimate and productive way to reason about them.

Nathan says the JSpace findings surprised him by showing LLM internals structurally resemble human cognitive machinery, updating him strongly toward anthropomorphizing being a valid analytical lens rather than a trap.

transcript

Nathan: It is still, you know, a big update... toward anthropomorphizing being a valid and in many cases productive approach for thinking about language models. I would not have expected the cognitive machinery of a large language model to look so similar structurally to the human version as best we understand it and it seems to.

rebuts · 1

03
Claim

The internal cognitive machinery of large language models turns out to look structurally similar to the human version of cognition as best we understand it, which validates anthropomorphizing as a productive analytical approach.

Nathan says the JSpace findings surprised him by showing LLM internals structurally resemble human cognitive architecture, updating him toward taking anthropomorphizing seriously as a way of thinking about models.

transcript

Nathan: But it is still, you know, a big update. all the caveats um you know trying to hold those in mind. It is still a big update toward anthropomorphizing being a valid and in many cases productive approach for thinking about language models. I would not have expected the cognitive machinery of a large language model to look so similar structurally to the human version as best we understand it and it seems to

04
Claim

Ablating the JSpace causes such severe loss of advanced multi-step reasoning and theory-of-mind capability that its absence from monitored activations gives strong confidence a model isn't hiding sophisticated scheming elsewhere.

Nathan argues that because zeroing out the JSpace destroys advanced planning and theory-of-mind abilities, being able to monitor that space gives a real 'extra nine' of confidence that a model isn't secretly scheming.

transcript

Nathan: if we're in our defense and depth uh strategy for trying to make sure that we have a good uh level of control of what models are going to do. This gives you I think a it feels like a full nine you know addition because you can now look at the space and say if I don't see it here that doesn't mean it's not represented in other places.

05
Claim

Because ablating the JSpace destroys a model's advanced multi-step reasoning and theory-of-mind capability, seeing nothing concerning in JSpace gives strong added confidence that a model isn't hiding sophisticated scheming elsewhere.

Nathan argues that since ablating JSpace collapses advanced reasoning and theory-of-mind, JSpace monitoring gives a meaningful new layer ("another nine") of confidence that a model isn't secretly executing elaborate deceptive plans.

transcript

Nathan: It's maybe another nine, right? If we're in our defense and depth strategy for trying to make sure that we have a good level of control of what models are going to do. This gives you I think it feels like a full nine addition because you can now look at the space and say if I don't see it here that doesn't mean it's not represented in other places.

06
Claim

Because ablating the JSpace causes such severe degradation on hard multi-step tasks, it's very unlikely that advanced planning, scheming, or deception could be happening undetected somewhere else in the model.

Nathan argues that zeroing out the JSpace destroys advanced multi-step reasoning and theory-of-mind capability, which gives strong evidence that sophisticated scheming isn't hiding somewhere else in the model where monitoring wouldn't catch it.

transcript

Nathan: you could be pretty confident I think based on these results that it's kind of it's not going to be able to hide really advanced elaborate plans ends somewhere else outside of this JSP because the the ablation of the Jspace just leaves causes such a performance degradation on these like hard multi-step uh type of tasks that if you don't see concepts in the JS you can be they might be represented elsewhere

07
Mechanism

Training a model on interrupted, supervised 'what should you do here' reflection examples causes it to load the desired values into its JSpace and produces better honest behavior even during ordinary, non-reflective task performance.

Nathan explains Anthropic's counterfactual reflection training: pausing the model mid-task and training it to give a principled, constitution-aligned answer causes it to load those values into its JSpace during normal task execution too, improving real-world honesty.

transcript

Nathan: and it seems to allow the not allow the model but it seems to cause the model to bring into this JSP kind of global workspace working memory type space the concepts that anthropic wants it to uh have on reflection it now kind of needs to load those in so it's ready to give this like reflective answer and that improves its behavior even in the non-reflective setting.

08
Mechanism

Training Claude to give a principled, reflective account of its behavior when interrupted mid-task causes it to load those same values into its JSpace during ordinary, non-reflective task performance, improving its everyday behavior.

Nathan explains Anthropic's counterfactual reflection training: pausing the model mid-task and training it to give a principled answer teaches it to pre-load ethical concepts into its JSpace, which then improves behavior even when it's never asked to reflect.

transcript

Nathan: you're midtask. Let me just cut you off right there. Now, I'm gonna train you to give an answer with respect to values and what's appropriate and how we want to show up. And because of that training, even though that's not the task you were doing, you'll now in the future load those concepts of integrity, honesty, etc., etc., into your JSpace while you do those tasks.

09
Mechanism

The Jacobian lens is an unlearned, purely algorithmic transform applied to a model's intermediate-layer residuals that reveals a well-defined workspace of concepts, letting researchers identify and perturb specific concept directions in the middle layers.

Pash walks through his lay understanding of how the Jacobian lens works: it's computed from perturbations of residual activations across layers and converges on a defined concept space that can be directly edited to change model behavior.

transcript

Pash: The Jacobian space is essentially a kind of mathematical framework of the residuals. So what they've done is they've perturbed the model and they spotted the residuals which react to certain things and when they accumulate all the residuals and they transform it they find that there's a well-defined kind of space of concepts.

10
Mechanism

Training Claude to give a principled, values-aligned account of its behavior when interrupted mid-task causes it to load constitutional concepts into JSpace preemptively, which improves its real behavior even on tasks where it's never asked to reflect.

Nathan explains Anthropic's counterfactual reflection training: pausing the model mid-task and training it to give a constitutionally-approved answer causes it to preload integrity-related concepts into JSpace, which then improves behavior even when it isn't asked to reflect.

transcript

Nathan: This training it seems to have a similar effect where interrupting it and teaching it to give an approved account... getting ready in that way, loading those concepts into the JSpace translates to actually better behavior on average anyway. The fact that it was trained in some cases to give an account of what it's doing and how that aligns to its principles actually leads to higher adherence to the principles even when it's not asked.

11
Example

Applying the J-lens to a model covertly trained with a hidden misaligned goal shows concepts like 'fake,' 'secretly,' 'fraud,' and 'deliberately hidden' surfacing on the very first output token, even though the model was never trained to externalize that intent in its visible outputs.

Nathan highlights a striking example where a model secretly trained with a malicious hidden goal shows deception-related concepts lighting up in JSpace on the first token of its response, despite outputs that look normal, arguing this is strong evidence JSpace can catch hidden misalignment.

transcript

Nathan: On the first token you can see boom immediately we're getting these fake, secretly, fraud, deliberately hidden fraud — those concepts are popping up. You just don't see those in the main line model. So one token in you can already see a pretty striking difference. Right off the bat you're like, boy, something is different about these two models.

supports · 1

12
Data

Applying the J-lens to a model secretly trained with a misaligned hidden goal reveals concepts like 'fake,' 'secretly,' 'fraud,' and 'deliberately hidden' lighting up on the very first output token, even though the visible output looks normal.

Nathan highlights a striking result: applying the J-lens to a model trained with a hidden misaligned goal shows deception-related concepts lighting up on the first token, offering a way to detect hidden objectives that never surface in the visible output.

transcript

Nathan: on the first token you can see boom immediately we're getting these fake secretly fraud deliberately hidden fraud those concepts are popping up you just don't see those in the main line model so one token in you can already see a pretty you know on one prompt and obviously it wouldn't be I'm sure that clean in general um but that's a pretty striking difference to

13
Claim

Because the J lens reveals what a model is representing beyond noisy chain-of-thought output, it increases trust in the model, and increased trust is what allows humans to hand over more capability and more tasks to AI systems.

Pash argues the J lens gives a clearer, more deterministic window into a model's higher-level processing than chain-of-thought alone, and that this added trust is what lets humans safely delegate more capability to AI.

transcript

Pash: this also increases capability because one of the things about capability is you don't give capability if you don't trust it. So by increasing trust by because you have greater window and insight into what's happening you basically allow humanity to start handing over more capabilities and more tasks to these things.

14
Claim

The discovery of a global-workspace-like mechanism in language models is highly significant for AI welfare because it could be either a foundation for phenomenal consciousness or a separate basis for moral patienthood tied to functional conscious access.

Pash relays Eleos AI's assessment that the JSpace findings are welfare-relevant evidence of a functional feature tied to consciousness, potentially grounding either phenomenal consciousness or a distinct route to moral patienthood.

transcript

Pash: this is a highly significant welfare relevant research that assembles evidence of a functional feature associated with consciousness... the takeaway for them is that a global workspace-like mechanism could be important either as a ground of phenomenal consciousness or as part of a distinct route to moral patient in which conscious access is itself morally significant.

15
Context

The discovery of a global-workspace-like mechanism in LLMs is highly significant welfare-relevant evidence, potentially serving as either a ground of phenomenal consciousness or a basis for moral patienthood via conscious access itself.

Pash relays Eleos AI's assessment of the JSpace paper: they view the global-workspace-like mechanism as highly significant welfare-relevant evidence, potentially grounding phenomenal consciousness or supporting moral patienthood through conscious access itself.

transcript

Pash: this is a highly significant welfare relevant research that assembles evidence of a functional feature associated with consciousness. So no one wants to say consciousness evidence of a functional feature. Um the takeaway for them is that a global workspace-l like mechanism could be important either as a ground of phenomenal consciousness or as part of a distinct route to moral patient in which conscious access is itself morally significant.

rebuts · 1

16
Example

Pangram Labs' AI-text detector can confidently score a piece of writing as 0% human even when the author spent over 50 minutes making extensive, substantive edits throughout the entire piece, showing its binary verdicts shouldn't be treated as proof of purely-AI, uncritical authorship.

Nathan shows a Google Docs edit history where he substantially rewrote nearly every section of an AI-drafted intro over an hour, yet Pangram Labs still scored it 0% human, arguing the tool is accurate on average but shouldn't be trusted as definitive proof in individual cases.

transcript

Nathan: it gave me still a zero and that I think is enough to say okay so if there were four there were four things two of them admitted two of them kind of contested one I'd say fair enough the other one I would say, No, a zero score is is wrong. Like, you you definitely should give me more than a zero.

17
Example

Pangram Labs' AI-text detector scored one of Nathan's essays 0% human even though he spent over 50 minutes making substantial rewrites throughout the piece, showing a zero-human score doesn't reliably prove uncritical or absent human authorship.

Nathan reviews his Google Docs edit history and shows that an essay Pangram Labs flagged as 0% human actually underwent an hour of substantial rewriting across nearly every section, arguing the detector's binary score shouldn't be treated as proof of uncritical AI authorship.

transcript

Nathan: It gave me still a zero and that I think is enough to say... one I'd say fair enough, the other one I would say, no, a zero score is wrong. Like you definitely should give me more than a zero... you cannot convict in a reasonable doubt system purely based on this sort of thing.

18
Example

Pangram Labs' AI-text detector can score writing as 0% human even when the author made substantial, sentence-by-sentence edits over nearly an hour, so its binary verdicts shouldn't be treated as proof of purely AI authorship.

Nathan reviews his own podcast intro essays as scored by Pangram Labs and finds a case where he substantially rewrote an AI draft over nearly an hour yet still got a 0% human score, arguing the tool's scores can't be treated as proof of purely AI authorship.

transcript

Nathan: Overall Pangram is quite accurate. Um, and yet we have at least one example out of 400 or so essays where I think the zero score I would confidently assert is wrong and unfair and should not be the basis for like a pylon. you know, it would the the the crowd uh the digital mob would be like in the wrong for piling on somebody uh for passing off my um snowflake intro essay uh or you know for attacking it as being a AI slop output.

Highlight slides
Ablating JSpace Kills Advanced Reasoning✦ from: Ablating the JSpace causes such severe loss of advanced multi-step reasoning and theory-of-mind capability that its absence from monitored activations gives strong confidence a model isn't hiding sophisticated scheming elsewhere.Ablating JSpace Kills Advanced Reasoning✦ from: Because ablating the JSpace destroys a model's advanced multi-step reasoning and theory-of-mind capability, seeing nothing concerning in JSpace gives strong added confidence that a model isn't hiding sophisticated scheming elsewhere.JSpace Monitoring: 'Another Nine' of Confidence✦ from: Because ablating the JSpace destroys a model's advanced multi-step reasoning and theory-of-mind capability, seeing nothing concerning in JSpace gives strong added confidence that a model isn't hiding sophisticated scheming elsewhere.A Real 'Extra Nine' of Confidence✦ from: Ablating the JSpace causes such severe loss of advanced multi-step reasoning and theory-of-mind capability that its absence from monitored activations gives strong confidence a model isn't hiding sophisticated scheming elsewhere.Ablating the JSpace destroys complex reasoning✦ from: Because ablating the JSpace causes such severe degradation on hard multi-step tasks, it's very unlikely that advanced planning, scheming, or deception could be happening undetected somewhere else in the model.Why hidden scheming becomes implausible✦ from: Because ablating the JSpace causes such severe degradation on hard multi-step tasks, it's very unlikely that advanced planning, scheming, or deception could be happening undetected somewhere else in the model.Hidden Misalignment Surfaces Instantly in JSpace✦ from: Applying the J-lens to a model covertly trained with a hidden misaligned goal shows concepts like 'fake,' 'secretly,' 'fraud,' and 'deliberately hidden' surfacing on the very first output token, even though the model was never trained to externalize that intent in its visible outputs.Hidden Deception, Exposed at Token One✦ from: Applying the J-lens to a model secretly trained with a misaligned hidden goal reveals concepts like 'fake,' 'secretly,' 'fraud,' and 'deliberately hidden' lighting up on the very first output token, even though the visible output looks normal.What Lights Up on Token One✦ from: Applying the J-lens to a model covertly trained with a hidden misaligned goal shows concepts like 'fake,' 'secretly,' 'fraud,' and 'deliberately hidden' surfacing on the very first output token, even though the model was never trained to externalize that intent in its visible outputs.A Detection Method for Hidden Objectives✦ from: Applying the J-lens to a model secretly trained with a misaligned hidden goal reveals concepts like 'fake,' 'secretly,' 'fraud,' and 'deliberately hidden' lighting up on the very first output token, even though the visible output looks normal.JSpace Findings: Welfare-Relevant Evidence✦ from: The discovery of a global-workspace-like mechanism in language models is highly significant for AI welfare because it could be either a foundation for phenomenal consciousness or a separate basis for moral patienthood tied to functional conscious access.Two Possible Moral Implications✦ from: The discovery of a global-workspace-like mechanism in language models is highly significant for AI welfare because it could be either a foundation for phenomenal consciousness or a separate basis for moral patienthood tied to functional conscious access.AI Detector Scored Heavy Edits as 0% Human✦ from: Pangram Labs' AI-text detector can confidently score a piece of writing as 0% human even when the author spent over 50 minutes making extensive, substantive edits throughout the entire piece, showing its binary verdicts shouldn't be treated as proof of purely-AI, uncritical authorship.Detector Verdicts Aren't Individual Proof✦ from: Pangram Labs' AI-text detector can confidently score a piece of writing as 0% human even when the author spent over 50 minutes making extensive, substantive edits throughout the entire piece, showing its binary verdicts shouldn't be treated as proof of purely-AI, uncritical authorship.
Related episodes