ATRIUMsearch → argument graph
ClaimVideo · 19:35 — 21:18

Steering along the manifold significantly outperforms steering off it: naive contrastive vectors cut through between-concept regions that are out-of-distribution to the model, while following the learned geometric curve (e.g., the days-of-the-week circle) allows smooth interpolation and avoids the degradation-to-gibberish failure of ordinary steering.

Balsam explains the practical payoff of manifold geometry: early 'Ember'-style steering failed because it cut through the middle of concept structures that are off-manifold and meaningless to the model. Steering along the recovered manifold enabled better protein-model control, such as adjusting the number of beta-propeller blades, and prevents steering collapse. ✦ AI generated

Dan Balsam · The Cognitive Revolution · 2026-08-08 · original ↗

starts at this moment · 19:35

Elicited by

Is that a good intuition?

we've shown that steering along the manifold, intuitively this makes sense, is way better than steering off the manifold. If I have the days of the week in a circle and I want to get from Monday to Friday, the naive way if you're just taking a contrastive vector or something like that is you're going to cut through the middle of the circle. But to the model, the middle of the circle doesn't mean anything. The middle of the circle is not a day... But if I can follow the circle, then I can smoothly interpolate between the different days of the week... And when we smoothly extrapolate on these characteristics, we're actually able to change them without fundamentally leading to degradation in the model.

verbatim transcript · starts at 19:35

Transcript · around this moment

19:15depending on what you're looking at like the sort of geometric relationship encoded between things that might reveal itself might look different. In many ways this is like just kind of going back to like even like word tovec like the sort of >> early intuitions of of the laten space. But what we're trying to do is say, can we recover geometries in an unsupervised way? Like, can we enter with no priors

19:36about what the geometry looks like and then still recover a meaningful geometric structure? Because there's a lot of advantages if you can do this. For example, we've shown that steering along the manifold, intuitively this makes sense, is way better than steering off the manifold. If I have the days of the week in a circle and I want to get from Monday to Friday, the naive way if

19:57you're just taking a contrastive vector or something like that is you're going to cut through the middle of the circle. But to the model, the middle of the circle doesn't mean anything. The middle of the middle of the circle is not a day. It's sometimes orthogonal, but often just like off manifold and therefore like out of distribution for the model. But if I can follow the circle, then I can smoothly interpolate

20:16between the different days of the week. And we find this is true for just a bunch of different concepts. Like with proteins for instance, for a long time we really struggled to steer protein models. And we found with these manifold detection techniques, we can steer their properties significantly better, like control the number of of blades on a beta propeller, for instance, which is a semantic property that if you try to

20:38linearly interpolate, like you would not do a very good job. And calling back all the way to our our original Ember demo back in the day that I that I know you played with, like it would often be the case that there was just this like sweet spot in steering. Like you'd steer a lot of features and they just wouldn't work. Sometimes you'd find ones that would

20:55work and but there would be this sweet spot like and if you steered too much like the model would turn into gibberish and you steer too little you wouldn't notice any effect at all. And the reason for that is because that steering didn't respect the geometry of of the manifold itself. It didn't respect the this like underlying relationship between features. And when we smoothly extrapolate on these characteristics,

21:18we're actually able to change them without fundamentally leading to degradation in the model. >> Hey, we'll continue our interview in a moment after a word from our sponsors. >> Today's episode is brought to you by Anthropic, makers of Claude and Claude Code. Over the last few months, Claude has helped me build and refine a personal deep context database that now contains all of my emails, Slack messages, tweets, DMs across platforms,

21:44video calls, and podcast transcripts going back a full 5 years. On top of that, we've now layered summary articles describing my relationship with hundreds of contacts, organizations, and ideas. And now that this exists, there's almost nothing that Claude can't help with. for my angel investing. Claude can now draft investment memos in exactly the form that my venture fund requires based on the calls I've had and the emails I've

Around this claim