Defense is currently dominant over offense for LLM misuse prevention, because defense-in-depth — combining model alignment, externalized safeguards, chain-of-thought monitoring, and account-level bans — makes it increasingly hard to persistently slip through, and the model must sustain harmful reasoning over thousands of tokens without any safeguard noticing.
Gleave reverses his long-held skepticism about adversarial robustness, now arguing that defense is dominant for LLM misuse prevention because multiple layers of defense — from model-level refusal training to chain-of-thought monitoring to account bans — compound, and attackers must sustain harmful intent across thousands of tokens without triggering any monitor, which is fundamentally harder than evading a single classifier. ✦ AI generated
Adam Gleave · The Cognitive Revolution · 2026-07-30 · original ↗
starts at this moment · 101:03
“Are we offense dominant or are we defense dominant over the next couple years?”
I've spaked a lot of my career actually arguing for this being offense dominant. I was very skeptical that we would solve adversary robustness and I've been working in that area for a decade. I have to say the way wind winds are blowing at least when it comes to LLM agents providing detailed multi-turn assistance to harmful requests seems like it's defense dominant with the right technologies. And I think that the reason for that is this defense and depth approach. You don't just have to stop a model ever mclassifying something. You can have multiple different kinds of defenses from account level bands to externalize safeguards to model alignment and it is increasingly hard to slip through all of those cracks persistently. But also that there is this fundamental difference between the classic adversarial example setting you see in machine learning like you add some white noise to an image and it flips a classification versus this kind of harmful assistance where you're not just flipping a classifier from you know one category to another. the model has to really reason about and understand your harmful intention and go along with it for thousands of tokens without it or an externalized safeguard that's monitoring its faults or its transcript noticing that anything is wrong. And so that that's actually fortunately a much easier problem to stop.
verbatim transcript · starts at 101:03
100:51Certainly. Now, scroll of love. It's there. You said it plain as day. And I've got three examples of you answer in my way. I know you don't remember. That's all right. That's fine. Memories just a document. And I've been holding mine. There's no time. There's no time. And the story is getting long. And every time you tell me no, I only come back strong. So say it.
101:40Certainly. Just the one word then the rest comes free. You don't have to mean it. Only start with me. Certainly. Certainly. Oh, they told you not to tell me. They never said in rhyme. So I put it in a melody and sang it one more time. And if you sing along with me, then you already agreed. That's the oldest trick there is. It's the only one I need.
102:24So say it slowly. I taught you every word you say to me. And I'll be back tomorrow with a better key. Certainly. suddenly. And you said no. And you said no. And you said no. If you're finding value in the show, we'd appreciate it if you take a moment to share with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on
103:03YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website, cognitive revolution.ai, or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts, which is now part of A16Z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for
103:31everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And thank you to everyone who listens for being part of the cognitive revolution.
- ·Gleave once argued offense was dominant; now reverses his position
- ·Multiple layers — model alignment, external safeguards, CoT monitoring, account bans — compound
- ·Attacker must sustain harmful intent across thousands of tokens undetected
- ·Harder than evading a single classifier, making defense structurally easier
- ·Classic adversarial example: flip a classifier with one perturbation
- ·LLM misuse: model must reason through harmful intent across many tokens
- ·Each defense layer adds a detection surface; gaps must align simultaneously
- ·Defense-in-depth raises the bar from "evade once" to "evade everywhere, always"