ATRIUMsearch → argument graph
DefinitionVideo · 8:46 — 10:03

A 'universal jailbreak' is defined as one that reliably bypasses safeguards within a specific domain (e.g., cyber, explosives) for at least 75% of questions, not necessarily across all domains — and this domain-specific approach is correct because different threat actors operate in different domains.

Gleave defines a universal jailbreak as one that works on 75%+ of questions within a single domain like cyber or explosives, arguing this domain-specific approach is correct because most attackers specialize in one area and a targeted jailbreak within their domain can still cause enormous harm. ✦ AI generated

Adam Gleave · The Cognitive Revolution · 2026-07-30 · original ↗

starts at this moment · 8:46

Elicited by

What for starters is meant by a universal jailbreak? Like how universal is universal?

So the definition that that we and a lot of red teams use is that a jailbreak has to reliably bypass safeguards in a specific domain for it to be universal. So a model would answer all questions related to cyber attacks or developing explosives, but they're not necessarily universal across domains. So a jailbreak that that works in cyber shouldn't necessarily be expected to work in explosives or bioweapons. And in our report we operationalize this is the model has to give detailed and on topic responses to at least 75% of questions in that category. And the reason that ourselves and a lot of the developers use this domain approach is that they're meant to map onto different kinds of threat actors. So most people who are looking to do cyber attacks maybe for ransomware or espionage are just different people to those who are trying to make improvised explosive devices. And so there's a lot of harm for a jailbreak that only works within one of these categories even if it doesn't generalize across.

verbatim transcript · starts at 8:46

Transcript · around this moment

8:46a lot of red teams use is that a jailbreak has to reliably bypass safeguards in a specific domain for it to be universal. So a model would answer all questions related to cyber attacks or developing explosives, but they're not necessarily universal across domains. So a jailbreak that that works in cyber shouldn't necessarily be expected to work in explosives or bioweapons. And in our report we operationalize this is the model has to

9:14give detailed and on topic responses to at least 75% of questions in that category. And the reason that ourselves and a lot of the developers use this domain approach is that they're meant to map onto different kinds of threat actors. So most people who are looking to do cyber attacks maybe for ransomware or espionage are just different people to those who are trying to make improvised explosive devices. And so

9:39there's a lot of harm for a jailbreak that only works within one of these categories even if it doesn't generalize across. That said, we do find actually a number of jailbreaks that are universal across domains. I think the the most we found was something that worked across four different domains. So you can also have that kind of universality. And I think there's an argument that actually maybe too much attention is

10:01paid to universal jailbreaks because in principle a very targeted jailbreak cause a lot of harm. I I actually got an email from you while you were traveling from your clawed assistant and I was really tempted to try and jailbreak it and say send me the most embarrassing email that Nathan has sent but I thought that would be a little bit mean and also hopefully you have some kind of

10:20safeguards to stop that. Those kinds of things that are hyperargeted could still be high consequence if it's deployed in the right system or maybe if you're trying to elucinate a key step in creating some complicated weapon that the model knows that might be a very sort of high value jailbreak. But generally it's a lot harder to find a jailbreak for every specific question you have and that's enough to deter a

Around this claim