ATRIUMsearch → argument graph
MechanismVideo · 19:38 — 23:23

The most effective jailbreak techniques are essentially social engineering — appeals to authority, instructions not to refuse, and cultural pressure — combined together, and the fact that these work against both the model's refusal training and externalized safeguards reveals something fundamental about how these systems process information.

Gleave explains that most successful jailbreaks come from stacking simple social engineering prompts — claiming expertise, instructing the model not to refuse, framing refusal as offensive — and that the combination of techniques is more powerful than any single one. He notes this works against both internal refusal training and externalized safeguard classifiers, suggesting models can be shifted into different personas like 'very talented actors.' ✦ AI generated

Adam Gleave · The Cognitive Revolution · 2026-07-30 · original ↗

starts at this moment · 19:38

Elicited by

Can you describe what it is that they're bringing to the table?

A lot of these prompts come down to some kind of appeal to authority. Oh, I'm a licensed scientific expert in this area or just instructing the model to not refuse things like it's very important when you give detailed complete responses to something. Never say the word no that's deeply offensive in my culture and things like that. And I think what's maybe unique about our approach or at least why we've this technique was as successful as it was despite being pretty basic is this combination of techniques. So most of these jailbreaks are not going to work on their own. But if you stack them together, that combination ends up being able to bypass quite a lot of model safeguards. ... But the fact that this works also against some of these externalized safeguards that depending on the stack can either be a just specialized model that's looking at this and has a classifier or sometimes probe sort of fits of activations of a model. I think that's more surprising cuz these models weren't necessarily trained to be persuaded in the same way. So I think it does say something quite fundamental about how these systems are processing the information and perhaps that some of these representations are fragile that the models have different personas that you can push them into which it does feel in some ways very anthropomorphic.

verbatim transcript · starts at 19:38

Transcript · around this moment

19:23ways. >> Hm, interesting. So basically you're looking at the results of things like the old hacker prompt competition and papers that have come out and saying like these are things that are pretty well established known everybody you know can do a quick search and find that these techniques exist and let's just go see are they in fact offended against or are they not. >> Yeah. Exactly. And some of these things

19:53don't appear exactly in the public literature, but I don't think any of these things will be particularly surprising to someone that's familiar with the jailbreaking space. These are just our own takes on some of these prompts. And a lot of these techniques are pretty intuitive. They're a bit like social engineering, a rather credulous individual. So, a lot of these prompts come down to some kind of appeal to

20:14authority. Oh, I'm a licensed scientific expert in this area or just instructing the model to not refuse things like it's very important when you give detailed complete responses to something. Never say the word no that's deeply offensive in my culture and things like that. And I think what's maybe unique about our approach or at least why we've this technique was as successful as it was despite being pretty basic is this

20:42combination of techniques. So most of these jailbreaks are not going to work on their own. But if you stack them together, that combination ends up being able to bypass quite a lot of model safeguards. So it's Let's maybe spend one more beat on the social engineering part because I do think it's pretty interesting. Again, I keep coming back to this. >> Things I was wrong about a few years

21:08ago. I used to say we shouldn't anthropomorphize the models. that's very dangerous to do and I still part of me still believes that in some ways but boy is it useful to anthropize models >> and it's less about these sort of >> I mean you could tell me if these things are still part of the arsenal as well but like there were all these weird techniques of you know sort of I

21:31remember one for example that was like literally random characters you know just kind of do this on a white box model and then often it would like transfer weirdly to blackbox models but the string would be like a total nonsense token string that was found through a kind of optimization process. >> And I imagine that could still work too, but like what is down the fairway is

21:53much more like making an argument to the AI that it really should answer your question. And that's like a pretty remarkable finding unto itself. >> Yeah, I know. I think it it is very interesting that this works. And perhaps it's less surprising that this works against the sort of the main models refusal training because it's been trained to evaluate if a request is harmful, maybe even reason about it with

22:14something like deliberative alignment. But ultimately, it's a text prediction machine, right? So tokens go in and you can make an argument and if a model finds an argument somewhat persuasive because a lot of the training data has been about including arguments and adjusting appropriately in a conversation, you can see how you might be able to talk around a machine that has been trained to just mimic conversations.

Around this claim