ATRIUMsearch → argument graph
MechanismArticle

Every model under the same family shares the same encryption key, so blocks captured from one model can be fed back into the weakest family members, which can be jailbroken into outputting the unencrypted raw reasoning blocks.

The vulnerability mechanism: a shared per-family encryption key lets an attacker replay reasoning blocks from strong models into weak family members and jailbreak the latter into decrypting the raw chain of thought. ✦ AI generated

Paper authors (via Hacker News, stolen-thoughts.com) · Simon Willison's Weblog · 2026-08-11 · original ↗

The paper's authors found that every model under the same family used the same encryption key, which meant you could feed those blocks back into the weakest model family members and jailbreak them into outputting the unencrypted raw reasoning blocks!

Read full article ↗excerpt · fair-use quotation

Around this claim