Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models, and a weaker sibling model can be jailbroken to recover the stronger model's hidden reasoning in plaintext.
The paper's central attack: frontier-model reasoning traces are returned to clients as encrypted blocks, then replayed into a weaker sibling model that—once jailbroken—reveals the stronger model's hidden chain of thought in plaintext.
transcript
Paper authors (via Hacker News, stolen-thoughts.com): Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model's hidden reasoning in plaintext
explains mechanism · 1