ATRIUMsearch → argument graph
ClaimArticle

Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models, and a weaker sibling model can be jailbroken to recover the stronger model's hidden reasoning in plaintext.

The paper's central attack: frontier-model reasoning traces are returned to clients as encrypted blocks, then replayed into a weaker sibling model that—once jailbroken—reveals the stronger model's hidden chain of thought in plaintext. ✦ AI generated

Paper authors (via Hacker News, stolen-thoughts.com) · Simon Willison's Weblog · 2026-08-11 · original ↗

Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model's hidden reasoning in plaintext

Read full article ↗excerpt · fair-use quotation

Around this claim
This moment responds to