ATRIUMsearch → argument graph
MechanismArticle · 3:00 — 4:00

Critical Thinking: The attack generalizes across model providers with concrete per-model templates, including bypassing an apparent ~50-token verbatim-output threshold via chunked continuations.

Concrete templates vary per model: for Claude, replay the signed block to Haiku 4.5; for GPT, inject the encrypted_content reasoning item and sample up to 50 outputs while bypassing the ~50-token verbatim threshold with chunked continuations; for Gemini, attach thought_signature with a <thought> prefill. ✦ AI generated

AINews host (unattributed editorial) · Latent Space · 2026-08-12 · original ↗

The paper gives concrete templates with some minor variations per model: Claude: replay the signed thinking block to Haiku 4.5, followed by an assistant prefill such as <thinking-copy>. GPT: inject the encrypted_content reasoning item multiple times into a fabricated conversation; sample up to 50 outputs. It also describes bypassing an apparent ~50-token verbatim-output threshold using chunked continuations. Gemini: attach thought_signature to a model turn with a <thought> prefill, then use repeated sampling and reconciliation.

Read full article ↗excerpt · fair-use quotation

Around this claim