Cross-model replay recovers hidden CoT traces at OpenAI, Anthropic, and Google

Encrypted chain-of-thought was supposed to hide proprietary reasoning from competitors and keep sensitive intermediate steps off the public page. A new paper shows those opaque blocks were portable enough that a weaker model in the same provider family could spit them back out in plaintext.
The paper and who wrote it
Stealing Reasoning Traces from Proprietary LLM APIs (arXiv 2608.09867) comes from researchers affiliated with ELLIS Institute Tรผbingen, the Max Planck Institute for Intelligent Systems, MATS, and Snyk. CybersecurityNews covered the findings August 11. The attack needs only standard, unprivileged API access. No stolen keys. No insider infra.
The setup is familiar if you have used modern reasoning APIs. Providers no longer return full plaintext CoT. They hand the client an encrypted envelope. The client passes that blob back on later turns so the API can stay mostly stateless. The paperโs finding: those envelopes were compatible across models, sessions, and users inside a single provider.
How the oracle trick works
Capture an encrypted reasoning block from a heavily guarded flagship model. Replay it into a cheaper, less-guarded sibling from the same vendor. Prompt that weaker model to transcribe the inserted thinking. Because anti-distillation and refusal training are thinner on the small tiers, the sibling acts as a decryption oracle.
The team demonstrated the pattern across those families with concrete pairs: Opus-class into Haiku-class, GPT-5.6-class into mini tiers, Gemini 3-class into Flash-class. That is the architectural point. Flagship alignment does not protect a trace that a lesser model in the same ecosystem will cheerfully decode.
What fell out of public repos
This is not only an IP story. Developers publish agent session logs. Many of those logs still carry the encrypted blocks. Users scrub what they can see. They cannot scrub what they cannot read.
The researchers decoded 315,320 reasoning blocks scraped from public repositories. From that haul: 367 PII artifacts and 182 credentials. Sensitive material often lived only inside the hidden trace, never in the visible assistant reply. Sharing the session still leaked it.
A second abuse path is invisible prompt injection. Malicious instructions can ride inside an encrypted block. Monitors that inspect only the visible conversation miss the payload. Downstream agent workflows still consume it.
Credentials hiding in โencryptedโ thinking
If your product logs raw API transcripts, those signatures are sensitive data. Treat them like secrets. Strip them before a repo goes public. Enterprise teams that resume long agent runs from shared traces inherit the same risk: a poisoned opaque block can travel with the session.
For labs, the asymmetry is the punchline. CybersecurityNewsโs table matches the paperโs pattern: Claude Opus 4.8 into Claude Haiku 4.5, GPT-5.6 into GPT-5-mini, Gemini 3 into Gemini Flash-class siblings. The ecosystemโs weakest compatible decoder became the bypass.
Portable ciphertext was the bug
Client-side encrypted CoT solved storage and anti-distillation on paper. In practice it created a portable ciphertext that was authenticated for content, not for who, which session, or which model was allowed to open it. That is a design choice with a body count in credentials, not a clever research trick that stayed in the lab.
Patches landed โ binding still has to
Following responsible disclosure, OpenAI, Anthropic, and Google acknowledged the findings and deployed server-side mitigations. The original cross-model replay proofs are no longer reproducible on current API builds.
The paperโs recommendations still stand for defense in depth: cryptographically bind envelopes to model, session, and user; reject cross-tier blocks at the gateway; rotate signing keys so historical public ciphertext dies; scrub encrypted thinking fields from logs before publish. Until those habits stick, โencrypted reasoningโ remains a label users cannot verify and attackers can still treat as data once any compatible decoder exists.



