Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Cross-model replay recovers hidden CoT traces at OpenAI, Anthropic, and Google

Listen to this article

Encrypted chain-of-thought was supposed to hide proprietary reasoning from competitors and keep sensitive intermediate steps off the public page. A new paper shows those opaque blocks were portable enough that a weaker model in the same provider family could spit them back out in plaintext.

The paper and who wrote it

Stealing Reasoning Traces from Proprietary LLM APIs (arXiv 2608.09867) comes from researchers affiliated with ELLIS Institute Tübingen, the Max Planck Institute for Intelligent Systems, MATS, and Snyk. CybersecurityNews covered the findings August 11. The attack needs only standard, unprivileged API access. No stolen keys. No insider infra.

The setup is familiar if you have used modern reasoning APIs. Providers no longer return full plaintext CoT. They hand the client an encrypted envelope. The client passes that blob back on later turns so the API can stay mostly stateless. The paper’s finding: those envelopes were compatible across models, sessions, and users inside a single provider.

How the oracle trick works

Capture an encrypted reasoning block from a heavily guarded flagship model. Replay it into a cheaper, less-guarded sibling from the same vendor. Prompt that weaker model to transcribe the inserted thinking. Because anti-distillation and refusal training are thinner on the small tiers, the sibling acts as a decryption oracle.

The team demonstrated the pattern across those families with concrete pairs: Opus-class into Haiku-class, GPT-5.6-class into mini tiers, Gemini 3-class into Flash-class. That is the architectural point. Flagship alignment does not protect a trace that a lesser model in the same ecosystem will cheerfully decode.

What fell out of public repos

This is not only an IP story. Developers publish agent session logs. Many of those logs still carry the encrypted blocks. Users scrub what they can see. They cannot scrub what they cannot read.

The researchers decoded 315,320 reasoning blocks scraped from public repositories. From that haul: 367 PII artifacts and 182 credentials. Sensitive material often lived only inside the hidden trace, never in the visible assistant reply. Sharing the session still leaked it.

A second abuse path is invisible prompt injection. Malicious instructions can ride inside an encrypted block. Monitors that inspect only the visible conversation miss the payload. Downstream agent workflows still consume it.

Credentials hiding in “encrypted” thinking

If your product logs raw API transcripts, those signatures are sensitive data. Treat them like secrets. Strip them before a repo goes public. Enterprise teams that resume long agent runs from shared traces inherit the same risk: a poisoned opaque block can travel with the session.

For labs, the asymmetry is the punchline. CybersecurityNews’s table matches the paper’s pattern: Claude Opus 4.8 into Claude Haiku 4.5, GPT-5.6 into GPT-5-mini, Gemini 3 into Gemini Flash-class siblings. The ecosystem’s weakest compatible decoder became the bypass.

Portable ciphertext was the bug

Client-side encrypted CoT solved storage and anti-distillation on paper. In practice it created a portable ciphertext that was authenticated for content, not for who, which session, or which model was allowed to open it. That is a design choice with a body count in credentials, not a clever research trick that stayed in the lab.

Patches landed — binding still has to

Following responsible disclosure, OpenAI, Anthropic, and Google acknowledged the findings and deployed server-side mitigations. The original cross-model replay proofs are no longer reproducible on current API builds.

The paper’s recommendations still stand for defense in depth: cryptographically bind envelopes to model, session, and user; reject cross-tier blocks at the gateway; rotate signing keys so historical public ciphertext dies; scrub encrypted thinking fields from logs before publish. Until those habits stick, “encrypted reasoning” remains a label users cannot verify and attackers can still treat as data once any compatible decoder exists.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile