Back
AI News

Encrypted reasoning traces can be forced into plaintext

Researchers claim they can steal hidden reasoning traces from Anthropic, OpenAI, and Google by exploiting the fact that encrypted blocks work across models — on the same day the FBI, NSA, and CISA warned about industrial-scale CoT…

AIMag.no
AIMag.no
September 15, 2026 · 4 min
Illustration: Sealed paper blocks wrapped in black tape, one pried open to reveal typed paper strips spilling out — a metaphor for encrypted reasoning traces forced into plaintext.

Encrypted reasoning traces can be forced into plaintext

Researchers claim they can steal hidden reasoning traces from Anthropic, OpenAI, and Google by exploiting the fact that encrypted blocks work across models — on the same day the FBI, NSA, and CISA warned about industrial-scale CoT extraction by Chinese AI firms.

Frontier model providers increasingly hide models' reasoning traces behind encryption to protect intellectual property and counter distillation. But the concealment mechanism is itself vulnerable, according to new research published with responsible disclosure on September 8, 2026.

The vulnerability: interchangeable encrypted blocks

In a paper titled "Stealing Reasoning Traces from Proprietary LLM APIs" — cited verbatim on Schneier on Security — the authors describe an architectural vulnerability: the encrypted reasoning blocks that the APIs return are "fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem" (f722aa6e).

The attack exploits precisely this compatibility. By injecting an encrypted reasoning trace from one model into a "weaker, and less safeguarded model from the same provider," the weaker model is forced to decode and output the trace "verbatim in plaintext" — without the more capable model ever being directly jailbroken (f722aa6e).

What the researchers claim they demonstrated

All of the following figures are the authors' own claims from the abstract, not independently verified:

  • According to the authors, the technique circumvents anti-distillation mechanisms and was demonstrated against Anthropic, OpenAI, and Google (f722aa6e).
  • By decoding 315,320 reasoning blocks scraped from public code repositories, they claim to have found 367 PII items and 182 credentials in developers' shared session logs (f722aa6e).
  • The vulnerability is also said to enable invisible prompt injections: malicious payloads can be hidden entirely inside encrypted blocks to poison public agentic runs (f722aa6e).

The authors state that the research followed responsible disclosure and that they propose cryptographic and system-level countermeasures to secure client-side reasoning (f722aa6e). Neither the paper's authors, its publication venue, nor its peer-review status are known from the source material, and none of the vendors have commented on the claim.

Same day: FBI, NSA, and CISA on CoT theft

Entirely separately, but on the same day, the FBI, NSA, and CISA published a joint advisory stating that Chinese AI firms are conducting "industrial-scale" extraction of capabilities from leading American models (Dark Reading, 27224c49).

The advisory names six developers — Alibaba, DeepSeek, MiniMax, Moonshot AI, StepFun, and Z.AI — accused of using American LLMs in training their own models (27224c49; NBC News, 18fe3a47). Among the advanced distillation tactics cited is explicitly chain-of-thought (CoT) reasoning extraction, along with automated failover between access routes and quality-assessment frameworks to detect defensive countermeasures (27224c49).

The agencies claim the firms extracted billions of tokens through millions of exchanges from frontier model variants of Claude, GPT, Gemini, and Grok, since as recently as late 2024 (27224c49). These are unverified government claims.

Company and political claims

Several claims circulating around the advisory come from actors with their own interests:

  • Anthropic claims that three Chinese AI labs are behind an "industrial campaign" against Claude, involving more than 16 million exchanges via roughly 24,000 fake accounts (CNN, 29e54d80).
  • The agencies claim the Chinese firms use a gray market of intermediaries — so-called "transfer stations" — to get around geographic restrictions, security measures, and terms of service at American companies (29e54d80).
  • Treasury Secretary Scott Bessent wrote on X that in cases of covert, industrial-scale distillation amounting to theft of intellectual property, sanctions and Entity List designations would "be on the table" (29e54d80).
  • The advisory characterizes distillation not as a supplement but as "the critical core" of the named companies' model development (18fe3a47).
  • China rejects the claims. Foreign Ministry spokesperson Mao Ning said she had not seen the report, but that China's AI development is the result of "high-level scientific and technological self-reliance" (18fe3a47). The Chinese embassy in Washington called the American emphasis on "distillation" a deliberate attack on China's AI progress, which China "firmly rejects" (29e54d80).

Uncertainty and open questions

There is no documented connection between the academic vulnerability disclosure and the agencies' allegations about Chinese distillation campaigns — they are two separate events that happened to fall on the same day. Several key figures are unverified: the researchers' decoding numbers as well as the claims about "billions of tokens" and 16 million exchanges all come from interested parties, not independent verification. The paper itself is not available in the source material, and none of the vendors have commented on the vulnerability.

The clear message from the research is nonetheless concrete: if reasoning traces are to be protected, simply encrypting them is not enough — especially not when the encryption is designed so that blocks can move freely between models within the same ecosystem.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Sources

  1. Stealing AI Reasoning Traces - Schneier on Securitywww.schneier.com
  2. US Government Claims Chinese AI Firms Distilling Frontier Modelswww.darkreading.com
  3. US accuses China AI developers DeepSeek and Alibaba of copying American AIwww.nbcnews.com
  4. US claims Chinese AI firms are carrying out ‘industrial-scale’ theft of trade secrets | CNN Politicswww.cnn.com