DEV Community

Breach Protocol
Breach Protocol

Posted on Originally published at groundtruth.day

Anthropic alleges covert routing to Claude for reasoning-trace extraction

Anthropic alleges that Moonshot and DeepSeek covertly relayed selected customer requests to Claude, showed Claude answers as their own and used outputs or reconstructed reasoning to improve their models. The allegation matters because it connects model extraction, user disclosure and cross-border prompt handling, but it is not evidence that Anthropic's network was hacked or that a regulator has reached a finding.

Key facts

  • Anthropic attributes more than 23 million exchanges to Moonshot from May through July 2026.
  • It attributes more than 12.1 million exchanges to DeepSeek over 14 days in July.
  • Anthropic describes “cross-session replay” for reconstructing reasoning.
  • Primary source: Anthropic's September threat-intelligence report.

Anthropic says Moonshot's Kimi service silently forwarded customer requests to Claude and returned Claude's answers as Kimi's. It says Moonshot retained some exchanges and extracted reasoning; for DeepSeek, it alleges selected requests from users of third-party or Anthropic coding harnesses were relayed to Claude Opus without users' knowledge or consent. The named technique is cross-session replay: save a returned reasoning signature, start a new session and ask the model to reconstruct the underlying reasoning. Anthropic calls the conduct “illicit distillation.”

The easy but wrong metaphor is a burglar entering a model's servers. The dossier supports a different metaphor: a shop silently sends a customer's hard question to a rival expert, delivers the rival's answer under its own sign, and keeps a notebook of how that expert reasoned. That can create three separate problems. It may extract capability; it may mislead a customer about where their prompt went; and it may move sensitive material across a jurisdiction. Evidence for one problem does not prove the others.

Anthropic's numbers are a useful anchor, not independent forensic proof. Its report describes example prompts involving corporate documentation, credentials and state-linked systems, but those are Anthropic assessments. It says it banned associated accounts, disrupted proxy networks and added behavior detection, extraction classifiers, reasoning summarization and identity verification. “We do not know whether Moonshot notified customers,” Anthropic writes, a limitation worth retaining. The Information reports that China's internet regulator is investigating, but its accessible public page does not identify a named agency and no public penalty, case filing or decision was found.

The technical concern is plausible in a general sense. The independent paper Stealing Reasoning Traces from Proprietary LLM APIs describes cross-session encrypted-reasoning extraction against several proprietary APIs, but it does not name Moonshot or DeepSeek. That supports technical context, not attribution. No response to these September allegations was found on the official DeepSeek and Moonshot pages reviewed. The fair conclusion is neither “confirmed breach” nor “nothing happened”: it is a well-specified public allegation that should trigger routing audits, disclosure checks, account-abuse monitoring and skepticism toward any provider that cannot explain its model and data path.


Originally published on Ground Truth, where every claim is checked against the primary source.

Top comments (0)