News · 2026-08-02
Beijing says U.S. firms distilled Chinese models, and names none of them
China's Ministry of Commerce accused "many U.S. AI companies" of distilling Chinese models during research and training in a written spokesperson statement published on 27 July, without naming a single company, model or training run, and without offering any technical evidence. It is a direct counter to a U.S. accusation made five days earlier — one that named two companies and also published nothing auditable.
Key facts
- The statement: a written spokesperson question-and-answer published 27 July 2026 by China's Ministry of Commerce — not a technical report, case filing or Foreign Ministry briefing. Original statement
- What it names: no U.S. company, no Chinese model, no U.S. model, no training run, no dates, no volumes.
- The accusation it answers: on 22 July, White House OSTP Director Michael Kratsios named Moonshot AI and Anthropic's Fable, alleging an internal platform for large-scale distillation that switched access methods to evade detection. Post
- Primary sources: the Ministry of Commerce statement, the Kratsios post, and Anthropic's February distillation disclosure.
Distillation is the technical heart of the dispute, and it is worth being precise about it because the word is doing two very different jobs. In its ordinary sense, distillation means training a smaller model to imitate a larger one — standard practice, usually applied to a company's own models, and the reason most cheap fast models exist. In the sense the accusations intend, it means systematically querying someone else's model to harvest its outputs as training data, which is closer to model extraction and is generally prohibited by terms of service.
Beijing's statement never says which it alleges. Nor does it say whether unauthorised access is claimed at all. Its supporting points are that model releases from the two countries have come close together in time, that Chinese capability is high, and that unnamed parts of U.S. industry oppose restrictions on access to Chinese models. None of that establishes distillation. The reference to nearly 200 U.S. startups opposing a cutoff from Chinese open-weight models goes to access policy, not to whether any of those startups trained on Chinese outputs. Observing that distillation is common industry practice does not identify improper conduct by anyone.
The U.S. accusation it answers is more specific and no better evidenced. Kratsios named a company, a target model and an alleged mechanism, including deliberate switching of access methods to avoid detection. He released no logs, no account identities, no outputs, no provenance trail and no model-forensics result.
There is one genuine asymmetry, and it should not be flattened. It would be too strong to say the U.S. side has never produced anything. Anthropic's disclosure on 23 February described what it says it observed on its own systems: large-scale extraction through fraudulent accounts, request metadata matching the public profiles of senior staff at the accused company, and attempts to reconstruct the model's reasoning traces. Anthropic is describing its own telemetry, which is a materially different kind of statement than a government asserting a conclusion. But Anthropic also published no logs, no metadata, no prompts and no attribution analysis, and it does not publicly demonstrate that any extracted output entered a specific training run. Its account of what it saw is verified as Anthropic's published claim. The attribution and the training conclusion are not independently verifiable from public material.
So the symmetry is real at the level of government statements. Washington: named company, named target, asserted mechanism, no public proof attached. Beijing: unnamed companies, no models, no mechanism, no public proof attached.
Why this matters beyond the diplomacy: the underlying question is genuinely answerable. Whether outputs from one model entered the training of another is a technical claim that logs, request metadata, statistical fingerprinting of outputs and training-data forensics can address. Researchers do this work. Neither government is showing any of it, which turns a checkable dispute into a matter of assertion — at a moment when Chinese open models have passed U.S. models in OpenRouter token share and China has been building governance structures around its model releases.
Reception among researchers has been appropriately unimpressed. Wharton's Ethan Mollick explicitly noted he had "no inside information" and framed the episode as escalating tension over open weights rather than a factual development. Researcher Eric W. Tramel listed the specifics a real allegation would contain: token volume, dates, what was input, which training stage, to what purpose, with what effect, and a clear definition of distillation. Those are reactions, not confirmation of either claim.
The honest caveat is that absence of published evidence is not evidence of absence. Both governments may hold material they are unwilling to release for intelligence, legal or commercial reasons, and companies rarely publish the forensics behind an accusation. But readers are entitled to be told the difference between a demonstrated finding and a stated position, and today both of these are stated positions.
Key questions
Which U.S. companies did China accuse?
What evidence did either government publish?
Is distillation illegitimate?
Cite this
APA
Ground Truth. (2026, August 2). Beijing says U.S. firms distilled Chinese models, and names none of them. Ground Truth. https://groundtruth.day/news/beijing-says-us-firms-distilled-chinese-models-and-names-none.html
BibTeX
@misc{groundtruth:beijing-says-us-firms-distilled-chinese-models-and-names-none,
title = {Beijing says U.S. firms distilled Chinese models, and names none of them},
author = {{Ground Truth}},
year = {2026},
month = {aug},
url = {https://groundtruth.day/news/beijing-says-us-firms-distilled-chinese-models-and-names-none.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.