News · 2026-07-28
Chinese open models passed US models in OpenRouter token share
OpenRouter's analysis of its own logged usage shows Chinese models overtaking U.S. models in token share in early June, with DeepSeek roughly doubling from about 9% to 18% of token flow between January and June. The driver is not a broad collapse in model prices. It is that token-hungry agent workloads - which OpenRouter says consume roughly 15 times the tokens of human requests - moved onto the cheapest capable endpoints.
Key facts
- DeepSeek rose from about 9% to 18% of OpenRouter token flow, January to June 2026.
- Chinese models overtook U.S. models in token share in early June.
- Agentic requests use about 15x the tokens of human requests; V4 Flash became 70% of DeepSeek's agentic flow by late May.
- Primary source: OpenRouter's own usage analysis.
OpenRouter is a routing layer: developers send requests to it, and it forwards them to whichever model provider they have chosen. That makes its logs an unusually direct view of what people actually run, as opposed to what they say in surveys or what vendors announce.
What those logs show is a shift with a specific mechanism, and the mechanism is the interesting part - because it is being widely reported as something it is not.
The claim in circulation is that AI is getting cheaper. The data supports something narrower: the average price paid per token fell, because the composition of traffic changed. This is the difference between a store cutting prices and its customers switching to the cheaper aisle. Both lower the average receipt. Only one is a price cut.
The composition shift has a clear cause. Agent workloads are structurally different from chat. A person asks a question, reads an answer, and thinks for a while. An agent runs a loop - plan, call a tool, read the result, revise, call another tool - and each cycle re-sends accumulated context. OpenRouter puts the difference at roughly 15 times the tokens per request. So when agents became a mainstream way to use models, the token market gained a new dominant buyer with a completely different sensitivity to price.
For a person paying per conversation, a threefold price difference between models is noise against the value of a better answer. For a loop burning fifteen times the tokens and running unattended, it is the entire budget. That buyer optimizes hard on cost, and it will accept a meaningfully weaker model to do it.
DeepSeek's V4 Flash is what that buyer selected. By late May it accounted for 70% of DeepSeek's agentic flow on the platform - an inexpensive, fast, mixture-of-experts model with a small active-parameter count, which is exactly the profile that wins when a loop is paying the bill. The same model turned up this week in a local benchmark hitting 32 tokens per second on one desktop, which is the other half of the same story: cheap to call, and increasingly cheap to self-host.
There is a further complication for anyone trying to read prices out of this. Prompt caching can cut effective input costs to 60-80% below list price, and agent loops - which re-send the same long context repeatedly - are precisely the workload caching helps most. So even the list prices that did not change may not describe what anyone actually paid.
The honest caveats are substantial and worth stating plainly. OpenRouter is one router, with a developer-heavy and price-sensitive user base that is not the market. Its public data does not expose a historical catalog-price series detailed enough to separate mix shift from genuine price movement, so any "AI got cheaper" claim built on this number is unsupported as stated. And token share is not revenue share - cheap tokens are, by construction, worth less per token, so a model can dominate volume while representing a small fraction of spending.
What survives all of that is a structural observation rather than a headline. The marginal buyer of inference tokens is now a loop, not a person. Loops are patient, unattended, price-sensitive, and indifferent to brand. That is a fundamentally different competitive environment than the one frontier labs built their positioning for, and it favors whoever can serve adequate quality at the lowest cost per token - which, right now, disproportionately means open-weight models from Chinese labs.
Whether that holds depends on things the data cannot yet answer: whether agentic workloads keep growing as a share of all inference, whether cheap models stay adequate as agent tasks get harder, and whether frontier labs respond with cheap tiers of their own. A second independent router publishing comparable numbers would turn this from a datapoint into a trend.
See also: open weights become an insurance policy and training vs inference.
Key questions
Does this mean AI got cheaper across the board?
Why do agents change the economics so much?
Is OpenRouter representative of the whole market?
Cite this
APA
Ground Truth. (2026, July 28). Chinese open models passed US models in OpenRouter token share. Ground Truth. https://groundtruth.day/news/chinese-open-models-passed-us-models-in-openrouter-token-share.html
BibTeX
@misc{groundtruth:chinese-open-models-passed-us-models-in-openrouter-token-share,
title = {Chinese open models passed US models in OpenRouter token share},
author = {{Ground Truth}},
year = {2026},
month = {jul},
url = {https://groundtruth.day/news/chinese-open-models-passed-us-models-in-openrouter-token-share.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.