In July, Chinese models captured all five of the top positions on OpenRouter, the neutral routing service that tracks AI usage worldwide. Xiaomi's MiMo V2.5 led by token volume, followed by models from DeepSeek, MiniMax, Alibaba's Qwen family and Moonshot's Kimi. The platform now processes more than 20 trillion tokens a week, with Chinese models accounting for over 60% of that traffic.
Why the shift matters
A year ago, US models handled roughly 70% of OpenRouter traffic; today they are down to about 30%. Even American firms are routing a record 58% of their tokens through Chinese services, not because of mandates but because the price-performance balance is hard to ignore.
The gap between frontier capability and commodity efficiency is widening. While US labs still lead on the most demanding tasks, with GPT-5.5, Claude Fable 5 and Gemini 3.x topping benchmarks, Chinese offerings are dramatically cheaper. DeepSeek's V4-Pro costs roughly one-twelfth of GPT-5.5 at comparable performance, and its V4 Flash is priced at $0.14 per million input tokens versus $5.00 for GPT-5.5. OpenRouter analysts note that Chinese open models run 60% to 90% cheaper than leading American products.
Impact on corporate AI strategies
Developers gravitate to models they can download and integrate. Alibaba's Qwen family has surpassed one billion cumulative downloads, overtaking Meta's Llama as the most downloaded open model family. Llama now represents less than 1% of routed volume. The data shows a clear split: premium providers like Anthropic hold about 12% of token share but capture roughly half of total spending, while cheap open models dominate the bulk of token traffic.
Most large enterprises are caught in the middle, using models that are neither the best nor the cheapest. This "death zone" squeezes budgets, as companies pay frontier prices for work that could be handled by efficient open models.
What executives can do now
Four practical steps can help organisations navigate the new landscape:
- Adopt hybrid routing. Direct high-risk, regulated tasks to frontier models and route high-volume, cost-sensitive workloads to efficient open models.
- Prioritise efficiency. Invest in inference optimisation, quantisation and hardware-model co-design to reduce costs without sacrificing quality.
- Build above the model layer. Leverage proprietary data, domain-specific fine-tuning and robust evaluation to create lasting competitive advantage.
- Avoid the middle. Choose either a capability-led or cost-led path this year; the hybrid middle is unlikely to survive beyond 2027.
Policy implications for the United States
US policymakers are debating restrictions on Chinese AI models for security reasons. While data-sovereignty concerns are valid for sensitive sectors, a blanket ban could hinder domestic developers who rely on affordable, high-quality open weights. The market shows that open-weight models are the primary vehicle for exporting standards and safety norms worldwide.
To remain competitive, the US should accelerate the release of credible open-weight models, backed by procurement incentives and clear commitments from leading labs. Matching both frontier capability and radical efficiency will be essential for shaping the next generation of global software.
The decisive question for boardrooms and regulators alike is whose models will underpin the AI-driven products of the future. Current download numbers suggest the balance is already tipping away from the US-preferred ecosystem.

