Chinese artificial intelligence models have pulled ahead of their U.S. competitors in global usage, according to fresh data from OpenRouter, the world's largest aggregation platform for large language models. The figures, reported across multiple outlets on June 20, 2026, show that systems built by Chinese labs now consume more tokens worldwide than those from OpenAI, Google, or Anthropic — a shift driven by aggressive pricing, rising AI-agent workloads, and a wave of efficiency gains forced by U.S. chip export controls.

DeepSeek V4 Flash leads a Chinese-dominated ranking For more context on this story, see our ongoing AI industry coverage.

According to OpenRouter data cited by The News International and corroborated by CGTN and Firstpost, Chinese models now occupy the top of the global token-consumption leaderboard. DeepSeek V4 Flash ranked first with 4.63 trillion tokens processed, followed by MiniMax M3 at 4.13 trillion tokens and Xiaomi's MiMo-V2.5 at 3.8 trillion tokens.

The broader top 10 also included Tencent's Hy3, DeepSeek V4 Pro, and Owl Alpha from Chinese developers, alongside Anthropic's Claude Opus 4.7, Claude Sonnet 4.6, and Claude Opus 4. Anthropic was the only U.S.-based company to crack the top five. Google's Gemini 2 and OpenAI's GPT-5.5 fared considerably worse, landing in 12th and 13th place respectively.

The result is striking because OpenRouter is not a Chinese platform. Founded by Alex Atallah, the former CTO of the NFT marketplace OpenSea, it offers a single API gateway to more than 400 models from over 60 providers, and its transparent usage data is widely treated as an industry barometer of which models developers and enterprises actually choose to run.

Why tokens have become the new battleground

A token is the smallest unit an AI model processes — a whole word, a fragment, or even a single character — and providers bill on the volume of tokens handled for both input and output. That makes token consumption a direct proxy for how heavily a model is being used, and it has become the central pricing front in the industry.

The stakes are rising fast because of AI agents. A simple chatbot might burn around 30,000 tokens summarizing Shakespeare's Hamlet, but an autonomous agent executing a coding task can consume up to 20 million tokens. Goldman Sachs has predicted that the spread of AI agents will drive a 24-fold increase in token consumption by 2030, contributing to an expected chip-supply crunch over the next 12 to 18 months.

Chinese models win on cost

The core reason for China's lead is price. Reporting from Trending Topics, based on OpenRouter pricing data, shows Chinese models are roughly 10 to 20 times cheaper than leading U.S. alternatives. MiniMax M2.5, for example, charges around $0.30 per million input tokens and $1.10 per million output tokens, compared with about $5.00 and $25.00 for Anthropic's Claude Opus 4.6.

That gap matters enormously for agentic workloads that process huge token volumes continuously. Several factors underpin the cost advantage: cheaper energy bolstered by heavy state investment in renewables, more efficient "Mixture-of-Experts" architectures, and the discipline imposed by U.S. export restrictions on advanced chips, which forced Chinese developers to do more with less compute. China's government has explicitly linked energy policy and AI competitiveness as a national priority in its 2026 work plan.

Performance closes the gap

Price alone would not be enough if the models lagged in quality. But the technical gap has narrowed sharply. MiniMax M2.5 scored 80.2% on SWE-Bench Verified, a standard benchmark for software-engineering capability, only narrowly behind Claude Opus 4.6's 80.8%. According to the analytics platform Artificial Analysis, several Chinese models have now reached the global top tier across programming, multimodal understanding, and long-context processing.

The adoption pattern extends to U.S. startups themselves. Martin Casado of the venture firm Andreessen Horowitz has estimated that around 80% of young AI companies building on open-source stacks now use Chinese models.

U.S. firms retreat to token caps and budget controls

The OpenRouter figures land at a moment when major American companies are pulling back on unconstrained AI spending. Several U.S. firms are moving away from flat-rate subscriptions toward token-based billing, and enterprises are introducing hard caps to contain costs. Uber, Amazon, Microsoft, and Meta have all reportedly limited token consumption in recent weeks as the cost of deploying AI at scale outstrips their budgets.

Microsoft's retrenchment has been particularly visible. The company recently decided to revoke its internal Claude code license for its Experiences and Devices division, ending access by June 30, 2026 — a signal that even the largest enterprise buyers are scrutinizing per-token economics.

A changing competitive map

The OpenRouter snapshot does not capture the entire global market — it reflects only the share of usage flowing through one aggregator. But the direction of travel is consistent with a broader realignment in which Chinese labs, once dismissed as fast followers, are increasingly setting the pace on the metrics that developers care about most: price, speed, and good-enough capability.

Whether U.S. labs respond by cutting prices, leaning harder into premium performance, or pushing policy levers remains to be seen. For now, the token data tells an unambiguous story: when developers vote with their API calls, China is winning.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →