Chinese AI labs have stunned the world, again. In the space of about a month, Chinese AI start-up labs Z.ai and Moonshot have each launched a model that is nearly as intelligent as competitors from OpenAI and Anthropic but far cheaper. Silicon Valley start-ups are already using Z.ai’s GLM-5.2, and Moonshot’s Kimi-K3 likely isn’t far behind in that market.
The success of lower-cost Chinese AI models highlights how high prices are likely to put Western models at a major disadvantage in the coming era of agentic AI, which is more expensive to use than large language models (LLMs) by several orders of magnitude. A future dominated by AI agents is one where companies value the cheapest intelligence possible. Right now, it’s advantage China.
AI prices are calculated per token – a computing-resources unit that Nvidia CEO Jensen Huang has called ‘the new commodity’ due to its growing importance in AI economics. AI companies use tokens to meter AI usage, roughly measured by the number of words that users put into an AI model and the number of words it then generates. While ordinary internet users pay for a monthly subscription, developers, who are using agentic models to build the future of AI, are charged per token. Buying 1 million tokens for, say, ChatGPT, buys a developer access to that model for a certain number of tasks. The price of 1 million tokens, some for users’ input and others for the AI output, varies between models.
GLM-5.2 tokens are cheap: Z.ai charges US$1.92 (A$2.74) per million output tokens for GLM-5.2, whereas Anthropic charges $25 per million for its Opus 4.8.
AI agents dramatically compound these cost differences. These models don’t just generate content but can also perform tasks online. The number of tokens needed to use agentic AI is exponentially higher than for conventional chatbot-style LLMs. Recent estimates from a team from MIT, Stanford and Google DeepMind suggest that agentic tasks use up to 3,500 times as many tokens as a simple reasoning task. A single prompt and response for information on a normal LLM usually requires a couple of thousand tokens. A single agentic coding or research task often runs into the millions, with some agent users at the extreme end already using tens or hundreds of millions of tokens per day for tasks that take days to complete. The emergence of agentic AI is one of the reasons that demand for tokens is going through the roof. One major AI provider, OpenRouter, recorded that its customers burned through more tokens in one week in June than in an entire three-month period the previous year.
Cheapness is attracting institutions with high token usage. High token payments factored into Microsoft’s recent cancellation of its subscription to Claude Code and could well be a factor behind its sudden interest in using DeepSeek to power its agentic AI office helper, Copilot Cowork. AI agent pioneer Azeem Azhar, who augments his own research firm’s work with AI agents, has started using MiMo-2.5-Pro, a model from Chinese conglomerate Xiaomi. MiMo’s token costs are dramatically cheaper than Claude’s, making all the difference when Azhar is using more than 100 million tokens daily. ‘At the scale agents operate, even small cost differences compound into meaningful budget gaps,’ wrote Azhar on his Substack, Exponential View.
The international spotlight has focused mainly on the most intelligent AI models, with Silicon Valley moguls such as OpenAI CEO Sam Altman and Anthropic’s Dario Amodei aiming for super-intelligent AI. Amodei in 2024 hailed this as AI’s future, ‘a country of geniuses in a data center’ capable of dramatic scientific discoveries. But when choosing a model, a customer may be looking not for the smartest model but for something more practical, something that’s good enough for the task at hand and affordable. For example, a company looking for AI agents that can draft and reply to emails doesn’t need a genius in a datacentre. A few percentage points lower in intelligence may matter less than a few hundred dollars more in savings.
Western and Chinese AI labs are wrestling with the question of how to keep prices down even as their costs go up. Anthropic and OpenAI have heavily subsidised their user subscriptions, with sovereign wealth funds and venture capital firms footing the bill. The rise in token usage within China has created a corresponding rise in computing costs, as demand for AI services outstrips domestic supplies of AI chips. In March, state broadcaster CCTV reported that cloud providers including Alibaba Cloud, Huawei Cloud and Tencent Cloud had all raised their computing prices by around 30 percent in just 10 days. This came at a time when market competition was also encouraging China’s AI enterprises to cut prices and offer free service schemes to attract users, as it still is. In 2025, SiliconFlow, a Chinese third-party provider of AI models, spent 24 percent more on marketing than its entire revenue.
But even as both sides contend with the economics of AI, current prices from OpenAI and Anthropic will likely be unsustainable if the world moves into AI agents. Chinese AI companies are already well-positioned to undercut the US companies. Take, for example, DeepSeek, which announced a permanent 75 percent cut to token prices for its application programming interface. That means that for any number of tokens (in fact, a very large number), an agent driven by DeepSeek can cost as little as 1/34 as much as one using competitors from OpenAI or Anthropic. Research company Artificial Analysis this month made leading AI models perform 657 agentic AI tasks, meaning high token usage. These tasks were all based around office and administrative work, tasks that agents may be performing for companies in the not too distant future. The total cost across these tasks for Anthropic’s Opus 4.8 was nearly $1,000. For Z.ai, it was $270. Despite burning the highest number of tokens of any model tested (1 billion), DeepSeek-V4 Flash cost just $14 to do all these tasks.
There are still disadvantages to Chinese models. The Economist has reported that GLM-5.2 uses more tokens when performing complex tasks than a Western equivalent, meaning the savings aren’t always as big as the sticker price per 1 million tokens implies. Developers also take performance into consideration when choosing a model – that is, whether it can do a job well. Chinese models don’t always perform well: they can break when they are given demanding tasks that take several days to complete. DeepSeek-V4 Flash scored poorly on performance, completing fewer tasks correctly for Artificial Analysis than both Western and Chinese competitors.
However, this may now be changing. Released just last week, Kimi-K3 completed more of Artificial Analysis’s tasks correctly than any other model tested, at around a third of the cost of Anthropic’s Fable 5. Once a Chinese AI model manages to combine cheapness with good performance, Western AI will be in trouble.
Lower prices for Western AI are already on the horizon. Meta announced on July 9 that its new AI model, Muse Spark 1.1, would be available at $4.25 per 1 million output tokens. This is a step away from the company’s earlier free open-source plan but is still a price that undercuts those of Anthropic and OpenAI. OpenAI has launched several versions of its new model, GPT-5.6, including Luna, a ‘cost-efficient’ model that is cheaper to run agentically than Kimi-K3 and GLM-5.2. However, a cheaper model sacrifices some of its performance. Kimi-K3 scores nearly 10 percentage points higher than Luna for performance on Artificial Analysis’s agentic test.
Another solution to high token costs is to host an agent within your own computer hardware. Nvidia is already positioning itself to cut the costs of AI agents, releasing a new laptop that can run agentic workflows locally, without paying additional costs through an application programming interface. Prices have not yet been announced but are said to be around US$2,000, making them within reach of enterprises in the West.
But if the cost of running AI agents is a problem even for well-financed Western companies, it will be an insurmountable obstacle for many in the Global South. Governments of developing countries seek to harness AI for their own development but chafe against multiple resource constraints and costs. China’s leadership is targeting the Global South for deployment of its AI products. In a speech last week at the World AI Conference in Shanghai, President Xi Jinping said China would launch AI application cooperation centres within six regional intergovernmental organisations that together cover almost the whole Global South. An op-ed in the People’s Daily last year argued that the low cost and open weights of Chinese AI made the technology more accessible to developing countries in the face of ‘hegemonic’ Western AI. An expansion of Chinese AI products in the Global South allows the technology to become part of the foundations of the AI ecosystems of developing countries going forward, with all the increased geopolitical leverage that would entail.
If Western companies don’t want to be left behind by their cheaper Chinese counterparts, they must find a balance between intelligence and affordability.
Facts Only
* Z.ai launched GLM-5.2 and Moonshot launched Kimi-K3 within approximately one month.
* AI prices are calculated per token, which Nvidia CEO Jensen Huang calls ‘the new commodity.’
* GLM-5.2 costs US$1.92 per million output tokens; Anthropic charges $25 per million for Opus 4.8.
* Agentic tasks require exponentially more tokens than conventional LLM tasks.
* Agentic tasks can use up to 3,500 times as many tokens as simple reasoning tasks.
* A single agentic coding or research task can involve millions of tokens.
* DeepSeek-V4 Flash cost $14 for a set of agentic tasks, compared to Anthropic's Opus 4.8 costing nearly $1,000 across similar agentic tests.
* Meta’s Muse Spark 1.1 model is priced at $4.25 per 1 million output tokens.
* China’s cloud providers raised computing prices by around 30 percent in ten days in March.
Executive Summary
Chinese AI labs have released models, such as Z.ai's GLM-5.2 and Moonshot's Kimi-K3, that offer near-competitive intelligence to models from OpenAI and Anthropic at significantly lower costs. This cost disparity is magnified by the emerging field of agentic AI, where token usage scales exponentially higher than standard chatbot interactions. The economic structure of AI has shifted toward valuing the cheapest possible intelligence in the age of AI agents.
The escalating demand for tokens drives up computing costs; for instance, cloud providers in China raised prices significantly following increased demand. This cost pressure is driving Western companies to seek cheaper alternatives and is fueling competition between Chinese and Western developers. While the pursuit of super-intelligent AI remains a goal for some entities, practical application often favors affordability over peak intelligence when deploying agentic systems.
Full Take
The narrative emerging from this data highlights a fundamental tension between optimizing for peak intelligence and optimizing for deployable economics, especially as AI evolves into agentic systems. The argument that cheaper models will dominate the future is less about sheer capability and more about access; when token costs become a multiplier of operational expense in agentic workflows, cost efficiency becomes a prerequisite for competitive advantage rather than an optional feature. This dynamic creates a path for geopolitical leverage, where nations prioritizing deployment—like China targeting the Global South—can offer accessible, low-cost infrastructure that bypasses the high capital barriers associated with Western super-models.
The observed performance trade-offs between models further complicate the cost-intelligence equation. The fact that Chinese models might require more tokens for complex reasoning, or conversely, that some agents demonstrate superior performance relative to cost (as seen with Kimi-K3), suggests that a purely cost-based comparison is insufficient. This implies a systemic shift where the viability of a model depends on its context-specific utility rather than an absolute benchmark of "genius." The move toward localized agent execution, evidenced by Nvidia’s hardware focus, signals a technological response to this economic pressure, suggesting solutions may lie in distributed, localized computation rather than purely centralized model superiority.
When examining the geopolitical implications, the cost differential is not merely an economic variable but a mechanism for establishing new technological dependencies. If agents become the primary interface for complex tasks, controlling the cheapest operational layer—the token economy—grants significant leverage over global AI development trajectories. The drive among Western entities to match or surpass this cost-efficiency suggests a necessary re-evaluation of where value is placed: in centralized intelligence research or in scalable, accessible deployment frameworks.
Bridge Questions: If performance and cost can be decoupled, does the pursuit of lower costs fundamentally alter the definition of "intelligent" in an agentic context? How will the trend toward localized hardware execution reshape international AI governance and data sovereignty agreements? What institutional incentives are required to shift development focus from raw model size to verifiable, practical task completion benchmarks across diverse operational environments?
Sentinel — Human
The text appears to be a sophisticated synthesis of real industry reports regarding AI economics, presenting a complex argument about cost competition and global strategy.
