ByteDance is reportedly training a massive artificial intelligence model with as many as 10 trillion parameters, a scale that could place it alongside some of the world's most advanced AI systems. The Financial Times, citing people familiar with the project, reported on Friday that the model could approach the size of Anthropic's flagship Mythos system.
If the reported figure is accurate, the model would be more than three times larger than Chinese startup Moonshot AI's Kimi K3, which is built with approximately 2.8 trillion parameters.
Parameters are the internal numerical values an AI model learns during training to identify patterns, understand information, generate responses and perform various tasks. While parameter count is often viewed as an indicator of a model's scale, it does not necessarily reflect overall performance or intelligence.
Before the launch of Kimi K3, Meituan's LongCat-2.0 and DeepSeek's V4-Pro were among China's largest AI models, each featuring around 1.6 trillion parameters. Several other Chinese developers have also crossed the one-trillion-parameter milestone as competition in the country's AI sector intensifies.
Comparing Chinese models with leading US systems remains difficult because companies such as Anthropic and OpenAI have not publicly disclosed the parameter counts for their latest models, including Fable, Mythos and GPT-5.5.
According to the Financial Times, industry estimates suggest Anthropic's flagship Mythos 5 model contains roughly 8 trillion parameters, while Fable 5 is believed to have around 5 trillion. If those estimates are accurate, ByteDance's upcoming model would be comparable in scale to Mythos, making it one of the largest AI systems under development.
The report comes as Chinese technology companies continue to accelerate AI development in response to intensifying global competition. Firms are racing to build increasingly capable foundation models while also trying to keep training and operating costs under control as they compete with major US developers.
According to the report, ByteDance's model is currently in the pre-training phase, during which it learns from vast amounts of data before being fine-tuned for specific applications. This stage typically lasts between three and six months before the model is prepared for public release or commercial deployment.
Chinese artificial intelligence startup DeepSeek is preparing to significantly increase the prices of its AI services, marking a notable departure from the aggressive low-cost strategy that has reshaped competition in the global AI industry.
The Hangzhou-based company has recently attracted widespread attention for offering high-performance AI models at a fraction of the cost charged by many rivals. Its V4 Flash model has demonstrated performance comparable to some of the world's leading AI systems while costing only a few cents, compared with several dollars charged by major U.S. competitors.
In a notice issued to users on Thursday, the company said it plans to introduce substantial price increases across its AI services. However, it did not disclose the exact extent of the hikes, instead advising customers to prepare for the changes.
At present, DeepSeek charges just $0.14 per million input tokens and $0.28 per million output tokens.
Catch all the Business News, Market News, Breaking News Events and Latest News Updates on Live Mint. Download The Mint News App to get Daily Market Updates.
Oops! Looks like you have exceeded the limit to bookmark the image. Remove some to bookmark this image.
Facts Only
* ByteDance is training an AI model with up to 10 trillion parameters.
* The model is currently in the pre-training phase.
* Pre-training typically lasts between three and six months.
* Moonshot AI's Kimi K3 has approximately 2.8 trillion parameters.
* Meituan's LongCat-2.0 and DeepSeek's V4-Pro each have around 1.6 trillion parameters.
* Anthropic's Mythos 5 is estimated at 8 trillion parameters.
* Anthropic's Fable 5 is estimated at 5 trillion parameters.
* OpenAI has not disclosed parameter counts for Fable, Mythos, or GPT-5.5.
* DeepSeek is increasing prices for its AI services.
* DeepSeek's current pricing is $0.14 per million input tokens and $0.28 per million output tokens.
* DeepSeek is based in Hangzhou.
Executive Summary
ByteDance is developing a massive AI model potentially reaching 10 trillion parameters, which would place it at a scale comparable to leading U.S. systems like Anthropic's Mythos. This development occurs amidst intense competition within the Chinese AI sector, where several firms, including Moonshot AI, Meituan, and DeepSeek, have already surpassed the one-trillion-parameter milestone. While parameter count is a common metric for scale, it is not a definitive proxy for a model's actual intelligence or performance.
Significant uncertainty exists regarding the exact capabilities of these systems because major U.S. developers like OpenAI and Anthropic do not publicly disclose their parameter counts; current comparisons rely on industry estimates. Parallel to these scaling efforts, the market is seeing a shift in pricing strategies. DeepSeek, previously known for aggressive low-cost offerings, is preparing to substantially increase its service prices, signaling a transition away from the low-cost disruption strategy that characterized its initial market entry.
Full Take
The strongest version of this narrative is that the global AI race has entered a "brute force" era, where scaling parameters is the primary vector for achieving parity between Chinese and American frontier models. It frames the current landscape as a high-stakes arms race of computational scale and economic endurance.
The narrative relies heavily on "parameter counting" as a proxy for progress. This creates a mental shortcut for the reader: larger equals better. However, by mentioning that scale does not necessarily reflect intelligence, the text maintains a veneer of objectivity while still centering the entire value proposition on the 10-trillion figure. The reliance on "people familiar with the project" and "industry estimates" for the most significant claims (the 10 trillion and 8 trillion figures) means the central tension of the piece is built on unverifiable data.
Patterns detected: ARC-0062 Authority Game
This narrative is driven by the "Scaling Hypothesis"—the belief that more data and more parameters inevitably lead to emergent intelligence. It echoes the Cold War space race, where the goal was often a visible milestone (reaching the moon/reaching 10 trillion parameters) rather than a specific functional utility. The second-order consequence is an immense increase in energy and capital expenditure, where the "winner" is decided by infrastructure capacity rather than algorithmic elegance.
If this were an influence campaign, the playbook would involve inflating the perceived scale of domestic AI to project technological dominance and deter competitors or attract investment, using precise-sounding but unverifiable numbers to create an illusion of transparency. The content matches this pattern slightly by emphasizing trillion-parameter milestones without providing the underlying benchmarks.
Bridge Questions:
1. If parameter count is not a direct indicator of intelligence, why has it become the primary metric for reporting "progress" in AI?
2. How does DeepSeek's price hike affect the accessibility of high-performance AI for smaller developers?
3. What functional capabilities would a 10-trillion parameter model provide that a 2-trillion parameter model cannot?
Sentinel — Human
This text functions as an aggregation of reported industry developments and comparative statistics, exhibiting the balanced structure typical of financial news reporting.
