Skip to main content
ByteDance AI model pretraining: server racks, glowing circuits, data streams, digital innovation in China.

Editorial illustration for ByteDance Develops China's Largest AI Model, Currently in Pretraining

ByteDance Builds China's Largest AI Model at 10T Parameters

ByteDance Develops China's Largest AI Model, Currently in Pretraining

4 min read

ByteDance is building an AI model with as many as ten trillion parameters, according to the Financial Times, which would make it the largest system under development in China by a wide margin. For comparison, Moonshot's Kimi K3, currently the country's biggest model, tops out at roughly a third of that size. A model in that range would also put ByteDance closer to Anthropic's Mythos 5, which outside estimates peg at around eight trillion parameters, though Anthropic has never confirmed the figure.

Three people familiar with the project told the FT the model is still in pretraining, the stage where a system absorbs raw data before fine-tuning begins. That phase usually runs three to six months, meaning ByteDance's model likely won't be ready for release anytime soon. Parameter count alone doesn't decide how good a model is, since data quality and training technique matter just as much, but the scale here is notable given how few companies, in China or elsewhere, have attempted models this large. ByteDance's TikTok ownership and deep pockets make the effort one to watch closely as the pretraining window plays out.

Bytedance is training an AI model with up to ten trillion parameters, according to the Financial Times. That's three times the size of Moonshot's Kimi K3, currently the largest Chinese model. It would put the TikTok parent company in the same ballpark as Anthropic's top system Mythos 5, which industry estimates place at around eight trillion parameters.

Why this matters

A ten-trillion-parameter model tells us less about capability than about intent. ByteDance already runs Doubao to hundreds of millions of users and owns TikTok's recommendation stack, so a model this size isn't a research flex, it's infrastructure for products that already exist at scale. The comparison to Anthropic's Mythos 5 matters less than the detail about skipping distillation.

Training from scratch at this size is expensive and slow, and it signals ByteDance wants a foundation model it controls end to end rather than one bootstrapped from a rival's outputs. For developers and founders building on Chinese AI infrastructure, that's the number to watch: whose stack are you actually inheriting biases and limitations from. Three to six months of pretraining means nothing shippable soon, and parameter count alone won't tell us if this thing reasons better than Kimi K3 or just costs more to run.

We'd treat this as a marker of ByteDance's ambition, not yet evidence of a leap in what Chinese models can do.

Common Questions Answered

How does ByteDance's new AI model compare in size to other Chinese AI models?

ByteDance's model with up to ten trillion parameters is approximately three times larger than Moonshot's Kimi K3, which is currently the largest AI model developed in China at roughly one-third the size. This makes ByteDance's model by far the largest system under development in China by a significant margin.

What is the estimated parameter count of Anthropic's Mythos 5 compared to ByteDance's model?

Industry estimates place Anthropic's Mythos 5 at around eight trillion parameters, which puts ByteDance's ten trillion parameter model in a similar ballpark but slightly larger. However, Anthropic has never officially confirmed the exact parameter count for Mythos 5.

Why is ByteDance training a ten-trillion-parameter model from scratch instead of using distillation?

Training from scratch at this massive scale is expensive and slow, but ByteDance's decision to skip distillation signals that the company wants to build a foundational model for infrastructure supporting its existing products at scale. This approach differs from a typical research demonstration, as ByteDance already operates Doubao to hundreds of millions of users and controls TikTok's recommendation system.

What does ByteDance's ten-trillion-parameter model reveal about the company's strategic intentions?

The model size indicates ByteDance's intent to build critical infrastructure for its existing products rather than simply showcase research capabilities. Since ByteDance already serves hundreds of millions of users through Doubao and manages TikTok's recommendation algorithm, this massive model represents a strategic investment in scaling its existing platform capabilities.

LIVE15:57OpenAI partners with psychologists to address AI's mental health shortcomings