Skip to main content
Nvidia Nemotron 4 AI chip with one trillion parameters, surpassing China's current large language models.

Editorial illustration for Nvidia's Nemotron 4 Aims for One Trillion Parameters, a Scale China Already Passed

Nvidia's Trillion-Parameter Nemotron 4 Chases China's Lead

Nvidia's Nemotron 4 Aims for One Trillion Parameters, a Scale China Already Passed

4 min read

Nvidia is building its biggest open-weight model yet, and by the time it ships, it still won't be the biggest model around. The Information reports that Nemotron 4, the successor to June's Nemotron 3 Ultra, will pack at least one trillion parameters in its largest version, twice the size of its predecessor. Nvidia has backed the effort with real money, tripling its cloud spending on in-house model training to $28 billion through 2031, with a release possible as early as this fall.

The problem is timing. Chinese labs have already blown past the trillion-parameter mark Nvidia is aiming for. Moonshot AI's Kimi K3 sits at 2.8 trillion parameters, and DeepSeek V4 Pro runs at 1.6 trillion. Nemotron 3 Ultra held the title of strongest open US model on the Artificial Analysis Intelligence Index at launch, but it never caught Kimi K2.6, and the gap has only widened since.

That scoreboard sets up an awkward position for Nvidia, which sells the chips that power both its own models and its rivals' training runs, all while lobbying against restrictions on the open models fueling that competition.

Even at one trillion parameters, Nvidia would only be reaching a scale that Chinese labs already occupy. Moonshot AI's Kimi K3 has 2.8 trillion parameters, and DeepSeek V4 Pro has 1.6 trillion.

Why this matters

Nvidia's $28 billion bet on Nemotron 4 tells us more about the open-weight race than the model itself will. Parameter count isn't the whole story, but the gap is stark: Kimi K3 at 2.8 trillion, DeepSeek V4 Pro at 1.6 trillion, and Nvidia's flagship still aiming for a fraction of that. For developers building on open models, this matters less as a scoreboard and more as a signal about who's setting the pace.

Chinese labs are shipping frontier-scale weights faster than Nvidia can train them, even with a cloud budget that tripled overnight. If Nemotron 4 lands and it's merely competitive with what DeepSeek and Moonshot already shipped months earlier, that's a data point worth sitting with: the "strongest open US model" title Nemotron 3 Ultra held in June is a moving target, and the movement isn't coming from Silicon Valley. Watch whether Nvidia ships on schedule or slips, and whether raw parameter count still predicts benchmark position once Nemotron 4 actually hits the Artificial Analysis leaderboard.

Common Questions Answered

What is the expected parameter count for Nvidia's Nemotron 4 model?

Nvidia's Nemotron 4 will pack at least one trillion parameters in its largest version, which is twice the size of its predecessor Nemotron 3 Ultra. The model could be released as early as fall of this year, representing a significant scale-up in Nvidia's open-weight model offerings.

How does Nemotron 4's one trillion parameters compare to Chinese AI models?

Even at one trillion parameters, Nemotron 4 will lag behind Chinese labs that have already surpassed this scale. Moonshot AI's Kimi K3 has 2.8 trillion parameters, while DeepSeek V4 Pro has 1.6 trillion parameters, demonstrating that Chinese laboratories are currently leading in model scale.

How much is Nvidia investing in cloud spending for model training through 2031?

Nvidia has tripled its cloud spending on in-house model training to $28 billion through 2031 to support the development of Nemotron 4. This substantial investment demonstrates Nvidia's commitment to competing in the open-weight model space despite the scale advantages held by Chinese competitors.

Why does parameter count matter less than the pace of development in the open-weight model race?

While parameter count isn't the whole story of model capability, the key signal is about who's setting the development pace in the industry. Chinese labs are shipping frontier-scale weights faster than Nvidia can, which matters more to developers building on open models than absolute parameter counts alone.

LIVE16:14Anthropic Hires Legal Startup Founder Robert Mahari to Lead Claude's Law Push