Editorial illustration for US benchmark shows China lagging; Deepseek model underperforms private tests
US benchmark shows China lagging; Deepseek model...
The US government has declared China's best AI model is falling behind. It's a political verdict.
DeepSeek V4 is the current Chinese champion. The US National Institute of Standards and Technology’s Center for AI Safety and Innovation, CAISI, ran the numbers. Its private testing shows the model underperforms its own maker's claims.
Deepseek says V4 is close to US leaders like Opus 4.6 and GPT-5.4. CAISI found it's closer to the older GPT-5.
It especially lags in abstract reasoning, cybersecurity, and software development. Math is the only bright spot, where it nearly matches the top tier.
CAISI tested performance across cybersecurity, software development, math, natural sciences, and abstract reasoning. CAISI calls Deepseek V4 the most capable Chinese AI model to date. But in private testing, it reportedly performs worse than Deepseek's own technical report suggests.
Deepseek pitches the model as roughly on par with current US models like Opus 4.6 and GPT-5.4. CAISI says it's actually closer to the older GPT-5 - especially on abstract reasoning, cybersecurity, and software development. Math is the one area where Deepseek V4 nearly matches the top US models.
The center, which likely has its own political agenda, sits within the National Institute of Standards and Technology (NIST). Its report paints a picture of a widening gap between US and Chinese models. Independent measurements tell a different story, showing the gap has stayed roughly constant.
Independent measurements disagree with CAISI's core conclusion. They show the performance gap has held steady, not widened. So you have a US government body inside NIST framing the race as one America is winning by a larger margin. And you have other data suggesting a persistent, stable rivalry.
This tells us more about benchmarks than about models. A test administered by one geopolitical actor will find what that actor needs to find. The math scores prove Chinese models can compete.
The reasoning scores show where they falter. The stalemate continues, measured differently depending on who owns the ruler.
Common Questions Answered
How does DeepSeek V4's performance compare to US models according to CAISI's testing?
According to the US National Institute of Standards and Technology's Center for AI Safety and Innovation (CAISI), DeepSeek V4 performs closer to the older GPT-5 rather than the newer US leaders like Opus 4.6 and GPT-5.4 as DeepSeek claims. The private testing shows the model underperforms its own maker's claims about its capabilities. This suggests a larger performance gap between Chinese and US AI models than DeepSeek's marketing suggests.
What specific area does DeepSeek V4 lag behind in according to CAISI's findings?
DeepSeek V4 especially lags in abstract reasoning capabilities compared to US models according to CAISI's benchmark testing. This weakness in abstract reasoning represents one of the most significant performance gaps identified in their evaluation.
Why do independent measurements disagree with CAISI's conclusions about the AI performance gap?
Independent measurements show that the performance gap between Chinese and US models has held steady rather than widened as CAISI suggests, indicating a persistent and stable rivalry rather than increasing American dominance. This disagreement highlights how benchmarks administered by one geopolitical actor may find results that align with that actor's interests, making it difficult to determine objective model performance differences.
What does the conflicting data between CAISI and independent measurements reveal about AI benchmarking?
The disagreement between CAISI's findings and independent measurements demonstrates that benchmark results tell us more about the testing methodology and the geopolitical interests of those administering tests than about the actual capabilities of AI models. A test administered by one geopolitical actor will tend to find what that actor needs to find, making truly objective comparisons challenging.
Further Reading
- Stanford: China 'effectively' closes AI model performance gap to US — Silicon Republic
- The U.S.-China AI Model Performance Gap Has Effectively Closed — Rich Turrin Substack
- US-China AI Technology Gap Narrows While Responsible AI Development Continues to Lag Behind — AI.cc
- Chinese AI models have lagged the US frontier by 7 months — Epoch AI