Skip to main content
Deepseek AI model matches GPT-5.6, 60% lower cost. AI technology, cost-effective, innovation, machine learning.

Editorial illustration for Deepseek's New AI Model Matches GPT-5.6 at 60% Lower Cost

Deepseek's New AI Model Matches GPT-5.6 at 60% Lower Cost

3 min read

Deepseek pushed out a new version of its budget model on July 31, calling it V4 Flash "0731," and the numbers put it within striking distance of OpenAI's cheaper offering. The Artificial Analysis Intelligence Index gives it a score of 50, a jump of ten points from the V4 Flash that launched back in April 2026. That's one point shy of GPT-5.6 Luna, OpenAI's budget model, but Deepseek's version runs about 60 percent cheaper per task. That gap holds even after OpenAI slashed its own prices by 80 percent.

The cost difference comes down partly to caching. Deepseek offers a 98 percent cache discount, well above the 90 percent that's become standard across the industry. The new model also trims token use by 12 percent compared to its predecessor, which adds up fast at scale.

Under the hood, nothing structural changed: 284 billion total parameters, 13 billion active at any given time, and a context window stretching to one million tokens. Deepseek released the weights under an MIT license on Hugging Face, so anyone can pull them down and run the model themselves.

Deepseek has released V4 Flash "0731," a major upgrade to its budget AI model. According to the Artificial Analysis Intelligence Index, the new version scores 50 points, ten more than the previous V4 Flash that launched in April 2026. That puts it just one point behind OpenAI's budget model GPT-5.6 Luna, but it costs about 60 percent less per task, even after OpenAI's 80 percent price cut.

Why this matters

For teams shipping AI products on tight margins, a one-point gap on the Artificial Analysis Intelligence Index at 60 percent lower cost is the kind of trade most engineering leads will take without blinking. Deepseek's 98 percent cache discount, well above the 90 percent norm, is doing a lot of the work here, and it's worth asking how that pricing holds up once usage scales past whatever promotional window Deepseek is currently running. The 12 percent token reduction is the more durable signal: it points to real efficiency gains rather than just aggressive discounting, and it's the kind of improvement that compounds across millions of API calls.

OpenAI already cut GPT-5.6 Luna's price by 80 percent before this comparison was even made, which tells us the budget-tier price war is already underway and Deepseek just landed the harder punch. For founders building agentic workflows specifically, the note that gains concentrated there matters more than the aggregate score. Watch whether OpenAI answers with another price cut or a model refresh, and whether Deepseek's cache economics survive contact with heavier production traffic.

LIVE20:44Chinese AI Researchers Turn to X for Technical Audience