Chinese artificial intelligence lab DeepSeek released V4-Flash "0731" on July 31, 2026, a retrained version of its budget model that nearly matches OpenAI's GPT-5.6 Luna on third-party benchmarks at roughly 60 percent lower cost. The release, which puts the model's public beta API into general availability, is the latest salvo in an intensifying price war between Chinese open-weight providers and Western frontier labs. For ongoing coverage of the model wars, check our latest AI news.
According to The Decoder, the new version scores 50 points on the Artificial Analysis Intelligence Index — a ten-point jump over the previous V4-Flash that launched in April 2026. That places it just one point behind OpenAI's budget-tier GPT-5.6 Luna, despite costing about 60 percent less per task, even after OpenAI's recent 80 percent price cut.
Built for Agentic Work
The biggest gains came in agentic tasks, the ability to autonomously complete multi-step real-world work. TechNode reported that the model scores 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE, two benchmarks designed to measure how well an AI can use tools, write code, and solve engineering problems without hand-holding.
On GDPval, a benchmark that tests models on complex office work, V4-Flash climbed from 1,189 to 1,559 Elo points. DeepSeek also said the model hallucinates less often than its predecessor and uses roughly 12 percent fewer tokens to complete the same tasks.
Same Architecture, Smarter Weights
The upgrade keeps the same underlying structure: 284 billion total parameters with 13 billion active at any one time, and a one-million-token context window. Rather than enlarging the model, DeepSeek retrained it — squeezing more performance out of the same footprint. The model weights are released under the permissive MIT license on Hugging Face, continuing DeepSeek's open-weight approach.
DeepSeek was careful to scope the release. The update applies only to the V4-Flash API; the higher-tier V4-Pro API and the models served through DeepSeek's consumer app and website remain unchanged. The new Flash API also adds support for the Responses API and is adapted for use with Codex, broadening its appeal to developers building coding agents.
The Cache Discount Edge
A significant part of DeepSeek's cost advantage comes from its pricing mechanics. The company offers a 98 percent cache discount — well above the industry-standard 90 percent — meaning that repeated or predictable portions of a request are billed at a small fraction of the normal rate. Combined with fewer tokens per task, the effective cost for many workloads drops dramatically.
That pricing structure is a direct challenge to OpenAI, which has itself been cutting prices to defend market share. The result is a market where near-frontier quality is becoming a commodity, and the competition is increasingly decided by infrastructure efficiency and unit economics rather than raw capability.
A Price War With No Floor
The release intensifies what analysts now describe as a full-blown price war. OpenAI cut GPT-5.6 Luna prices by 80 percent to defend its developer base, only for DeepSeek to answer with a model that is comparably capable and still far cheaper. Each round pushes the effective cost of near-frontier intelligence lower, and the gap between the cheapest and most expensive capable models keeps narrowing.
The dynamic is unusual in software. In most markets, a 60 percent cost advantage at near-equal quality would rapidly capture share. Here, the advantage is being matched and re-matched within weeks, leaving developers with a constantly shifting leaderboard. Benchmark parity, once the exclusive province of frontier labs, is increasingly available from open-weight challengers willing to publish both their weights and their pricing.
What It Means for Developers
For developers, the practical implication is clear: capable agentic models are getting cheaper fast. A model that can autonomously navigate a terminal, write and debug code, and handle complex office tasks — all at a fraction of a cent per request — changes the economics of building AI-powered products.
The open-weight release also means teams can run the model on their own hardware, avoiding per-token API fees entirely for sensitive workloads. As DeepSeek, OpenAI, and others continue to trade blows on price and performance, the winners are the builders who now have a growing menu of capable, affordable models to choose from.
Stay Ahead of the AI Curve
The artificial intelligence landscape is moving faster than ever, and the firms that track these shifts early gain a lasting edge. For continuous, source-checked coverage of model releases, funding rounds, and policy moves, follow our latest AI news.
Read more AI news →