BTC/USD $68,420 +2.8%
ETH/USD $3,540 +1.4%
SOL/USD $142.80 -0.6%
BNB/USD $605.20 +0.9%
XRP/USD $0.62 -1.2%
DOGE/USD $0.18 +5.4%
BTC/USD $68,420 +2.8%
ETH/USD $3,540 +1.4%
SOL/USD $142.80 -0.6%
BNB/USD $605.20 +0.9%
XRP/USD $0.62 -1.2%
DOGE/USD $0.18 +5.4%
Markets

DeepSeek’s Cheap Model Just Beat Its Own Flagship on Nine Benchmarks

DeepSeek moved its V4-Flash model out of preview and into official public beta on July 31, 2026, and the upgraded version now outscores the company’s own larger, more expensive V4-Pro-Preview

AnonymousCryptoCompass newsroom
August 2, 2026
4 min read
NEWS
DeepSeek’s Cheap Model Just Beat Its Own Flagship on Nine Benchmarks
CryptoCompass editorial visual for markets coverage.

DeepSeek moved its V4-Flash model out of preview and into official public beta on July 31, 2026, and the upgraded version now outscores the company’s own larger, more expensive V4-Pro-Preview model across every agent benchmark DeepSeek published. The catch: it’s the same architecture as before. DeepSeek only retrained it.

V4-Flash-0731 keeps the identical 284-billion-parameter mixture-of-experts design and 13-billion active parameters as the preview build — DeepSeek’s changelog is explicit that this is a post-training upgrade, not a new model. On Terminal-Bench 2.1, a measure of complex command-line agent work, it scored 82.7, up from 61.8 for the preview and ahead of V4-Pro-Preview’s 72.1, landing within a few points of Anthropic’s Claude Opus 4.8 at 85.0, according to a detailed benchmark breakdown. DeepSWE jumped from 7.3 to 54.4, and Cybergym rose from 38.7 to 76.7.

The comparison DeepSeek wants you to make, and the one it doesn’t

DeepSeek published a nine-benchmark comparison against Opus 4.8. Opus 4.8 leads on all nine, though several gaps are narrow — Agents’ Last Exam sits at 25.2 versus 25.7, essentially a rounding difference. Independent evaluator Artificial Analysis measured the pre-0731 endpoint at 61.8 on Terminal-Bench 2.1 in a July 27 snapshot, matching DeepSeek’s own preview figure to the decimal — solid independent confirmation of the baseline, even if the eye-catching 25.8-point jump some outlets cited compares two different benchmark versions rather than a clean before-and-after.

Pricing that undercuts everyone, including OpenAI’s new cuts

V4-Flash-0731 costs $0.14 per million input tokens on a cache miss (just $0.0028 on a cache hit) and $0.28 per million output tokens — pricing unchanged from the preview, and dramatically below even OpenAI’s newly discounted GPT-5.6 Luna at $0.20/$1.20. The timing is hard to ignore: DeepSeek shipped this release one day after OpenAI cut Luna’s price 80%, positioning V4-Flash-0731 as an immediate counter-move in the same price war rather than a coincidence.

What’s still unverified

As of release, the -0731 weights had not yet appeared on Hugging Face, and DeepSeek’s announced peak/off-peak pricing tiers weren’t live — meaning some details in the official changelog describe a commitment rather than a currently usable fact. Flash also supports 2,500 concurrent requests versus Pro’s 500, a fivefold headroom advantage that likely matters more to production agent fleets than any single benchmark score.

The migration cost for existing users is effectively zero: developers already calling the deepseek-v4-flash endpoint get the upgraded checkpoint automatically, under the same model name and API key, with no code changes required. DeepSeek framed this as a deliberate design choice — since the architecture didn’t change, only the post-training pass, there was no technical reason to force a version bump that would fragment the user base across old and new endpoints.

What to watch next

  • Whether the -0731 weights land on Hugging Face as promised, enabling self-hosted deployment.
  • Whether DeepSeek’s separate V4-Pro official release, still pending, closes more of the gap with Opus 4.8.
  • Whether rival labs respond with their own retrained, price-unchanged upgrades rather than net-new model launches.

Sources

Disclaimer: This content is meant to inform and should not be considered financial advice. The views expressed in this article may include the author’s personal opinions and do not represent Times Tabloid’s opinion. Readers are advised to conduct thorough research before making any investment decisions. Any action taken by the reader is strictly at their own risk. Times Tabloid is not responsible for any financial losses.

The post DeepSeek’s Cheap Model Just Beat Its Own Flagship on Nine Benchmarks appeared first on Times Tabloid.