Coin World reports:
On August 13, the official version of DeepSeek V4 Pro was updated to API. The model call name remains unchanged, but the version number has been updated to DeepSeek-V4-Pro-0813; the new version enhances Agent capabilities and supports Responses API and Codex integration.
The most eye-catching aspect remains the pricing: the cost for a million-token input with cache hit is 0.025 yuan, while a cache miss costs 3 yuan, and the output for a million tokens is 6 yuan. According to evaluations released by DeepSeek, the new version is close to Anthropic's Fable 5 in several tests. Decrypt calculated that the gap between the two in some capability metrics is about 5%, while the difference in calling prices could reach dozens of times.
These figures easily lead the discussion to a familiar phrase: domestic models have lowered prices again. However, this time, what is truly worth observing is not "how much cheaper can it get," but how application teams will reassess their calculations now that low-cost models have begun to fill in the development interface.
In the past year, cheap models have often been relegated to the end of processes: summarizing, rewriting, classifying, or handling high-frequency but non-critical customer service inquiries. Complex tasks were still assigned to more expensive models, as the latter were more reliable in tool invocation, long-task stability, and coding capabilities.
The V4 Pro emphasizes Agent and Responses API, targeting this boundary. For developers, whether a model can reliably call searches, databases, and internal tools is often more important than scoring two points higher in a single evaluation. If a task requires multiple steps, any formatting error, wrong tool selection, or loss of context in between can turn the saved token costs into retry expenses.
Thus, 6 yuan does not equate to a task costing only 6 yuan. What enterprises truly pay for also includes engineering integration, log monitoring, failure reruns, manual reviews, and data governance. However, when the base calling price drops to this level, previously "not worth automating" small tasks also begin to have economic viability. Bulk reading of contracts, organizing sales leads, generating test cases—each with low single value but occurring thousands of times daily—are precisely where low-cost models can easily make an impact.
The cache hit price is particularly noteworthy. The difference between 0.025 yuan and 3 yuan is 120 times, meaning whether a company can design an effective cache may impact the bill more than the choice of model. Stable system prompts, reusable product documentation, and standard operating procedures are all suitable for inclusion in reusable contexts; if every request reconstructs the information, the published lowest price is merely a number on a poster.
The support for Responses API and Codex integration reveals another layer of competition: model vendors not only want to sell tokens but also aim to reduce the friction of migration for developers. The closer the interface is to the workflow teams are already using, the more replacing a model resembles changing a configuration rather than rewriting an entire application.
This will give application companies stronger bargaining power. Previously, a set of Agents built around a single model's proprietary interface could easily lead to lock-in, no matter how effective; now, teams can incorporate model selection into the routing layer: using more stable models for high-risk tasks and cheaper models for bulk tasks, with automatic upgrades in case of failure. The evaluation differences between models will ultimately be converted into success rates and costs for each type of task.
However, "close to Fable 5" should still be understood with caution. Public evaluations are facets provided by vendors and do not equate to all real business scenarios being only 5% apart. Differences in codebase size, specialized Chinese materials, and tool invocation chain lengths can lead to completely different results. If enterprises want to migrate, the most reliable method is not to copy the overall rankings but to conduct small-scale blind tests using their own failure samples: the same inputs, the same tool permissions, and the same acceptance criteria, comparing success rates and complete task costs.
Therefore, the significance of this update from DeepSeek is not merely the emergence of another cheap model. It challenges the notion that "cheap models can only handle trivial tasks" and forces application teams to answer a more specific question: which tasks genuinely require the most expensive intelligence, and which tasks only need sufficiently stable, affordable, and replaceable basic capabilities?
Model prices will continue to fluctuate, and rankings will keep changing. What can truly remain in applications will not be a single evaluation sheet but a workflow that can switch suppliers at any time and accurately calculate failure costs.
This also poses more challenging requirements for model vendors. Low prices can quickly bring in call volumes but do not necessarily translate into long-term revenue; once developers can switch at low costs, service stability, throttling strategies, documentation quality, and fault response will directly affect retention. The more meaningful comparison moving forward will not be who offers the lowest unit price on a given day, but who can enable enterprises to run continuously for months while keeping price, performance, and interfaces predictable. Ultimately, model competition will return from press releases to bills and production logs.
Development teams should also view this round of price cuts as an architectural check rather than merely a procurement opportunity. If prompts, tool protocols, and business rules are all tightly bound to a single model, the savings today may turn into technical debt during migration tomorrow. Standardizing records of task inputs, model outputs, and acceptance results is essential for genuinely comparing different models. The greatest value of low prices is not to allow teams to relax cost management but to provide more real business opportunities for small-scale experiments, using results to decide whether to scale up.
Source: Geek Park August 13 Morning Report; Decrypt's report on DeepSeek V4 Pro official version; DeepSeek's public API pricing information. Cover image source: Decrypt.
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.











Today’s WEEX TradFi Daily Brief covers the cooler-than-feared July CPI print that reduced rate-hike expectations, the sharp rally in storage and AI infrastructure names, and the after-hours earnings focus on Applied Materials, helping you quickly capture stock-token trading opportunities.


















