Featured image of post DeepSeek V4.1 Flash Officially Launched: Surpasses V4 Pro in All Metrics, Legacy Service to be Retired on September 14

DeepSeek V4.1 Flash Officially Launched: Surpasses V4 Pro in All Metrics, Legacy Service to be Retired on September 14

DeepSeek launches V4.1 Flash with revised pricing and extended V4 Pro decommission timeline.

Key Event Summary

Key Event Summary
Key Event Summary|News screenshot

DeepSeek officially launched the V4.1 Flash model and implementation of the new Flash pricing at 12:00 Beijing time on September 10, 2026. The V4 Pro service will be retired on September 14 at 12:00 Beijing time—a delay from the previously announced timeline.

Key facts:

  • Launch date: September 10, 2026 (Beijing time)
  • New version: DeepSeek V4.1 Flash
  • Pricing effective: September 10, 2026, 12:00 Beijing time
  • V4 Pro retirement: September 14, 2026, 12:00 Beijing time
  • Migration logic: V4 Pro requests will be automatically routed to V4.1 Flash after retirement and billed under the new pricing tier
  • Weight availability: API users can now access V4.1 Flash; model is live for API integration

Performance and Pricing Details

V4.1 Flash has comprehensively surpassed V4 Pro across four metrics—performance, cost, speed, and total latency—according to multi-party internal and external testing. DeepSeek did not disclose specific benchmark figures but emphasized the “comprehensive superiority” of the new model.

The revised pricing scheme distinguishes between time periods:

  • Off-peak hours (effective from September 10, 12:00 Beijing time):
    • Input cache hit: ¥0.02 per 1k tokens
    • Input cache miss: ¥1 per 1k tokens
    • Output: ¥4 per 1k tokens
  • Peak hours: Twice the off-peak rates

The DeepSeek Open Platform also sent identical notifications to users. This decommission aligns with DeepSeek’s standard product iteration pipeline: phasing out legacy models as newer, more efficient alternatives become available.

Cost-Performance Edge: A Surprising Contrast

Cost-Performance Edge: A Surprising Contrast
Cost-Performance Edge: A Surprising Contrast|News screenshot

Despite the performance enhancements, V4.1 Flash’s cache-hit input price stands at ¥0.02 per 1k tokens, among the lowest in the industry where comparable models often charge several to ten-plus yuan per 1k tokens for input. Combined with the claimed performance superiority, the unit-performance cost drops significantly—the most(counterintuitive) data point in this update.

Recommendations for Developers

  • Migrate now: Developers currently using V4 Pro or evaluating DeepSeek API; V4.1 Flash delivers better performance at lower cost, with longer service lifecycle ahead.
  • Wait and test: Teams with heavy V4 Pro reliance and strict cache-hit dependencies should complete migration testing before September 14 to avoid configuration mismatches during the cutover.

Final Thoughts

With V4.1 Flash, DeepSeek signals a pivot toward efficiency-focused model lifecycle management—an industry-wide trend where competitive advantage shifts from raw scale to cost-optimized delivery.