Key Event Summary

DeepSeek officially launched the V4.1 Flash model and implementation of the new Flash pricing at 12:00 Beijing time on September 10, 2026. The V4 Pro service will be retired on September 14 at 12:00 Beijing time—a delay from the previously announced timeline.
Key facts:
- Launch date: September 10, 2026 (Beijing time)
- New version: DeepSeek V4.1 Flash
- Pricing effective: September 10, 2026, 12:00 Beijing time
- V4 Pro retirement: September 14, 2026, 12:00 Beijing time
- Migration logic: V4 Pro requests will be automatically routed to V4.1 Flash after retirement and billed under the new pricing tier
- Weight availability: API users can now access V4.1 Flash; model is live for API integration
Performance and Pricing Details
V4.1 Flash has comprehensively surpassed V4 Pro across four metrics—performance, cost, speed, and total latency—according to multi-party internal and external testing. DeepSeek did not disclose specific benchmark figures but emphasized the “comprehensive superiority” of the new model.
The revised pricing scheme distinguishes between time periods:
- Off-peak hours (effective from September 10, 12:00 Beijing time):
- Input cache hit: ¥0.02 per 1k tokens
- Input cache miss: ¥1 per 1k tokens
- Output: ¥4 per 1k tokens
- Peak hours: Twice the off-peak rates
The DeepSeek Open Platform also sent identical notifications to users. This decommission aligns with DeepSeek’s standard product iteration pipeline: phasing out legacy models as newer, more efficient alternatives become available.
Cost-Performance Edge: A Surprising Contrast

Despite the performance enhancements, V4.1 Flash’s cache-hit input price stands at ¥0.02 per 1k tokens, among the lowest in the industry where comparable models often charge several to ten-plus yuan per 1k tokens for input. Combined with the claimed performance superiority, the unit-performance cost drops significantly—the most(counterintuitive) data point in this update.
Recommendations for Developers
- Migrate now: Developers currently using V4 Pro or evaluating DeepSeek API; V4.1 Flash delivers better performance at lower cost, with longer service lifecycle ahead.
- Wait and test: Teams with heavy V4 Pro reliance and strict cache-hit dependencies should complete migration testing before September 14 to avoid configuration mismatches during the cutover.
Final Thoughts
With V4.1 Flash, DeepSeek signals a pivot toward efficiency-focused model lifecycle management—an industry-wide trend where competitive advantage shifts from raw scale to cost-optimized delivery.
