DeepSeek Tests V4.1 Flash Internally, Probing Full Replacement of Pro Version
DeepSeek launched a limited internal test of an intermediate V4.1 Flash checkpoint in early September 2026 and explicitly asked participants in its feedback form whether this new model could fully replace the current online V4 Pro version. Concurrently, the company announced a price cut for the Flash series, effective Beijing time September 10 at noon. A successful test would enable DeepSeek to restructure its model deployment strategy: migrating baseline and mid-complexity workloads to the lower-cost Flash tier, while reserving Pro for genuinely hard problems.
Core Information Summary

- Test launch time: Early September 2026, with the test model name valid until September 10
- New version: V4.1 Flash intermediate checkpoint—not the final commercial release
- Key features: New architecture, native multimodal support, improved capability, speed, and cost efficiency over prior Flash builds
- Pricing adjustment: Effective 12:00 Beijing time on September 10; testing usage billed at existing V4 Flash rates
- Usage limits: Maximum 20 concurrent requests per account during the test
- Weight release: The source material does not indicate whether weights are open-sourced
Technical Details and Performance Improvements

Internal testing has shown V4.1 Flash delivering notably faster generation speed, with early developer reports emphasizing significantly improved performance on coding and retrieval tasks. Technically, the new version introduces a revised structure and native multimodal support—meaning the model processes images, text, and other inputs directly without requiring separate adapter layers. Although the test is restricted to internal users, early feedback contains clear positive reports on performance gains.
In tandem, DeepSeek announced the Flash series price adjustments. The largest reduction—60%—applies to cache-hit input tokens. Specific rate changes are:
| Metric | Cache-hit input | Cache-miss input | Output |
|---|---|---|---|
| Off-peak rate (CNY per million tokens) | 0.02 | 1.00 | 4.00 |
| Peak rate (CNY per million tokens) | 0.04 | 2.00 | 8.00 |
These rates apply to both the main Flash model and the vision-experimental variant. Notably, the cache-hit input price for Flash has dropped to 0.02 yuan per million tokens, creating a historically substantial price gap compared to alternative tiers—Flash input is now just 2% of the former peak output price.
The Economics of Intelligence Density

DeepSeek repeatedly emphasizes “intelligence density”—the capacity to deliver effective answers faster and with fewer intermediate computation steps—within fewer tokens and less wall-clock time. For Agent-style workloads, a single user request may trigger multiple model calls: file reading, tool invocation, intermediate result validation, and multi-step progression. At such scale, low per-token pricing does not guarantee low task cost; step inflation can easily offset token savings. DeepSeek’s open-source Harness framework operates at this layer: the model makes decisions, while Harness provides tool integration, session state management, and execution environments.
Industry peers exhibit similar recalibration strategies: when a mid-tier model approaches the performance of a previous flagship, the premium tier must justify its cost by handling longer-horizon, higher-stakes tasks. Flagship models increasingly compete on the ability to sustain verifiable multi-hour (or multi-day) agentic workflows rather than on individual benchmark scores.
Product Suitability Guidance
- Choose Flash if you: Run routine reasoning, code generation, document retrieval, or mid-complexity dialogue tasks; prioritize low latency and cost-per-inference
- Consider waiting or sticking with Pro if: You require strictly auditable multi-step decision support, ultra-long-context consistency, or traceable reasoning chains in regulated domains (e.g., financial modeling, legal argumentation)
Final note: DeepSeek’s move signals a strategic pivot from “capability accessibility” to “efficiency affordability.” As strong reasoning becomes widely available, the next competitive axis is intelligence density—the amount of effective intelligence packed into each millisecond of latency and each unit of cost. Whether Flash can truly absorb the mainstream workload may reshape pricing tiers across both open and closed model ecosystems.
