DeepSeek Flash Price Cut, V4.1 Flash Enters Beta

DeepSeek slashed flash model pricing and launched V4.1 Flash beta on September 10, with the following key details:
- Launch time: September 10, 12:00 onward
- New version: DeepSeek V4.1 Flash (short-term beta, model ID includes “expires-on-0910”)
- Price changes: Input fell from ¥1.5 to ¥1 per million tokens (off-peak); output from ¥4.5 to ¥4; cache from ¥0.05 to ¥0.02; peak hours at 2× off-peak rates
- Access: API users only need to switch model name to “deepseek-v4.1-flash-expires-on-0910”; no base_url change required
- Limits: Single account up to 20 concurrent requests
- Weight status: Closed-beta; weights not open-sourced
Replacing the previous architecture, the beta natively supports multimodal inputs, while billing follows the V4 Flash schedule. DeepSeek states本轮测试 emphasizes capability, generation speed, and inference cost optimization.
Navier–Stokes Candidate Proof: OpenAI Claims AI Achievement

OpenAI announced on September 8 its candidate solution to the Navier–Stokes existence and smoothness problem—an official Clay Mathematics Institute Millennium Challenge—accompanying a paper and Lean formal proof. This marks a rare instance where an AI system claims to have completed a rigorous pure-mathematical proof.
Key facts:
- Proof generated by an internal model “superior to GPT-6 Astra”, coordinating ~10,000 concurrent agents
- ~88 hours from start to candidate output; followed by 17 hours of Lean formalization by GPT-6 Astra
- Total: ~2.7 million agent messages, 13 billion output tokens
- OpenAI explicitly stated no intention to claim the prize, awaiting independent mathematical review
A curveball emerged when NYU’s Tristan Buckmaster questioned potential data leakage; OpenAI admitted it “cannot entirely exclude that de-identified data contributed to model upgrades”, though no end-user data was directly accessed during-solving.
AI Agent Ecosystem Accelerates: Meta, WhatsApp, Xiaomi

AI agent commercialization is picking up pace:
- Meta Muse: Available via standalone app and WhatsApp, driven by Muse Spark; can decompose goals, draft plans, send emails, book trips, shop. For security, Meta provisions each Muse a dedicated cloud VM and a standalone Sentinel module to审查 outbound requests; sensitive actions require user confirmation
- WhatsApp third-party Agent test: Android beta allows up to 5 agents per account, each assigned a unique API key; agents may only read messages sent to the dedicated agent session and cannot join group chats; note: end-to-end encryption is disabled for agent sessions
- Xiaomi MiMo Desktop beta: Accepts multiple file formats, delegates tasks, and delivers editable output (PPT, Office, web, light game); interactivity preview within session; supports MCP for Figma control; complex jobs can split across multiple sessions
Comparative Overview

| Model Version | Billing Baseline | Input (Off-peak) | Output (Off-peak) | Multimodal | Weights Open |
|---|---|---|---|---|---|
| DeepSeek V4 Flash | Standard | ¥1.5 | ¥4.5 | No | No |
| DeepSeek V4.1 Flash | Beta | ¥1.0 | ¥4.0 | Yes | No |
Who should jump in: API developers needing low-latency multimodal inference may request V4.1 Flash beta access; enterprises will benefit from V4 Flash’s steep cost drop.
Best wait: The Navier–Stokes proof is unverified and industrially unusable for now; Muse’s launch is U.S.-only at launch.
Final Note
AI is evolving from isolated capability to system-level infrastructure: falling prices, tiered models, and agent orchestration dominate current trends. The synchronized moves by DeepSeek and OpenAI confirm that continuous improvement in reasoning cost and engineering efficiency is forcing the industry to re-calculate “usability thresholds”.
