Key Event Snapshot

- Launch time: Anthropic and OpenAI both unveiled new models on September 24, 2026
- New versions: Claude Opus 5.5 (Anthropic), GPT-6 Sol & Luna (OpenAI)
- Pricing changes: GPT-6 Sol charges $2/M input, $10/M output; Claude Opus 5.5 reduces cost by 40%
- Availability: Not specified in source material
- Weight openness: Not mentioned
GPT-6 Sol effectively serves as a “shadow replacement” for GPT-5.6 Terra: OpenAI removed the mid-tier Terra档位, positioning Sol to inherit Terra’s performance tier and pricing range.
Tier Realignment and Pricing Mechanics

OpenAI’s previous tiers were: flagship Sol, mid-tier Terra, lightweight Luna; now Astra, Sol, Luna. The key discrepancy—GPT-6 Sol and GPT-5.6 Terra show nearly identical capabilities. Terra billed $2/M input and $12/M output, while Sol charges $2/M input and $10/M output. Only output price dropped 17%; input remained unchanged.
| Model | Input ($/M) | Output ($/M) | vs. Previous |
|---|---|---|---|
| GPT-5.6 Terra | 2 | 12 | — |
| GPT-6 Sol | 2 | 10 | Output -17%, input flat |
| Claude Opus 5.5 | N/A | N/A | Cost down 40% |
Compression and caching form the real substance of OpenAI’s price cut: Caching-hits input drops to $0.20/M (10% of original), with cache validity extended to 30 minutes. Pre-warmed caching allows pre-computing KV cache for system prompts, tool definitions, and reference docs before requests arrive. Real-world tests show enterprise Agent scenarios achieving 90%+ stable cache-hit rates, up from 85%.
Competitive Landscape: Three Distinct Paths

Anthropic: Authentic discount, strong performance Claude Opus 5.5 targets daily tasks with near-Fable 5.1 capabilities, only lagging in deep logical reasoning. Savings come not just from lower list price, but 60% reduction in “thinking chain” overhead for long tasks—developers can now afford frequent high-level推理 calls.
DeepSeek: Extreme compression, off-peak advantage V4.1-Flash runs on a 552B MoE architecture yet activates only ~16B parameters. Through YOCO’s “cache once”, CSA2 compressed sparse attention, and FP4 quantization, KV cache reaches ~890 bytes per token (2% of standard Transformer). However, low pricing requires off-peak usage and high prefix repetition.
OpenAI: Pragmatic refactoring, rule-based control Avoiding fine-grained MoE to preserve training stability, OpenAI employs parameter distillation and sparse reconstruction. Over-length contexts (>272K tokens) face automatic 2× input and 1.5× output pricing, revealing efficiency limitations in long-context processing.
Practical Guidance

- Best to adopt now: Enterprises with short prompts (≤272K tokens) and high-repetition workflows (e.g., Copilot coding, fixed-structure Agents)—cache reuse can slash costs by ~50%
- Worth waiting: Applications requiring super-long contexts (>500K) or budgeted off-peak usage—DeepSeek’s time-offset pricing remains more economical; OpenAI’s over-length surcharges may offset actual costs
Final Thoughts
With the “Li Sheng” successor race heating up, OpenAI’s “cache economy” is less about raw efficiency gains and more about ecological lock-in: the 10% cache-hit discount creates significant migration friction. Technologically it may trail DeepSeek’s transparent pricing—yet commercially, OpenAI’s strategy proves more sophisticated than either Anthropic’s pure-performance discount or DeepSeek’s technical elegance.
