Xiaomi Unveils HySparse2 for MiMo-V3: Two-Tier KV Sharing, Half Prefill Compute
Fuli Luo, head of Xiaomi’s MiMo team, recently announced the new inference architecture HySparse2 for MiMo-V3 on X platform, followed by a detailed release via the official WeChat account. Designed specifically for long-context multi-turn Agent scenarios, HySparse2 targets three direct objectives: reduced Prefill computation, smaller KV Cache footprint, and improved long-context retrieval accuracy.
Key specifications:
- Release status: HySparse2 architecture confirmed, specific launch timeline not disclosed
- Primary use case: Long-context multi-turn Agent interactions
- Foundation: Built on MiMo-V2.6 Hybrid SWA with established relative performance coefficients
Technical Breakthroughs: KV Sharing and Prefill Optimization
The core innovation of HySparse2 lies in its upgraded KV Cache handling. Compared to MiMo-V2.6, the KV sharing mechanism has been elevated to two tiers, significantly enhancing cache utilization efficiency. In million-token long-context scenarios, the architecture achieves approximately 50% reduction in Prefill computation—dramatically lowering the initial processing burden during model inference.
A surprising metric: Despite halving Prefill compute workload and shrinking KV Cache size, long-context retrieval recall accuracy actually improves. This contradicts the industry assumption that “cache compression inevitably sacrifices precision,” validating the effectiveness of HySparse2’s architectural design.
Evolution Path
MiMo-V3 is not a ground-up rebuild but an evolution of the existing V2.6 Hybrid SWA stack. Hybrid SWA (Hybrid Sliding Window Attention) combines attention range adjustments to balance accuracy and efficiency. HySparse2’s two-tier sharing allows different attention layers to reuse partial key-value information, maintaining inter-layer differentiation while preventing semantic confusion that full sharing would cause.
KV Cache stores attention key-value pairs during inference, growing linearly with context length—a major bottleneck for long text processing. By sharing cache across layers, HySparse2 delivers performance gains without additional computational resources.
Deployment Recommendations
Consider immediate adoption:
- Enterprises developing long-dialogue Agent applications handling 100k+ token conversation history
- Model service providers sensitive to inference latency and able to accommodate initial adaptation period
Best to wait:
- Existing MiMo-V2.6 users with context lengths under 100k tokens—upgrade ROI limited
- Production systems requiring extreme stability that cannot absorb adjustments during early version adoption
Final Thoughts
HySparse2’s two-tier KV sharing delivers fresh insights for long-context models, proving that efficiency gains need not compromise quality. As Agent applications demand more memory capacity, efficient cache management is poised to become a critical frontier in large model engineering.
