Featured image of post MiMo-V3 Architecture Revealed: HySparse2 Reduces Prefill Compute by 80% vs. Hybrid SWA via Two-Level KV Sharing

MiMo-V3 Architecture Revealed: HySparse2 Reduces Prefill Compute by 80% vs. Hybrid SWA via Two-Level KV Sharing

Xiaomi MiMo team pre-releases HySparse2, core component of MiMo-V3, optimizing compute and cache for long-context agent tasks.

Core Announcement

Core Announcement
Core Announcement|News screenshot

On September 23, Luo Fuli, head of Xiaomi’s MiMo large model team, pre-released the core component of MiMo-V3—the HySparse2 architecture—while 15-author technical paper became publicly available. This follows merely two days after MiMo-V2.6’s release on September 22, marking an accelerated iteration pace from Xiaomi’s LLM team.

Key facts:

  • Announcement date: September 23, 2024 (pre-release, not official launch)
  • New architecture: HySparse2
  • Version association: Core component of MiMo-V3 (not yet officially named)
  • Paper status: Publicly available; 15 authors; Luo Fuli as corresponding author
  • Weight release: Not mentioned; only paper disclosed

Architectural Innovation: Two-Level KV Sharing for Efficiency

Architectural Innovation: Two-Level KV Sharing for Efficiency
Architectural Innovation: Two-Level KV Sharing for Efficiency|News screenshot

HySparse2 addresses three key bottlenecks in multi-turn Agent workflows: excessive Prefill compute, high KV Cache consumption, and imprecise long-context retrieval. Its core is the two-level KV sharing mechanism.

Level 1: KV Bridging The model splits into Self-Decoder and Cross-Decoder components. Self-Decoder processes new long inputs, and its full-attention layer’s K/V can be directly reused by Cross-Decoder, eliminating redundant processing of existing tokens.

Level 2: KV Reuse Within Cross-Decoder, KV Cache and token selection indices from one full-attention layer are shared across multiple subsequent sparse-attention layers. This enables Prefill to exit after processing only the first half of the network.

▲Prefill compute and KV Cache comparison across architectures (Prefill FLOPs shown as percentage of Hybrid SWA baseline)

ArchitecturePrefill FLOPs (relative)KV Cache UsageAgentPPLRULER-v2 (256K context)
Hybrid SWA (MiMo-V2.6)100% (baseline)12.09GBHigher35.74
HySparse (previous gen)34%6.72GBHigher32.61
HySparse2 (MiMo-V3)20%2.69GBLowest58.45

An unexpected contrast: deleting the independent SWA branch in Cross-Decoder slightly reduced math reasoning scores but eliminated extra projection parameters and local KV Cache, achieving architectural simplification—including enabling early Prefill exit.

Token-Level Sparse Selection and Deployment Optimization

Unlike HySparse’s Block-level sparse selection (selecting contiguous token chunks), HySparse2 uses Token-level sparse selection, directly extracting relevant tokens from anywhere in the context. This better suits multi-turn Agent scenarios where useful information may be scattered across early tool returns.

Ablation studies confirm Token-level selection scores higher on RULER-v2, MRCR-v2, and GraphWalks.

For deployment, a 49-layer model example shows Prefill serving nodes need only the first 25 layers (nearly half), executing just one full-attention layer during Prefill—substantially reducing hardware requirements.

Technical Sources and Citations

Technical Sources and Citations
Technical Sources and Citations|News screenshot

The HySparse2 paper cites four DeepSeek works: DeepSeek-V2, DeepSeek-V3.2, DeepSeek-V4, and DeepSeek-V4.1-Flash. OpenAI’s gpt-oss model card and GPT-4.1 evaluation work are also referenced, demonstrating Xiaomi’s broad technical reference approach.

Deployment Recommendations

Deployment Recommendations
Deployment Recommendations|News screenshot

  • Best for teams: Deploying long-context Agent systems where inference latency and VRAM are constrained budget. Enterprises building multi-tool-call Agents suffering from context bloat should track HySparse2 developments.
  • Wait if: You require stability or production guarantees. MiMo-V3 is not yet officially launched; actual model specs, open-source plans, and real-world performance remain unconfirmed.

Final Thoughts

From MiMo-V2.6’s large-scale RL training to MiMo-V3’s architectural innovation, Xiaomi demonstrates a dual-track strategy—fast-following combined with innovation. HySparse2’s real value lies beyond numbers: it validates the feasibility of combining sparse attention with KV sharing, paving technical pathways for lighter, efficient Agent models going forward.