Featured image of post Xiaomi Unveils HySparse2 Architecture for MiMo-V3: Designed to Improve Long-Context Agent Inference with 1/5 Compute

Xiaomi Unveils HySparse2 Architecture for MiMo-V3: Designed to Improve Long-Context Agent Inference with 1/5 Compute

Xiaomi's new long-context attention architecture cuts Prefill compute by 80% and KV Cache to 2.7GB

Xiaomi Unveils HySparse2 Architecture for MiMo-V3: Designed to Improve Long-Context Agent Inference

Xiaomi Unveils HySparse2 Architecture for MiMo-V3: Designed to Improve Long-Context Agent Inference
Xiaomi Unveils HySparse2 Architecture for MiMo-V3: Designed to Improve Long-Context Agent Inference|News screenshot

Xiaomi officially unveiled HySparse2, the core architectural framework for MiMo-V3, on September 24. Designed specifically for long-context, multi-turn Agent scenarios, this architecture addresses performance bottlenecks in processing web pages, files, code, and tool call records by fundamentally rethinking attention mechanisms.

Key specifications:

  • Release date: September 24, 2024
  • Target model: MiMo-V3 (based on 80B-A3B MoE foundation)
  • Key optimizations: Prefill compute reduced to 1/5, KV Cache shrunk from 12GB to 2.7GB
  • Predecessor: HySparse was already applied to MiMo-V2 series
  • Weight status: Architecture design is public; specific model weights not specified

Two-Level KV Sharing: Enabling Prefill Early Exit

HySparse2’s core innovation is its two-level KV sharing design, inspired by YOCO’s architecture, dividing the model into Self-Decoder (first half) and Cross-Decoder (second half).

KV Bridging enables提前 KV Cache construction across halves: The KV Cache for the second half’s Full Attention layers no longer waits for input to pass through layer by layer, but is directly generated from the hidden states of corresponding layers in the first half. Each target layer maintains independent K/V projections to construct different KV caches. This makes KV construction info available during the Self-Decoder phase.

KV Reuse复用 KV and selection results within the same Hybrid Block: The Full Attention layer selects important positions upon completion, and subsequent Sparse Attention layers directly reuse its KV Cache and selection indices, avoiding separate training of an independent selector.

Token-Level Sparse Selection: Precision and Efficiency Balance

Token-Level Sparse Selection: Precision and Efficiency Balance
Token-Level Sparse Selection: Precision and Efficiency Balance|News screenshot

While the first-generation HySparse used block-level sparse selection, HySparse2 upgrades to token-level selection. This change allows the same attention budget to be allocated more precisely to critical information points scattered across different rounds.

The design retains a local window forcing the most recent 128 tokens to be selected, plus an additional 1,024 global tokens. Since local and global information jointly use the shared KV Cache prepared during Self-Decoder, sparse layers no longer require an independent SWA branch or local KV Cache,解除 key computation dependency.

Ablation experiments prove the local window maintains competitiveness on multiple long-text tasks while saving projection parameters and KV Cache overhead of the independent branch.

Performance Metrics: Dual Breakthrough in Cost and Efficiency

ArchitecturePrefill Compute (80B-A3B MoE)KV CacheMRCR-v2 (avg gain)RULER-v2 (avg gain)
Hybrid SWA1× (baseline)12GBbaselinebaseline
HySparse1/36.7GB+11.30 pts+19.81 pts
HySparse21/52.7GBhigherhigher

In million-token cost analysis, HySparse2 reduces Prefill compute to 1/5 and KV Cache to 2.7GB versus Hybrid SWA; versus first-gen HySparse, Prefill compute drops to 1/3.

Post-pretraining, HySparse2 maintains comparable general capabilities and shows improved scores across all tested lengths up to 256k context: higher MRCR-v2 and RULER-v2, lower AgentPPL and LongPPL, demonstrating simultaneous improvements in multi-round Agent trajectory modeling and long-range information prediction.

Practical Application Guidance

Practical Application Guidance
Practical Application Guidance|News screenshot

Worth trying now: Agent projects requiring long document retrieval and multi-turn tool calls can significantly reduce deployment costs through HySparse2’s efficient Prefill and small KV Cache.

Wait for now: Those prioritizing real-time response should await YOCOHySparse2’s system-level synergy with MiMo-UltraSpeed, as these remain parallel explorations per the description.

Final Thoughts

HySparse2’s breakthrough proves that less KV Cache and shorter Prefill can coexist with better long-text retrieval. As model scales swell, this architectural innovation opens new efficiency pathways for practical long-context Agent deployment.