Xiaomi MiMo-7B Series Open-Sourced: Small Model RL Training Surpasses 32B-Class Reasoning Performance

Xiaomi's MiMo-7B series achieves 32B-class reasoning performance through meticulously designed pretraining and posttraining strategies.

Xiaomi MiMo-7B Series正式 Open-Sourced: A Small Model Breakthrough

On May 30, 2025, Xiaomi officially released and fully open-sourced the MiMo-7B series of large language models. The series comprises four key versions: the Base model, the SFT model, the RL model trained from the Base model, and the RL model trained from the SFT model. All checkpoints are publicly available on **Hugging Face and ModelScope **. Evaluations were conducted with temperature=0.6.

A standout achievement is MiMo-7B-RL-0530’s performance on the AIME24 benchmark, which **ultimately surpassed DeepSeek R1 with a score of 79.8 **. This model, despite its compact 7-billion-parameter scale, delivers reasoning capability previously thought to require models of 32B or larger.

Technical Innovation: Synergy Between Pretraining and Posttraining

MiMo team’s core philosophy is that unlocking reasoning potential requires both pretraining-stage potential mining and **posttraining-stage strategy optimization **, challenging the industry norm of relying on large base models for RL training.

For pretraining, the team employed a three-stage data mixture strategy, training MiMo-7B-Base on approximately 25 trillion tokens. Technical enhancements include optimized text extraction pipelines, multi-dimensional data filtering to increase reasoning pattern density, and generation of massive diverse synthetic reasoning data. Notably, Multiple-Token Prediction (MTP) was introduced as an additional training objective to boost performance and accelerate inference. The MTP layers are tuned during pretraining and SFT, then frozen during RL, achieving about 90% acceptance rate in speculative decoding.

For posttraining, the team designed a novel RL approach:

  • RL training data comprises 130K curated mathematics and code problems verified by rule-based verifiers, avoiding reward hacking
  • A test difficulty driven code reward addresses sparse rewards for hard problems by assigning fine-grained scores for test cases
  • A data re-sampling strategy for easy problems enhances rollout efficiency and stabilizes policy updates in later RL phases

Dedicated RL Infrastructure: Accelerated Training and Validation

The MiMo team developed the Seamless Rollout Engine for efficient RL training, achieving 2.29× faster training and 1.96× faster validation through continuous rollout, asynchronous reward computation, and early termination. The engine also improves robustness of the vLLM inference engine in RL systems and integrates MTP support.

MiMo-7B-RL series models are officially supported by SGLang for speculative decoding acceleration with parameters like --speculative-num-draft-tokens 2.

Model Comparison and Practical Advice

The following table summarizes key features (based on publicly available data):

Model SeriesParametersRL Training Starting PointAIME24 ScoreDistinction
MiMo-7B-RL-05307BRL from Base/SFT79.8Small-model RL reaching top performance; cold-start SFT achieves strong results
DeepSeek R1--79.8 (baseline)Industry reference model
OpenAI o1-mini--matchedPerformance parity
Typical RL models32B+32B base-Mainstream RL requires large base models

Deployment Guidance:

  1. Ideal for resource-constrained teams seeking strong reasoning capabilities: use MiMo-7B base for further SFT fine-tuning, or deploy MiMo-7B-RL-0530 directly for math/code reasoning tasks.
  2. For speculative decoding acceleration, MiMo’s official vLLM 0.7.3 fork is recommended, or SGLang integration with appropriate parameters.
  3. MTP functionality requires setting num_speculative_tokens=1. Note that MTP parameters are frozen during RL training and not trainable thereafter.

In Conclusion

MiMo-7B demonstrates that carefully designed pretraining data and posttraining strategies can enable small models to rival larger ones in reasoning capability, opening a new path toward lightweight, efficient reasoning systems.