Core Announcement

GitHub open-source project Colibrì has recently surged in popularity, reaching 32k Stars. It enables running ultra-large MoE (Mixture of Experts) models on consumer-grade laptops without any GPU. Developed independently by JustVugg, Colibrì is a zero-dependency pure-C implementation requiring no deep learning frameworks.
- Release Status: Open-sourced; pre-built binaries for Linux/macOS/Windows available
- Model Support: 9 model families including GLM-5.2/5.3, DeepSeek V4 Flash, Qwen, Inkling, and Kimi K3
- Quantization: Uniform int4 precision
- Minimal Hardware Requirement: GLM-5.2 runs on 16GB RAM (24GB recommended); 128GB RAM setups achieve ~1.8 token/s
- Weight Availability: Framework open-source; models must be acquired separately (conversion tools provided)
Technical Innovation: Multi-tier Loading as Weight JIT

Colibrì’s breakthrough lies in decoupling model inference from memory constraints, specifically designed for MoE architectures like GLM-5.2.
MoE models have massive total parameters but activate only a small expert subset per token. Colibrì splits models into two layers:
- Resident Layer: Attention, Embedding, shared Dense components—~9.9GB int4, always in RAM
- On-Demand Layer: 19,456 routed experts—~370GB int4, stored entirely on NVMe SSD
During inference, the router identifies required experts. If not in RAM cache, they are loaded from SSD on-demand. After processing, experts are retained or evicted based on usage patterns. Developers describe this as a “JIT for model weights"—loading only hot components when needed.
To optimize I/O, Colibrì implements three-tier scheduling (VRAM/RAM/SSD):
- Frequent experts reside in faster tiers; infrequent ones stay on SSD
- LRU caching + frequency tracking for dynamic priority adjustment
- Predictive preloading: Based on 71.6% predictability of router correlation between adjacent layers, background loading occurs during computation
- Multi-SSD Parallelism: Duplicate model copies across two SSDs for throughput scaling
Key Metrics & Model Range

Colibrì continuously expands supported models, revealing counterintuitive ratios:
| Model (int4) | Total Params | Active/Token | Storage Usage | RAM Min | GPU Required |
|---|---|---|---|---|---|
| GLM-5.2 | 744B | ~40B | ~372GB | 16GB | No |
| Kimi K3 | 2.8T | 104B | ~1.6TB | 32GB | No |
Note: Specific activation parameters and hardware requirements for DeepSeek V4 Flash, Inkling, and other models were not detailed in the original source.
Surprising Contrast: Initial inference on a pure-CPU machine (12-core + 25GB RAM) achieves only 0.05–0.1 token/s with cold cache, yet software-only optimizations push speeds to 5.8–6.8 token/s with full VRAM residency (6× RTX 5090), proving that the scheduling strategy contributes orders-of-magnitude performance gains.
Practical Recommendations

- Ready to Try: Laptop users with 24GB+ RAM seeking to experiment with GLM/Qwen models without GPU; preference for zero-engine dependency and open-source tools
- Wait Longer: Users requiring >5 token/s throughput, <500GB free disk space, or only 8–16GB RAM—current setup may be suboptimal
- Pro tip: Web Dashboard and “Brain” visualization show real-time expert locations, memory tiers, and usage heatmaps
Final Thoughts
Colibrì demonstrates that next-gen LLM deployment hinges less on raw hardware and more on intelligent scheduling. The MoE + multi-tier storage path offers a concrete blueprint for accessible local inference.
