Core Announcement

Zhipu held an investor call on September 16, 2026, revealing strategic progress on deploying RSI (Recursive System Improvement) at the infrastructure layer. The system is already live, supporting GLM-5.3-Flash (the anonymous model Ox-Alpha) on domestic chip clusters:
- Deployment Scale: 100,000 domestic chips
- Real-World Impact: Processed over 62 trillion tokens in 6 days; 3.2x end-to-end throughput improvement
- Speed: Full deployment from first successful run to full traffic cutover in just two weeks
- Technical Driver: GLM-5.3-powered Infra Agent as the system foundation
Key Facts and Contrasting Stats

Zhipu overcame three major infra-level challenges. First, domestic chips inherently lag behind NVIDIA in memory bandwidth and data transfer rates, making direct replication of NVIDIA optimization paths infeasible. Second, the software stack remains immature: critical operators are missing, and low-level documentation often requires reverse engineering. Third, GLM-5.3-Flash features a new architecture supporting multimodal (text + image) processing and 1 million token context, causing explosive growth in KV cache pressure.
The most striking contrast lies in deployment cycle. Previously, similar-scale infrastructure deployments required experienced teams weeks or months. Zhipu completed the task in two weeks. Core solutions included using communication to substitute memory (intra-node tensor parallelism + layer splitting), using computation to substitute memory (ReplaySSM), using precision to substitute capacity (W8A8 quantization + dynamically mixed-precision KV caching), and decoupling encoding/prefill/decoding phases for scheduling flexibility.
Equally notable is the “sparse feedback” problem. Traditional systems only report aggregate metrics (e.g., “20% throughput drop”) without root-cause tracking. Zhipu’s “Dense Feedback” framework addresses three layers: correctness (accuracy), system behavior (bottleneck location), and performance (optimal solution), enabling AI to emulate experienced engineers’ diagnostic intuition.
Dense Feedback in Action
Case 1: Long-context precision drift. AI identified KDA kernel default precision causing rounded-error accumulation, auto-matched a solution from open-source libraries, and upgraded computation to a higher-precision combo algorithm.
Case 2: Prefill-stage concurrency bottleneck. System automatically pulled full-chain timing diagrams, cross-checked Python/C++ language boundaries for GIL lock status, pinpointed communication thread blocking, reducing troubleshooting time from days to seconds.
Case 3: Decoding kernel optimization. Leveraging an “optimization template library” (extracted from SGLang, Flash Linear Attention, DeepGEMM), AI quickly matched and applied Tiling patterns, closed-loop refined and fed improvements back into the library.
Why Infra Became Strategic Priority

As disclosed in the investor call, Zhipu’s core growth constraint is compute supply. After GLM-5’s February 2026 launch, traffic spiked 10x overnight, forcing temporary suspension of Coding Plan for six months. Per acquisition, Coding Plan resumed July 2026 after $4B funding, selling 15x more units than before—confirming the “more compute = more revenue” equation.
Zhipu’s 300B RMB investment math: 40% for model training, 60% for inference; with 80% gross margin on GLM-5.3 inference, break-even occurs within one year. Strategy includes: 1GW data center construction, domestic chip reliance, and cloud revenue-sharing (GLM open models as cloud APIs). Revenue-sharing agreements with domestic and international cloud providers are signed, with revenue recognition starting October 2026.
Reader Recommendations

- Suitable for: Enterprise clients with strong domestic compute dependencies (government, finance, telecom); use cases needing ultra-long context (1M tokens) at cost-sensitive banks/insurance companies;
- Wait first: Customers sensitive to extreme latency (e.g., high-frequency trading) or tightly bound to NVIDIA-specific toolchains (H100/V100)—monitor stability in upcoming releases.
Final Note
Zhipu’s fusion of model capability with system engineering marks a shift from raw compute wars to soft-hardware co-optimization. If RSI scaling reflex gains compound, it could redefine how China pursues AGI breakthroughs under hardware constraints.
