Core Announcement

On September 17, the Zhipu GLM team officially disclosed the first engineering implementation of Recurrent Self-Improvement (RSI) in domestic large models. This marks the first public deployment of AI self-betterment capabilities in a production environment among Chinese AI model companies.
- Technical Driver: Infra Agent powered by GLM-5.3
- Target System: GLM-5.3-Flash inference infrastructure
- Hardware Platform: A cluster comprising over 100,000 domestic chips
- Performance Outcome: End-to-end throughput increased to 3X the baseline within under two weeks
- Live Deployment: Functions as anonymous model Ox-Alpha on OpenCode and OpenRouter
- Usage Metrics: Over 62 trillion tokens processed in 6 days
- Status: Actively serving production traffic, not a prototype
Technical Details and Breakthrough
Recursive Self-Improvement (RSI) enables AI models to autonomously identify system bottlenecks, generate optimized code, perform debugging, and execute performance tuning—all as part of a closed-loop engineering workflow. In this implementation, the Infra Agent did not merely generate code fragments; it fully dominated the construction and iterative optimization of the inference service infrastructure.
A key technical nuance: the inference model being improved is GLM-5.3-Flash, while the model acting as the ‘engineer’ is GLM-5.3. This creates a reverse operation where the latter serves as a self-improvement agent for the former’s runtime environment. Within less than two weeks, the system achieved a 3X end-to-end throughput improvement over the initial baseline, with hardware utilization efficiency and per-token cost reaching parity with mainstream NVIDIA GPU systems.
This advancement moves beyond the conventional role of large models as code-generation assistants. The Infra Agent now actively participates in and drives the complete engineering lifecycle—including architecture design, performance tuning, and stability validation—without requiring manual code review intervention. Its inference scheduling logic and deployment scripts are dynamically generated and iterated by the agent itself.
Product and Platform Summary

No detailed performance comparison table was provided in the original disclosure, though per-token cost and hardware efficiency parity with NVIDIA GPUs was explicitly stated. Based on verified facts:
| Metric | Infra Agent Result | Official Benchmark |
|---|---|---|
| Hardware Platform | Over 100,000 domestic chips | Chinese-made chips, model not disclosed |
| End-to-End Throughput | 3X baseline | Achieved in under two weeks |
| Per-Token Cost | Comparator level to NVIDIA GPU | Stated in official blog |
| Hardware Efficiency | Matches NVIDIA GPU level | Stated in official blog |
| Token Volume (6 days) | >62 trillion | Via Ox-Alpha on OpenCode/OpenRouter |
Note: The specific国产芯片 (domestic chip) model and vendor remain undisclosed in the source material.
Practical Recommendations for Users
- Ready to adopt now: Mid-to-large AI application teams with strong cost sensitivity and frequent infrastructure iteration needs; organizations already operating domestic chip clusters should monitor Zhipu’s upcoming API or SDK integration paths.
- Recommended to wait: Enterprise users expecting public web dashboards or fine-grained control interfaces; teams requiring engineering specifics like chip model, driver version, or scheduling strategy details, as these remain unavailable.
The system is currently exposed as the anonymous model Ox-Alpha. GLM-5.3 and the Infra Agent interface are not yet open for direct third-party integration. Users can only consume Ox-Alpha’s output via API. Organizations planning procurement should watch for Zhipu’s future open-source release or enterprise partnership documentation.
Final Note
This engineering breakthrough confirms that AI self-improvement has transitioned from theoretical concept to active infrastructure optimization. When models no longer just write code—but directly manage the “soil” in which that code runs—the boundary of model iteration expands beyond accuracy gains to encompass operational efficiency. It signals a pivotal milestone in the maturation of large model engineering practices.
