Featured image of post Zhipu AI Launches GLM-5.3-FlashX: Inference Speed on Domestic Chips Reaches 200 Tokens/s

Zhipu AI Launches GLM-5.3-FlashX: Inference Speed on Domestic Chips Reaches 200 Tokens/s

Zhipu AI releases FlashX variant achieving 200 tokens/s inference on domestic chips.

Core Event Summary

Zhipu AI has officially launched the GLM-5.3-FlashX, a high-speed inference variant of the GLM-5.3 series, achieving up to 200 tokens/s on domestic chip infrastructure. Prior to its formal release, the GLM-5.3 series operated anonymously as “Ox Alpha” and ranked as the most invoked model on both OpenCode and OpenRouter platforms.

  • Release Date: Around September 18, 2026
  • New Version: GLM-5.3-FlashX (accelerated variant of GLM-5.3)
  • Core Metric: Up to 200 tokens/s inference speed
  • Hardware Foundation: 100,000 Chinese-made chip units
  • License Status: Weight availability not specified in source
  • Availability: Live on OpenCode and OpenRouter platforms

Technical Details and Platform Performance

GLM-5.3-FlashX represents another significant milestone in Zhipu AI’s domestic chip infrastructure optimization. The team previously built its inference service from scratch on a 100,000-chip国产 (domestic) cluster during the last GLM deployment. This FlashX version achieved a substantial throughput jump to 200 token/s through deep engineering optimization.

A notable反差 (counterintuitive aspect) is that the anonymous Ox Alpha version ranked as the top-called model on both OpenCode and OpenRouter before official launch, indicating strong developer validation. The newly released FlashX now delivers proven high-performance inference on domestic hardware—a performance benchmark previously associated mainly with NVIDIA GPU infrastructure.

OpenCode and OpenRouter are internationally recognized platforms for open-source model testing and distribution. This means GLM-5.3-FlashX passed both domestic engineering validation and international pressure testing via real-world API calls.

Model Version Comparison (Public Information Only)

FeatureGLM-5.3-FlashXGLM-5.3 (Ox Alpha)
Inference SpeedUp to 200 tokens/sNot disclosed, slower than FlashX
Deployed PlatformsOpenCode, OpenRouterOpenCode, OpenRouter
Call Volume RankNot disclosed#1 on both platforms
Hardware Infrastructure100,000 domestic chips100,000 domestic chips

Note: Only information explicitly mentioned in the source is included.

Practical Recommendations

  • Early Adopters Should Try: Teams requiring high-throughput inference on domestic chips, especially small-to-medium businesses building API services—existing integrations on OpenCode/OpenRouter can quickly benchmark performance gains.
  • Wait-and-See if You: Need detailed model specs (context length, multilingual accuracy, task-specific metrics) not yet published; or rely on commercial API billing models pending FlashX pricing clarification.

Final Thoughts

Domestic large model inference has matured from “just functional” to “highly efficient.” GLM-5.3-FlashX’s 200 tokens/s throughput confirms that China-developed infrastructure can now support commercially viable LLM services at scale.