Skip to content
AI IntelligenceSep 18, 2026AI Intelligence
Article

Zhipu AI launched GLM-5.3-FlashX, claiming nearly 200 tokens per second inference on ~100,000 domestic accelerators.

The base Flash model is a 320B MoE with 18B active parameters, optimized by an Infra Agent serving stack.

AI-generated: summaries written by AI from the linked sources. How we use AI

AI-generatedSource: Pandaily
01

Source Brief

Zhipu AI launched GLM-5.3-FlashX, claiming nearly 200 tokens per second inference on ~100,000 domestic accelerators. The base Flash model is a 320B MoE with 18B active parameters, optimized by an Infra Agent serving stack.