Skip to content
Intelligence IASep 18, 2026Intelligence IA
Article

Zhipu AI launched GLM-5.3-FlashX, claiming nearly 200 tokens per second inference on ~100,000 domestic accelerators.

The base Flash model is a 320B MoE with 18B active parameters, optimized by an Infra Agent serving stack.

Généré par IA : résumés rédigés par une IA à partir des sources citées. Notre usage de l'IA

Généré par IASource: Pandaily
01

Brief source

Zhipu AI launched GLM-5.3-FlashX, claiming nearly 200 tokens per second inference on ~100,000 domestic accelerators. The base Flash model is a 320B MoE with 18B active parameters, optimized by an Infra Agent serving stack.