Skip to content
AI情报Sep 18, 2026AI情报
文章

Zhipu AI launched GLM-5.3-FlashX, claiming nearly 200 tokens per second inference on ~100,000 domestic accelerators.

The base Flash model is a 320B MoE with 18B active parameters, optimized by an Infra Agent serving stack.

AI 生成:摘要由 AI 根据所链接的来源撰写。 我们如何使用 AI

AI 生成来源: Pandaily
01

来源简报

Zhipu AI launched GLM-5.3-FlashX, claiming nearly 200 tokens per second inference on ~100,000 domestic accelerators. The base Flash model is a 320B MoE with 18B active parameters, optimized by an Infra Agent serving stack.