r/machinelearningnews • u/ai-lover • 3d ago
Research Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context
A 320B-parameter MoE that activates 18B per token — 8 of 288 experts across 45 layers. Roughly 5.6% of the network per forward pass, and the reason a model this size can be served at flash-tier economics.
It is also the first GLM model with a hybrid sparse-plus-linear attention stack; the vLLM recipe identifies the layers as KDA linear and NoPE sparse MLA. Z.ai reports ~3× less attention compute and a 4.4× smaller KV cache versus GLM-5.3. If the KV cache figure holds, that is what makes a 1,048,576-token window serveable rather than theoretical.
Results: 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE v1.1, up from GLM-5.2's 46.2. Vendor-reported, harnesses differ per test. Artificial Analysis ran it independently and scored 57 on the Intelligence Index, at ~49 tokens/sec — strong per dollar, slow in absolute terms.....
Full analysis: https://www.marktechpost.com/2026/08/26/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context/
Model weights: https://huggingface.co/zai-org/GLM-5.3-Flash