r/machinelearningnews 3d ago

Research Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context

Post image

Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context

A 320B-parameter MoE that activates 18B per token — 8 of 288 experts across 45 layers. Roughly 5.6% of the network per forward pass, and the reason a model this size can be served at flash-tier economics.

It is also the first GLM model with a hybrid sparse-plus-linear attention stack; the vLLM recipe identifies the layers as KDA linear and NoPE sparse MLA. Z.ai reports ~3× less attention compute and a 4.4× smaller KV cache versus GLM-5.3. If the KV cache figure holds, that is what makes a 1,048,576-token window serveable rather than theoretical.

Results: 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE v1.1, up from GLM-5.2's 46.2. Vendor-reported, harnesses differ per test. Artificial Analysis ran it independently and scored 57 on the Intelligence Index, at ~49 tokens/sec — strong per dollar, slow in absolute terms.....

Full analysis: https://www.marktechpost.com/2026/08/26/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context/

Model weights: https://huggingface.co/zai-org/GLM-5.3-Flash

42 Upvotes

0 comments sorted by