← /pulse/glm-5-3-flash-open-weights-mit
INFOMODEL RELEASE·2026-10-07

GLM-5.3-Flash: 320B total, 18B active, MIT license

▼ WHAT HAPPENED

Z.ai published GLM-5.3 and GLM-5.3-Flash on Hugging Face on August 25, 2026. GLM-5.3-Flash is a multimodal Mixture-of-Experts model under the MIT license; its model card states 320B total parameters with 18B active. The full GLM-5.3 is listed at 753B parameters under its own `glm-5.3` license. Community GGUF conversions of GLM-5.3-Flash exist, and Ollama lists `glm-5.3-flash`.

▼ OPERATOR ANGLE

Only 18B parameters are active per token, but all 320B have to be resident, so memory rather than compute decides whether you can run it: think large unified-memory machines or multi-GPU servers, not a single consumer card. Check the license per model: Flash is MIT, the full GLM-5.3 is not.
[pulse item] · runlocalai.co/pulse/glm-5-3-flash-open-weights-mit