docs/source/en/model_doc/glm5_next.md
This model was contributed to Hugging Face Transformers on 2026-08-26.
The GLM-5.3-Flash model use this class. The implementation in transformers does not include an MTP layer.
GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.
GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute.
[[autodoc]] Glm5NextConfig
[[autodoc]] Glm5NextTextConfig
[[autodoc]] Glm5NextVisionConfig
[[autodoc]] Glm5NextPreTrainedModel - forward
[[autodoc]] Glm5NextTextModel - forward
[[autodoc]] Glm5NextVisionModel - forward
[[autodoc]] Glm5NextModel - forward
[[autodoc]] Glm5NextForConditionalGeneration - forward
[[autodoc]] Glm5NextProcessor
[[autodoc]] Glm5NextImageProcessor
[[autodoc]] Glm5NextImageProcessorPil
[[autodoc]] Glm5NextVideoProcessor