GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...
Details
| Field | Value |
|---|---|
| id | z-ai/glm-5.3-flashx |
| family | glm |
| modality | text+image+video->text |
| context_length | 1048576 |
| max_output_tokens | 131072 |
| release_date | 2026-09-18 |
| is_open_weights | no |
| supports_tool_call | yes |
| supports_reasoning | yes |
| supports_structured_output | no |
| supports_attachment | yes |
| input_modalities | text, image, video |
| output_modalities | text |