GLM-5.3-Flash
AvailableZ.ai's first natively multimodal GLM-5 (text, image, and video understanding in one stack), released Aug 26 2026 and stealth-tested beforehand as 'ox-alpha'. A 320B-total / 18B-active MoE (45 layers, hybrid linear + sparse attention) with a 1M-token context window and MIT-licensed open weights (zai-org/GLM-5.3-Flash). It is a distinct model from the text-only flagship GLM-5.3 (753B, $1.40/$4.40) and is priced roughly 10x cheaper on input: list $0.15 / $0.03 cached / $0.50 per Mtok, with a 50% promo through 2026-09-09. Vision sits inside the coding/agent loop (self-visual judgment) rather than as a bolted-on VL head. Self-reported vs GLM-5.2: DeepSWE 63.4 vs 46.2, AutomationBench 48.8 vs 26.2, Terminal-Bench 2.1 84.3 vs 81.0; vision self-reports include CharXiv Reasoning 89.4 and Chartography 78.0, though BabyVision 53.4 trails Gemini 3.7 Flash — all vendor numbers, unverified at launch. Thinking is always on and cannot be disabled.
Specifications
- License
- Open source · MIT
- Weights
- Downloadable
- Architecture
- Mixture-of-Experts
- Parameters
- 320B · 18B active
- Context window
- 1M tokens
- Max output
- —
- Knowledge cutoff
- —
- Price (in / out, $/M)
- $0.15 / $0.5
- Modalities
- TextVisionCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.