GLM-5.2 Turbo
AvailableA speed-optimized, hosted "Turbo" serving tier of Z.ai's GLM-5.2, surfaced Aug 17 2026 (API id glm-5.2-fast). It targets latency-sensitive coding and agent workloads at the premium fast tier and carries GLM-5.2's 1M-token context. Served through Z.ai and SCX.ai with list pricing around $1.99 input / $6.16 output per Mtok โ well above standard GLM-5.2 ($0.55/$1.78), reflecting the dedicated fast-serving tier. Z.ai has not published a separate parameter count, architecture detail, or open-weight release for the Turbo variant, so size and weights are recorded undisclosed pending confirmation; the underlying GLM-5.2 base is an open-weight (MIT) ~753B MoE. Added this sweep to resolve a catalog gap prior runs flagged for lacking a solid spec source (now sourced from the LLM Gateway model page); fields beyond context, pricing, and providers are conservative and unverified.
Specifications
- License
- Proprietary
- Weights
- Not released
- Architecture
- Mixture-of-Experts
- Parameters
- Undisclosed
- Context window
- 1M tokens
- Max output
- โ
- Knowledge cutoff
- โ
- Price (in / out, $/M)
- $1.99 / $6.16
- Modalities
- TextCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.