Lab release history
Last updated Aug 26, 2026
Z.ai (Zhipu AI) model releases
Tsinghua-spun lab behind the GLM family, rebranded internationally as Z.ai. This page collects the lab's model releases, lifecycle events, source links, and model metadata in one crawlable record.
18 models
GLM-5.3-Flash
AvailableZ.ai's first natively multimodal GLM-5 (text, image, and video understanding in one stack), released Aug 26 2026 and stealth-tested beforehand as 'ox-alpha'. A 320B-total / 18B-active MoE (45 layers, hybrid linear + sparse attention) with a 1M-token context window and MIT-licensed open weights (zai-org/GLM-5.3-Flash). It is a distinct model from the text-only flagship GLM-5.3 (753B, $1.40/$4.40) and is priced roughly 10x cheaper on input: list $0.15 / $0.03 cached / $0.50 per Mtok, with a 50% promo through 2026-09-09. Vision sits inside the coding/agent loop (self-visual judgment) rather than as a bolted-on VL head. Self-reported vs GLM-5.2: DeepSWE 63.4 vs 46.2, AutomationBench 48.8 vs 26.2, Terminal-Bench 2.1 84.3 vs 81.0; vision self-reports include CharXiv Reasoning 89.4 and Chartography 78.0, though BabyVision 53.4 trails Gemini 3.7 Flash — all vendor numbers, unverified at launch. Thinking is always on and cannot be disabled.
GLM-5.2 Turbo
AvailableA speed-optimized, hosted "Turbo" serving tier of Z.ai's GLM-5.2, surfaced Aug 17 2026 (API id glm-5.2-fast). It targets latency-sensitive coding and agent workloads at the premium fast tier and carries GLM-5.2's 1M-token context. Served through Z.ai and SCX.ai with list pricing around $1.99 input / $6.16 output per Mtok — well above standard GLM-5.2 ($0.55/$1.78), reflecting the dedicated fast-serving tier. Z.ai has not published a separate parameter count, architecture detail, or open-weight release for the Turbo variant, so size and weights are recorded undisclosed pending confirmation; the underlying GLM-5.2 base is an open-weight (MIT) ~753B MoE. Added this sweep to resolve a catalog gap prior runs flagged for lacking a solid spec source (now sourced from the LLM Gateway model page); fields beyond context, pricing, and providers are conservative and unverified.
GLM-5.3
AvailableZ.ai's 2026-08-14 coding model, pitched as the strongest open-weights coder on the market. At 743B parameters, it reuses the GLM-5.2 base model with all gains coming from expanded post-training ("more environments, more diverse tasks, more compute"), and keeps the 1M-token context and 128K output ceiling while consuming far fewer tokens per task. Live at launch through the GLM Coding Plan subscription and ZCode, with API access following. Open weights shipped 2026-08-28 on Hugging Face (zai-org/GLM-5.3) after a roughly two-week safety review that Z.ai attributed to unexpectedly strong multi-stage exploit-chaining behavior surfaced during evaluation (2,436 vulnerabilities found across 269 open-source projects). Unlike GLM-5.3-Flash's plain MIT, the flagship weights carry a bespoke "GLM-5.3 License": MIT-equivalent grants (use, modify, distribute, sell without restriction) for most users, with one divergence — companies whose aggregate revenue exceeds $10B over any consecutive 12 months must pass Z.AI's security review before using the weights or derivatives commercially as a Model-as-a-Service.
Z.ai Fable-class model
RumoredA speculative Z.ai frontier model tracked after Z.ai founder Jie Tang responded to Elon Musk's prediction of a Chinese Fable 5-class model by saying it would not take that long. Name, architecture, weights, and launch timing remain unconfirmed.
GLM-5.2
AvailableZ.ai's latest open flagship for long-horizon coding, agentic engineering, and million-token workflows, adding IndexShare sparse-attention reuse over GLM-5.1.
GLM-5.1
AvailableZ.ai agentic-engineering follow-up to GLM-5, with stronger coding performance and better long-horizon tool-use behavior.
GLM-5V-Turbo
AvailableZ.ai's native-multimodal vision agent: the first GLM model designed from the start as a multimodal agent, taking image, video, and text input and producing agent-oriented output (tool calling, task decomposition, and GUI interaction). Served via API with a ~203K-token context.
GLM-5
AvailableZ.ai flagship for complex systems engineering and long-horizon agentic tasks, scaling the GLM line to 744B total / 40B active parameters.
GLM-4.7
AvailableCoding-focused GLM release with improved multilingual agentic coding, terminal tasks, tool use, and interface generation.
GLM-4.6V
AvailableOpen 106B-class vision-language model with native multimodal function calling for visual agents.
GLM-4.6
AvailableAgentic reasoning and coding upgrade over GLM-4.5, expanding the text context window from 128K to 200K tokens.
GLM-4.5V
AvailableVision-language GLM based on GLM-4.5-Air, covering image, video, document, grounding, and GUI-agent tasks.
GLM-4.5
AvailableOpen agentic, reasoning, and coding foundation model that marked Z.ai international rebrand and MIT-licensed GLM push.
GLM-4.5-Air
AvailableCompact GLM-4.5 companion with 106B total / 12B active parameters for efficient agentic reasoning and coding.
GLM-4-9B
AvailableOpen GLM-4 9B model family, covering chat, long-context, and code-oriented variants.
GLM-4
AvailableZhipu's GLM-4 flagship generation, launched as the successor to ChatGLM3 with stronger tool use and multimodal variants.
ChatGLM2-6B
AvailableSecond open ChatGLM generation, improving long context, inference efficiency, and bilingual chat quality.
ChatGLM-6B
AvailableZhipu AI and Tsinghua KEG's first widely used open bilingual ChatGLM checkpoint.