Z.ai’s GLM (General Language Model) family began with language-model pretraining research at Tsinghua University. GLM-130B expanded that work into a large bilingual model, and the ChatGLM releases of 2023 made Chinese-English conversation models available for local deployment.
Later GLM generations introduced tool use, reasoning, coding, image and speech models, and long-running agent tasks. This release timeline follows the major models from the original 2021 research through GLM-5.3 and GLM-5.3-Flash, with their launch dates and defining changes.
Last updated: October 9, 2026.
GLM model release timeline
2026: GLM-5 and native multimodal models
| Date | Milestone |
|---|---|
| Aug. 26, 2026 | GLM-5.3-Flash launches with open weights and native multimodal understanding. This newly trained mixture-of-experts model has 320 billion total and 18 billion active parameters, uses hybrid sparse and linear attention, and supports a one-million-token context window. |
| Aug. 25, 2026 | GLM-5.3 model weights are published on Hugging Face and ModelScope after the earlier security review. The release uses the GLM-5.3 model license, distinct from the MIT license of GLM-5.2 and GLM-5.3-Flash. |
| Aug. 14, 2026 | GLM-5.3 is announced with stronger complex coding, long-running agent tasks, and vulnerability analysis. It retains the GLM-5.2 base model and improves it through post-training. Model weights are withheld at announcement pending additional security evaluation. |
| June 16, 2026 | GLM-5.2 launches with a one-million-token context window and adjustable reasoning effort. Its IndexShare architecture shares an indexer across groups of sparse-attention layers to reduce computation on long inputs. |
| Apr. 7, 2026 | GLM-5.1 launches with stronger coding and sustained task execution. Its post-training emphasizes repeated planning, experimentation, and revision across extended tool-use sessions. Model weights are available for local deployment. |
| Apr. 2, 2026 | The GLM-V team announces GLM-5V-Turbo, a hosted multimodal model for visual coding and agent tasks. It accepts images, video, and text, and combines visual perception with planning and tool use. |
| Mar. 15, 2026 | GLM-5-Turbo is announced as a hosted model optimized for OpenClaw workflows. Training targets tool invocation, instruction following, and persistent tasks across long sequences of actions. |
| Feb. 12, 2026 | GLM-5 launches for complex software engineering and long-running agent tasks. The open-weight mixture-of-experts model has 744 billion total parameters and 40 billion active parameters, and incorporates DeepSeek Sparse Attention. |
| Feb. 3, 2026 | GLM-OCR launches as a compact document-recognition model. Its CogViT visual encoder and GLM language decoder extract text, tables, and formulas from document images, with weights available for local use. |
| Jan. 19, 2026 | GLM-4.7-Flash launches as a smaller open-weight reasoning and coding model. Its mixture-of-experts architecture uses 30 billion total parameters and 3 billion active parameters to reduce deployment requirements. |
| Jan. 14, 2026 | GLM-Image launches with publicly available weights for image generation. It combines an autoregressive model with a diffusion decoder and emphasizes text rendering and knowledge-intensive images, including diagrams and educational graphics. |
2025: Reasoning, coding, and open-weight expansion
| Date | Milestone |
|---|---|
| Dec. 22, 2025 | GLM-4.7 launches with improvements in multilingual coding, terminal tasks, tool use, and mathematical reasoning. Its thinking controls support reasoning before individual actions and retaining reasoning across turns in extended agent workflows. |
| Dec. 11, 2025 | GLM-TTS is released with open model weights and inference code. The text-to-speech model supports zero-shot voice cloning, expressive speech, and streaming generation. |
| Dec. 10, 2025 | GLM-ASR-2512 launches for speech recognition. The GLM-ASR-Nano-2512 weights extend the GLM audio line with a compact transcription model for Mandarin, English, and Cantonese, including accented and noisy speech. |
| Dec. 8, 2025 | GLM-4.6V and GLM-4.6V-Flash launch with open weights at 106 billion and 9 billion parameters. The vision-language models support a 128K-token context and native multimodal tool calling for images, screenshots, and document pages. |
| Sept. 30, 2025 | GLM-4.6 launches as the successor to GLM-4.5. It expands the context window from 128K to 200K tokens and improves coding, reasoning, and tool use. Weights are available for local deployment. |
| Aug. 11, 2025 | GLM-4.5V launches as an open-weight vision-language model based on GLM-4.5-Air. It supports visual reasoning, video understanding, object localization, and graphical interface tasks, with selectable thinking and direct-response modes. |
| July 28, 2025 | GLM-4.5 and GLM-4.5-Air launch with open weights. These mixture-of-experts models combine reasoning, coding, and agent capabilities, with 355 billion and 106 billion total parameters. Both support thinking and non-thinking modes. |
| July 2, 2025 | GLM-4.1V-9B-Thinking receives publicly available weights, following its technical report on July 1. The 9-billion-parameter vision-language model uses reinforcement learning to improve reasoning about images, documents, video, and graphical interfaces. |
| Apr. 14, 2025 | The GLM-4-32B-0414 series introduces larger open GLM-4 models for dialogue and tool use. Related GLM-Z1 reasoning models add extended thinking for mathematics, code, and logic. GLM-Z1-Rumination explores deeper reasoning on open-ended tasks. |
| Jan. 8, 2025 | GLM-Realtime is introduced as an end-to-end real-time audio model with interactive speech and function-calling capabilities. |
2024: GLM-4, open models, and speech
| Date | Milestone |
|---|---|
| Dec. 20, 2024 | GLM-Zero-Preview becomes available as a reasoning-model preview through ChatGLM and the BigModel API. Extended reinforcement learning targets mathematical reasoning, code, and complex problems. |
| Oct. 25, 2024 | GLM-4-Voice is released with open weights for spoken dialogue in Mandarin and English. Built on GLM-4-9B, it understands and generates speech and supports instructions that change speaking style, emotion, and speed. |
| Aug. 12, 2024 | Zhipu AI introduces GLM-4-Plus and makes the model available through its API. The update improves instruction following and long-text processing. GLM-4V-Plus is also announced for image and video understanding. |
| June 5, 2024 | The GLM-4-9B family is released with open weights, alongside GLM-4V-9B for image understanding. The language releases include a 128K-context chat model and an experimental one-million-token chat variant. |
| Jan. 16, 2024 | GLM-4 becomes available through the BigModel API. GLM-4 All Tools is accessible through ChatGLM services, with model support for selecting and using a browser, Python interpreter, image-generation tool, and user-defined functions. |
2023: Three ChatGLM generations
| Date | Milestone |
|---|---|
| Oct. 27, 2023 | ChatGLM3-6B is released with open weights. A new prompt format adds native function calling, code execution, and agent-task support. The release also includes a base model and a separate 32K-context chat model. |
| June 25, 2023 | ChatGLM2-6B is released as the second open ChatGLM generation. It extends the context length from 2K to 32K tokens and uses multi-query attention for more efficient inference, alongside improvements in reasoning and coding. |
| Mar. 14, 2023 | ChatGLM-130B goes live as a bilingual dialogue model, and ChatGLM-6B is released with open weights. The smaller 6.2-billion-parameter model supports local Chinese-English conversation and quantized deployment on consumer graphics cards. |
2022: GLM-130B
| Date | Milestone |
|---|---|
| Aug. 20, 2022 | GLM-130B is released with publicly available model checkpoints. The 130-billion-parameter model is trained for English and Chinese. INT4 quantization follows on Aug. 24, reducing the hardware required for inference. |
2021: The original GLM research
| Date | Milestone |
|---|---|
| September 2021 | GLM-10B launches publicly after its training milestone earlier in the year, making a ten-billion-parameter GLM model available to researchers. |
| June 2021 | The research team completes training of GLM-10B, its first 10-billion-parameter model, built on the original GLM pretraining framework. |
| Mar. 18, 2021 | The original GLM research paper is submitted to arXiv. It introduces autoregressive blank infilling, a training objective that reconstructs missing text and supports language understanding, conditional generation, and unconditional generation in a shared framework. |
FAQs
Do GLM version numbers indicate model size?
Generation numbers and parameter counts describe different things. GLM-4 and GLM-5 identify model generations. Names such as GLM-130B and ChatGLM-6B include an approximate parameter count. For mixture-of-experts models, total parameters and active parameters also differ. Only a subset of the model’s experts processes each token.
Are all GLM models open source?
No. Some GLM models are available only through hosted services, and other releases have downloadable weights. GLM-4.5, GLM-4.5-Air, and GLM-5.3-Flash have MIT-licensed weights. GLM-5.3 weights are publicly available under a separate GLM-5.3 model license. Repository code licenses and model-weight licenses can differ.
What is the difference between GLM, ChatGLM, and Z.ai?
GLM is the model family and the name of its original pretraining research. ChatGLM refers to the earlier conversational models and chatbot products built on that research. Z.ai is the current international brand of the company previously known as Zhipu AI.
Related resources
- AI Model Release Calendar: Latest AI Model Releases by Date
- ZCode: Z.ai’s Open-Source AI Coding Workspace with Goal Mode
- AI LLM API Pricing: GPTl, Gemini, Claude Opus & More










