GLM Release Timeline: Every Model From GLM-4.5 to GLM-5.3 (and What’s Next)

Every GLM release date, newest first, with what changed in each model and what Z.ai has said about the next one.

The newest GLM release is GLM-5.3-Flash, released on August 26, 2026, one day after the open weights of the flagship GLM-5.3 went up on Hugging Face. GLM-5.3 itself was announced on August 14, 2026 and reached Z.ai’s API release notes on August 18. This page lists the GLM release date of every model from GLM-4.5 (July 28, 2025) to today, newest first, with what changed in each release and a link to the full model page.

Z.ai (formerly Zhipu AI) has shipped a new flagship roughly every two months in 2026: GLM-5 in February, GLM-5.1 in April, GLM-5.2 in June and GLM-5.3 in August. Dates below come from Z.ai’s release notes at docs.z.ai unless a note says otherwise. Want to try the models while you read? Every current one is in the free GLM chat.

GLM release date timeline from GLM-4.5 in 2025 to GLM-5.3 and GLM-5.3-Flash in 2026
Every GLM release from GLM-4.5 to GLM-5.3-Flash.

GLM release dates at a glance

DateReleaseWhat it isWeights
Aug 26, 2026GLM-5.3-Flash320B-A18B natively multimodal model, 1M contextMIT
Aug 25, 2026GLM-5.3 weightsFP8 and BF16 repos on Hugging FaceGLM-5.3 License
Aug 14 / 18, 2026GLM-5.3Flagship, coding and cyber gains from post-trainingGLM-5.3 License
Jun 16, 2026GLM-5.21M-token context, IndexShare, effort levelsMIT
Apr 7, 2026GLM-5.1Long-horizon agent model, up to 8 hours per taskMIT
Mar 2026GLM-5-TurboGLM-5 variant tuned for OpenClawNot published
Feb 12, 2026GLM-5744B-A40B, first GLM with sparse attentionMIT
Feb 3, 2026GLM-OCRCompact OCR model–
Jan 19, 2026GLM-4.7-Flash30B-A3B free-tier modelMIT
Jan 14, 2026GLM-ImageImage generation, strong text renderingMIT
Dec 22, 2025GLM-4.7Coding flagship, turn-level thinkingMIT
Dec 11, 2025AutoGLM-Phone-MultilingualMobile automation framework–
Dec 10, 2025GLM-ASR-2512Speech recognition–
Dec 8, 2025GLM-4.6VVision model, 128K contextMIT
Sep 30, 2025GLM-4.6Context 128K to 200KMIT
Aug 11, 2025GLM-4.5V100B-scale open vision reasoning modelOpen
Aug 8, 2025GLM Slide/Poster Agent (beta)Slides and posters from prompts–
Jul 28, 2025GLM-4.5 and GLM-4.5-Air355B-A32B and 106B-A12B agentic modelsMIT
Jul 15, 2025CogVideoX-3Video generation upgrade–
GLM release dates, newest first. A dash means the release is an API product or our sources list no weights.
Timeline chart of GLM model release dates from July 2025 to August 2026
The GLM flagship line on one axis: in 2026, a new flagship roughly every two months.

The latest GLM model version

If you searched for the latest Zhipu GLM model version, the answer has two parts:

  • Most capable: GLM-5.3 (model ID glm-5.3), a 744B-parameter mixture-of-experts model with 40B active parameters, a 1M-token context and 128K maximum output. It costs $1.40 per 1M input tokens and $4.40 per 1M output tokens on the Z.ai API.
  • Newest: GLM-5.3-Flash (model ID glm-5.3-flash), a smaller 320B-A18B model built on a new base, the first natively multimodal model of the GLM-5 series. It costs $0.15 in and $0.50 out, and it is the default model in this site’s chat.

Everything older is still callable on the API, but on the GLM Coding Plan requests for GLM-5.2 and GLM-5.1 are now routed to GLM-5.3, and GLM-4.7 requests go to GLM-5.3-Flash. The GLM models comparison helps you pick between them, and the GLM pricing page lists what each one costs.

2026 releases: the GLM-5 line

GLM-5 line release dates in 2026 with key features of each model
Four flagships and a Flash model in seven months.

GLM-5.3-Flash: August 26, 2026

GLM-5.3-Flash starts from a newly trained base model rather than a cut-down GLM-5.3. It has 320B total and 18B active parameters across 45 layers and is the first GLM model to combine linear and sparse attention. Z.ai says this cuts attention compute by 3.0x and the KV cache by 4.4x compared with GLM-5.3, which is why it can serve a 1M-token context cheaply. It accepts video, image, text and file input and produces text.

Before launch it ran anonymously as “ox-alpha” on OpenCode and OpenRouter, where it became the most popular model of that week. Weights are MIT-licensed (FP8 and BF16), the API price is $0.15 in and $0.50 out, and a faster API-only variant, GLM-5.3-FlashX, runs at about 200 tokens per second for $0.37 in and $1.25 out. Z.ai reports it beats GLM-5.2 on six coding and agentic benchmarks, including DeepSWE v1.1 (63.4 vs 46.2). Details: GLM-5.3-Flash.

GLM-5.3: August 14 and 18, 2026 (weights August 25)

Z.ai announced GLM-5.3 on its blog on August 14 under the title “GLM-5.3: Frontier Coding with Emergent Cyber Capabilities”, and the API release-notes entry followed on August 18. The blog promised weights two weeks after launch, once safety evaluation was complete; they appeared on Hugging Face on August 25 as zai-org/GLM-5.3 (FP8) and zai-org/GLM-5.3-BF16.

GLM-5.3 uses the same base model as GLM-5.2. Every gain comes from post-training: Z.ai reports a 50% improvement over GLM-5.2 on its in-house Z.ai Code Bench, open-source state of the art on Terminal Bench 3.0 and Agents’ Last Exam, and 84.5 on CyberGym. In real-world runs with security teams it found 2,436 vulnerabilities across 269 projects, 1,097 of them medium-to-high severity. Two breaking changes matter for developers: thinking can no longer be disabled (send reasoning_effort: "low" instead of thinking: disabled), and the weights ship under a new custom GLM-5.3 License rather than MIT. The GLM license guide explains that license clause by clause.

GLM-5.2: June 16, 2026

GLM-5.2 (“Built for Long-Horizon Tasks”) took the context window from 200K to 1M tokens, which Z.ai calls “solid 1M lossless context”. The architecture added IndexShare (one lightweight indexer shared by every four sparse-attention layers, cutting per-token FLOPs by 2.9x at 1M context) and an improved multi-token prediction layer that raises speculative-decoding acceptance length by up to 20%. It introduced effort levels (High and Max). Weights went up the same day under MIT. In Z.ai’s table it scores 81.0 on Terminal Bench 2.1 and 62.1 on SWE-bench Pro. See GLM-5.2.

GLM-5.1: April 7, 2026

GLM-5.1 kept the GLM-5 architecture (744B total, 40B active, 200K context) and focused on long-horizon work: Z.ai says it can work on one task independently for up to 8 hours. It scored 58.4 on SWE-Bench Pro, which Z.ai claimed as state of the art at the time, and Z.ai described it as aligned with Claude Opus 4.6 overall. Weights are MIT. See GLM-5.1.

GLM-5-Turbo: March 2026

GLM-5-Turbo is a GLM-5 variant trained for OpenClaw workflows: tool calling, complex instruction following, scheduled and persistent tasks, and long execution chains. It has a 200K context and 128K output. It has no entry in Z.ai’s release notes; the March date comes from its OpenRouter listing (z-ai/glm-5-turbo). Z.ai published ZClawBench alongside it and says Turbo beats GLM-5 there. Weights are not published. A vision sibling, GLM-5V-Turbo, has its own docs page. Coverage is on the GLM-5 and GLM-5-Turbo page.

GLM-5: February 12, 2026

GLM-5 opened the fifth generation with a big jump in scale: 744B total and 40B active parameters (up from GLM-4.5’s 355B and 32B) and 28.5T pre-training tokens (up from 23T). It was the first GLM to integrate DeepSeek Sparse Attention and was post-trained with Z.ai’s asynchronous RL framework, slime. Z.ai reported 77.8 on SWE-bench Verified, 56.2 on Terminal Bench 2.0 and a final balance of $4,432 on Vending Bench 2, first among open models. The weights appeared on Hugging Face a day early, on February 11, under MIT, with a technical report on arXiv (2602.15763).

GLM-OCR: February 3, 2026

A compact optical character recognition model built on a CogViT encoder and a GLM-0.5B decoder. It costs $0.03 per 1M tokens for input and output on the API.

GLM-4.7-Flash: January 19, 2026

The “free-tier version of GLM-4.7”: a 30B-A3B mixture-of-experts model with a 200K context and 128K output, free on the Z.ai API for input, cached input and output. It runs locally with vLLM and SGLang and is the most-downloaded GLM repository on Hugging Face, with more than 1.8 million downloads a month. See GLM-4.7-Flash.

GLM-Image: January 14, 2026

An image generation model that pairs a 9B autoregressive model with a 7B diffusion decoder and a glyph encoder for text rendering. It is strong at text inside images, such as posters, slides and diagrams (Z.ai reports 0.9116 word accuracy on CVTG-2K). It costs $0.015 per image, and weights are MIT. You can try it in this site’s chat with the GLM-Image generator.

Other GLM news in 2026

  • January 8, 2026: the company listed on the Main Board of HKEX (stock code 2513) as Knowledge Atlas Technology. The listed company changed its name to Z.AI Co., Ltd. in July 2026. More on the Zhipu AI company page.
  • July 30, 2026: the GLM Coding Plan moved from prompt-based to credits-based limits, with new Lite, Pro and Max tiers at $18, $80 and $168 per month.
  • September 3 to October 7, 2026: a GLM-5.3-Flash campaign gives paid Coding Plan users unlimited GLM-5.3-Flash in ZCode and AutoClaw from 23:00 to 09:00 UTC+8, and doubled quota on other agents.
  • September 25 to October 7, 2026: all Coding Plan usage is charged at the off-peak (50%) credit rate around the clock.

2025 releases: the GLM-4.x line

GLM-4.7: December 22, 2025

The last GLM-4 flagship kept the 355B-class size of GLM-4.5 (32B active) with a 200K context and 128K output, and brought large coding gains: Z.ai reports 73.8% on SWE-bench Verified (+5.8 points over GLM-4.6), 66.7% on SWE-bench Multilingual and 41% on Terminal Bench 2.0. It introduced turn-level thinking and preserved thinking. MIT weights. See GLM-4.7.

AutoGLM-Phone-Multilingual, GLM-ASR-2512 and GLM-4.6V: December 8–11, 2025

Three launches in four days. GLM-4.6V (December 8) is a vision model with a 128K context, MIT weights and a price of $0.30 in and $0.90 out, with a free sibling, GLM-4.6V-Flash. GLM-ASR-2512 (December 10) is a speech recognition model with a reported character error rate of 0.0717, priced at $0.03 per 1M tokens. AutoGLM-Phone-Multilingual (December 11) is a mobile automation framework that reads the screen and performs real actions through ADB across more than 50 mainstream apps.

GLM-4.6: September 30, 2025

GLM-4.6 expanded the context window from 128K to 200K and raised maximum output to 128K. Z.ai says it is more than 30% more token-efficient than GLM-4.5, and it published 74 real coding test trajectories run in Claude Code. There is no “GLM-4.6-Air”; the lightweight open model of this era remained GLM-4.5-Air. See GLM-4.6.

GLM-4.5V and GLM Slide/Poster Agent: August 2025

GLM-4.5V (August 11) is a 100B-scale open-source vision reasoning model covering video understanding, visual grounding and GUI agents, with a thinking mode. The GLM Slide/Poster Agent beta (August 8) generates slides and posters from natural-language instructions and costs $0.70 per 1M tokens.

GLM-4.5 and GLM-4.5-Air: July 28, 2025

The GLM-4.5 series is where the modern GLM line begins. GLM-4.5 has 355B total and 32B active parameters; GLM-4.5-Air has 106B total and 12B active. Both have a 128K context, 96K maximum output and hybrid reasoning (a thinking mode and a non-thinking mode), and both are MIT-licensed. At launch Z.ai ranked GLM-4.5 second overall and first among open models on the average of 12 benchmarks, and the release notes highlight one-click compatibility with Claude Code. The series also includes the faster paid variants GLM-4.5-X and GLM-4.5-AirX and the free GLM-4.5-Flash. See GLM-4.5 and GLM-4.5-Air.

CogVideoX-3: July 15, 2025

An incremental upgrade to Z.ai’s video generation model with better quality and start-and-end-frame synthesis. It costs $0.20 per video on the API.

GLM news for developers: what changed in 2026

Release dates tell you when a model arrived. These are the changes that actually break or improve existing code, in the order they happened:

  1. A bigger flagship (GLM-5, February). A new model ID, glm-5, on the same Chat Completions endpoint, priced at $1.00 in and $3.20 out against GLM-4.7’s $0.60 and $2.20.
  2. One flagship price (GLM-5.1, April onward). GLM-5.1, GLM-5.2 and GLM-5.3 are all listed at $1.40 in and $4.40 out, with cached input at $0.26, so upgrading costs nothing extra.
  3. 1M-token context and effort levels (GLM-5.2, June). reasoning_effort accepts high and max (default max); none or minimal skip thinking.
  4. Coding Plan becomes credits-based (July 30). Usage is now measured in credits per 5 hours and per week, with off-peak hours at half rate.
  5. Forced thinking and a new license (GLM-5.3, August). Sending thinking: {"type": "disabled"} to glm-5.3 makes the request fail; use reasoning_effort: "low", which joins high and max. The weights use the GLM-5.3 License instead of MIT.
  6. Native image input in the GLM-5 series (GLM-5.3-Flash, August). Add image_url parts to messages[].content; thinking is always on here too.
  7. Coding Plan routing. Requests for GLM-5.2 and GLM-5.1 on the plan now run on GLM-5.3, and GLM-4.7 requests on GLM-5.3-Flash, so the model name in your tool config no longer pins an older model there.

If a request that worked last month now returns a 400 error, the forced-thinking change is the usual cause; GLM API error codes lists the fixes, and GLM thinking mode shows the right settings per model.

Earlier history: from GLM to GLM-Zero

GLM is older than most people think. Z.ai’s company page traces the line back to the company’s founding on March 15, 2019, originating from Tsinghua University research. The milestones before GLM-4.5:

DateMilestone
Mar 10, 2021GLM, a new pre-training paradigm
Aug 20, 2022GLM-130B released and open-sourced (100B+ parameters)
Mar 15, 2023ChatGLM; the open ChatGLM-6B passed 20 million downloads
Jan 25, 2024GLM-4 foundation model
Jun 18, 2024GLM-4-9B and GLM-4V-9B open-sourced
Aug 12, 2024GLM-4-Plus
Oct 28, 2024GLM-4-Voice and AutoGLM
Nov 15, 2024GLM-PC
Dec 20, 2024GLM-Zero-Preview, the first GLM reasoning model
Jan 8, 2025GLM-Realtime
Mar 22, 2025AutoGLM Reflection
GLM milestones before GLM-4.5, from Z.ai’s company page.

The “ChatGLM” name from 2023 still shows up in searches and old tutorials. The current models are simply called GLM. Z.ai’s older developer platform, bigmodel.cn, still exists and uses a separate account system from the z.ai platform most readers use today.

What’s next: GLM-5.5 and GLM-6

Z.ai has not announced a release date, specifications or even a confirmed name for GLM-5.5 or GLM-6. Any page quoting a date or a parameter count for either model is guessing. The only forward-looking statement from Z.ai is at the end of the GLM-5.3-Flash announcement:

We are now scaling this recipe to larger models — GLM-5.3-Flash pushes the cost-performance frontier, and the lessons from building it are already shaping our next frontier model.

Z.ai, GLM-5.3-Flash announcement

“This recipe” refers to the GLM-5.3-Flash design: a newly trained base model, hybrid linear and sparse attention, Manifold-Constrained Hyper-Connections and a 30T-token multimodal pre-training corpus. Z.ai has not said which model the recipe goes into next, what that model will be called, or when it will ship. That is the full official record.

The track record is the other clue. In 2026 the gaps between flagships were 54 days (GLM-5 to GLM-5.1), 70 days (GLM-5.1 to GLM-5.2) and 59 days (GLM-5.2 to GLM-5.3 announcement). That pattern tells you how often to look, not when the next one lands. When it does, this page and the GLM models hub are updated with the date, specs and price.

Where to follow GLM news

New GLM models usually appear in several official places within a day or two. Watch these:

  • Z.ai blog: the long-form launch posts with benchmark tables. GLM-5.2, GLM-5.3 and GLM-5.3-Flash all launched there; the GLM-5.3 launch post is a good example.
  • Release notes at docs.z.ai: the dated API changelog this page is built on.
  • GitHub at github.com/zai-org/GLM-5: the README with download links and serving notes for every GLM-5 series model.
  • Hugging Face at huggingface.co/zai-org: new repos often appear shortly before or on launch day (GLM-5 weights were up a day before its release note).
  • Z.ai on X and Discord: both are linked in the z.ai site footer, and the Discord community is also linked from the GitHub README.
  • OpenRouter: anonymous test models sometimes appear there first, as GLM-5.3-Flash did under the name “ox-alpha”.

For a list of every genuine Z.ai domain (useful when a “GLM-6 download” link looks suspicious), see Z.ai official website links.

GLM release date FAQ

What is the latest GLM model?

GLM-5.3-Flash, released August 26, 2026, is the newest. GLM-5.3, announced August 14, 2026, is the most capable flagship.

When was GLM-5.3 released?

Z.ai announced GLM-5.3 on August 14, 2026, listed it in the API release notes on August 18, and published its weights on Hugging Face on August 25.

When was GLM-5 released?

February 12, 2026, according to Z.ai’s release notes. The weights were published on Hugging Face one day earlier, on February 11.

When was GLM-5.2 released?

June 16, 2026. The API, the chat.z.ai app and the MIT weights on Hugging Face all went live the same day.

When is GLM-5.5 coming out?

Z.ai has not announced GLM-5.5 or a date for it. Z.ai has said only that the lessons from GLM-5.3-Flash are shaping its next frontier model.

Is GLM-6 released?

No. There is no GLM-6 model, release date or specification from Z.ai. Treat any “GLM-6” download or benchmark you find as fake until it appears on z.ai, docs.z.ai or the zai-org Hugging Face page.

When was GLM-4.5 released?

July 28, 2025, together with GLM-4.5-Air. The Hugging Face repositories were created on July 20, 2025.

How often does Z.ai release new GLM models?

In 2026, a new flagship arrived roughly every two months (February, April, June, August), with smaller models such as GLM-4.7-Flash, GLM-OCR, GLM-Image and GLM-5.3-Flash in between.

The fastest way to feel the difference between releases is to ask the same question twice: once to GLM-5.3 in the chat and once to GLM-4.7 in the chat.