GLM Models Compared: Every Z.ai GLM Model (2026)

Every Z.ai GLM model in one place: specs, prices, licenses and where to use each, plus a straight answer on which GLM model fits your job.

The latest GLM model is GLM-5.3, Z.ai’s flagship for coding and long agent runs, announced on August 14, 2026. Alongside it sits GLM-5.3-Flash (August 26, 2026), a cheaper multimodal model that reads images and video for $0.15 per million input tokens. Every other GLM model on this page is older, more specialised or both. This hub lists all GLM models Z.ai (formerly Zhipu AI) currently offers, with release dates, parameter counts, context windows, prices, licenses and where you can use each one.

Short version: use GLM-5.3 when quality matters most, GLM-5.3-Flash for everything else, and GLM-4.7-Flash when you need a model that costs nothing on the API. The rest of this page explains why, and when an older model is still the better choice.

GLM models compared: GLM-5.3, GLM-5.3-Flash, GLM-5.2 and GLM-4.7-Flash specs and prices
Every current GLM model, from the 744B flagships to the free 30B Flash.

All GLM models compared in one table

This master table covers the language models you can call through the Z.ai API. Prices are pay-as-you-go per million tokens (input / output). “Chat” shows where you can talk to the model without writing code: GLM Chat is this site, chat.z.ai is Z.ai’s own app.

ModelReleasedParameters (total / active)Context / max outputInputPrice in / outWeightsChat
GLM-5.3Aug 14, 2026744B / 40B1M / 128KText$1.40 / $4.40GLM-5.3 LicenseGLM Chat, chat.z.ai
GLM-5.3-FlashAug 26, 2026320B / 18B1M / 128KText, image, video, file$0.15 / $0.50MITGLM Chat (default)
GLM-5.3-FlashXAug 26, 2026—1M / —Text, image, video, file$0.37 / $1.25API only—
GLM-5.2Jun 16, 2026744B / 40B1M / 128KText$1.40 / $4.40MITGLM Chat, chat.z.ai
GLM-5.1Apr 7, 2026744B / 40B200K / 128KText$1.40 / $4.40MITGLM Chat
GLM-5Feb 12, 2026744B / 40B200K / 128KText$1.00 / $3.20MIT—
GLM-5-TurboMar 2026—200K / 128KTextOpenRouter: $1.20 / $4.00Not published—
GLM-4.7Dec 22, 2025~358B / 32B200K / 128KText$0.60 / $2.20MITGLM Chat
GLM-4.7-FlashJan 19, 202631B / 3B200K / 128KTextFreeMITGLM Chat (unlimited)
GLM-4.7-FlashX——200K / 128KText$0.07 / $0.40API only—
GLM-4.6Sep 30, 2025357B (355B class)200K / 128KText$0.60 / $2.20MITGLM Chat
GLM-4.5Jul 28, 2025355B / 32B128K / 96KText$0.60 / $2.20MITGLM Chat
GLM-4.5-AirJul 28, 2025106B / 12B128K / 96KText$0.20 / $1.10MIT—
GLM-4.5-XJul 28, 2025—128K / 96KText$2.20 / $8.90API only—
GLM-4.5-AirXJul 28, 2025—128K / 96KText$1.10 / $4.50API only—
GLM-4.5-Flash———TextFreeAPI only—
GLM-4-32B-0414-128K—32B128K / —Text$0.10 / $0.10——
GLM language models on the Z.ai API. Dates from Z.ai release notes and announcements (GLM-5-Turbo: OpenRouter listing). A dash means Z.ai publishes no figure.

Three patterns stand out. First, the GLM-5 flagships (GLM-5.1, GLM-5.2, GLM-5.3) all cost the same $1.40 / $4.40, so upgrading within the line never raises your bill. Second, only GLM-5.2, GLM-5.3 and GLM-5.3-Flash read 1M tokens; everything from GLM-5.1 back tops out at 200K or 128K. Third, every model with its own page on this site has open weights, and all of them except GLM-5.3 use the plain MIT license. Cached input is cheaper on every model: $0.26 per million on the GLM-5.3, 5.2 and 5.1 tier, $0.03 on GLM-5.3-Flash. The GLM pricing page has the full cached-input column and worked cost examples.

Which GLM model should you use?

Pick by job, not by version number. The newest model is not always the best fit: GLM-5.3 cannot switch its reasoning off, GLM-5.3-Flash is paid on the API, and the older models have simpler thinking controls that some apps rely on.

Which GLM models to use: GLM-5.3 for hard coding, GLM-5.3-Flash as default, GLM-4.7-Flash for free API use
The short answer for most people.

Hard coding and long agent sessions: GLM-5.3

GLM-5.3 is the strongest GLM for software work. Z.ai reports a 50% gain over GLM-5.2 on its in-house Z.ai Code Bench, 88.2 on Terminal Bench 2.1 and 66.9 on DeepSWE v1.1, and open-source state-of-the-art results on Terminal Bench 3.0. It also uses fewer tokens per task than GLM-5.2 (about 75K vs 96K output tokens per Code Bench task at Max effort). Use it for multi-file refactors, debugging sessions, security review and agents that run for hours. Try GLM-5.3 in the chat.

Default for chat, apps and anything with images: GLM-5.3-Flash

At $0.15 / $0.50, GLM-5.3-Flash costs roughly a tenth of GLM-5.3 for input and output, keeps the 1M context and 128K output, and is the first natively multimodal model of the GLM-5 series, reading images, video and files. Z.ai says it beats GLM-5.2 across six coding and agent benchmarks, for example 63.4 vs 46.2 on DeepSWE v1.1. Make it your default and escalate to GLM-5.3 only when it falls short. Chat with GLM-5.3-Flash.

Free API use and prototypes: GLM-4.7-Flash

GLM-4.7-Flash is free on the Z.ai API for input, cached input and output; only your account’s rate limits apply. With 200K context, 128K output and a model-card score of 59.2 on SWE-bench Verified, it handles summaries, classification, extraction and routine code well. It is also the easiest GLM to run yourself: 31B total parameters, 3B active.

Fast answers with thinking switched off: GLM-5.2 or GLM-4.7

GLM-5.3 and GLM-5.3-Flash always reason. If your app needs instant, non-reasoning replies, use GLM-5.2 with reasoning_effort set to none (it then skips thinking) or GLM-4.7 with thinking disabled. GLM-5.2 gives you the 1M context; GLM-4.7 costs roughly half as much.

Self-hosting with a plain MIT license

For MIT weights at frontier scale, take GLM-5.2 (744B-A40B, 1M context). For a single-server deployment, GLM-4.5-Air (106B-A12B) or GLM-4.7-Flash (31B-A3B) are the practical choices. GLM-5.3-Flash (320B-A18B) sits in between and needs 4.4x less KV cache than GLM-5.3, according to Z.ai. GLM-5.3 itself is open too, under the GLM-5.3 License, which only adds a condition for very large Model-as-a-Service companies.

OpenClaw agents: GLM-5-Turbo

GLM-5-Turbo is a GLM-5 variant tuned for OpenClaw workflows: tool calling, complex instructions, scheduled and persistent tasks. Z.ai says it beats GLM-5 on ZClawBench, its public OpenClaw benchmark. See the GLM and OpenClaw setup guide for how it compares with GLM-5.3 there.

Use caseFirst choiceBudget choice
Agentic codingGLM-5.3GLM-5.3-Flash
General chat and writingGLM-5.3-FlashGLM-4.7-Flash
Screenshots, photos, videoGLM-5.3-FlashGLM-4.6V-Flash (free)
1M-token documents or reposGLM-5.3GLM-5.3-Flash
Non-reasoning, low latencyGLM-5.2 (effort none)GLM-4.7-Flash (thinking off)
Math and scienceGLM-5.3GLM-5.2
Image generationGLM-ImageCogView-4
Local deploymentGLM-5.2 or GLM-5.3-FlashGLM-4.7-Flash
GLM model picks by use case.

The GLM-5 line: GLM-5 to GLM-5.3

The GLM-5 generation started in February 2026 and has produced a new flagship roughly every two months. GLM-5, GLM-5.1, GLM-5.2 and GLM-5.3 share one architecture: 744B total parameters, 40B active, with DeepSeek Sparse Attention. GLM-5.3-Flash breaks the pattern with a new, smaller base model.

GLM-5.3

Same base model as GLM-5.2, with every gain from post-training. Announced August 14, 2026, listed in the API release notes on August 18 and published on Hugging Face on August 25 after a safety evaluation. Text-only input, 1M context, 128K output, forced reasoning with low, high or max effort. Besides coding, Z.ai highlights an emergent cyber capability: 84.5% on CyberGym, the best result on that benchmark, and 2,436 real vulnerabilities found across 269 projects with partner security teams. Weights come in FP8 (zai-org/GLM-5.3) and BF16 (zai-org/GLM-5.3-BF16). Full details on the GLM-5.3 page.

GLM-5.3-Flash and GLM-5.3-FlashX

A newly trained 320B-total, 18B-active model with 45 layers, the first GLM with hybrid linear and sparse attention, plus IndexPool and Manifold-Constrained Hyper-Connections. It was pre-trained on a 30T-token multimodal corpus and is the first natively multimodal GLM-5 model. Before launch it ran anonymously as “ox-alpha” on OpenCode and OpenRouter. GLM-5.3-FlashX is the API-only speed variant at about 200 tokens per second. Read the GLM-5.3-Flash guide.

GLM-5.2

Released June 16, 2026 with MIT weights on the same day. It brought the 1M-token context to the GLM line, IndexShare (one attention indexer shared by every four sparse layers, 2.9x fewer per-token FLOPs at 1M context) and an improved multi-token prediction layer with up to 20% longer speculative-decoding acceptance. Z.ai reports 62.1 on SWE-bench Pro and 81.0 on Terminal Bench 2.1. GLM-5.2 specs, benchmarks and OpenRouter access.

GLM-5.1

Released April 7, 2026 for long-horizon agent work, able to run one task for up to 8 hours. Z.ai describes it as broadly aligned with Claude Opus 4.6 and reports 58.4 on SWE-bench Pro, a top score at launch. 200K context, thinking on by default but switchable. GLM-5.1 specs and pricing.

GLM-5 and GLM-5-Turbo

GLM-5 (February 12, 2026) roughly doubled the parameter count of GLM-4.5, from 355B to 744B, raised pre-training data to 28.5T tokens, adopted DeepSeek Sparse Attention and was trained with Z.ai’s asynchronous RL framework, slime. It reports 77.8 on SWE-bench Verified and 56.2 on Terminal Bench 2.0. At $1.00 / $3.20 it remains the cheapest 744B GLM. GLM-5-Turbo, its OpenClaw-tuned variant, is API-only. Both are covered on the GLM-5 page.

How the GLM-5 models compare on Z.ai’s benchmarks

Z.ai publishes a benchmark table with each launch. Putting the overlapping rows side by side shows how fast the line has moved. Scores for GLM-5.3 and GLM-5.3-Flash come from the GLM-5.3 and GLM-5.3-Flash announcements; GLM-5.1 scores come from the GLM-5.2 table. All are Z.ai’s reported numbers.

BenchmarkGLM-5.3GLM-5.3-FlashGLM-5.2GLM-5.1
Z.ai Code Bench (max effort)34.5%29.0%23.4%—
DeepSWE v1.166.963.446.218.0
AutomationBench48.248.826.2—
Terminal Bench 2.188.2—81.063.5
NL2Repo58.0—48.942.7
HLE with tools62.5—54.752.3
Z.ai-reported scores from its GLM-5.2, GLM-5.3 and GLM-5.3-Flash announcements.

Two things jump out. GLM-5.3-Flash lands close to GLM-5.3 on agentic coding while costing about a tenth as much, which is why it is the default in GLM Chat. And GLM-5.2 more than doubled GLM-5.1’s DeepSWE score, from 18.0 to 46.2, before GLM-5.3 added about 21 more points on the same base model. These are vendor benchmarks, so test your own tasks too: each model page includes a five-task run in GLM Chat’s own test harness.

The GLM-4.x line: GLM-4.5 to GLM-4.7-Flash

The GLM-4.5 series (July 2025) introduced the Mixture-of-Experts design and hybrid reasoning that every later GLM builds on. These models are cheaper than the GLM-5 line and have simpler, fully switchable thinking, which keeps them useful in production.

GLM-4.7 and GLM-4.7-FlashX

GLM-4.7 (December 22, 2025) keeps GLM-4.5’s 32B-active design and focuses on coding: Z.ai reports 73.8% on SWE-bench Verified, 66.7% on SWE-bench Multilingual and 84.9 on LiveCodeBench v6. It added turn-level thinking and preserved thinking. GLM-4.7-FlashX is a paid, faster sibling of the free Flash model at $0.07 / $0.40. GLM-4.7 coding performance.

GLM-4.7-Flash

Released January 19, 2026 as the free-tier version of GLM-4.7. A 30B-A3B MoE model, MIT-licensed, with 200K context. Its model card shows 91.6 on AIME 25, 75.2 on GPQA and 79.5 on τ²-Bench. It is the most downloaded GLM repository on Hugging Face, at more than 1.8 million downloads a month. How to use GLM-4.7-Flash for free.

GLM-4.6

Released September 30, 2025. It raised the context from 128K to 200K, kept 128K output and, per Z.ai, is over 30% more token-efficient than GLM-4.5. Hybrid thinking is on by default. Note that there is no “GLM-4.6-Air”: the lightweight open model of that period is GLM-4.5-Air, later joined by GLM-4.7-Flash. GLM-4.6 specs and open weights.

GLM-4.5, GLM-4.5-Air, X, AirX and Flash

Released July 28, 2025. GLM-4.5 is 355B total / 32B active; GLM-4.5-Air is 106B / 12B. Both have 128K context, 96K output, MIT weights and a default temperature of 0.6 (the newer models default to 1.0). At launch Z.ai ranked GLM-4.5 second overall and first among open models on the average of 12 benchmarks. GLM-4.5-X and GLM-4.5-AirX are high-speed paid variants, and GLM-4.5-Flash is free. Read how GLM-4.5 aged and the GLM-4.5-Air guide.

Vision, image and other GLM models

Beyond the text line, Z.ai runs a set of specialist models on the same API and account. Prices are per million tokens unless noted.

ModelReleasedWhat it doesContextPrice
GLM-5V-Turbo—Multimodal coding model for agents200K—
GLM-4.6VDec 8, 2025Image, video, file understanding with function calling128K$0.30 / $0.90
GLM-4.6V-FlashX—Fast, cheap vision128K$0.04 / $0.40
GLM-4.6V-Flash—Free vision128KFree
GLM-4.5VAug 11, 2025100B-scale open vision reasoning64K$0.60 / $1.80
GLM-ImageJan 14, 2026Text-to-image, strong text rendering—$0.015 per image
GLM-OCRFeb 3, 2026Document parsing and extraction—$0.03 / $0.03
GLM-ASR-2512Dec 10, 2025Speech recognition—$0.03 (about $0.0024 per minute)
CogView-4—Image generation—$0.01 per image
CogVideoX-3Jul 15, 2025Video generation—$0.20 per video
Specialist models on the Z.ai API.

GLM-Image deserves a closer look. It combines a 9B autoregressive model with a 7B diffusion decoder and a Glyph Encoder, and is built for images with text in them: posters, slides, diagrams. Z.ai reports a CVTG-2K word accuracy of 0.9116. Weights are MIT. On the API it uses POST /images/generations with model glm-image, a default size of 1280×1280 (custom sizes from 1024 to 2048 pixels, divisible by 32), and returns a URL that expires after 30 days. In GLM Chat you can generate images with GLM-Image, three a day.

For vision, GLM-5.3-Flash now covers most needs on its own. The older GLM-4.6V family remains useful when you want a free vision model (GLM-4.6V-Flash) or the very low input price of GLM-4.6V-FlashX.

GLM model names explained: Flash, FlashX, Air, X, AirX, Turbo, V

GLM model names follow a pattern: a version number, then a suffix that tells you the size or speed tier. Once you know the suffixes, the model list reads easily.

SuffixMeaningExamples
(none)Full-size model of that versionGLM-5.3, GLM-4.7
FlashLightweight model; free on the API through GLM-4.7, cheap paid from GLM-5.3GLM-4.7-Flash, GLM-5.3-Flash
FlashXPaid, faster variant of a Flash model, API onlyGLM-5.3-FlashX, GLM-4.7-FlashX
AirSmaller open sibling of a flagshipGLM-4.5-Air
XHigh-speed paid variant of the full modelGLM-4.5-X
AirXHigh-speed paid variant of AirGLM-4.5-AirX
TurboVariant tuned for agent workflowsGLM-5-Turbo, GLM-5V-Turbo
VVision model (reads images and video)GLM-4.6V, GLM-4.5V
-FP8 / -BF16Weight precision on Hugging FaceGLM-5.2-FP8, GLM-5.3-BF16
GLM naming conventions.

Two naming traps catch people. First, “Flash” no longer means free: GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash are free on the API, but GLM-5.3-Flash costs $0.15 / $0.50. Second, the precision suffix flips between releases: for GLM-5.2, GLM-5.1 and GLM-5 the plain repository holds BF16 weights and the -FP8 repository the FP8 build, while for GLM-5.3 and GLM-5.3-Flash the plain repository is FP8 and the -BF16 repository holds BF16. API model IDs are always lowercase (glm-5.3-flash, glm-4.5-air).

The older name ChatGLM referred to Z.ai’s chat-tuned models from 2023, such as the open-source ChatGLM-6B. Current models drop the “Chat” prefix.

Thinking and reasoning controls by model

Every current GLM model can reason before answering, but the controls differ, and getting them wrong is the most common reason a request fails after switching models. Reasoning tokens are billed as output tokens, so these settings also drive your cost.

ModelDefaultCan disable?reasoning_effort
GLM-5.3, GLM-5.3-FlashOn (forced)No (request fails)low / high / max (default max)
GLM-5.2OnYes, effort none or minimalhigh / max (default max)
GLM-5.1, GLM-5OnYes—
GLM-4.7 seriesOn, turn-levelYes—
GLM-4.6, GLM-4.5 seriesHybrid (dynamic)Yes—
Thinking defaults on the Z.ai API. Thinking is set with thinking.type enabled or disabled.

On GLM-5.2, low and medium map to high, and xhigh maps to max, so code written for other providers keeps working. Full examples are in the GLM thinking mode guide.

Routing, migration and older GLM models

Z.ai has not published retirement dates for older GLM models on the pay-as-you-go API; GLM-4.5 through GLM-5.2 are all still listed with prices. The changes that matter today are on the Coding Plan and in the GLM-5.3 migration.

Coding Plan routing

  • Every Coding Plan tier (Lite, Pro, Max) includes GLM-5.3 and GLM-5.3-Flash.
  • Requests for glm-5.2 or glm-5.1 are automatically routed to GLM-5.3.
  • Requests for glm-4.7 are routed to GLM-5.3-Flash.
  • GLM-5.3-Flash gets 3x the quota of GLM-5.3 on every tier. GLM-5.3-FlashX is not on the Coding Plan yet.

So if your Claude Code or Cline config still says GLM-5.2, you are already running GLM-5.3 on the plan. The plan works only inside supported coding tools; for your own apps, use the pay-as-you-go API, where every model ID in the master table still resolves to that model. The GLM Coding Plan guide explains tiers and credits.

Moving an app to GLM-5.3

  1. If your code sends "thinking": {"type": "disabled"}, change it to enabled and set reasoning_effort to low before switching the model ID. Otherwise the request fails.
  2. Change the model to glm-5.3. Defaults are temperature 1.0 and top_p 0.95; Z.ai recommends tuning only one of them.
  3. Handle delta.reasoning_content separately from delta.content when streaming.
  4. For streamed tool calls, set both stream and tool_stream to true and concatenate the argument fragments.
  5. Test latency and cost: max effort, the default, produces the most reasoning tokens.

The GLM API quickstart has working curl, Python and Node code for every step, and the GLM release timeline tracks each model change as it happens.

Where each GLM model is available

You can reach GLM through five channels. OpenRouter prices change often and vary by provider, so treat them as OpenRouter’s listed price and confirm on the model’s OpenRouter page before you commit.

ModelGLM ChatZ.ai APIOpenRouter (listed, in / out)Coding PlanHugging Face
GLM-5.3Yesglm-5.3$1.40 / $4.40Yeszai-org/GLM-5.3
GLM-5.3-FlashYes (default)glm-5.3-flash$0.15 / $0.50Yes, 3x quotazai-org/GLM-5.3-Flash
GLM-5.2Yesglm-5.2~$0.65 / $2.04; free variantRoutes to 5.3zai-org/GLM-5.2
GLM-5.1Yesglm-5.1$0.97 / $3.04Routes to 5.3zai-org/GLM-5.1
GLM-5Noglm-5$0.60 / $1.92—zai-org/GLM-5
GLM-5-TurboNoglm-5-turbo$1.20 / $4.00—Not published
GLM-4.7Yesglm-4.7$0.40 / $1.75Routes to 5.3-FlashYes
GLM-4.7-FlashYes (unlimited)glm-4.7-flash (free)$0.06 / $0.40—Yes
GLM-4.6Yesglm-4.6$0.43 / $1.75—Yes
GLM-4.5Yesglm-4.5$0.60 / $2.20—Yes
GLM-4.5-AirNoglm-4.5-air$0.13 / $0.85—Yes
GLM availability by channel. OpenRouter prices are OpenRouter’s listed prices per 1M tokens.

OpenRouter also lists z-ai/glm-5.2:free, a free, rate-limited GLM-5.2 with a 32,768-token context, and z-ai/glm-5.3:batch at $0.45 / $2.00. Z.ai’s own chat app, chat.z.ai, offers GLM-5.3 and other GLM models for free. On Hugging Face every open GLM lives under zai-org, with ModelScope mirrors under ZhipuAI; supported local serving stacks include vLLM, SGLang, Transformers, KTransformers and Unsloth. For licensing details per model, see is GLM open source.

In GLM Chat, the paid models allow 40 messages a day per visitor, GLM-4.7-Flash is unlimited at up to one request every 3 seconds, and each message can run to 8,000 characters. Deep links open a specific model straight away: chat with GLM-5.2, chat with GLM-4.7, chat with GLM-4.6 or chat with GLM-4.5. That makes it easy to put the same prompt to two generations and see the difference for yourself.

GLM models FAQ

What is the latest GLM model?

GLM-5.3 is the latest flagship (announced August 14, 2026). GLM-5.3-Flash, released August 26, 2026, is the newest model overall and the cheaper multimodal option. Z.ai has announced no date or specs for GLM-5.5 or GLM-6; it has said only that it is scaling the GLM-5.3-Flash recipe to larger models.

How many GLM models are there?

Z.ai’s docs cover 17 GLM language models (from GLM-4-32B-0414-128K to GLM-5.3), five vision models (GLM-5V-Turbo, GLM-4.6V, GLM-4.6V-FlashX, GLM-4.6V-Flash, GLM-4.5V) and GLM-Image, GLM-OCR and GLM-ASR-2512. For most people, four matter: GLM-5.3, GLM-5.3-Flash, GLM-5.2 and GLM-4.7-Flash.

Which GLM model is best for coding?

GLM-5.3. Z.ai reports a 50% improvement over GLM-5.2 on its in-house coding benchmark and open-source state-of-the-art results on Terminal Bench 3.0. For a cheaper option, GLM-5.3-Flash beats GLM-5.2 on six coding and agent benchmarks according to Z.ai, at about a tenth of the price.

Which GLM model is free?

On the Z.ai API, GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash are free for input and output. In GLM Chat, every model on the model chips is free to use within daily limits, and GLM-4.7-Flash has no daily cap.

What is the difference between GLM-5.3 and GLM-5.2?

They share the same 744B base model, context (1M) and price ($1.40 / $4.40). GLM-5.3 adds post-training that lifts coding and cyber scores sharply, but it always reasons, while GLM-5.2 can skip thinking. GLM-5.2 has a plain MIT license; GLM-5.3 uses the GLM-5.3 License.

Which GLM models have a 1M context window?

GLM-5.3, GLM-5.3-Flash (and FlashX) and GLM-5.2. GLM-5.1, GLM-5, GLM-4.7 and GLM-4.6 have 200K; the GLM-4.5 series has 128K.

Is there a GLM-4.6 Air model?

No. Z.ai never released a GLM-4.6-Air. If you want a lightweight open GLM, use GLM-4.5-Air (106B total, 12B active) or the newer GLM-4.7-Flash (31B total, 3B active, free on the API).

Can I try every GLM model before paying?

Most of them. The free GLM chat on this site covers GLM-5.3, GLM-5.3-Flash, GLM-5.2, GLM-5.1, GLM-4.7, GLM-4.6, GLM-4.5, GLM-4.7-Flash and GLM-Image with no account. For GLM-5 and GLM-4.5-Air, use their closest chat siblings (GLM-5.1 and GLM-4.5) or call the API.