The latest GLM model is GLM-5.3, Z.ai’s flagship for coding and long agent runs, announced on August 14, 2026. Alongside it sits GLM-5.3-Flash (August 26, 2026), a cheaper multimodal model that reads images and video for $0.15 per million input tokens. Every other GLM model on this page is older, more specialised or both. This hub lists all GLM models Z.ai (formerly Zhipu AI) currently offers, with release dates, parameter counts, context windows, prices, licenses and where you can use each one.
Short version: use GLM-5.3 when quality matters most, GLM-5.3-Flash for everything else, and GLM-4.7-Flash when you need a model that costs nothing on the API. The rest of this page explains why, and when an older model is still the better choice.

All GLM models compared in one table
This master table covers the language models you can call through the Z.ai API. Prices are pay-as-you-go per million tokens (input / output). “Chat” shows where you can talk to the model without writing code: GLM Chat is this site, chat.z.ai is Z.ai’s own app.
| Model | Released | Parameters (total / active) | Context / max output | Input | Price in / out | Weights | Chat |
|---|---|---|---|---|---|---|---|
| GLM-5.3 | Aug 14, 2026 | 744B / 40B | 1M / 128K | Text | $1.40 / $4.40 | GLM-5.3 License | GLM Chat, chat.z.ai |
| GLM-5.3-Flash | Aug 26, 2026 | 320B / 18B | 1M / 128K | Text, image, video, file | $0.15 / $0.50 | MIT | GLM Chat (default) |
| GLM-5.3-FlashX | Aug 26, 2026 | — | 1M / — | Text, image, video, file | $0.37 / $1.25 | API only | — |
| GLM-5.2 | Jun 16, 2026 | 744B / 40B | 1M / 128K | Text | $1.40 / $4.40 | MIT | GLM Chat, chat.z.ai |
| GLM-5.1 | Apr 7, 2026 | 744B / 40B | 200K / 128K | Text | $1.40 / $4.40 | MIT | GLM Chat |
| GLM-5 | Feb 12, 2026 | 744B / 40B | 200K / 128K | Text | $1.00 / $3.20 | MIT | — |
| GLM-5-Turbo | Mar 2026 | — | 200K / 128K | Text | OpenRouter: $1.20 / $4.00 | Not published | — |
| GLM-4.7 | Dec 22, 2025 | ~358B / 32B | 200K / 128K | Text | $0.60 / $2.20 | MIT | GLM Chat |
| GLM-4.7-Flash | Jan 19, 2026 | 31B / 3B | 200K / 128K | Text | Free | MIT | GLM Chat (unlimited) |
| GLM-4.7-FlashX | — | — | 200K / 128K | Text | $0.07 / $0.40 | API only | — |
| GLM-4.6 | Sep 30, 2025 | 357B (355B class) | 200K / 128K | Text | $0.60 / $2.20 | MIT | GLM Chat |
| GLM-4.5 | Jul 28, 2025 | 355B / 32B | 128K / 96K | Text | $0.60 / $2.20 | MIT | GLM Chat |
| GLM-4.5-Air | Jul 28, 2025 | 106B / 12B | 128K / 96K | Text | $0.20 / $1.10 | MIT | — |
| GLM-4.5-X | Jul 28, 2025 | — | 128K / 96K | Text | $2.20 / $8.90 | API only | — |
| GLM-4.5-AirX | Jul 28, 2025 | — | 128K / 96K | Text | $1.10 / $4.50 | API only | — |
| GLM-4.5-Flash | — | — | — | Text | Free | API only | — |
| GLM-4-32B-0414-128K | — | 32B | 128K / — | Text | $0.10 / $0.10 | — | — |
Three patterns stand out. First, the GLM-5 flagships (GLM-5.1, GLM-5.2, GLM-5.3) all cost the same $1.40 / $4.40, so upgrading within the line never raises your bill. Second, only GLM-5.2, GLM-5.3 and GLM-5.3-Flash read 1M tokens; everything from GLM-5.1 back tops out at 200K or 128K. Third, every model with its own page on this site has open weights, and all of them except GLM-5.3 use the plain MIT license. Cached input is cheaper on every model: $0.26 per million on the GLM-5.3, 5.2 and 5.1 tier, $0.03 on GLM-5.3-Flash. The GLM pricing page has the full cached-input column and worked cost examples.
Which GLM model should you use?
Pick by job, not by version number. The newest model is not always the best fit: GLM-5.3 cannot switch its reasoning off, GLM-5.3-Flash is paid on the API, and the older models have simpler thinking controls that some apps rely on.

Hard coding and long agent sessions: GLM-5.3
GLM-5.3 is the strongest GLM for software work. Z.ai reports a 50% gain over GLM-5.2 on its in-house Z.ai Code Bench, 88.2 on Terminal Bench 2.1 and 66.9 on DeepSWE v1.1, and open-source state-of-the-art results on Terminal Bench 3.0. It also uses fewer tokens per task than GLM-5.2 (about 75K vs 96K output tokens per Code Bench task at Max effort). Use it for multi-file refactors, debugging sessions, security review and agents that run for hours. Try GLM-5.3 in the chat.
Default for chat, apps and anything with images: GLM-5.3-Flash
At $0.15 / $0.50, GLM-5.3-Flash costs roughly a tenth of GLM-5.3 for input and output, keeps the 1M context and 128K output, and is the first natively multimodal model of the GLM-5 series, reading images, video and files. Z.ai says it beats GLM-5.2 across six coding and agent benchmarks, for example 63.4 vs 46.2 on DeepSWE v1.1. Make it your default and escalate to GLM-5.3 only when it falls short. Chat with GLM-5.3-Flash.
Free API use and prototypes: GLM-4.7-Flash
GLM-4.7-Flash is free on the Z.ai API for input, cached input and output; only your account’s rate limits apply. With 200K context, 128K output and a model-card score of 59.2 on SWE-bench Verified, it handles summaries, classification, extraction and routine code well. It is also the easiest GLM to run yourself: 31B total parameters, 3B active.
Fast answers with thinking switched off: GLM-5.2 or GLM-4.7
GLM-5.3 and GLM-5.3-Flash always reason. If your app needs instant, non-reasoning replies, use GLM-5.2 with reasoning_effort set to none (it then skips thinking) or GLM-4.7 with thinking disabled. GLM-5.2 gives you the 1M context; GLM-4.7 costs roughly half as much.
Self-hosting with a plain MIT license
For MIT weights at frontier scale, take GLM-5.2 (744B-A40B, 1M context). For a single-server deployment, GLM-4.5-Air (106B-A12B) or GLM-4.7-Flash (31B-A3B) are the practical choices. GLM-5.3-Flash (320B-A18B) sits in between and needs 4.4x less KV cache than GLM-5.3, according to Z.ai. GLM-5.3 itself is open too, under the GLM-5.3 License, which only adds a condition for very large Model-as-a-Service companies.
OpenClaw agents: GLM-5-Turbo
GLM-5-Turbo is a GLM-5 variant tuned for OpenClaw workflows: tool calling, complex instructions, scheduled and persistent tasks. Z.ai says it beats GLM-5 on ZClawBench, its public OpenClaw benchmark. See the GLM and OpenClaw setup guide for how it compares with GLM-5.3 there.
| Use case | First choice | Budget choice |
|---|---|---|
| Agentic coding | GLM-5.3 | GLM-5.3-Flash |
| General chat and writing | GLM-5.3-Flash | GLM-4.7-Flash |
| Screenshots, photos, video | GLM-5.3-Flash | GLM-4.6V-Flash (free) |
| 1M-token documents or repos | GLM-5.3 | GLM-5.3-Flash |
| Non-reasoning, low latency | GLM-5.2 (effort none) | GLM-4.7-Flash (thinking off) |
| Math and science | GLM-5.3 | GLM-5.2 |
| Image generation | GLM-Image | CogView-4 |
| Local deployment | GLM-5.2 or GLM-5.3-Flash | GLM-4.7-Flash |
The GLM-5 line: GLM-5 to GLM-5.3
The GLM-5 generation started in February 2026 and has produced a new flagship roughly every two months. GLM-5, GLM-5.1, GLM-5.2 and GLM-5.3 share one architecture: 744B total parameters, 40B active, with DeepSeek Sparse Attention. GLM-5.3-Flash breaks the pattern with a new, smaller base model.
GLM-5.3
Same base model as GLM-5.2, with every gain from post-training. Announced August 14, 2026, listed in the API release notes on August 18 and published on Hugging Face on August 25 after a safety evaluation. Text-only input, 1M context, 128K output, forced reasoning with low, high or max effort. Besides coding, Z.ai highlights an emergent cyber capability: 84.5% on CyberGym, the best result on that benchmark, and 2,436 real vulnerabilities found across 269 projects with partner security teams. Weights come in FP8 (zai-org/GLM-5.3) and BF16 (zai-org/GLM-5.3-BF16). Full details on the GLM-5.3 page.
GLM-5.3-Flash and GLM-5.3-FlashX
A newly trained 320B-total, 18B-active model with 45 layers, the first GLM with hybrid linear and sparse attention, plus IndexPool and Manifold-Constrained Hyper-Connections. It was pre-trained on a 30T-token multimodal corpus and is the first natively multimodal GLM-5 model. Before launch it ran anonymously as “ox-alpha” on OpenCode and OpenRouter. GLM-5.3-FlashX is the API-only speed variant at about 200 tokens per second. Read the GLM-5.3-Flash guide.
GLM-5.2
Released June 16, 2026 with MIT weights on the same day. It brought the 1M-token context to the GLM line, IndexShare (one attention indexer shared by every four sparse layers, 2.9x fewer per-token FLOPs at 1M context) and an improved multi-token prediction layer with up to 20% longer speculative-decoding acceptance. Z.ai reports 62.1 on SWE-bench Pro and 81.0 on Terminal Bench 2.1. GLM-5.2 specs, benchmarks and OpenRouter access.
GLM-5.1
Released April 7, 2026 for long-horizon agent work, able to run one task for up to 8 hours. Z.ai describes it as broadly aligned with Claude Opus 4.6 and reports 58.4 on SWE-bench Pro, a top score at launch. 200K context, thinking on by default but switchable. GLM-5.1 specs and pricing.
GLM-5 and GLM-5-Turbo
GLM-5 (February 12, 2026) roughly doubled the parameter count of GLM-4.5, from 355B to 744B, raised pre-training data to 28.5T tokens, adopted DeepSeek Sparse Attention and was trained with Z.ai’s asynchronous RL framework, slime. It reports 77.8 on SWE-bench Verified and 56.2 on Terminal Bench 2.0. At $1.00 / $3.20 it remains the cheapest 744B GLM. GLM-5-Turbo, its OpenClaw-tuned variant, is API-only. Both are covered on the GLM-5 page.
How the GLM-5 models compare on Z.ai’s benchmarks
Z.ai publishes a benchmark table with each launch. Putting the overlapping rows side by side shows how fast the line has moved. Scores for GLM-5.3 and GLM-5.3-Flash come from the GLM-5.3 and GLM-5.3-Flash announcements; GLM-5.1 scores come from the GLM-5.2 table. All are Z.ai’s reported numbers.
| Benchmark | GLM-5.3 | GLM-5.3-Flash | GLM-5.2 | GLM-5.1 |
|---|---|---|---|---|
| Z.ai Code Bench (max effort) | 34.5% | 29.0% | 23.4% | — |
| DeepSWE v1.1 | 66.9 | 63.4 | 46.2 | 18.0 |
| AutomationBench | 48.2 | 48.8 | 26.2 | — |
| Terminal Bench 2.1 | 88.2 | — | 81.0 | 63.5 |
| NL2Repo | 58.0 | — | 48.9 | 42.7 |
| HLE with tools | 62.5 | — | 54.7 | 52.3 |
Two things jump out. GLM-5.3-Flash lands close to GLM-5.3 on agentic coding while costing about a tenth as much, which is why it is the default in GLM Chat. And GLM-5.2 more than doubled GLM-5.1’s DeepSWE score, from 18.0 to 46.2, before GLM-5.3 added about 21 more points on the same base model. These are vendor benchmarks, so test your own tasks too: each model page includes a five-task run in GLM Chat’s own test harness.
The GLM-4.x line: GLM-4.5 to GLM-4.7-Flash
The GLM-4.5 series (July 2025) introduced the Mixture-of-Experts design and hybrid reasoning that every later GLM builds on. These models are cheaper than the GLM-5 line and have simpler, fully switchable thinking, which keeps them useful in production.
GLM-4.7 and GLM-4.7-FlashX
GLM-4.7 (December 22, 2025) keeps GLM-4.5’s 32B-active design and focuses on coding: Z.ai reports 73.8% on SWE-bench Verified, 66.7% on SWE-bench Multilingual and 84.9 on LiveCodeBench v6. It added turn-level thinking and preserved thinking. GLM-4.7-FlashX is a paid, faster sibling of the free Flash model at $0.07 / $0.40. GLM-4.7 coding performance.
GLM-4.7-Flash
Released January 19, 2026 as the free-tier version of GLM-4.7. A 30B-A3B MoE model, MIT-licensed, with 200K context. Its model card shows 91.6 on AIME 25, 75.2 on GPQA and 79.5 on τ²-Bench. It is the most downloaded GLM repository on Hugging Face, at more than 1.8 million downloads a month. How to use GLM-4.7-Flash for free.
GLM-4.6
Released September 30, 2025. It raised the context from 128K to 200K, kept 128K output and, per Z.ai, is over 30% more token-efficient than GLM-4.5. Hybrid thinking is on by default. Note that there is no “GLM-4.6-Air”: the lightweight open model of that period is GLM-4.5-Air, later joined by GLM-4.7-Flash. GLM-4.6 specs and open weights.
GLM-4.5, GLM-4.5-Air, X, AirX and Flash
Released July 28, 2025. GLM-4.5 is 355B total / 32B active; GLM-4.5-Air is 106B / 12B. Both have 128K context, 96K output, MIT weights and a default temperature of 0.6 (the newer models default to 1.0). At launch Z.ai ranked GLM-4.5 second overall and first among open models on the average of 12 benchmarks. GLM-4.5-X and GLM-4.5-AirX are high-speed paid variants, and GLM-4.5-Flash is free. Read how GLM-4.5 aged and the GLM-4.5-Air guide.
Vision, image and other GLM models
Beyond the text line, Z.ai runs a set of specialist models on the same API and account. Prices are per million tokens unless noted.
| Model | Released | What it does | Context | Price |
|---|---|---|---|---|
| GLM-5V-Turbo | — | Multimodal coding model for agents | 200K | — |
| GLM-4.6V | Dec 8, 2025 | Image, video, file understanding with function calling | 128K | $0.30 / $0.90 |
| GLM-4.6V-FlashX | — | Fast, cheap vision | 128K | $0.04 / $0.40 |
| GLM-4.6V-Flash | — | Free vision | 128K | Free |
| GLM-4.5V | Aug 11, 2025 | 100B-scale open vision reasoning | 64K | $0.60 / $1.80 |
| GLM-Image | Jan 14, 2026 | Text-to-image, strong text rendering | — | $0.015 per image |
| GLM-OCR | Feb 3, 2026 | Document parsing and extraction | — | $0.03 / $0.03 |
| GLM-ASR-2512 | Dec 10, 2025 | Speech recognition | — | $0.03 (about $0.0024 per minute) |
| CogView-4 | — | Image generation | — | $0.01 per image |
| CogVideoX-3 | Jul 15, 2025 | Video generation | — | $0.20 per video |
GLM-Image deserves a closer look. It combines a 9B autoregressive model with a 7B diffusion decoder and a Glyph Encoder, and is built for images with text in them: posters, slides, diagrams. Z.ai reports a CVTG-2K word accuracy of 0.9116. Weights are MIT. On the API it uses POST /images/generations with model glm-image, a default size of 1280×1280 (custom sizes from 1024 to 2048 pixels, divisible by 32), and returns a URL that expires after 30 days. In GLM Chat you can generate images with GLM-Image, three a day.
For vision, GLM-5.3-Flash now covers most needs on its own. The older GLM-4.6V family remains useful when you want a free vision model (GLM-4.6V-Flash) or the very low input price of GLM-4.6V-FlashX.
GLM model names explained: Flash, FlashX, Air, X, AirX, Turbo, V
GLM model names follow a pattern: a version number, then a suffix that tells you the size or speed tier. Once you know the suffixes, the model list reads easily.
| Suffix | Meaning | Examples |
|---|---|---|
| (none) | Full-size model of that version | GLM-5.3, GLM-4.7 |
| Flash | Lightweight model; free on the API through GLM-4.7, cheap paid from GLM-5.3 | GLM-4.7-Flash, GLM-5.3-Flash |
| FlashX | Paid, faster variant of a Flash model, API only | GLM-5.3-FlashX, GLM-4.7-FlashX |
| Air | Smaller open sibling of a flagship | GLM-4.5-Air |
| X | High-speed paid variant of the full model | GLM-4.5-X |
| AirX | High-speed paid variant of Air | GLM-4.5-AirX |
| Turbo | Variant tuned for agent workflows | GLM-5-Turbo, GLM-5V-Turbo |
| V | Vision model (reads images and video) | GLM-4.6V, GLM-4.5V |
| -FP8 / -BF16 | Weight precision on Hugging Face | GLM-5.2-FP8, GLM-5.3-BF16 |
Two naming traps catch people. First, “Flash” no longer means free: GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash are free on the API, but GLM-5.3-Flash costs $0.15 / $0.50. Second, the precision suffix flips between releases: for GLM-5.2, GLM-5.1 and GLM-5 the plain repository holds BF16 weights and the -FP8 repository the FP8 build, while for GLM-5.3 and GLM-5.3-Flash the plain repository is FP8 and the -BF16 repository holds BF16. API model IDs are always lowercase (glm-5.3-flash, glm-4.5-air).
The older name ChatGLM referred to Z.ai’s chat-tuned models from 2023, such as the open-source ChatGLM-6B. Current models drop the “Chat” prefix.
Thinking and reasoning controls by model
Every current GLM model can reason before answering, but the controls differ, and getting them wrong is the most common reason a request fails after switching models. Reasoning tokens are billed as output tokens, so these settings also drive your cost.
| Model | Default | Can disable? | reasoning_effort |
|---|---|---|---|
| GLM-5.3, GLM-5.3-Flash | On (forced) | No (request fails) | low / high / max (default max) |
| GLM-5.2 | On | Yes, effort none or minimal | high / max (default max) |
| GLM-5.1, GLM-5 | On | Yes | — |
| GLM-4.7 series | On, turn-level | Yes | — |
| GLM-4.6, GLM-4.5 series | Hybrid (dynamic) | Yes | — |
On GLM-5.2, low and medium map to high, and xhigh maps to max, so code written for other providers keeps working. Full examples are in the GLM thinking mode guide.
Routing, migration and older GLM models
Z.ai has not published retirement dates for older GLM models on the pay-as-you-go API; GLM-4.5 through GLM-5.2 are all still listed with prices. The changes that matter today are on the Coding Plan and in the GLM-5.3 migration.
Coding Plan routing
- Every Coding Plan tier (Lite, Pro, Max) includes GLM-5.3 and GLM-5.3-Flash.
- Requests for
glm-5.2orglm-5.1are automatically routed to GLM-5.3. - Requests for
glm-4.7are routed to GLM-5.3-Flash. - GLM-5.3-Flash gets 3x the quota of GLM-5.3 on every tier. GLM-5.3-FlashX is not on the Coding Plan yet.
So if your Claude Code or Cline config still says GLM-5.2, you are already running GLM-5.3 on the plan. The plan works only inside supported coding tools; for your own apps, use the pay-as-you-go API, where every model ID in the master table still resolves to that model. The GLM Coding Plan guide explains tiers and credits.
Moving an app to GLM-5.3
- If your code sends
"thinking": {"type": "disabled"}, change it toenabledand setreasoning_efforttolowbefore switching the model ID. Otherwise the request fails. - Change the model to
glm-5.3. Defaults are temperature 1.0 and top_p 0.95; Z.ai recommends tuning only one of them. - Handle
delta.reasoning_contentseparately fromdelta.contentwhen streaming. - For streamed tool calls, set both
streamandtool_streamto true and concatenate the argument fragments. - Test latency and cost:
maxeffort, the default, produces the most reasoning tokens.
The GLM API quickstart has working curl, Python and Node code for every step, and the GLM release timeline tracks each model change as it happens.
Where each GLM model is available
You can reach GLM through five channels. OpenRouter prices change often and vary by provider, so treat them as OpenRouter’s listed price and confirm on the model’s OpenRouter page before you commit.
| Model | GLM Chat | Z.ai API | OpenRouter (listed, in / out) | Coding Plan | Hugging Face |
|---|---|---|---|---|---|
| GLM-5.3 | Yes | glm-5.3 | $1.40 / $4.40 | Yes | zai-org/GLM-5.3 |
| GLM-5.3-Flash | Yes (default) | glm-5.3-flash | $0.15 / $0.50 | Yes, 3x quota | zai-org/GLM-5.3-Flash |
| GLM-5.2 | Yes | glm-5.2 | ~$0.65 / $2.04; free variant | Routes to 5.3 | zai-org/GLM-5.2 |
| GLM-5.1 | Yes | glm-5.1 | $0.97 / $3.04 | Routes to 5.3 | zai-org/GLM-5.1 |
| GLM-5 | No | glm-5 | $0.60 / $1.92 | — | zai-org/GLM-5 |
| GLM-5-Turbo | No | glm-5-turbo | $1.20 / $4.00 | — | Not published |
| GLM-4.7 | Yes | glm-4.7 | $0.40 / $1.75 | Routes to 5.3-Flash | Yes |
| GLM-4.7-Flash | Yes (unlimited) | glm-4.7-flash (free) | $0.06 / $0.40 | — | Yes |
| GLM-4.6 | Yes | glm-4.6 | $0.43 / $1.75 | — | Yes |
| GLM-4.5 | Yes | glm-4.5 | $0.60 / $2.20 | — | Yes |
| GLM-4.5-Air | No | glm-4.5-air | $0.13 / $0.85 | — | Yes |
OpenRouter also lists z-ai/glm-5.2:free, a free, rate-limited GLM-5.2 with a 32,768-token context, and z-ai/glm-5.3:batch at $0.45 / $2.00. Z.ai’s own chat app, chat.z.ai, offers GLM-5.3 and other GLM models for free. On Hugging Face every open GLM lives under zai-org, with ModelScope mirrors under ZhipuAI; supported local serving stacks include vLLM, SGLang, Transformers, KTransformers and Unsloth. For licensing details per model, see is GLM open source.
In GLM Chat, the paid models allow 40 messages a day per visitor, GLM-4.7-Flash is unlimited at up to one request every 3 seconds, and each message can run to 8,000 characters. Deep links open a specific model straight away: chat with GLM-5.2, chat with GLM-4.7, chat with GLM-4.6 or chat with GLM-4.5. That makes it easy to put the same prompt to two generations and see the difference for yourself.
GLM models FAQ
What is the latest GLM model?
GLM-5.3 is the latest flagship (announced August 14, 2026). GLM-5.3-Flash, released August 26, 2026, is the newest model overall and the cheaper multimodal option. Z.ai has announced no date or specs for GLM-5.5 or GLM-6; it has said only that it is scaling the GLM-5.3-Flash recipe to larger models.
How many GLM models are there?
Z.ai’s docs cover 17 GLM language models (from GLM-4-32B-0414-128K to GLM-5.3), five vision models (GLM-5V-Turbo, GLM-4.6V, GLM-4.6V-FlashX, GLM-4.6V-Flash, GLM-4.5V) and GLM-Image, GLM-OCR and GLM-ASR-2512. For most people, four matter: GLM-5.3, GLM-5.3-Flash, GLM-5.2 and GLM-4.7-Flash.
Which GLM model is best for coding?
GLM-5.3. Z.ai reports a 50% improvement over GLM-5.2 on its in-house coding benchmark and open-source state-of-the-art results on Terminal Bench 3.0. For a cheaper option, GLM-5.3-Flash beats GLM-5.2 on six coding and agent benchmarks according to Z.ai, at about a tenth of the price.
Which GLM model is free?
On the Z.ai API, GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash are free for input and output. In GLM Chat, every model on the model chips is free to use within daily limits, and GLM-4.7-Flash has no daily cap.
What is the difference between GLM-5.3 and GLM-5.2?
They share the same 744B base model, context (1M) and price ($1.40 / $4.40). GLM-5.3 adds post-training that lifts coding and cyber scores sharply, but it always reasons, while GLM-5.2 can skip thinking. GLM-5.2 has a plain MIT license; GLM-5.3 uses the GLM-5.3 License.
Which GLM models have a 1M context window?
GLM-5.3, GLM-5.3-Flash (and FlashX) and GLM-5.2. GLM-5.1, GLM-5, GLM-4.7 and GLM-4.6 have 200K; the GLM-4.5 series has 128K.
Is there a GLM-4.6 Air model?
No. Z.ai never released a GLM-4.6-Air. If you want a lightweight open GLM, use GLM-4.5-Air (106B total, 12B active) or the newer GLM-4.7-Flash (31B total, 3B active, free on the API).
Can I try every GLM model before paying?
Most of them. The free GLM chat on this site covers GLM-5.3, GLM-5.3-Flash, GLM-5.2, GLM-5.1, GLM-4.7, GLM-4.6, GLM-4.5, GLM-4.7-Flash and GLM-Image with no account. For GLM-5 and GLM-4.5-Air, use their closest chat siblings (GLM-5.1 and GLM-4.5) or call the API.