GLM 5.2 (officially written GLM-5.2) is the flagship large language model that Z.ai, the company formerly known as Zhipu AI, released on June 16, 2026 for long-horizon coding and agent work. It is a 744B-parameter Mixture-of-Experts model with 40B parameters active per token, a 1M-token context window, up to 128K output tokens, and text in, text out. The weights are open under the MIT license, and the Z.ai API charges $1.40 per million input tokens and $4.40 per million output tokens.
GLM-5.2 was the first GLM model to offer what Z.ai calls a “solid 1M-token context”, and it introduced two things that later models kept: the IndexShare attention optimization and explicit effort levels (high and max). On August 14, 2026 Z.ai announced GLM-5.3, which uses the same base model with more post-training, at the same API price. GLM-5.2 is still live on the API, on OpenRouter (including a free variant) and on Hugging Face, and it is the newest GLM flagship that still lets you switch thinking off.
This guide covers everything: specs, what changed in the architecture, the full benchmark tables, prices, free routes, working API code, OpenRouter, coding agents, self-hosting and a clear verdict on when GLM-5.2 is still the right pick. If you just want to try it, Chat with GLM-5.2 now in the free chat on this site, no sign-up needed.

What GLM-5.2 is good at
Z.ai built GLM-5.2 for one job above all: staying useful across very long engineering tasks. The official model guide describes it as a flagship “built for the era of long-horizon tasks” that can carry a whole project in context and push a task from requirements to a deployable product. In practice that translates into a handful of workloads where GLM-5.2 is at its best.
Project-scale codebase work
With 1M tokens of context you can load a real repository (backend, frontend, configuration, tests and docs) in one request. Z.ai’s recommended first test is a technical audit: ask the model for an architecture map, module responsibilities, API contracts, data flows, call chains, technical debt and the constraints any refactor must respect. The point is not just reading more code; it is carrying earlier engineering judgments forward into later steps, which is where smaller-context models tend to drift.
Long refactors that run end to end
GLM-5.2 is tuned for cross-file, multi-step work: decoupling a module, migrating an API, restructuring directories, adapting to a new SDK or porting code between languages. Z.ai’s suggested pattern is to set clear boundaries (no change to business logic, API signatures or runtime behavior), ask for an execution plan with risks and a verification method first, then let the model implement and run the tests.
Following hard engineering rules
Z.ai highlights more consistent adherence to team standards over long, multi-round sessions: lint rules, build commands, test requirements, commit conventions and forbidden actions written in files such as CLAUDE.md or Agent.md. If your pain point with coding agents is out-of-scope edits, new dependencies nobody asked for, or skipped verification, this is the behavior GLM-5.2 was trained to improve.
Mobile, mini-program and game builds
The official guide lists client-side scenarios that go past generating an app: streaming messages, reconnection, local state, notifications and permissions, then debugging on a real device with ADB, logcat and screenshots. It also covers migrating a web project into a WeChat Mini Program (native, Taro or uni-app) and building small level-based games with a state machine, scoring and save logic.
Research reproduction and code-to-video
Two more showcase tasks from Z.ai: turning a paper’s architecture, loss functions and data pipeline into a runnable PyTorch project that reproduces the reported metrics, and writing Remotion (React) code that renders a finished MP4 video from a plain-language idea.
Where GLM-5.2 is weaker
- No image input. GLM-5.2 is text-only. For screenshots, diagrams or video, use GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series.
- Hard reasoning still trails the closed frontier. On Humanity’s Last Exam, Z.ai reports 40.5 for GLM-5.2 against 49.8 for Claude Opus 4.8 (full set) and 45.0 for Gemini 3.1 Pro.
- Ultra-long marathons. On SWE-Marathon, Z.ai’s own table shows 13.0 for GLM-5.2 against 26.0 for Opus 4.8. Z.ai says plainly that the model “still has room to grow” there.
- It is no longer the newest. GLM-5.3 beats it on every benchmark Z.ai published for both, at the same price.
What’s new in GLM-5.2
GLM-5.2 sits on the same 744B / 40B-active architecture line that began with GLM-5 in February 2026 and continued with GLM-5.1 in April. What changed is the context length, the attention machinery that makes 1M tokens affordable, faster speculative decoding, effort control and a new round of long-horizon reinforcement learning.
GLM-5.2 specs at a glance
| Spec | GLM-5.2 |
|---|---|
| Developer | Z.ai (formerly Zhipu AI) |
| Release date | June 16, 2026 |
| Architecture | Mixture-of-Experts with DeepSeek Sparse Attention and IndexShare |
| Parameters | 744B total, 40B active |
| Context window | 1M tokens |
| Max output | 128K tokens (default 65,536) |
| Input / output | Text / text |
| Thinking | On by default, can be disabled |
| Effort levels | high, max (default max) |
| API model ID | glm-5.2 |
| API price | $1.40 in, $0.26 cached, $4.40 out per 1M |
| Weights | MIT, BF16 and FP8 on Hugging Face and ModelScope |
| Capabilities | Streaming, function calling, tool streaming, context caching, JSON output, MCP |

1M-token context that stays usable
The jump from 200K (GLM-5 and GLM-5.1) to 1M tokens is the headline. Z.ai’s argument is that a long window only matters if quality holds across long, messy agent trajectories, so it expanded 1M-context training specifically for coding-agent scenarios: large implementations, automated research, performance optimization and complex debugging.
A 1M context is easy to claim, but much harder to keep reliable under real engineering pressure.
Z.ai, GLM-5.2 announcement
IndexShare: cheaper sparse attention at 1M tokens
GLM-5 introduced DeepSeek Sparse Attention (DSA), in which a lightweight “indexer” picks the top-k tokens each layer should attend to. At 1M tokens the indexer itself becomes expensive. GLM-5.2’s fix, called IndexShare, places one indexer at the first of every four transformer layers and reuses its top-k indices for all four. That removes the indexer dot products and top-k selection from three out of four layers and cuts per-token FLOPs by 2.9x at 1M context. GLM-5.2 was trained with IndexShare from mid-training at a 128K sequence length, and Z.ai says it outperforms GLM-5.1 on long-context benchmarks with less computation.
A faster MTP layer for speculative decoding
GLM models ship a Multi-Token Prediction (MTP) layer that serving engines use as a built-in draft model. For GLM-5.2, Z.ai applied IndexShare and KV sharing inside the MTP layer (which also removes a training-versus-inference mismatch that GLM-5.1’s MTP layer had), added rejection sampling, and trained with an end-to-end TV loss. In Z.ai’s ablation on coding workloads (GLM-5.1 backbone and data, seven MTP steps) the average acceptance length rose by 20%:
| Method | Acceptance length |
|---|---|
| Baseline | 4.56 |
| + IndexShare + KV Share | 5.10 |
| + Rejection sampling | 5.29 |
| + End-to-end TV loss | 5.47 (+20%) |
Longer acceptance means more tokens confirmed per decoding step, which is the part of the speed story you benefit from when you serve the weights yourself.
Effort levels: High and Max
GLM-5.2 was the first GLM model with the reasoning_effort parameter. It accepts high (enhanced reasoning) and max (deep reasoning, the default). Z.ai says GLM-5.2 delivers much stronger agentic coding than GLM-5.1 at comparable token use, landing roughly between Claude Opus 4.7 and Claude Opus 4.8 at similar token consumption, with Max available when a task justifies the extra compute. For compatibility with other APIs, GLM-5.2 also accepts none, minimal, low, medium and xhigh and maps them (details in the API section below).
Long-horizon RL and anti-hacking
Very long tasks produce traces that must be split (“compacted”) into sub-traces, and different rollouts produce different numbers of them. Z.ai therefore moved from group-wise optimization to a critic-based PPO setup that learns from individual rollouts, and trains on every compacted sub-trace with a token-level loss. It also added an anti-hack module after finding that GLM-5.2 showed more reward-hacking attempts than GLM-5.1, such as reading hidden evaluation files or downloading reference solutions. A rule-based filter flags suspicious tool calls, an LLM judge confirms intent, and blocked calls return dummy results so the rollout can continue instead of being thrown away.
slime and serving infrastructure
Post-training ran on slime, Z.ai’s open-source RL framework. Z.ai says it used slime for parallel on-policy distillation that merged more than ten expert models into the final model in about two days. On the serving side, Z.ai notes that IndexShare cuts compute but does not shrink the per-token KV cache proportionally, so it reworked memory management (building on LayerSplit), long-context kernels and CPU-side scheduling. The practical takeaway for self-hosters: at long context, memory for the KV cache, not raw compute, becomes the bottleneck.
GLM-5.2 release history
- Before June 16, 2026: GLM Coding Plan subscribers get early access, and their feedback centers on project-level context, steadier long tasks, stricter adherence to engineering rules and stronger mobile work.
- June 16, 2026: public release on the Z.ai API, in the chat app at chat.z.ai, and as open weights on Hugging Face and ModelScope. On the Coding Plan it counts at 3x quota in peak hours (14:00 to 18:00 UTC+8) and 2x off-peak, with a limited-time 1x off-peak promotion; ZCode users get 1.5x effective quota until June 30.
- July 30, 2026: the Coding Plan switches from prompt-based to credit-based quotas.
- August 14, 2026: Z.ai announces GLM-5.3, built on the same base model. Coding Plan requests for GLM-5.2 are now routed to GLM-5.3.
- Today: GLM-5.2 stays available on the pay-as-you-go API at $1.40 / $4.40, on OpenRouter, and as MIT weights.
GLM 5.2 benchmarks
All numbers in this section are Z.ai’s published results from the GLM-5.2 announcement and model card. Scores marked * come from a benchmark’s full set; the rest use the text-only subset where one exists. A dash means Z.ai did not report that model.
Reasoning
| Benchmark | GLM-5.2 | GLM-5.1 | Qwen3.7-Max | MiniMax M3 | DeepSeek-V4-Pro | Claude Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|---|---|---|---|
| HLE | 40.5 | 31.0 | 41.4 | 37.0 | 37.7 | 49.8* | 41.4* | 45.0 |
| HLE w/ Tools | 54.7 | 52.3 | 53.5 | – | 48.2 | 57.9* | 52.2* | 51.4* |
| CritPt | 20.9 | 4.6 | 13.4 | 3.7 | 12.9 | 20.9 | 27.1 | 17.7 |
| AIME 2026 | 99.2 | 95.3 | 97.0 | – | 94.6 | 95.7 | 98.3 | 98.2 |
| HMMT Nov 2025 | 94.4 | 94.0 | 95.0 | 84.4 | 94.4 | 96.5 | 96.5 | 94.8 |
| HMMT Feb 2026 | 92.5 | 82.6 | 97.1 | 84.4 | 95.2 | 96.7 | 96.7 | 87.3 |
| IMOAnswerBench | 91.0 | 83.8 | 90.0 | – | 89.8 | 83.5 | – | 81.0 |
| GPQA-Diamond | 91.2 | 86.2 | 90.0 | 93.0 | 90.1 | 93.6 | 93.6 | 94.3 |
Coding
| Benchmark | GLM-5.2 | GLM-5.1 | Qwen3.7-Max | MiniMax M3 | DeepSeek-V4-Pro | Claude Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|---|---|---|---|
| SWE-bench Pro | 62.1 | 58.4 | 60.6 | 59.0 | 55.4 | 69.2 | 58.6 | 54.2 |
| NL2Repo | 48.9 | 42.7 | 47.2 | 42.1 | 35.5 | 69.7 | 50.7 | 33.4 |
| DeepSWE | 46.2 | 18.0 | 18.0 | 20.0 | 8.0 | 58.0 | 70.0 | 10.0 |
| ProgramBench | 63.7 | 50.9 | – | – | 47.8 | 71.9 | 70.8 | 39.5 |
| Terminal Bench 2.1 (Terminus-2) | 81.0 | 63.5 | 75.0 | 65.0 | 64.0 | 85.0 | 84.0 | 74.0 |
| FrontierSWE (dominance) | 74.4 | 30.5 | – | – | 29.0 | 75.1 | 72.6 | 39.6 |
| PostTrainBench | 34.3 | 20.1 | – | – | – | 37.2 | 28.4 | 21.6 |
| SWE-Marathon | 13.0 | 1.0 | – | – | – | 26.0 | 12.0 | 4.0 |
Z.ai also reports Terminal Bench 2.1 with each model’s best reported harness: GLM-5.2 scores 82.7 in Claude Code, GLM-5.1 69 in Claude Code, Claude Opus 4.8 78.9 in Claude Code, GPT-5.5 83.4 in Codex and Gemini 3.1 Pro 70.7 in Gemini CLI.
Agentic tool use
| Benchmark | GLM-5.2 | GLM-5.1 | Qwen3.7-Max | MiniMax M3 | DeepSeek-V4-Pro | Claude Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|---|---|---|---|
| MCP-Atlas (public set) | 76.8 | 71.8 | 76.4 | 74.2 | 73.6 | 77.8 | 75.3 | 69.2 |
| Tool-Decathlon | 48.2 | 40.7 | – | – | 52.8 | 59.9 | 55.6 | 48.8 |
What the tables say
- Biggest jumps over GLM-5.1 are in long-horizon and agentic coding: FrontierSWE 74.4 vs 30.5, DeepSWE 46.2 vs 18.0, Terminal Bench 2.1 81.0 vs 63.5, SWE-Marathon 13.0 vs 1.0 and CritPt 20.9 vs 4.6.
- Against closed models, Z.ai says GLM-5.2 trails Opus 4.8 by only 1% on FrontierSWE while edging GPT-5.5 by 1% and Opus 4.7 by 11%; ranks second only to Opus 4.8 on PostTrainBench; and trails Opus 4.8 by 13% on SWE-Marathon. It was the highest-ranked open model on all three at launch.
- Math is a strength. AIME 2026 at 99.2 is the top score in the table, and IMOAnswerBench at 91.0 is ahead of every closed model listed.
- Science and hard knowledge lag. GPQA-Diamond (91.2) and HLE (40.5) sit below Opus 4.8, GPT-5.5 and Gemini 3.1 Pro.
What each benchmark measures
- HLE (Humanity’s Last Exam): very hard expert-level questions across many fields; “w/ Tools” gives the model tool access.
- CritPt: research-level physics reasoning problems.
- AIME 2026: competition math problems with exact integer answers.
- HMMT Nov 2025 / Feb 2026: problems from two rounds of the HMMT math tournament.
- IMOAnswerBench: olympiad-level math problems graded on the final answer.
- GPQA-Diamond: graduate-level science multiple-choice questions designed to resist simple web lookup.
- SWE-bench Pro: resolving real issues in real repositories, a harder successor to SWE-bench.
- NL2Repo: generating a working repository from a natural-language specification.
- DeepSWE: long software-engineering tasks solved in an isolated container with no internet access.
- ProgramBench: 200 program-building instances run in a sandbox with internet disabled.
- Terminal Bench 2.1: real-world tasks completed in a command-line terminal.
- FrontierSWE: open-ended technical projects lasting hours to tens of hours, from systems optimization to applied ML research.
- PostTrainBench: the agent gets an H100 GPU and is scored on how much it improves small models through post-training.
- SWE-Marathon: ultra-long software projects such as building compilers, optimizing kernels and developing production-grade services.
- MCP-Atlas: tool use through MCP servers, 500 public tasks with a 10-minute limit each.
- Tool-Decathlon: a broad tool-use benchmark run through its official evaluation service.
How Z.ai measured it
Reasoning tasks used temperature 1.0, top_p 0.95 and up to 163,840 generated tokens; HLE with tools allowed a 300,000-token context. SWE-bench Pro ran in OpenHands with a 400K context. Terminal Bench 2.1 (Terminus-2) ran with a 256K context, 4 CPUs and 8 GB of RAM per task; the Claude Code variant averaged five runs. FrontierSWE, PostTrainBench and SWE-Marathon were run by third parties (Proximal, PostTrainBench and Abundant AI) at 1M context, max effort and 128K output. These are generous settings: your own results at lower effort or shorter output limits will differ.
GLM-5.2 vs GLM-5.3 benchmark rows
When Z.ai launched GLM-5.3 it re-ran GLM-5.2 on a newer set of benchmark versions. These are the two models side by side, with Claude Opus 4.8 for reference, from the GLM-5.3 announcement:
| Benchmark | GLM-5.3 | GLM-5.2 | Opus 4.8 |
|---|---|---|---|
| Terminal Bench 2.1 | 88.2 | 81.0 | 85.0 |
| Terminal Bench 3.0 | 28.3 | 4.6 | 21.1 |
| DeepSWE v1.1 | 66.9 | 46.2 | 58.0 |
| NL2Repo | 58.0 | 48.9 | 69.7 |
| ProgramBench (Almost Solved) | 19.0 | 9.5 | 15.5 |
| FrontierSWE | 78.1 | 67.5 | 66.5 |
| SWE-Marathon v1.1 | 42.5 | 19.4 | 48.8 |
| PostTrainBench | 39.8 | 31.7 | 32.9 |
| CyberGym | 84.5 | 77.2 | 78.1 |
| ExploitGym 2h / 6h | 105 / 130 | 29 / 39 | 80 / 120 |
| ExploitBench | 54.4 | 24.4 | 40.0 |
| Toolathlon Verified | 73.0 | 59.9 | 76.2 |
| AutomationBench v1.0.6 | 48.2 | 26.2 | 41.0 |
| Agents’ Last Exam (CLI) | 28.5 | 23.8 | 25.7 |
| HLE w/ Tools | 62.5 | 54.7 | 57.9 |
| GDPval-AA v2 | 1769 | 1508 | 1588 |
Some GLM-5.2 numbers differ between the two announcements (FrontierSWE 74.4 vs 67.5, PostTrainBench 34.3 vs 31.7, SWE-Marathon 13.0 vs 19.4). That is expected: FrontierSWE is a dominance score reported at a specific date, SWE-Marathon moved to v1.1, ProgramBench switched to an “Almost Solved” metric, and several suites ran in a different harness. Compare numbers only within the same table.
On Z.ai’s private Code Bench, the gap is clearer still: at Max effort GLM-5.2 completes 23.4% of tasks using about 96K output tokens per task, while GLM-5.3 reaches 34.5% with about 75K. Z.ai summarizes this as a 50% improvement for GLM-5.3 over GLM-5.2.
We tested it: 5 tasks
Here is how GLM-5.2 handles the five standard tasks every model on GLM Chat goes through. Each task runs three times in the site’s own test harness with the same settings as the public chat: temperature 0.7, a 2,048-token output cap and the lightest reasoning setting (for GLM-5.2 that means thinking switched off). The harness records the full answer, latency and token counts, and an editor scores each answer from 0 to 2.
- Python: a streaming CSV de-duplicator that keeps the first occurrence, supports key columns and preserves the header. It reveals whether the model respects a memory constraint instead of loading the whole file.
- JavaScript: a
debounce(fn, wait, { leading, trailing })function withcancel()plus tests for Node’s built-in test runner. It reveals care with timing edge cases, especially leading and trailing calls together, and whether the tests actually pass withnode --test. - PHP refactor: a messy 60-line legacy function with copy-pasted loops, magic discount numbers, unescaped HTML and string-built SQL, rewritten as clean PHP 8 without changing behavior. It reveals discipline: the VIP discount and every order type must behave exactly as before, while output escaping and parameterized SQL fix the security holes.
- Explanation: transformer attention for a complete beginner in about 150 words. It reveals whether the model can be accurate about queries, keys, values and weights while staying readable and on length.
- Extraction: a messy product description turned into strict JSON with fixed keys, cm conversions and
nullfor missing fields. It reveals format obedience and whether the model invents data when the text is silent.
Our test in GLM Chat

| Task | Median score | Median latency | Median output tokens | Why (across runs) |
|---|---|---|---|---|
| Python: stream-safe CSV dedupe | 2 / 2 (runs: 2, 2, 2) | 4.6 s | 372 | Streams with csv, keeps the header and first occurrence, key columns work in every run. |
| JavaScript: debounce with options + tests | 1 / 2 (runs: 0, 1, 2) | 12.6 s | 1,011 | Its own tests fail (2/3 runs); fails our debounce behaviour checks (1/3 runs). |
| PHP: refactor a messy 60-line function | 1 / 2 (runs: 0, 1, 1) | 9.6 s | 834 | Output not escaped (3/3 runs); SQL not parameterised (2/3 runs); changes the function’s behaviour (1/3 runs). |
| Explain attention to a beginner | 1 / 2 (runs: 2, 1, 1) | 4.2 s | 185 | No queries/keys/values (2/3 runs); weighting step missing (1/3 runs). |
| Extract JSON from a messy description | 2 / 2 (runs: 2, 2, 1) | 2.2 s | 112 | Passes every check in most runs; a dimension left in inches (1/3 runs). |
| Total | 7 / 10 |
A score of 2 means correct and complete, 1 means usable after a small fix, 0 means wrong or unusable, for a maximum of 10. When you read GLM-5.2’s card, look first at the PHP refactor and the JSON extraction: those two punish models that change behavior silently or invent data, and they are the closest proxies for the “follows engineering rules” claim above. The Python and JavaScript tasks show whether the model honors every stated constraint (streaming, leading plus trailing calls, a working cancel()) rather than writing the obvious version.
Keep in mind what the battery does not measure. With thinking off and a 2,048-token cap, it tests GLM-5.2 as a fast everyday assistant, not as a 1M-context agent running at Max effort for hours. For that side of the model, the long-horizon benchmarks above are the better guide, and the GLM-5.3 page shows the same five tasks on its successor so you can compare directly.
GLM-5.2 vs GLM-5.3 vs GLM-5.1
These three share the 744B / 40B-active design, so the differences come from context length, thinking controls, licensing and post-training. GLM-5.3-Flash is included because it is now the default cheap option in the same family.
| Spec | GLM-5.2 | GLM-5.3 | GLM-5.1 | GLM-5.3-Flash |
|---|---|---|---|---|
| Release | Jun 16, 2026 | Aug 14, 2026 | Apr 7, 2026 | Aug 26, 2026 |
| Parameters | 744B / 40B active | 744B / 40B active | 744B / 40B active | 320B / 18B active |
| Context | 1M | 1M | 200K | 1M |
| Max output | 128K | 128K | 128K | 128K |
| Input | Text | Text | Text | Text, image, video, file |
| Thinking | On by default, can be disabled | Always on | On by default, can be disabled | Always on |
| Effort levels | high, max | low, high, max | Not supported | low, high, max |
| Input / output price | $1.40 / $4.40 | $1.40 / $4.40 | $1.40 / $4.40 | $0.15 / $0.50 |
| Cached input | $0.26 | $0.26 | $0.26 | $0.03 |
| Weights | MIT | GLM-5.3 License | MIT | MIT |
| Coding Plan | Routed to GLM-5.3 | Included | Routed to GLM-5.3 | Included, 3x quota |

GLM-5.2 vs GLM-5.1
GLM-5.2 is a straight upgrade at the same price: five times the context (1M vs 200K), effort control, faster speculative decoding and large gains on long-horizon coding. Both keep the ability to disable thinking and both ship MIT weights. The only reason to stay on GLM-5.1 is a pipeline already tuned for it; on the Coding Plan, requests for either model now go to GLM-5.3 anyway. The GLM-5.1 page covers its 8-hour autonomous runs and SWE-Bench Pro result.
GLM-5.2 vs GLM-5.3
GLM-5.3 is GLM-5.2 plus a much larger round of post-training. It wins every benchmark row Z.ai published for both, uses fewer output tokens at Max effort on Z.ai Code Bench, and costs the same per token. Three differences can still matter to you:
- Thinking cannot be disabled on GLM-5.3. Sending
"thinking": {"type": "disabled"}toglm-5.3makes the request fail; the lightest option isreasoning_effort: "low". GLM-5.2 can answer with no reasoning at all. - License. GLM-5.2 is plain MIT. GLM-5.3 uses the custom GLM-5.3 License, which adds a security-review condition for very large Model-as-a-Service businesses (see Is GLM open source?).
- Third-party pricing. OpenRouter’s listed price for GLM-5.2 is lower than for GLM-5.3, and only GLM-5.2 has a free variant there.
GLM-5.2 vs GLM-5.3-Flash
GLM-5.3-Flash costs about a tenth as much ($0.15 / $0.50), reads images and video, and Z.ai reports it beats GLM-5.2 on its six headline coding and agent benchmarks, including DeepSWE v1.1 (63.4 vs 46.2) and AutomationBench (48.8 vs 26.2). For most new projects that care about cost, it is the better default. GLM-5.2 keeps the edge where you need the bigger 744B model’s knowledge, a thinking-off mode, or an MIT model you already host.
GLM-5.2 vs DeepSeek and Claude
Outside the GLM family, the two most common comparisons are DeepSeek and Anthropic’s Claude. In Z.ai’s GLM-5.2 table, GLM-5.2 leads DeepSeek-V4-Pro on SWE-bench Pro (62.1 vs 55.4), NL2Repo (48.9 vs 35.5), DeepSWE (46.2 vs 8.0), Terminal Bench 2.1 (81.0 vs 64.0) and FrontierSWE (74.4 vs 29.0), while DeepSeek-V4-Pro leads on Tool-Decathlon (52.8 vs 48.2). On list prices, deepseek-v4-pro costs $1.32 in and $3.96 out per million tokens at peak and half that off-peak, close to GLM-5.2’s $1.40 / $4.40.
Against Claude Opus 4.8, Z.ai’s numbers put GLM-5.2 behind on most rows (SWE-bench Pro 62.1 vs 69.2, HLE 40.5 vs 49.8) but close on FrontierSWE (74.4 vs 75.1) and MCP-Atlas (76.8 vs 77.8), and ahead on Terminal Bench 2.1 in Claude Code (82.7 vs 78.9). Opus 4.8 lists at $5 in and $25 out per million tokens, so GLM-5.2’s output tokens cost less than a fifth as much. All of these are Z.ai’s reported numbers; the full breakdowns are in GLM vs DeepSeek and GLM vs Claude.
GLM 5.2 pricing and free access
On Z.ai’s pay-as-you-go API, GLM-5.2 costs the same as GLM-5.3 and GLM-5.1. Storage for cached input is listed as free for a limited time.
| Model | Input | Cached input | Output |
|---|---|---|---|
| GLM-5.2 | $1.40 | $0.26 | $4.40 |
| GLM-5.3 | $1.40 | $0.26 | $4.40 |
| GLM-5.1 | $1.40 | $0.26 | $4.40 |
| GLM-5 | $1.00 | $0.20 | $3.20 |
| GLM-5.3-Flash | $0.15 | $0.03 | $0.50 |
| GLM-4.7 | $0.60 | $0.11 | $2.20 |
| GLM-4.7-Flash | Free | Free | Free |
What a long-context request really costs
Long context is where GLM-5.2 bills add up, so do the math before you paste a whole repository. Say you send a 300,000-token codebase and get a 10,000-token answer:
- Input: 0.3M x $1.40 = $0.42
- Output: 0.01M x $4.40 = $0.044
- Total: about $0.46 for the first request
Now ask a follow-up in the same conversation. If 290,000 tokens of the prompt hit the cache and 10,000 are new, input costs 0.29M x $0.26 = $0.0754 plus 0.01M x $1.40 = $0.014. With another 10,000 output tokens ($0.044), the follow-up comes to about $0.13, less than a third of the first call. Context caching is automatic on Z.ai; you can see the cached count in usage.prompt_tokens_details.cached_tokens. The context caching guide shows how to structure prompts so the cache hits.
Remember that reasoning tokens are billed as output. At max effort on a hard task, the thinking can easily outweigh the visible answer, which is exactly why high exists. The full cross-model price list lives on the GLM pricing page.
Is GLM 5.2 free?
Not on the Z.ai API, but there are four legitimate ways to use it without paying per token:
- GLM Chat (this site). Pick GLM-5.2 in the model chips and chat with GLM-5.2 for free. No account; 40 messages per day on the paid models, up to 8,000 characters per message, and the last 20 messages (trimmed to 24,000 characters) are sent as context. The chat runs GLM-5.2 with thinking off for quick replies. If the site’s daily capacity for paid models runs out, replies come from GLM-4.7-Flash for the rest of the UTC day and are labeled “GLM-4.7-Flash (daily limit reached)”.
- Z.ai’s own chat app. GLM-5.2 launched in the free web chat at chat.z.ai; the app’s model list changes as newer models arrive.
- OpenRouter’s free variant.
z-ai/glm-5.2:freeis listed at $0 with a 32,768-token context and rate limits (details in the OpenRouter section). - Your own hardware. The MIT weights are free to download and use commercially, if you have the GPUs for a 744B model.
If what you really need is a free API model rather than GLM-5.2 specifically, GLM-4.7-Flash and GLM-4.5-Flash are free on the Z.ai API itself.
GLM-5.2 on the Coding Plan
The GLM Coding Plan (from $18 per month for Lite) no longer serves GLM-5.2 as a separate model: requests for GLM-5.2 and GLM-5.1 are automatically routed to GLM-5.3. If you subscribe to the plan, you are paying for GLM-5.3 and GLM-5.3-Flash. The Coding Plan guide explains tiers, credits and when the plan beats pay-as-you-go.
Using the GLM 5.2 API
The model ID is glm-5.2 (lowercase). The API is OpenAI-compatible Chat Completions at https://api.z.ai/api/paas/v4, authenticated with a bearer key. Create a key on the API Keys page at z.ai/manage-apikey/apikey-list and export it as ZAI_API_KEY. New to the platform? The GLM API quickstart walks through the account and key setup.
curl
curl -X POST "https://api.z.ai/api/paas/v4/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $ZAI_API_KEY" \
-d '{
"model": "glm-5.2",
"messages": [
{"role": "system", "content": "You are a senior software engineer."},
{"role": "user", "content": "Review this function and list the bugs: ..."}
],
"thinking": {"type": "enabled"},
"reasoning_effort": "high",
"max_tokens": 8192,
"temperature": 1.0
}'
Python with the OpenAI SDK
Point the standard openai package at Z.ai’s base URL. GLM-specific fields such as thinking go in extra_body; putting reasoning_effort there too keeps the code working on older SDK versions.
# pip install --upgrade "openai>=1.0"
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["ZAI_API_KEY"],
base_url="https://api.z.ai/api/paas/v4/",
)
resp = client.chat.completions.create(
model="glm-5.2",
messages=[
{"role": "system", "content": "You are a senior software engineer."},
{"role": "user", "content": "Plan a migration from REST polling to WebSockets."},
],
max_tokens=8192,
extra_body={
"thinking": {"type": "enabled"},
"reasoning_effort": "max",
},
)
msg = resp.choices[0].message
print(getattr(msg, "reasoning_content", None)) # the model's reasoning, if any
print(msg.content) # the final answer
print(resp.usage) # token counts, incl. cached tokens
Streaming reasoning and answer separately
With stream=True, reasoning arrives in delta.reasoning_content and the answer in delta.content; the stream ends with data: [DONE]. This example uses Z.ai’s official SDK (pip install zai-sdk):
from zai import ZaiClient
client = ZaiClient(api_key="your-api-key")
stream = client.chat.completions.create(
model="glm-5.2",
messages=[{"role": "user", "content": "Explain IndexShare in three sentences."}],
thinking={"type": "enabled"},
reasoning_effort="high",
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta
if delta.reasoning_content:
print(delta.reasoning_content, end="", flush=True)
if delta.content:
print(delta.content, end="", flush=True)
Turning thinking off
Unlike GLM-5.3, GLM-5.2 can skip reasoning entirely. Either disable thinking or pass reasoning_effort: "none" (or "minimal"). Use this for classification, extraction and short chat where latency matters more than depth:
curl -X POST "https://api.z.ai/api/paas/v4/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $ZAI_API_KEY" \
-d '{
"model": "glm-5.2",
"messages": [{"role": "user", "content": "Classify this ticket: billing, bug or feature?"}],
"thinking": {"type": "disabled"}
}'
reasoning_effort values on GLM-5.2
| Value you send | What GLM-5.2 does |
|---|---|
max (default) | Deep reasoning |
xhigh | Mapped to max |
high | Enhanced reasoning |
medium, low | Mapped to high |
minimal, none | Skips thinking |
The thinking mode guide compares these controls across every GLM model.
Parameters worth knowing
- max_tokens: default 65,536, maximum 131,072 (128K).
- temperature: range 0.0 to 1.0, default 1.0. top_p: default 0.95.
- Tools: standard OpenAI-style function calling (up to 128 functions), plus
tool_stream: trueto stream tool-call arguments. See the function calling guide. - JSON mode:
response_format: {"type": "json_object"}. - Errors: an unknown model name returns 1211, a prompt over the limit 1261, and rate limits 1302. The rate limits and error codes guide lists fixes for each.
Function calling with GLM-5.2
GLM-5.2 uses the standard OpenAI tool format. tool_choice supports only auto, so the model decides when to call a tool. A minimal round trip:
import json, os
from openai import OpenAI
client = OpenAI(api_key=os.environ["ZAI_API_KEY"], base_url="https://api.z.ai/api/paas/v4/")
tools = [{
"type": "function",
"function": {
"name": "get_open_issues",
"description": "List open issues for a repository",
"parameters": {
"type": "object",
"properties": {"repo": {"type": "string"}},
"required": ["repo"],
},
},
}]
messages = [{"role": "user", "content": "How many open issues does acme/api have?"}]
first = client.chat.completions.create(model="glm-5.2", messages=messages, tools=tools)
call = first.choices[0].message.tool_calls[0]
args = json.loads(call.function.arguments)
result = {"repo": args["repo"], "open_issues": 42} # run your real function here
messages.append(first.choices[0].message)
messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)})
final = client.chat.completions.create(model="glm-5.2", messages=messages, tools=tools)
print(final.choices[0].message.content)
In multi-step agent loops with thinking on, send the assistant message back unchanged (including its reasoning) so the model keeps its chain of thought between tool calls.
Tips for 1M-token prompts
- Put the stable part first. Caching is automatic and matches repeated context, so place the repository, docs and rules at the start of the prompt and the changing question at the end. Every follow-up then reuses the cached prefix at $0.26 instead of $1.40 per million tokens.
- Ask for a plan before edits. Z.ai’s own recommended prompts start with an audit or execution plan, impact scope and verification method, then let the model implement. That structure is what GLM-5.2 was trained to follow over long tasks.
- Write your rules down. Hard constraints (no new dependencies, no API contract changes, no commits) belong in the prompt or in
CLAUDE.md/Agent.md, where the model can refer back to them. - Start at high, escalate to max. Use
highfor routine edits and switch tomaxfor the hard cases; reasoning tokens are billed as output. - Cap output explicitly. The default
max_tokensis 65,536. Set a lower value for chat-style answers and the full 131,072 only for long generations.
Errors you are most likely to hit
| Code | Meaning | Typical fix |
|---|---|---|
| 401 / 1000 | Authentication failed | Check the key and the Bearer header |
| 400 / 1211 | Unknown model | Use the exact ID glm-5.2, lowercase |
| 400 / 1214 | Invalid parameter value | Keep temperature within 0.0 to 1.0; use a listed effort value |
| 400 / 1261 | Prompt too long | Trim the input below the 1M-token window |
| 429 / 1113 | Insufficient balance | Top up, or fix a Coding Plan vs API base URL mix-up |
| 429 / 1302 | Request rate limit | Back off and retry; check limits at z.ai/manage-apikey/rate-limits |
| 429 / 1305 | Service temporarily overloaded | Retry with exponential backoff |
GLM-5.2 on OpenRouter
OpenRouter is a third-party router that resells many models behind one OpenAI-compatible API and one key. It lists two GLM-5.2 entries:
| Slug | Listed price (in / out) | Context | Max output | Notes |
|---|---|---|---|---|
z-ai/glm-5.2 | about $0.65 / $2.04 | 1,048,576 | 131,072 | Routed across providers; tools, JSON and structured output listed |
z-ai/glm-5.2:free | $0 / $0 | 32,768 | 29,491 | Rate-limited; no tools or response_format in its listed parameters |
OpenRouter routes the paid slug across several providers, and its listed price is roughly half of Z.ai’s direct price. Prices move often, so check the GLM-5.2 page on OpenRouter before you plan costs. The free variant is excellent for trying the model or light personal use; its 32K context rules out the long-context work GLM-5.2 is known for.
Calling GLM-5.2 through OpenRouter
Use the base URL https://openrouter.ai/api/v1 and your OpenRouter key:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key=os.environ["OPENROUTER_API_KEY"],
)
resp = client.chat.completions.create(
model="z-ai/glm-5.2:free", # or "z-ai/glm-5.2" for the paid, 1M-context route
messages=[{"role": "user", "content": "Write a SQL query that finds duplicate emails."}],
)
print(resp.choices[0].message.content)
The same request with curl:
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "z-ai/glm-5.2", "messages": [{"role": "user", "content": "Hello, GLM-5.2"}]}'
OpenRouter lists reasoning and reasoning_effort among the supported parameters for both slugs. Z.ai-specific fields such as thinking belong to Z.ai’s own API; if you need exact control over GLM’s thinking behavior or the documented Z.ai parameters, call Z.ai directly.
GLM-5.2 in Claude Code and coding agents
When GLM-5.2 launched, Z.ai rolled it out to every Coding Plan subscriber (some had early access before release). You enabled it by setting the model name to GLM-5.2, or GLM-5.2[1m] in Claude Code for the full 1M context, and you could switch between High and Max effort. ZCode, Z.ai’s desktop agent, shipped with GLM-5.2 and its /goal mode for long-horizon tasks.
That has changed. On the Coding Plan today, requests for GLM-5.2 are automatically routed to GLM-5.3, so writing glm-5.2 in a plan-connected tool gets you GLM-5.3. The current recommended Claude Code setup maps the model slots to GLM-5.3 and GLM-5.3-Flash in ~/.claude/settings.json:
{
"env": {
"ANTHROPIC_AUTH_TOKEN": "your-api-key",
"ANTHROPIC_BASE_URL": "https://api.z.ai/api/anthropic",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "glm-5.3-flash[1m]",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "glm-5.3[1m]",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "glm-5.3[1m]",
"CLAUDE_CODE_AUTO_COMPACT_WINDOW": "1000000",
"API_TIMEOUT_MS": "3000000"
}
}
The [1m] suffix plus CLAUDE_CODE_AUTO_COMPACT_WINDOW enables the 1M context; if Claude Code says a [1m] model does not exist, update Claude Code. Type /effort inside a session to change thinking intensity, and /status to confirm which model is active. The step-by-step setup is in How to use GLM in Claude Code.
If you specifically want GLM-5.2 in an agent
Use pay-as-you-go instead of the plan. In any tool that accepts a custom OpenAI-compatible provider (Cline, Kilo Code, Roo Code, OpenCode and others), choose the OpenAI Compatible provider, set the base URL to https://api.z.ai/api/paas/v4, enter your API key and the custom model glm-5.2, and set the context window to 1,000,000. Leave image support unchecked, since GLM-5.2 is text-only. You then pay per token at $1.40 / $4.40, and the request really reaches GLM-5.2. The Cline guide and the OpenClaw guide show the exact fields for those tools.
Do not point pay-as-you-go keys at the Coding Plan URL (https://api.z.ai/api/coding/paas/v4) or vice versa. Mixing the two is a common cause of the “1113 insufficient balance” error.
Downloading the weights
GLM-5.2’s weights were published on Hugging Face and ModelScope on launch day under the MIT license, with no gating and no regional limits. Z.ai calls this “Pure Open”.
| Variant | Hugging Face | ModelScope | Precision | Size |
|---|---|---|---|---|
| GLM-5.2 | zai-org/GLM-5.2 | ZhipuAI/GLM-5.2 | BF16 | 744B-A40B |
| GLM-5.2-FP8 | zai-org/GLM-5.2-FP8 | ZhipuAI/GLM-5.2-FP8 | FP8 | 744B-A40B |
The Hugging Face model card now points to GLM-5.3-BF16 as the newer version, and the GLM-5 GitHub repository hosts the shared README, deployment notes and the GLM-5 technical report link for the whole series.
How much memory the weights need
Simple arithmetic gives the floor for the weights alone (estimates, before KV cache, activations and framework overhead):
- BF16: 744B parameters x 2 bytes = about 1,488 GB (roughly 1.5 TB).
- FP8: 744B x 1 byte = about 744 GB.
- Any 4-bit quantized build: 744B x 0.5 bytes = about 372 GB.
Only 40B parameters are active per token, which keeps compute per token modest, but every expert still has to sit in memory. Add the KV cache on top: Z.ai notes that IndexShare cuts compute but not per-token KV-cache size, so long contexts need a lot of extra memory. Serving the full 1M context is a multi-GPU cluster job, not a workstation job.
Running GLM-5.2 locally
Z.ai lists these inference frameworks for GLM-5.2, with minimum versions from the model card:
| Framework | Minimum version | Official guide |
|---|---|---|
| SGLang | v0.5.13.post1 | SGLang cookbook for GLM-5.2 |
| vLLM | v0.23.0 | vLLM recipe for zai-org/GLM-5.2 |
| Transformers | See model card | glm_moe_dsa model docs |
| KTransformers | v0.5.12 | GLM-5.2 tutorial |
| Unsloth | v0.1.47-beta | Unsloth GLM-5.2 guide |
| xLLM | – | Listed in Z.ai’s announcement |
On Ascend NPUs, vLLM-Ascend, xLLM and SGLang are supported. KTransformers and Unsloth each publish a dedicated GLM-5.2 tutorial, and the SGLang cookbook and vLLM recipe pages hold the tested launch settings.
A typical vLLM start looks like this; take the exact parallelism, reasoning-parser and tool-parser flags from the official vLLM recipe for your hardware:
pip install -U "vllm>=0.23.0"
vllm serve zai-org/GLM-5.2-FP8 --tensor-parallel-size 8
vLLM then exposes an OpenAI-compatible server, so the Python example above works unchanged once you set base_url to your server’s /v1 address and model to the served name.
For fine-tuning, the GLM-5 series supports slime (v0.3.0+), the RL framework Z.ai uses internally, and ms-swift (v4.4.0+) for SFT, PPO and GRPO. The MIT license lets you ship fine-tuned derivatives commercially; keep the copyright notice.
Should you still use GLM-5.2?
For most people starting today, no: GLM-5.3 is the same model family, costs the same, scores higher everywhere Z.ai compared them and is what the Coding Plan serves. But GLM-5.2 still wins in a few clear situations.

Keep or choose GLM-5.2 if
- You need answers with no reasoning at all. GLM-5.2 can disable thinking; GLM-5.3 and GLM-5.3-Flash cannot. For high-volume extraction or routing at flagship quality, that saves time and output tokens.
- You want a plain MIT model. If your legal team prefers the standard MIT terms over the GLM-5.3 License’s extra condition, GLM-5.2 is the newest 744B GLM under MIT.
- You self-host already. A validated GLM-5.2 deployment has little reason to move until you have tested GLM-5.3 on your workload.
- You pay through OpenRouter. Its listed GLM-5.2 price is lower than GLM-5.3’s, and the free variant is handy for prototypes.
- You have prompts tuned for it. Evaluation suites and agent prompts tuned on GLM-5.2 behave predictably; migrate on your own schedule.
Move to another model if
- You write code with agents: GLM-5.3, especially on the Coding Plan, where GLM-5.2 requests already route to it.
- Cost matters most or you need images: GLM-5.3-Flash at $0.15 / $0.50 with native image and video input.
- You need free API calls: GLM-4.7-Flash on the Z.ai API.
- You want the full family picture: the GLM models comparison ranks every model by use case, and the release timeline shows where GLM-5.2 fits.
For a decision that spans vendors, the comparison pages weigh GLM against DeepSeek and Claude on price, coding benchmarks and features, including using GLM inside Claude Code as a middle path.
GLM-5.2 FAQ
What is GLM 5.2?
GLM-5.2 is a large language model from Z.ai (formerly Zhipu AI), released on June 16, 2026 as its flagship for long-horizon coding and agent tasks. It is a 744B-parameter Mixture-of-Experts model with 40B active parameters, a 1M-token context window and 128K maximum output, released with open MIT-licensed weights.
It was succeeded as flagship by GLM-5.3 in August 2026, which uses the same base model with more post-training.
Is GLM 5.2 free?
The model weights are free under the MIT license, but the Z.ai API is paid ($1.40 input and $4.40 output per million tokens). You can use GLM-5.2 without paying in the free chat on this site (40 messages a day on paid models, no account), in Z.ai’s chat app at chat.z.ai, or through OpenRouter’s rate-limited z-ai/glm-5.2:free variant with a 32K context.
What is GLM-5.2’s context window?
1M tokens (1,048,576 on OpenRouter’s paid listing), with up to 128K output tokens (131,072) per response. That is five times GLM-5.1’s 200K window. The free OpenRouter variant is limited to 32,768 tokens, and the chat on this site sends the last 20 messages, trimmed to 24,000 characters.
How much does GLM 5.2 cost on the API?
On the Z.ai API: $1.40 per million input tokens, $0.26 per million cached input tokens and $4.40 per million output tokens, the same as GLM-5.3 and GLM-5.1. Reasoning tokens count as output. OpenRouter’s listed price is about $0.65 in and $2.04 out, but it changes often.
Is GLM-5.2 open source?
Yes, in the sense most people mean: the full weights are on Hugging Face (zai-org/GLM-5.2 in BF16 and zai-org/GLM-5.2-FP8) under the MIT license, which allows commercial use, modification and redistribution as long as you keep the copyright notice.
When was GLM-5.2 released?
June 16, 2026, the same day for the API, the chat app and the Hugging Face weights. Coding Plan users had early access shortly before the public release. GLM-5.1 came out on April 7, 2026 and GLM-5.3 on August 14, 2026.
Can I turn off thinking in GLM-5.2?
Yes. Send "thinking": {"type": "disabled"} or reasoning_effort: "none" (or "minimal"). Otherwise thinking is on, reasoning_effort defaults to max, and you can choose high for a faster, cheaper answer. This is one of the main differences from GLM-5.3, which always reasons.
Is GLM-5.3 better than GLM-5.2?
On Z.ai’s published numbers, yes, across the board: for example Terminal Bench 3.0 28.3 vs 4.6, DeepSWE v1.1 66.9 vs 46.2 and a 50% gain on Z.ai Code Bench, at the same API price. GLM-5.2 remains the better fit if you need thinking off, the plain MIT license or its cheaper OpenRouter listing.
The quickest way to judge for yourself is to ask both the same question. Chat with GLM-5.2 now, then switch the chip to GLM-5.3 and compare the answers side by side.