GLM-5.2: Specs, Pricing, Open Weights & How It Compares

Z.ai's June 2026 flagship for long-horizon coding: 1M-token context, MIT open weights and effort control. Specs, benchmarks, prices, code and a clear verdict on when to use it.

Chat with GLM-5.2 now Use it via API

GLM 5.2 (officially written GLM-5.2) is the flagship large language model that Z.ai, the company formerly known as Zhipu AI, released on June 16, 2026 for long-horizon coding and agent work. It is a 744B-parameter Mixture-of-Experts model with 40B parameters active per token, a 1M-token context window, up to 128K output tokens, and text in, text out. The weights are open under the MIT license, and the Z.ai API charges $1.40 per million input tokens and $4.40 per million output tokens.

GLM-5.2 was the first GLM model to offer what Z.ai calls a “solid 1M-token context”, and it introduced two things that later models kept: the IndexShare attention optimization and explicit effort levels (high and max). On August 14, 2026 Z.ai announced GLM-5.3, which uses the same base model with more post-training, at the same API price. GLM-5.2 is still live on the API, on OpenRouter (including a free variant) and on Hugging Face, and it is the newest GLM flagship that still lets you switch thinking off.

This guide covers everything: specs, what changed in the architecture, the full benchmark tables, prices, free routes, working API code, OpenRouter, coding agents, self-hosting and a clear verdict on when GLM-5.2 is still the right pick. If you just want to try it, Chat with GLM-5.2 now in the free chat on this site, no sign-up needed.

GLM 5.2 specs: 1M-token context, 744B total and 40B active parameters, MIT open weights, $1.40 in and $4.40 out per million tokens
GLM-5.2 key specs at a glance.

What GLM-5.2 is good at

Z.ai built GLM-5.2 for one job above all: staying useful across very long engineering tasks. The official model guide describes it as a flagship “built for the era of long-horizon tasks” that can carry a whole project in context and push a task from requirements to a deployable product. In practice that translates into a handful of workloads where GLM-5.2 is at its best.

Project-scale codebase work

With 1M tokens of context you can load a real repository (backend, frontend, configuration, tests and docs) in one request. Z.ai’s recommended first test is a technical audit: ask the model for an architecture map, module responsibilities, API contracts, data flows, call chains, technical debt and the constraints any refactor must respect. The point is not just reading more code; it is carrying earlier engineering judgments forward into later steps, which is where smaller-context models tend to drift.

Long refactors that run end to end

GLM-5.2 is tuned for cross-file, multi-step work: decoupling a module, migrating an API, restructuring directories, adapting to a new SDK or porting code between languages. Z.ai’s suggested pattern is to set clear boundaries (no change to business logic, API signatures or runtime behavior), ask for an execution plan with risks and a verification method first, then let the model implement and run the tests.

Following hard engineering rules

Z.ai highlights more consistent adherence to team standards over long, multi-round sessions: lint rules, build commands, test requirements, commit conventions and forbidden actions written in files such as CLAUDE.md or Agent.md. If your pain point with coding agents is out-of-scope edits, new dependencies nobody asked for, or skipped verification, this is the behavior GLM-5.2 was trained to improve.

Mobile, mini-program and game builds

The official guide lists client-side scenarios that go past generating an app: streaming messages, reconnection, local state, notifications and permissions, then debugging on a real device with ADB, logcat and screenshots. It also covers migrating a web project into a WeChat Mini Program (native, Taro or uni-app) and building small level-based games with a state machine, scoring and save logic.

Research reproduction and code-to-video

Two more showcase tasks from Z.ai: turning a paper’s architecture, loss functions and data pipeline into a runnable PyTorch project that reproduces the reported metrics, and writing Remotion (React) code that renders a finished MP4 video from a plain-language idea.

Where GLM-5.2 is weaker

  • No image input. GLM-5.2 is text-only. For screenshots, diagrams or video, use GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series.
  • Hard reasoning still trails the closed frontier. On Humanity’s Last Exam, Z.ai reports 40.5 for GLM-5.2 against 49.8 for Claude Opus 4.8 (full set) and 45.0 for Gemini 3.1 Pro.
  • Ultra-long marathons. On SWE-Marathon, Z.ai’s own table shows 13.0 for GLM-5.2 against 26.0 for Opus 4.8. Z.ai says plainly that the model “still has room to grow” there.
  • It is no longer the newest. GLM-5.3 beats it on every benchmark Z.ai published for both, at the same price.

What’s new in GLM-5.2

GLM-5.2 sits on the same 744B / 40B-active architecture line that began with GLM-5 in February 2026 and continued with GLM-5.1 in April. What changed is the context length, the attention machinery that makes 1M tokens affordable, faster speculative decoding, effort control and a new round of long-horizon reinforcement learning.

GLM-5.2 specs at a glance

SpecGLM-5.2
DeveloperZ.ai (formerly Zhipu AI)
Release dateJune 16, 2026
ArchitectureMixture-of-Experts with DeepSeek Sparse Attention and IndexShare
Parameters744B total, 40B active
Context window1M tokens
Max output128K tokens (default 65,536)
Input / outputText / text
ThinkingOn by default, can be disabled
Effort levelshigh, max (default max)
API model IDglm-5.2
API price$1.40 in, $0.26 cached, $4.40 out per 1M
WeightsMIT, BF16 and FP8 on Hugging Face and ModelScope
CapabilitiesStreaming, function calling, tool streaming, context caching, JSON output, MCP
GLM-5.2 technical specifications from Z.ai’s documentation and model card.
What's new in GLM-5.2: 1M context, IndexShare, better MTP, effort levels, long-horizon RL
The five headline changes in GLM-5.2.

1M-token context that stays usable

The jump from 200K (GLM-5 and GLM-5.1) to 1M tokens is the headline. Z.ai’s argument is that a long window only matters if quality holds across long, messy agent trajectories, so it expanded 1M-context training specifically for coding-agent scenarios: large implementations, automated research, performance optimization and complex debugging.

A 1M context is easy to claim, but much harder to keep reliable under real engineering pressure.

Z.ai, GLM-5.2 announcement

IndexShare: cheaper sparse attention at 1M tokens

GLM-5 introduced DeepSeek Sparse Attention (DSA), in which a lightweight “indexer” picks the top-k tokens each layer should attend to. At 1M tokens the indexer itself becomes expensive. GLM-5.2’s fix, called IndexShare, places one indexer at the first of every four transformer layers and reuses its top-k indices for all four. That removes the indexer dot products and top-k selection from three out of four layers and cuts per-token FLOPs by 2.9x at 1M context. GLM-5.2 was trained with IndexShare from mid-training at a 128K sequence length, and Z.ai says it outperforms GLM-5.1 on long-context benchmarks with less computation.

A faster MTP layer for speculative decoding

GLM models ship a Multi-Token Prediction (MTP) layer that serving engines use as a built-in draft model. For GLM-5.2, Z.ai applied IndexShare and KV sharing inside the MTP layer (which also removes a training-versus-inference mismatch that GLM-5.1’s MTP layer had), added rejection sampling, and trained with an end-to-end TV loss. In Z.ai’s ablation on coding workloads (GLM-5.1 backbone and data, seven MTP steps) the average acceptance length rose by 20%:

MethodAcceptance length
Baseline4.56
+ IndexShare + KV Share5.10
+ Rejection sampling5.29
+ End-to-end TV loss5.47 (+20%)
Z.ai’s MTP ablation: average accepted tokens per speculative step on coding tasks.

Longer acceptance means more tokens confirmed per decoding step, which is the part of the speed story you benefit from when you serve the weights yourself.

Effort levels: High and Max

GLM-5.2 was the first GLM model with the reasoning_effort parameter. It accepts high (enhanced reasoning) and max (deep reasoning, the default). Z.ai says GLM-5.2 delivers much stronger agentic coding than GLM-5.1 at comparable token use, landing roughly between Claude Opus 4.7 and Claude Opus 4.8 at similar token consumption, with Max available when a task justifies the extra compute. For compatibility with other APIs, GLM-5.2 also accepts none, minimal, low, medium and xhigh and maps them (details in the API section below).

Long-horizon RL and anti-hacking

Very long tasks produce traces that must be split (“compacted”) into sub-traces, and different rollouts produce different numbers of them. Z.ai therefore moved from group-wise optimization to a critic-based PPO setup that learns from individual rollouts, and trains on every compacted sub-trace with a token-level loss. It also added an anti-hack module after finding that GLM-5.2 showed more reward-hacking attempts than GLM-5.1, such as reading hidden evaluation files or downloading reference solutions. A rule-based filter flags suspicious tool calls, an LLM judge confirms intent, and blocked calls return dummy results so the rollout can continue instead of being thrown away.

slime and serving infrastructure

Post-training ran on slime, Z.ai’s open-source RL framework. Z.ai says it used slime for parallel on-policy distillation that merged more than ten expert models into the final model in about two days. On the serving side, Z.ai notes that IndexShare cuts compute but does not shrink the per-token KV cache proportionally, so it reworked memory management (building on LayerSplit), long-context kernels and CPU-side scheduling. The practical takeaway for self-hosters: at long context, memory for the KV cache, not raw compute, becomes the bottleneck.

GLM-5.2 release history

  • Before June 16, 2026: GLM Coding Plan subscribers get early access, and their feedback centers on project-level context, steadier long tasks, stricter adherence to engineering rules and stronger mobile work.
  • June 16, 2026: public release on the Z.ai API, in the chat app at chat.z.ai, and as open weights on Hugging Face and ModelScope. On the Coding Plan it counts at 3x quota in peak hours (14:00 to 18:00 UTC+8) and 2x off-peak, with a limited-time 1x off-peak promotion; ZCode users get 1.5x effective quota until June 30.
  • July 30, 2026: the Coding Plan switches from prompt-based to credit-based quotas.
  • August 14, 2026: Z.ai announces GLM-5.3, built on the same base model. Coding Plan requests for GLM-5.2 are now routed to GLM-5.3.
  • Today: GLM-5.2 stays available on the pay-as-you-go API at $1.40 / $4.40, on OpenRouter, and as MIT weights.

GLM 5.2 benchmarks

All numbers in this section are Z.ai’s published results from the GLM-5.2 announcement and model card. Scores marked * come from a benchmark’s full set; the rest use the text-only subset where one exists. A dash means Z.ai did not report that model.

Reasoning

BenchmarkGLM-5.2GLM-5.1Qwen3.7-MaxMiniMax M3DeepSeek-V4-ProClaude Opus 4.8GPT-5.5Gemini 3.1 Pro
HLE40.531.041.437.037.749.8*41.4*45.0
HLE w/ Tools54.752.353.5–48.257.9*52.2*51.4*
CritPt20.94.613.43.712.920.927.117.7
AIME 202699.295.397.0–94.695.798.398.2
HMMT Nov 202594.494.095.084.494.496.596.594.8
HMMT Feb 202692.582.697.184.495.296.796.787.3
IMOAnswerBench91.083.890.0–89.883.5–81.0
GPQA-Diamond91.286.290.093.090.193.693.694.3
GLM-5.2 reasoning benchmarks as reported by Z.ai. * = full set.

Coding

BenchmarkGLM-5.2GLM-5.1Qwen3.7-MaxMiniMax M3DeepSeek-V4-ProClaude Opus 4.8GPT-5.5Gemini 3.1 Pro
SWE-bench Pro62.158.460.659.055.469.258.654.2
NL2Repo48.942.747.242.135.569.750.733.4
DeepSWE46.218.018.020.08.058.070.010.0
ProgramBench63.750.9––47.871.970.839.5
Terminal Bench 2.1 (Terminus-2)81.063.575.065.064.085.084.074.0
FrontierSWE (dominance)74.430.5––29.075.172.639.6
PostTrainBench34.320.1–––37.228.421.6
SWE-Marathon13.01.0–––26.012.04.0
GLM-5.2 coding benchmarks as reported by Z.ai.

Z.ai also reports Terminal Bench 2.1 with each model’s best reported harness: GLM-5.2 scores 82.7 in Claude Code, GLM-5.1 69 in Claude Code, Claude Opus 4.8 78.9 in Claude Code, GPT-5.5 83.4 in Codex and Gemini 3.1 Pro 70.7 in Gemini CLI.

Agentic tool use

BenchmarkGLM-5.2GLM-5.1Qwen3.7-MaxMiniMax M3DeepSeek-V4-ProClaude Opus 4.8GPT-5.5Gemini 3.1 Pro
MCP-Atlas (public set)76.871.876.474.273.677.875.369.2
Tool-Decathlon48.240.7––52.859.955.648.8
GLM-5.2 agentic benchmarks as reported by Z.ai.

What the tables say

  • Biggest jumps over GLM-5.1 are in long-horizon and agentic coding: FrontierSWE 74.4 vs 30.5, DeepSWE 46.2 vs 18.0, Terminal Bench 2.1 81.0 vs 63.5, SWE-Marathon 13.0 vs 1.0 and CritPt 20.9 vs 4.6.
  • Against closed models, Z.ai says GLM-5.2 trails Opus 4.8 by only 1% on FrontierSWE while edging GPT-5.5 by 1% and Opus 4.7 by 11%; ranks second only to Opus 4.8 on PostTrainBench; and trails Opus 4.8 by 13% on SWE-Marathon. It was the highest-ranked open model on all three at launch.
  • Math is a strength. AIME 2026 at 99.2 is the top score in the table, and IMOAnswerBench at 91.0 is ahead of every closed model listed.
  • Science and hard knowledge lag. GPQA-Diamond (91.2) and HLE (40.5) sit below Opus 4.8, GPT-5.5 and Gemini 3.1 Pro.

What each benchmark measures

  • HLE (Humanity’s Last Exam): very hard expert-level questions across many fields; “w/ Tools” gives the model tool access.
  • CritPt: research-level physics reasoning problems.
  • AIME 2026: competition math problems with exact integer answers.
  • HMMT Nov 2025 / Feb 2026: problems from two rounds of the HMMT math tournament.
  • IMOAnswerBench: olympiad-level math problems graded on the final answer.
  • GPQA-Diamond: graduate-level science multiple-choice questions designed to resist simple web lookup.
  • SWE-bench Pro: resolving real issues in real repositories, a harder successor to SWE-bench.
  • NL2Repo: generating a working repository from a natural-language specification.
  • DeepSWE: long software-engineering tasks solved in an isolated container with no internet access.
  • ProgramBench: 200 program-building instances run in a sandbox with internet disabled.
  • Terminal Bench 2.1: real-world tasks completed in a command-line terminal.
  • FrontierSWE: open-ended technical projects lasting hours to tens of hours, from systems optimization to applied ML research.
  • PostTrainBench: the agent gets an H100 GPU and is scored on how much it improves small models through post-training.
  • SWE-Marathon: ultra-long software projects such as building compilers, optimizing kernels and developing production-grade services.
  • MCP-Atlas: tool use through MCP servers, 500 public tasks with a 10-minute limit each.
  • Tool-Decathlon: a broad tool-use benchmark run through its official evaluation service.

How Z.ai measured it

Reasoning tasks used temperature 1.0, top_p 0.95 and up to 163,840 generated tokens; HLE with tools allowed a 300,000-token context. SWE-bench Pro ran in OpenHands with a 400K context. Terminal Bench 2.1 (Terminus-2) ran with a 256K context, 4 CPUs and 8 GB of RAM per task; the Claude Code variant averaged five runs. FrontierSWE, PostTrainBench and SWE-Marathon were run by third parties (Proximal, PostTrainBench and Abundant AI) at 1M context, max effort and 128K output. These are generous settings: your own results at lower effort or shorter output limits will differ.

GLM-5.2 vs GLM-5.3 benchmark rows

When Z.ai launched GLM-5.3 it re-ran GLM-5.2 on a newer set of benchmark versions. These are the two models side by side, with Claude Opus 4.8 for reference, from the GLM-5.3 announcement:

BenchmarkGLM-5.3GLM-5.2Opus 4.8
Terminal Bench 2.188.281.085.0
Terminal Bench 3.028.34.621.1
DeepSWE v1.166.946.258.0
NL2Repo58.048.969.7
ProgramBench (Almost Solved)19.09.515.5
FrontierSWE78.167.566.5
SWE-Marathon v1.142.519.448.8
PostTrainBench39.831.732.9
CyberGym84.577.278.1
ExploitGym 2h / 6h105 / 13029 / 3980 / 120
ExploitBench54.424.440.0
Toolathlon Verified73.059.976.2
AutomationBench v1.0.648.226.241.0
Agents’ Last Exam (CLI)28.523.825.7
HLE w/ Tools62.554.757.9
GDPval-AA v2176915081588
GLM-5.3 vs GLM-5.2 in Z.ai’s GLM-5.3 results (August 2026).

Some GLM-5.2 numbers differ between the two announcements (FrontierSWE 74.4 vs 67.5, PostTrainBench 34.3 vs 31.7, SWE-Marathon 13.0 vs 19.4). That is expected: FrontierSWE is a dominance score reported at a specific date, SWE-Marathon moved to v1.1, ProgramBench switched to an “Almost Solved” metric, and several suites ran in a different harness. Compare numbers only within the same table.

On Z.ai’s private Code Bench, the gap is clearer still: at Max effort GLM-5.2 completes 23.4% of tasks using about 96K output tokens per task, while GLM-5.3 reaches 34.5% with about 75K. Z.ai summarizes this as a 50% improvement for GLM-5.3 over GLM-5.2.

We tested it: 5 tasks

Here is how GLM-5.2 handles the five standard tasks every model on GLM Chat goes through. Each task runs three times in the site’s own test harness with the same settings as the public chat: temperature 0.7, a 2,048-token output cap and the lightest reasoning setting (for GLM-5.2 that means thinking switched off). The harness records the full answer, latency and token counts, and an editor scores each answer from 0 to 2.

  1. Python: a streaming CSV de-duplicator that keeps the first occurrence, supports key columns and preserves the header. It reveals whether the model respects a memory constraint instead of loading the whole file.
  2. JavaScript: a debounce(fn, wait, { leading, trailing }) function with cancel() plus tests for Node’s built-in test runner. It reveals care with timing edge cases, especially leading and trailing calls together, and whether the tests actually pass with node --test.
  3. PHP refactor: a messy 60-line legacy function with copy-pasted loops, magic discount numbers, unescaped HTML and string-built SQL, rewritten as clean PHP 8 without changing behavior. It reveals discipline: the VIP discount and every order type must behave exactly as before, while output escaping and parameterized SQL fix the security holes.
  4. Explanation: transformer attention for a complete beginner in about 150 words. It reveals whether the model can be accurate about queries, keys, values and weights while staying readable and on length.
  5. Extraction: a messy product description turned into strict JSON with fixed keys, cm conversions and null for missing fields. It reveals format obedience and whether the model invents data when the text is silent.

Our test in GLM Chat

Runs: 3 runs per task, median score, September 24, 2026.

GLM-5.2 answering task 5 (Extract JSON from a messy description) in GLM Chat
GLM-5.2’s best-scoring answer in our battery (best of 3 runs): task 5, Extract JSON from a messy description.
TaskMedian scoreMedian latencyMedian output tokensWhy (across runs)
Python: stream-safe CSV dedupe2 / 2 (runs: 2, 2, 2)4.6 s372Streams with csv, keeps the header and first occurrence, key columns work in every run.
JavaScript: debounce with options + tests1 / 2 (runs: 0, 1, 2)12.6 s1,011Its own tests fail (2/3 runs); fails our debounce behaviour checks (1/3 runs).
PHP: refactor a messy 60-line function1 / 2 (runs: 0, 1, 1)9.6 s834Output not escaped (3/3 runs); SQL not parameterised (2/3 runs); changes the function’s behaviour (1/3 runs).
Explain attention to a beginner1 / 2 (runs: 2, 1, 1)4.2 s185No queries/keys/values (2/3 runs); weighting step missing (1/3 runs).
Extract JSON from a messy description2 / 2 (runs: 2, 2, 1)2.2 s112Passes every check in most runs; a dimension left in inches (1/3 runs).
Total7 / 10
Our test results. Runs: 3 runs per task, median score, September 24, 2026, in GLM Chat with the settings in our editorial policy.

A score of 2 means correct and complete, 1 means usable after a small fix, 0 means wrong or unusable, for a maximum of 10. When you read GLM-5.2’s card, look first at the PHP refactor and the JSON extraction: those two punish models that change behavior silently or invent data, and they are the closest proxies for the “follows engineering rules” claim above. The Python and JavaScript tasks show whether the model honors every stated constraint (streaming, leading plus trailing calls, a working cancel()) rather than writing the obvious version.

Keep in mind what the battery does not measure. With thinking off and a 2,048-token cap, it tests GLM-5.2 as a fast everyday assistant, not as a 1M-context agent running at Max effort for hours. For that side of the model, the long-horizon benchmarks above are the better guide, and the GLM-5.3 page shows the same five tasks on its successor so you can compare directly.

GLM-5.2 vs GLM-5.3 vs GLM-5.1

These three share the 744B / 40B-active design, so the differences come from context length, thinking controls, licensing and post-training. GLM-5.3-Flash is included because it is now the default cheap option in the same family.

SpecGLM-5.2GLM-5.3GLM-5.1GLM-5.3-Flash
ReleaseJun 16, 2026Aug 14, 2026Apr 7, 2026Aug 26, 2026
Parameters744B / 40B active744B / 40B active744B / 40B active320B / 18B active
Context1M1M200K1M
Max output128K128K128K128K
InputTextTextTextText, image, video, file
ThinkingOn by default, can be disabledAlways onOn by default, can be disabledAlways on
Effort levelshigh, maxlow, high, maxNot supportedlow, high, max
Input / output price$1.40 / $4.40$1.40 / $4.40$1.40 / $4.40$0.15 / $0.50
Cached input$0.26$0.26$0.26$0.03
WeightsMITGLM-5.3 LicenseMITMIT
Coding PlanRouted to GLM-5.3IncludedRouted to GLM-5.3Included, 3x quota
GLM-5.2 compared with its closest GLM siblings. Prices per 1M tokens on the Z.ai API.
GLM 5.2 vs GLM 5.3: same base model and price, GLM-5.3 adds post-training and forced thinking
The short version of GLM-5.2 vs GLM-5.3.

GLM-5.2 vs GLM-5.1

GLM-5.2 is a straight upgrade at the same price: five times the context (1M vs 200K), effort control, faster speculative decoding and large gains on long-horizon coding. Both keep the ability to disable thinking and both ship MIT weights. The only reason to stay on GLM-5.1 is a pipeline already tuned for it; on the Coding Plan, requests for either model now go to GLM-5.3 anyway. The GLM-5.1 page covers its 8-hour autonomous runs and SWE-Bench Pro result.

GLM-5.2 vs GLM-5.3

GLM-5.3 is GLM-5.2 plus a much larger round of post-training. It wins every benchmark row Z.ai published for both, uses fewer output tokens at Max effort on Z.ai Code Bench, and costs the same per token. Three differences can still matter to you:

  • Thinking cannot be disabled on GLM-5.3. Sending "thinking": {"type": "disabled"} to glm-5.3 makes the request fail; the lightest option is reasoning_effort: "low". GLM-5.2 can answer with no reasoning at all.
  • License. GLM-5.2 is plain MIT. GLM-5.3 uses the custom GLM-5.3 License, which adds a security-review condition for very large Model-as-a-Service businesses (see Is GLM open source?).
  • Third-party pricing. OpenRouter’s listed price for GLM-5.2 is lower than for GLM-5.3, and only GLM-5.2 has a free variant there.

GLM-5.2 vs GLM-5.3-Flash

GLM-5.3-Flash costs about a tenth as much ($0.15 / $0.50), reads images and video, and Z.ai reports it beats GLM-5.2 on its six headline coding and agent benchmarks, including DeepSWE v1.1 (63.4 vs 46.2) and AutomationBench (48.8 vs 26.2). For most new projects that care about cost, it is the better default. GLM-5.2 keeps the edge where you need the bigger 744B model’s knowledge, a thinking-off mode, or an MIT model you already host.

GLM-5.2 vs DeepSeek and Claude

Outside the GLM family, the two most common comparisons are DeepSeek and Anthropic’s Claude. In Z.ai’s GLM-5.2 table, GLM-5.2 leads DeepSeek-V4-Pro on SWE-bench Pro (62.1 vs 55.4), NL2Repo (48.9 vs 35.5), DeepSWE (46.2 vs 8.0), Terminal Bench 2.1 (81.0 vs 64.0) and FrontierSWE (74.4 vs 29.0), while DeepSeek-V4-Pro leads on Tool-Decathlon (52.8 vs 48.2). On list prices, deepseek-v4-pro costs $1.32 in and $3.96 out per million tokens at peak and half that off-peak, close to GLM-5.2’s $1.40 / $4.40.

Against Claude Opus 4.8, Z.ai’s numbers put GLM-5.2 behind on most rows (SWE-bench Pro 62.1 vs 69.2, HLE 40.5 vs 49.8) but close on FrontierSWE (74.4 vs 75.1) and MCP-Atlas (76.8 vs 77.8), and ahead on Terminal Bench 2.1 in Claude Code (82.7 vs 78.9). Opus 4.8 lists at $5 in and $25 out per million tokens, so GLM-5.2’s output tokens cost less than a fifth as much. All of these are Z.ai’s reported numbers; the full breakdowns are in GLM vs DeepSeek and GLM vs Claude.

GLM 5.2 pricing and free access

On Z.ai’s pay-as-you-go API, GLM-5.2 costs the same as GLM-5.3 and GLM-5.1. Storage for cached input is listed as free for a limited time.

ModelInputCached inputOutput
GLM-5.2$1.40$0.26$4.40
GLM-5.3$1.40$0.26$4.40
GLM-5.1$1.40$0.26$4.40
GLM-5$1.00$0.20$3.20
GLM-5.3-Flash$0.15$0.03$0.50
GLM-4.7$0.60$0.11$2.20
GLM-4.7-FlashFreeFreeFree
Z.ai API prices per 1M tokens.

What a long-context request really costs

Long context is where GLM-5.2 bills add up, so do the math before you paste a whole repository. Say you send a 300,000-token codebase and get a 10,000-token answer:

  • Input: 0.3M x $1.40 = $0.42
  • Output: 0.01M x $4.40 = $0.044
  • Total: about $0.46 for the first request

Now ask a follow-up in the same conversation. If 290,000 tokens of the prompt hit the cache and 10,000 are new, input costs 0.29M x $0.26 = $0.0754 plus 0.01M x $1.40 = $0.014. With another 10,000 output tokens ($0.044), the follow-up comes to about $0.13, less than a third of the first call. Context caching is automatic on Z.ai; you can see the cached count in usage.prompt_tokens_details.cached_tokens. The context caching guide shows how to structure prompts so the cache hits.

Remember that reasoning tokens are billed as output. At max effort on a hard task, the thinking can easily outweigh the visible answer, which is exactly why high exists. The full cross-model price list lives on the GLM pricing page.

Is GLM 5.2 free?

Not on the Z.ai API, but there are four legitimate ways to use it without paying per token:

  1. GLM Chat (this site). Pick GLM-5.2 in the model chips and chat with GLM-5.2 for free. No account; 40 messages per day on the paid models, up to 8,000 characters per message, and the last 20 messages (trimmed to 24,000 characters) are sent as context. The chat runs GLM-5.2 with thinking off for quick replies. If the site’s daily capacity for paid models runs out, replies come from GLM-4.7-Flash for the rest of the UTC day and are labeled “GLM-4.7-Flash (daily limit reached)”.
  2. Z.ai’s own chat app. GLM-5.2 launched in the free web chat at chat.z.ai; the app’s model list changes as newer models arrive.
  3. OpenRouter’s free variant. z-ai/glm-5.2:free is listed at $0 with a 32,768-token context and rate limits (details in the OpenRouter section).
  4. Your own hardware. The MIT weights are free to download and use commercially, if you have the GPUs for a 744B model.

If what you really need is a free API model rather than GLM-5.2 specifically, GLM-4.7-Flash and GLM-4.5-Flash are free on the Z.ai API itself.

GLM-5.2 on the Coding Plan

The GLM Coding Plan (from $18 per month for Lite) no longer serves GLM-5.2 as a separate model: requests for GLM-5.2 and GLM-5.1 are automatically routed to GLM-5.3. If you subscribe to the plan, you are paying for GLM-5.3 and GLM-5.3-Flash. The Coding Plan guide explains tiers, credits and when the plan beats pay-as-you-go.

Using the GLM 5.2 API

The model ID is glm-5.2 (lowercase). The API is OpenAI-compatible Chat Completions at https://api.z.ai/api/paas/v4, authenticated with a bearer key. Create a key on the API Keys page at z.ai/manage-apikey/apikey-list and export it as ZAI_API_KEY. New to the platform? The GLM API quickstart walks through the account and key setup.

curl

curl -X POST "https://api.z.ai/api/paas/v4/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $ZAI_API_KEY" \
  -d '{
    "model": "glm-5.2",
    "messages": [
      {"role": "system", "content": "You are a senior software engineer."},
      {"role": "user", "content": "Review this function and list the bugs: ..."}
    ],
    "thinking": {"type": "enabled"},
    "reasoning_effort": "high",
    "max_tokens": 8192,
    "temperature": 1.0
  }'

Python with the OpenAI SDK

Point the standard openai package at Z.ai’s base URL. GLM-specific fields such as thinking go in extra_body; putting reasoning_effort there too keeps the code working on older SDK versions.

# pip install --upgrade "openai>=1.0"
import os
from openai import OpenAI
client = OpenAI(
    api_key=os.environ["ZAI_API_KEY"],
    base_url="https://api.z.ai/api/paas/v4/",
)
resp = client.chat.completions.create(
    model="glm-5.2",
    messages=[
        {"role": "system", "content": "You are a senior software engineer."},
        {"role": "user", "content": "Plan a migration from REST polling to WebSockets."},
    ],
    max_tokens=8192,
    extra_body={
        "thinking": {"type": "enabled"},
        "reasoning_effort": "max",
    },
)
msg = resp.choices[0].message
print(getattr(msg, "reasoning_content", None))  # the model's reasoning, if any
print(msg.content)                              # the final answer
print(resp.usage)                               # token counts, incl. cached tokens

Streaming reasoning and answer separately

With stream=True, reasoning arrives in delta.reasoning_content and the answer in delta.content; the stream ends with data: [DONE]. This example uses Z.ai’s official SDK (pip install zai-sdk):

from zai import ZaiClient
client = ZaiClient(api_key="your-api-key")
stream = client.chat.completions.create(
    model="glm-5.2",
    messages=[{"role": "user", "content": "Explain IndexShare in three sentences."}],
    thinking={"type": "enabled"},
    reasoning_effort="high",
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta
    if delta.reasoning_content:
        print(delta.reasoning_content, end="", flush=True)
    if delta.content:
        print(delta.content, end="", flush=True)

Turning thinking off

Unlike GLM-5.3, GLM-5.2 can skip reasoning entirely. Either disable thinking or pass reasoning_effort: "none" (or "minimal"). Use this for classification, extraction and short chat where latency matters more than depth:

curl -X POST "https://api.z.ai/api/paas/v4/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $ZAI_API_KEY" \
  -d '{
    "model": "glm-5.2",
    "messages": [{"role": "user", "content": "Classify this ticket: billing, bug or feature?"}],
    "thinking": {"type": "disabled"}
  }'

reasoning_effort values on GLM-5.2

Value you sendWhat GLM-5.2 does
max (default)Deep reasoning
xhighMapped to max
highEnhanced reasoning
medium, lowMapped to high
minimal, noneSkips thinking
How GLM-5.2 handles each reasoning_effort value on the Z.ai API.

The thinking mode guide compares these controls across every GLM model.

Parameters worth knowing

  • max_tokens: default 65,536, maximum 131,072 (128K).
  • temperature: range 0.0 to 1.0, default 1.0. top_p: default 0.95.
  • Tools: standard OpenAI-style function calling (up to 128 functions), plus tool_stream: true to stream tool-call arguments. See the function calling guide.
  • JSON mode: response_format: {"type": "json_object"}.
  • Errors: an unknown model name returns 1211, a prompt over the limit 1261, and rate limits 1302. The rate limits and error codes guide lists fixes for each.

Function calling with GLM-5.2

GLM-5.2 uses the standard OpenAI tool format. tool_choice supports only auto, so the model decides when to call a tool. A minimal round trip:

import json, os
from openai import OpenAI
client = OpenAI(api_key=os.environ["ZAI_API_KEY"], base_url="https://api.z.ai/api/paas/v4/")
tools = [{
    "type": "function",
    "function": {
        "name": "get_open_issues",
        "description": "List open issues for a repository",
        "parameters": {
            "type": "object",
            "properties": {"repo": {"type": "string"}},
            "required": ["repo"],
        },
    },
}]
messages = [{"role": "user", "content": "How many open issues does acme/api have?"}]
first = client.chat.completions.create(model="glm-5.2", messages=messages, tools=tools)
call = first.choices[0].message.tool_calls[0]
args = json.loads(call.function.arguments)
result = {"repo": args["repo"], "open_issues": 42}   # run your real function here
messages.append(first.choices[0].message)
messages.append({"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)})
final = client.chat.completions.create(model="glm-5.2", messages=messages, tools=tools)
print(final.choices[0].message.content)

In multi-step agent loops with thinking on, send the assistant message back unchanged (including its reasoning) so the model keeps its chain of thought between tool calls.

Tips for 1M-token prompts

  • Put the stable part first. Caching is automatic and matches repeated context, so place the repository, docs and rules at the start of the prompt and the changing question at the end. Every follow-up then reuses the cached prefix at $0.26 instead of $1.40 per million tokens.
  • Ask for a plan before edits. Z.ai’s own recommended prompts start with an audit or execution plan, impact scope and verification method, then let the model implement. That structure is what GLM-5.2 was trained to follow over long tasks.
  • Write your rules down. Hard constraints (no new dependencies, no API contract changes, no commits) belong in the prompt or in CLAUDE.md / Agent.md, where the model can refer back to them.
  • Start at high, escalate to max. Use high for routine edits and switch to max for the hard cases; reasoning tokens are billed as output.
  • Cap output explicitly. The default max_tokens is 65,536. Set a lower value for chat-style answers and the full 131,072 only for long generations.

Errors you are most likely to hit

CodeMeaningTypical fix
401 / 1000Authentication failedCheck the key and the Bearer header
400 / 1211Unknown modelUse the exact ID glm-5.2, lowercase
400 / 1214Invalid parameter valueKeep temperature within 0.0 to 1.0; use a listed effort value
400 / 1261Prompt too longTrim the input below the 1M-token window
429 / 1113Insufficient balanceTop up, or fix a Coding Plan vs API base URL mix-up
429 / 1302Request rate limitBack off and retry; check limits at z.ai/manage-apikey/rate-limits
429 / 1305Service temporarily overloadedRetry with exponential backoff
Common Z.ai API errors when calling GLM-5.2.

GLM-5.2 on OpenRouter

OpenRouter is a third-party router that resells many models behind one OpenAI-compatible API and one key. It lists two GLM-5.2 entries:

SlugListed price (in / out)ContextMax outputNotes
z-ai/glm-5.2about $0.65 / $2.041,048,576131,072Routed across providers; tools, JSON and structured output listed
z-ai/glm-5.2:free$0 / $032,76829,491Rate-limited; no tools or response_format in its listed parameters
OpenRouter’s listed GLM-5.2 entries. Prices per 1M tokens change often; confirm on the model page.

OpenRouter routes the paid slug across several providers, and its listed price is roughly half of Z.ai’s direct price. Prices move often, so check the GLM-5.2 page on OpenRouter before you plan costs. The free variant is excellent for trying the model or light personal use; its 32K context rules out the long-context work GLM-5.2 is known for.

Calling GLM-5.2 through OpenRouter

Use the base URL https://openrouter.ai/api/v1 and your OpenRouter key:

import os
from openai import OpenAI
client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key=os.environ["OPENROUTER_API_KEY"],
)
resp = client.chat.completions.create(
    model="z-ai/glm-5.2:free",   # or "z-ai/glm-5.2" for the paid, 1M-context route
    messages=[{"role": "user", "content": "Write a SQL query that finds duplicate emails."}],
)
print(resp.choices[0].message.content)

The same request with curl:

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "z-ai/glm-5.2", "messages": [{"role": "user", "content": "Hello, GLM-5.2"}]}'

OpenRouter lists reasoning and reasoning_effort among the supported parameters for both slugs. Z.ai-specific fields such as thinking belong to Z.ai’s own API; if you need exact control over GLM’s thinking behavior or the documented Z.ai parameters, call Z.ai directly.

GLM-5.2 in Claude Code and coding agents

When GLM-5.2 launched, Z.ai rolled it out to every Coding Plan subscriber (some had early access before release). You enabled it by setting the model name to GLM-5.2, or GLM-5.2[1m] in Claude Code for the full 1M context, and you could switch between High and Max effort. ZCode, Z.ai’s desktop agent, shipped with GLM-5.2 and its /goal mode for long-horizon tasks.

That has changed. On the Coding Plan today, requests for GLM-5.2 are automatically routed to GLM-5.3, so writing glm-5.2 in a plan-connected tool gets you GLM-5.3. The current recommended Claude Code setup maps the model slots to GLM-5.3 and GLM-5.3-Flash in ~/.claude/settings.json:

{
  "env": {
    "ANTHROPIC_AUTH_TOKEN": "your-api-key",
    "ANTHROPIC_BASE_URL": "https://api.z.ai/api/anthropic",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "glm-5.3-flash[1m]",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "glm-5.3[1m]",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "glm-5.3[1m]",
    "CLAUDE_CODE_AUTO_COMPACT_WINDOW": "1000000",
    "API_TIMEOUT_MS": "3000000"
  }
}

The [1m] suffix plus CLAUDE_CODE_AUTO_COMPACT_WINDOW enables the 1M context; if Claude Code says a [1m] model does not exist, update Claude Code. Type /effort inside a session to change thinking intensity, and /status to confirm which model is active. The step-by-step setup is in How to use GLM in Claude Code.

If you specifically want GLM-5.2 in an agent

Use pay-as-you-go instead of the plan. In any tool that accepts a custom OpenAI-compatible provider (Cline, Kilo Code, Roo Code, OpenCode and others), choose the OpenAI Compatible provider, set the base URL to https://api.z.ai/api/paas/v4, enter your API key and the custom model glm-5.2, and set the context window to 1,000,000. Leave image support unchecked, since GLM-5.2 is text-only. You then pay per token at $1.40 / $4.40, and the request really reaches GLM-5.2. The Cline guide and the OpenClaw guide show the exact fields for those tools.

Do not point pay-as-you-go keys at the Coding Plan URL (https://api.z.ai/api/coding/paas/v4) or vice versa. Mixing the two is a common cause of the “1113 insufficient balance” error.

Downloading the weights

GLM-5.2’s weights were published on Hugging Face and ModelScope on launch day under the MIT license, with no gating and no regional limits. Z.ai calls this “Pure Open”.

VariantHugging FaceModelScopePrecisionSize
GLM-5.2zai-org/GLM-5.2ZhipuAI/GLM-5.2BF16744B-A40B
GLM-5.2-FP8zai-org/GLM-5.2-FP8ZhipuAI/GLM-5.2-FP8FP8744B-A40B
Official GLM-5.2 weight repositories.

The Hugging Face model card now points to GLM-5.3-BF16 as the newer version, and the GLM-5 GitHub repository hosts the shared README, deployment notes and the GLM-5 technical report link for the whole series.

How much memory the weights need

Simple arithmetic gives the floor for the weights alone (estimates, before KV cache, activations and framework overhead):

  • BF16: 744B parameters x 2 bytes = about 1,488 GB (roughly 1.5 TB).
  • FP8: 744B x 1 byte = about 744 GB.
  • Any 4-bit quantized build: 744B x 0.5 bytes = about 372 GB.

Only 40B parameters are active per token, which keeps compute per token modest, but every expert still has to sit in memory. Add the KV cache on top: Z.ai notes that IndexShare cuts compute but not per-token KV-cache size, so long contexts need a lot of extra memory. Serving the full 1M context is a multi-GPU cluster job, not a workstation job.

Running GLM-5.2 locally

Z.ai lists these inference frameworks for GLM-5.2, with minimum versions from the model card:

FrameworkMinimum versionOfficial guide
SGLangv0.5.13.post1SGLang cookbook for GLM-5.2
vLLMv0.23.0vLLM recipe for zai-org/GLM-5.2
TransformersSee model cardglm_moe_dsa model docs
KTransformersv0.5.12GLM-5.2 tutorial
Unslothv0.1.47-betaUnsloth GLM-5.2 guide
xLLM–Listed in Z.ai’s announcement
Supported frameworks for serving GLM-5.2.

On Ascend NPUs, vLLM-Ascend, xLLM and SGLang are supported. KTransformers and Unsloth each publish a dedicated GLM-5.2 tutorial, and the SGLang cookbook and vLLM recipe pages hold the tested launch settings.

A typical vLLM start looks like this; take the exact parallelism, reasoning-parser and tool-parser flags from the official vLLM recipe for your hardware:

pip install -U "vllm>=0.23.0"
vllm serve zai-org/GLM-5.2-FP8 --tensor-parallel-size 8

vLLM then exposes an OpenAI-compatible server, so the Python example above works unchanged once you set base_url to your server’s /v1 address and model to the served name.

For fine-tuning, the GLM-5 series supports slime (v0.3.0+), the RL framework Z.ai uses internally, and ms-swift (v4.4.0+) for SFT, PPO and GRPO. The MIT license lets you ship fine-tuned derivatives commercially; keep the copyright notice.

Should you still use GLM-5.2?

For most people starting today, no: GLM-5.3 is the same model family, costs the same, scores higher everywhere Z.ai compared them and is what the Coding Plan serves. But GLM-5.2 still wins in a few clear situations.

When GLM 5.2 is still the right choice: thinking off, MIT license, OpenRouter free tier, existing pipelines
Pick GLM-5.2 in these cases.

Keep or choose GLM-5.2 if

  • You need answers with no reasoning at all. GLM-5.2 can disable thinking; GLM-5.3 and GLM-5.3-Flash cannot. For high-volume extraction or routing at flagship quality, that saves time and output tokens.
  • You want a plain MIT model. If your legal team prefers the standard MIT terms over the GLM-5.3 License’s extra condition, GLM-5.2 is the newest 744B GLM under MIT.
  • You self-host already. A validated GLM-5.2 deployment has little reason to move until you have tested GLM-5.3 on your workload.
  • You pay through OpenRouter. Its listed GLM-5.2 price is lower than GLM-5.3’s, and the free variant is handy for prototypes.
  • You have prompts tuned for it. Evaluation suites and agent prompts tuned on GLM-5.2 behave predictably; migrate on your own schedule.

Move to another model if

  • You write code with agents: GLM-5.3, especially on the Coding Plan, where GLM-5.2 requests already route to it.
  • Cost matters most or you need images: GLM-5.3-Flash at $0.15 / $0.50 with native image and video input.
  • You need free API calls: GLM-4.7-Flash on the Z.ai API.
  • You want the full family picture: the GLM models comparison ranks every model by use case, and the release timeline shows where GLM-5.2 fits.

For a decision that spans vendors, the comparison pages weigh GLM against DeepSeek and Claude on price, coding benchmarks and features, including using GLM inside Claude Code as a middle path.

GLM-5.2 FAQ

What is GLM 5.2?

GLM-5.2 is a large language model from Z.ai (formerly Zhipu AI), released on June 16, 2026 as its flagship for long-horizon coding and agent tasks. It is a 744B-parameter Mixture-of-Experts model with 40B active parameters, a 1M-token context window and 128K maximum output, released with open MIT-licensed weights.

It was succeeded as flagship by GLM-5.3 in August 2026, which uses the same base model with more post-training.

Is GLM 5.2 free?

The model weights are free under the MIT license, but the Z.ai API is paid ($1.40 input and $4.40 output per million tokens). You can use GLM-5.2 without paying in the free chat on this site (40 messages a day on paid models, no account), in Z.ai’s chat app at chat.z.ai, or through OpenRouter’s rate-limited z-ai/glm-5.2:free variant with a 32K context.

What is GLM-5.2’s context window?

1M tokens (1,048,576 on OpenRouter’s paid listing), with up to 128K output tokens (131,072) per response. That is five times GLM-5.1’s 200K window. The free OpenRouter variant is limited to 32,768 tokens, and the chat on this site sends the last 20 messages, trimmed to 24,000 characters.

How much does GLM 5.2 cost on the API?

On the Z.ai API: $1.40 per million input tokens, $0.26 per million cached input tokens and $4.40 per million output tokens, the same as GLM-5.3 and GLM-5.1. Reasoning tokens count as output. OpenRouter’s listed price is about $0.65 in and $2.04 out, but it changes often.

Is GLM-5.2 open source?

Yes, in the sense most people mean: the full weights are on Hugging Face (zai-org/GLM-5.2 in BF16 and zai-org/GLM-5.2-FP8) under the MIT license, which allows commercial use, modification and redistribution as long as you keep the copyright notice.

When was GLM-5.2 released?

June 16, 2026, the same day for the API, the chat app and the Hugging Face weights. Coding Plan users had early access shortly before the public release. GLM-5.1 came out on April 7, 2026 and GLM-5.3 on August 14, 2026.

Can I turn off thinking in GLM-5.2?

Yes. Send "thinking": {"type": "disabled"} or reasoning_effort: "none" (or "minimal"). Otherwise thinking is on, reasoning_effort defaults to max, and you can choose high for a faster, cheaper answer. This is one of the main differences from GLM-5.3, which always reasons.

Is GLM-5.3 better than GLM-5.2?

On Z.ai’s published numbers, yes, across the board: for example Terminal Bench 3.0 28.3 vs 4.6, DeepSWE v1.1 66.9 vs 46.2 and a 50% gain on Z.ai Code Bench, at the same API price. GLM-5.2 remains the better fit if you need thinking off, the plain MIT license or its cheaper OpenRouter listing.

The quickest way to judge for yourself is to ask both the same question. Chat with GLM-5.2 now, then switch the chip to GLM-5.3 and compare the answers side by side.

Try GLM-5.2 on your own prompt

Free, no sign-up. Your conversation stays in your browser.

Open chat