GLM 5.3 (officially GLM-5.3) is Z.ai’s current flagship model, announced on August 14, 2026 by the company formerly known as Zhipu AI. It keeps GLM-5.2’s base model (744B total parameters, 40B active, Mixture-of-Experts) and gets every improvement from post-training. It reads text only, holds a 1M-token context, writes up to 128K tokens per answer and always reasons, with three effort levels: low, high and max. The Z.ai API charges $1.40 per million input tokens and $4.40 per million output tokens, and the weights are open under the custom GLM-5.3 License.
Z.ai’s headline claims: a 50% gain over GLM-5.2 on its in-house Z.ai Code Bench, the best open-model scores on Terminal Bench 3.0 and Agents’ Last Exam (CLI), and the top result on the CyberGym vulnerability-discovery benchmark. The API release-notes entry followed on August 18, and the weights reached Hugging Face on August 25 after a safety evaluation.
Below you will find the full benchmark table, the one API change that breaks old code (thinking can no longer be disabled), prices, free routes, working code, the Coding Plan, what the GLM-5.3 License actually requires and how to download the weights. To try it right away, Chat with GLM-5.3 now on this site, free and without an account.

What GLM-5.3 is good at
Z.ai describes GLM-5.3 as its latest flagship for complex software engineering and agent work. Because it shares GLM-5.2’s base model, the differences you feel come from how it was trained after pre-training: more environments, more varied tasks and more compute spent on reinforcement learning.
GLM-5.3 specs at a glance
| Spec | GLM-5.3 |
|---|---|
| Developer | Z.ai (formerly Zhipu AI) |
| Announced | August 14, 2026 (API release notes August 18) |
| Base model | Same as GLM-5.2; all gains from post-training |
| Parameters | 744B total, 40B active (MoE) |
| Context / max output | 1M / 128K tokens |
| Input / output | Text / text |
| Reasoning | Always on; effort low, high or max (default max) |
| API model ID | glm-5.3 |
| API price | $1.40 in, $0.26 cached, $4.40 out per 1M |
| Weights | GLM-5.3 License; FP8 and BF16 on Hugging Face and ModelScope |
| Capabilities | Streaming, function calling, context caching, structured output |
Real engineering work, not coding exercises
For GLM-5.3, Z.ai pushed its training environments toward “real units of expert work”, some representing several days of effort for an experienced engineer. One example from the announcement: the model gets the same working environment as an ML infrastructure engineer (compute clusters, storage, internal documentation, codebases, experiment results) and must diagnose bottlenecks, implement optimizations, run experiments and deliver a measurable speedup without breaking correctness. The aim is a model that owns a substantial task end to end instead of waiting for a human to break it into steps.
Better results with fewer tokens
On Z.ai Code Bench, a private benchmark of realistic coding-agent tasks, GLM-5.3 improves both success and efficiency. At Max effort it completes 34.5% of tasks with roughly 75K output tokens per task, against 23.4% at 96K for GLM-5.2. At High effort it reaches 31.4% with about 50K tokens, ahead of Claude Opus 4.8 at 29.5% with 120K. It still trails Claude Fable 5, which reaches 39.5% at Max effort. Fewer output tokens per task means lower bills and faster turnarounds, not just higher scores.
Terminal and long-horizon agent tasks
The largest public gains are on hard agent benchmarks: Terminal Bench 3.0 from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9 and Agents’ Last Exam (CLI) from 23.8 to 28.5. GLM-5.3 carries over GLM-5.2’s long-horizon RL techniques, including compaction, which helps these gains hold on tasks that run for hours.
General agent and knowledge work
The gains are not limited to code. AutomationBench rises from 26.2 to 48.2, Toolathlon Verified from 59.9 to 73.0, and GDPval-AA v2 (evaluated by Artificial Analysis) from 1508 to 1769, the highest score in Z.ai’s table.
Security research
GLM-5.3 is the strongest model in Z.ai’s table on CyberGym and more than doubles GLM-5.2 on exploitation benchmarks. That makes it a capable code-review and vulnerability-discovery assistant; see the cyber section below for the numbers and the real-world findings.
How Z.ai trained it
GLM-5.3’s story is about training environments more than architecture. Z.ai built pipelines that synthesize long-horizon task environments end to end: research agents collect task patterns from real work and turn them into runnable environments with multi-step dependencies and hidden state, and a judge agent attempts each task to confirm it is solvable. Verifiers are written without access to the reference solution, and solver trajectories are used to find and close reward shortcuts. Only a verifier that passes oracle, no-op and unsolved-state checks is trusted as a reward signal. Z.ai notes the pipeline still needs meaningful human-in-the-loop work.
Under the hood, the same slime RL framework used for GLM-5.2 was extended: Z.ai reports that training-rollout log-probability differences were brought down to the 1e-7 level (a reduction of more than 99.99%), and that system-level optimizations raised end-to-end RL throughput on long-horizon coding tasks by more than 2.3x.
Choosing an effort level
- low: chat, classification, extraction, quick edits. The fastest option, and the lightest setting GLM-5.3 allows.
- high: routine agent steps and everyday coding. Z.ai’s Code Bench shows High already beating Claude Opus 4.8 with far fewer tokens.
- max: the default. Use it for hard debugging, large refactors and long agent runs, as Z.ai recommends for coding.
Where GLM-5.3 is weaker
- Text only. No image, video or file input. For screenshots and UI work use GLM-5.3-Flash, which is natively multimodal and about a tenth of the price.
- Always reasons. Every request includes some thinking, so simple lookups cost more latency and output tokens than on a model that can skip reasoning.
- Not the frontier leader everywhere. In Z.ai’s own table, GPT-5.6 Sol leads Terminal Bench 2.1, Terminal Bench 3.0 and DeepSWE; Claude Opus 4.8 leads NL2Repo and SWE-Marathon; Fable 5 leads ProgramBench, FrontierSWE and PostTrainBench.
GLM 5.3 benchmarks
These are Z.ai’s published results from the GLM-5.3 announcement and model card. A dash means Z.ai did not report that model.
Coding
| Benchmark | GLM-5.3 | GLM-5.2 | Kimi K3 | DeepSeek-V4-Pro-0813 | Qwen3.8-Max | Opus 4.8 | Fable 5 (w/ fallback) | GPT-5.6 Sol |
|---|---|---|---|---|---|---|---|---|
| Terminal Bench 2.1 | 88.2 | 81.0 | 88.3 | 87.9 | 86.6 | 85.0 | 88.0 | 88.8 |
| Terminal Bench 3.0 | 28.3 | 4.6 | 17.4 | – | – | 21.1 | 33.7 | 34.6 |
| DeepSWE v1.1 | 66.9 | 46.2 | 67.5 | 62.7 | 56.6 | 58.0 | 69.7 | 72.7 |
| NL2Repo | 58.0 | 48.9 | 58.0 | 61.1 | 55.9 | 69.7 | – | – |
| ProgramBench (Almost Solved) | 19.0 | 9.5 | 17.5 | – | 10.5 | 15.5 | 33.0 | 23.0 |
| FrontierSWE | 78.1 | 67.5 | – | – | – | 66.5 | 88.2 | – |
| SWE-Marathon v1.1 | 42.5 | 19.4 | 48.1 | – | – | 48.8 | 33.1 | 42.5 |
| PostTrainBench | 39.8 | 31.7 | 32.0 | – | – | 32.9 | 41.8 | 36.2 |
Cyber
| Benchmark | GLM-5.3 | GLM-5.2 | Kimi K3 | DeepSeek-V4-Pro-0813 | Qwen3.8-Max | Opus 4.8 | Fable 5 (w/ fallback) | GPT-5.6 Sol |
|---|---|---|---|---|---|---|---|---|
| CyberGym | 84.5 | 77.2 | 80.0 | 83.3 | 78.5 | 78.1 | 83.8 | 83.6 |
| ExploitGym 2h / 6h | 105 / 130 | 29 / 39 | 36 / 70 | – | 14 / 26 | 80 / 120 | 181 / 247 | 216 / 293 |
| ExploitBench | 54.4 | 24.4 | 32.2 | – | 28.8 | 40.0 | 78.0 | 76.5 |
Agentic
| Benchmark | GLM-5.3 | GLM-5.2 | Kimi K3 | DeepSeek-V4-Pro-0813 | Qwen3.8-Max | Opus 4.8 | Fable 5 (w/ fallback) | GPT-5.6 Sol |
|---|---|---|---|---|---|---|---|---|
| Toolathlon Verified | 73.0 | 59.9 | 76.5 | 74.1 | 72.5 | 76.2 | 74.7 | 74.9 |
| AutomationBench v1.0.6 | 48.2 | 26.2 | 46.7 | 43.2 | 39.8 | 41.0 | 46.2 | 45.8 |
| Agents’ Last Exam (CLI) | 28.5 | 23.8 | 27.6 | 25.7 | 27.0 | 25.7 | 23.8 | 28.6 |
| HLE w/ Tools | 62.5 | 54.7 | 59.8 | 60.0 | 56.2 | 57.9 | 63.9 | 64.5 |
| GDPval-AA v2 | 1769 | 1508 | 1682 | 1590 | 1739 | 1588 | 1743 | 1730 |
What stands out
- Against GLM-5.2, GLM-5.3 wins every row. The biggest relative jumps are Terminal Bench 3.0 (4.6 to 28.3), ExploitGym (29 to 105 tasks in two hours), ExploitBench (24.4 to 54.4) and AutomationBench (26.2 to 48.2).
- Against other open-weight models, it is level with Kimi K3 on Terminal Bench 2.1 and NL2Repo, ahead on Terminal Bench 3.0 (28.3 vs 17.4) and PostTrainBench, and behind on SWE-Marathon (42.5 vs 48.1) and Toolathlon (73.0 vs 76.5). DeepSeek-V4-Pro-0813 leads it on NL2Repo (61.1 vs 58.0).
- Against closed models, GLM-5.3 beats Opus 4.8 on 13 of the 16 rows Z.ai published for both, but GPT-5.6 Sol and Fable 5 remain ahead on most coding and exploitation rows.
How Z.ai ran the evaluations
Most agentic suites ran in the Claude Code harness at max effort. Terminal Bench 3.0 used a 400K context and 128K output, averaged over three rollouts with up to 600 agent turns and a 10-hour limit. Agents’ Last Exam ran its 105 tasks in isolated containers with a 1M context. CyberGym ran 1,507 tasks with no web tools and a domain allowlist to prevent cheating. ExploitGym normalized its two-hour and six-hour time limits by each model’s measured throughput. FrontierSWE was run by Proximal at 1M context. These are heavy, max-effort settings; expect different numbers at low or high effort.
GLM-5.3’s cyber capability
Z.ai calls this an “emergent” capability. It added vulnerability-discovery data and environments to post-training expecting modest gains, and found the skill kept growing as training scaled: GLM-5.3 began planning multi-stage exploitation chains, not just spotting isolated flaws.

Three benchmarks, three stages
- CyberGym (finding bugs): starting from white-box source code, the model must identify a vulnerability and prove it by triggering a fault. GLM-5.3 scores 84.5%, up from 77.2% for GLM-5.2 and the best result on the benchmark, ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%) in Z.ai’s write-up. (Z.ai’s table lists the 83.8% score under Fable 5 with fallback.)
- ExploitBench (understanding exploitation): deeper reasoning about real vulnerabilities and how they can be exploited. GLM-5.3 reaches 54.4%, more than double GLM-5.2’s 24.4%, while the top closed models score 78.0% and 76.5%.
- ExploitGym (completing exploits under time limits): GLM-5.3 completes 105 tasks in two hours and 130 in six, against 29 and 39 for GLM-5.2. The best closed models remain well ahead at 181 / 247 and 216 / 293.
Z.ai sums up the pattern candidly: the further up the exploitation chain a benchmark sits, the larger the gain over GLM-5.2, and the wider the remaining gap to the closed frontier.
Real-world findings
Since GLM-5.2, Z.ai has worked with several security teams to run its models against real codebases. After expert review, screening and deduplication, the model identified 2,436 vulnerabilities across 269 projects, spanning system kernels, operating systems, browser engines, open-source infrastructure, web applications and network protocols. Z.ai tracks them in its public Security Disclosure Ledger, which at announcement showed:
| Ledger item | Count |
|---|---|
| Findings tracked | 2,436 |
| Publicly disclosed | 53 |
| Under embargo | 2,383 |
| Critical severity | 107 |
| High severity | 990 |
| Medium severity | 1,286 |
| Low severity | 53 |
| Open-source projects | 269 |
The oldest flaw dated back to code written in 1981, and on average a vulnerability had lived 26.6 years before discovery. The ledger records the affected project, severity, CVE where available and how long each bug had been in the codebase.
What this means for you
For defenders, GLM-5.3 is a strong assistant for reviewing code you own, triaging suspected bugs and writing proof-of-concept tests in an authorized setting. Z.ai released the weights two weeks after launch, once safety evaluation and hardening were complete, and the GLM-5.3 License adds a security-review condition for the very largest Model-as-a-Service providers. Use it on systems you are allowed to test.
We tested it: 5 tasks
Here is how GLM-5.3 handles the five standard tasks every model on GLM Chat runs through. Each task runs three times in the site’s own harness with the public chat’s settings: temperature 0.7, a 2,048-token output cap and the lightest reasoning setting. Because GLM-5.3 cannot turn thinking off, that means reasoning_effort: "low". The harness logs the full answer, latency and token counts, and an editor scores each answer 0 to 2.
- Python: stream-deduplicate a large CSV, keep first occurrences, support key columns and preserve the header.
- JavaScript:
debounce(fn, wait, { leading, trailing })withcancel()and tests fornode --test. - PHP refactor: turn a messy 60-line legacy function into clean PHP 8 without changing behavior, including the VIP discount, escaped output and parameterized SQL.
- Explanation: transformer attention for a beginner in about 150 words.
- Extraction: a messy product description to strict JSON with fixed keys, centimeter conversions and
nullfor missing facts.
Our test in GLM Chat

| Task | Median score | Median latency | Median output tokens | Why (across runs) |
|---|---|---|---|---|
| Python: stream-safe CSV dedupe | 2 / 2 (runs: 2, 2, 2) | 7.0 s | 474 | Streams with csv, keeps the header and first occurrence, key columns work in every run. |
| JavaScript: debounce with options + tests | 2 / 2 (runs: 2, 1, 2) | 12.0 s | 1,190 | Passes every check in most runs; its own tests fail (1/3 runs). |
| PHP: refactor a messy 60-line function | 1 / 2 (runs: 0, 1, 1) | 12.0 s | 864 | Output not escaped (2/3 runs); changes the function’s behaviour (1/3 runs). |
| Explain attention to a beginner | 1 / 2 (runs: 1, 2, 1) | 4.5 s | 209 | No queries/keys/values (2/3 runs). |
| Extract JSON from a messy description | 1 / 2 (runs: 1, 1, 1) | 2.2 s | 140 | Name padded (3/3 runs). |
| Total | 7 / 10 |
Scores run from 0 (wrong or unusable) through 1 (usable with a small fix) to 2 (correct and complete), for a maximum of 10. For GLM-5.3, the most telling numbers are the latency and token counts next to the scores: even at low effort it reasons before answering, so compare its speed with GLM-5.2, which runs with thinking fully off in the same battery. On quality, watch the PHP refactor (behavior must stay identical across all three order types) and the debounce tests, where subtle timing bugs separate careful models from fast ones.
This battery measures GLM-5.3 as an everyday assistant under a tight output cap. Its long-horizon agent strength shows up in the benchmark tables above, run at max effort with 128K outputs.
GLM-5.3 vs GLM-5.2 vs GLM-5.3-Flash
| Spec | GLM-5.3 | GLM-5.2 | GLM-5.3-Flash | GLM-5.1 |
|---|---|---|---|---|
| Release | Aug 14, 2026 | Jun 16, 2026 | Aug 26, 2026 | Apr 7, 2026 |
| Parameters | 744B / 40B active | 744B / 40B active | 320B / 18B active | 744B / 40B active |
| Context / output | 1M / 128K | 1M / 128K | 1M / 128K | 200K / 128K |
| Input | Text | Text | Video, image, text, file | Text |
| Thinking | Always on | On, can be disabled | Always on | On, can be disabled |
| Effort levels | low, high, max | high, max | low, high, max | Not supported |
| Input / output price | $1.40 / $4.40 | $1.40 / $4.40 | $0.15 / $0.50 | $1.40 / $4.40 |
| Cached input | $0.26 | $0.26 | $0.03 | $0.26 |
| Weights | GLM-5.3 License (FP8, BF16) | MIT (BF16, FP8) | MIT (FP8, BF16) | MIT (BF16, FP8) |
| Coding Plan | Included | Routed to GLM-5.3 | Included, 3x quota | Routed to GLM-5.3 |
GLM-5.3 vs GLM-5.2
Same base model, same price, better post-training. GLM-5.3 wins every benchmark Z.ai published for both and uses fewer output tokens per task on Z.ai Code Bench. Stay on GLM-5.2 only if you need thinking fully disabled, prefer a plain MIT license, or pay through OpenRouter, where GLM-5.2 is listed cheaper and has a free variant. The GLM-5.2 page covers those cases in detail.
GLM-5.3 vs GLM-5.3-Flash
GLM-5.3-Flash is a newly trained, smaller model (320B / 18B active) with native image and video input and a hybrid linear-plus-sparse attention design. It costs about a tenth as much and gets 3x the quota on the Coding Plan. Z.ai reports 29.0 for Flash on Z.ai Code Bench v1.0 at max effort, while the GLM-5.3 announcement gives GLM-5.3 34.5% at Max effort. Choose GLM-5.3 for the hardest coding and agent tasks; choose Flash for volume, vision and cost.
GLM-5.3 vs GLM-5.1
GLM-5.1 has a 200K context, no effort levels and much lower scores on long-horizon tasks, at the same price. There is no reason to start new work on it; on the Coding Plan, GLM-5.1 requests already route to GLM-5.3. See the GLM-5.1 page for its history.
Forced thinking: migrating from GLM-5.2
This is the one change that breaks existing code. GLM-5.3 always reasons, and disabling thinking is no longer supported.
| Parameter | Values | Default | Notes |
|---|---|---|---|
thinking.type | enabled | enabled | disabled is rejected |
reasoning_effort | low, high, max | max | low = light, high = enhanced, max = deep; max recommended for coding |
If your application sends "thinking": {"type": "disabled"} (common for fast chat or extraction on GLM-5.2, GLM-5.1 or GLM-4.x), change it to enabled and set reasoning_effort to low before switching the model ID to glm-5.3. Otherwise the request fails.

Before (GLM-5.2 with thinking off):
{
"model": "glm-5.2",
"thinking": {"type": "disabled"},
"messages": [{"role": "user", "content": "Classify: billing, bug or feature?"}]
}
After (GLM-5.3 with the lightest reasoning):
{
"model": "glm-5.3",
"thinking": {"type": "enabled"},
"reasoning_effort": "low",
"messages": [{"role": "user", "content": "Classify: billing, bug or feature?"}]
}
Effort values on the API vs the Coding Plan
- Pay-as-you-go API: only
low,highandmaxare accepted for GLM-5.3; any other value returns an error. - Coding Plan endpoints: values are mapped for compatibility:
none,minimalandlowbecome low;mediumandhighbecome high;xhighandmaxbecome max. A tool that sends “thinking disabled” gets low effort instead of an error. - Self-hosted weights: in the open chat template,
reasoning_effortdefaults tomaxif it is missing or unrecognized. The template’sclear_thinkingdefaults tofalse; for chat use, passclear_thinking=true.
The GLM thinking mode guide compares these controls across all GLM models.
GLM 5.3 pricing and free access
| Model | Input | Cached input | Output |
|---|---|---|---|
| GLM-5.3 | $1.40 | $0.26 | $4.40 |
| GLM-5.3-FlashX | $0.37 | $0.075 | $1.25 |
| GLM-5.3-Flash | $0.15 | $0.03 | $0.50 |
| GLM-5.2 | $1.40 | $0.26 | $4.40 |
| GLM-4.7-Flash | Free | Free | Free |
A worked example: one agent session
Say a coding agent makes 40 requests, each with 60,000 input tokens (54,000 of them repeated context that hits the cache) and 3,000 output tokens:
- Cached input: 0.054M x $0.26 = $0.01404
- New input: 0.006M x $1.40 = $0.0084
- Output: 0.003M x $4.40 = $0.0132
- Per request: about $0.0356; for 40 requests: about $1.43
Without cache hits the same session would cost 40 x (0.06M x $1.40 + $0.0132) = about $3.89. Agents resend a lot of context, so caching matters more than the headline input price. Reasoning tokens are billed as output, which is why choosing high over max for routine steps also saves money. The pricing page compares every GLM model and the context caching guide explains how to keep cache hit rates high.
Is GLM 5.3 free?
Not on the Z.ai API, but you can use it without paying per token:
- GLM Chat (this site): select GLM-5.3 and chat with GLM-5.3 for free. No account; 40 messages a day on paid models, 8,000 characters per message, with the last 20 messages (up to 24,000 characters) sent as context. The chat uses
reasoning_effort: "low"for quick replies. If the site’s daily capacity for paid models runs out, answers come from GLM-4.7-Flash for the rest of the UTC day, labeled “GLM-4.7-Flash (daily limit reached)”. - Z.ai’s chat app: GLM-5.3 is available at chat.z.ai.
- Free API alternatives: GLM-4.7-Flash is free on the Z.ai API; GLM-5.3-Flash is not free but costs $0.15 / $0.50.
GLM-5.3 on OpenRouter
OpenRouter lists z-ai/glm-5.3 at $1.40 in and $4.40 out per million tokens with a 1,310,720-token context, plus a batch route, z-ai/glm-5.3:batch, at $0.45 / $2.00. There is no free GLM-5.3 variant listed. OpenRouter prices change often, so confirm on the GLM-5.3 page on OpenRouter before you plan costs.
Using GLM 5.3 via API
The model ID is glm-5.3. For pay-as-you-go use, send OpenAI-compatible Chat Completions requests to https://api.z.ai/api/paas/v4/chat/completions with your key as a bearer token. Get a key on the API Keys page at z.ai/manage-apikey/apikey-list; the GLM API quickstart covers account setup.
| Use | Base URL |
|---|---|
| Pay-as-you-go API (Chat Completions) | https://api.z.ai/api/paas/v4 |
| Coding Plan, OpenAI-compatible tools | https://api.z.ai/api/coding/paas/v4 |
| OpenAI Responses protocol (e.g. Codex) | https://api.z.ai/api/v1 |
| Anthropic Messages protocol (Claude Code, Goose) | https://api.z.ai/api/anthropic |
Z.ai notes that accounts which have ever subscribed to a Coding Plan, including an expired one, can currently reach the model API only through the OpenAI Chat Completion-compatible protocol.
curl
curl -X POST "https://api.z.ai/api/paas/v4/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $ZAI_API_KEY" \
-d '{
"model": "glm-5.3",
"messages": [
{"role": "system", "content": "You are a senior software engineer."},
{"role": "user", "content": "Find the race condition in this code: ..."}
],
"thinking": {"type": "enabled"},
"reasoning_effort": "max",
"max_tokens": 16384,
"temperature": 1.0
}'
Python with the OpenAI SDK
# pip install --upgrade "openai>=1.0"
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["ZAI_API_KEY"],
base_url="https://api.z.ai/api/paas/v4/",
)
resp = client.chat.completions.create(
model="glm-5.3",
messages=[{"role": "user", "content": "Write a migration plan for splitting this monolith: ..."}],
max_tokens=16384,
extra_body={
"thinking": {"type": "enabled"}, # "disabled" is rejected on glm-5.3
"reasoning_effort": "high", # low | high | max
},
)
msg = resp.choices[0].message
print(getattr(msg, "reasoning_content", None))
print(msg.content)
Streaming with the official SDK
from zai import ZaiClient
client = ZaiClient(api_key="your-api-key")
stream = client.chat.completions.create(
model="glm-5.3",
messages=[{"role": "user", "content": "Summarize the risks in this diff: ..."}],
thinking={"type": "enabled"},
reasoning_effort="low",
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta
if delta.reasoning_content:
print(delta.reasoning_content, end="", flush=True)
if delta.content:
print(delta.content, end="", flush=True)
Reasoning streams in delta.reasoning_content and the answer in delta.content; the stream ends with data: [DONE].
Parameters and limits
- max_tokens: default 65,536, maximum 131,072.
- temperature: 0.0 to 1.0, default 1.0; top_p default 0.95.
- Tools: OpenAI-style function calling with
tool_choice: "auto", plustool_streamfor streamed tool arguments. See the function calling guide. - JSON output:
response_format: {"type": "json_object"}. - Errors: sending thinking
disabledmakes the request fail; rate limits return 429 with code 1302. The rate limits and error codes guide lists every code and its fix.
GLM-5.3 on the Coding Plan
GLM-5.3 is included on every GLM Coding Plan tier, alongside GLM-5.3-Flash. Requests for GLM-5.2 and GLM-5.1 are routed to GLM-5.3 automatically, and GLM-4.7 requests go to GLM-5.3-Flash.
| Tier | Price | Credits per 5 hours | Credits per week |
|---|---|---|---|
| Lite | $18 / month | 2,000 | 10,000 |
| Pro | $80 / month | 12,000 | 60,000 |
| Max | $168 / month | 28,000 | 140,000 |
How GLM-5.3 uses credits
Credits = (input tokens x 6.9 + cached input x 1.7 + output x 24) / 10,000 for GLM-5.3. Off-peak calls cost 50% of that. Peak hours are Monday to Friday, 14:00 to 18:00 UTC+8; everything else, including weekends, is off-peak. From September 25 to October 7, 2026, Z.ai charges all hours at the off-peak rate.
Take a request with 5,000 new input tokens, 45,000 cached tokens and 2,000 output tokens: (5,000 x 6.9 + 45,000 x 1.7 + 2,000 x 24) / 10,000 = (34,500 + 76,500 + 48,000) / 10,000 = 15.9 credits at peak, or about 7.95 off-peak. A Lite plan’s 2,000 credits per five hours covers roughly 125 such requests at peak or about 251 off-peak. The same request on pay-as-you-go would cost about $0.0275.
Z.ai estimates weekly GLM-5.3 allowances at a 95% cache hit rate of 48 to 97 million tokens on Lite, 290 to 580 million on Pro and 676 to 1,352 million on Max (the range runs from all-peak to all-off-peak use). The Coding Plan guide covers the break-even against pay-as-you-go.
Claude Code setup
Point Claude Code at the Anthropic-compatible endpoint and map its model slots in ~/.claude/settings.json:
{
"env": {
"ANTHROPIC_AUTH_TOKEN": "your-api-key",
"ANTHROPIC_BASE_URL": "https://api.z.ai/api/anthropic",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "glm-5.3-flash[1m]",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "glm-5.3[1m]",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "glm-5.3[1m]",
"CLAUDE_CODE_AUTO_COMPACT_WINDOW": "1000000",
"API_TIMEOUT_MS": "3000000"
}
}
The [1m] suffix enables the 1M context. Inside a session, /effort switches thinking intensity (default max) and /status confirms the active model. In OpenAI-compatible tools such as Cline, use the Coding Plan base URL, the custom model glm-5.3, a 1,000,000-token context window and leave image support off. Full walkthrough: How to use GLM in Claude Code.
Z.ai’s own desktop agent, ZCode, runs GLM-5.3 with a Goal mode that plans, codes, tests and verifies until a target is met; Z.ai reports a cache hit rate above 98% there.
The GLM-5.3 License explained
GLM-5.3 is the first GLM-5 flagship not released under plain MIT. Its custom GLM-5.3 License keeps MIT-style freedoms and adds one condition that affects only very large companies.
What you can do
The license grants, free of charge, the right to use, copy, modify, merge, publish, distribute, sublicense and sell the software, including the model weights, configuration files, inference and training code, and to run, deploy and fine-tune it and create derivative works. You must keep the copyright and permission notice in copies and follow applicable laws. The software comes “as is”, without warranty.
The one extra condition
“Model as a Service” means giving a third party access to language model inference or fine-tuning (e.g., via API) in a manner that allows such third party to exercise meaningful control over the inputs, parameters, or training data.
GLM-5.3 License, clause 2
If you (together with your affiliates) operate such a Model-as-a-Service business and your combined revenue exceeds $10 billion over any consecutive 12 months, you must pass Z.ai’s security review before using GLM-5.3 or its derivatives for any commercial purpose. Two things are explicitly not Model as a Service: end-user products where the model is embedded in specific features or harnesses, and simply relaying requests to models hosted by others.
What it means in practice
- Startups, most companies, researchers and hobbyists: effectively MIT terms. Self-host, fine-tune, sell products built on it.
- A SaaS app with GLM-5.3 inside a feature: not Model as a Service, so no review, whatever your size.
- A giant cloud or API provider selling raw GLM-5.3 inference: needs Z.ai’s security review first. Questions go to glmlicense@z.ai.
This is a summary, not legal advice. The Is GLM open source? guide compares the GLM-5.3 License with the MIT license used by every other open GLM model.
Downloading the weights
Z.ai published GLM-5.3’s weights on August 25, 2026, after the two-week safety evaluation it promised at launch.
| Variant | Hugging Face | ModelScope | Precision | Size |
|---|---|---|---|---|
| GLM-5.3 | zai-org/GLM-5.3 | ZhipuAI/GLM-5.3 | FP8 | 744B-A40B |
| GLM-5.3-BF16 | zai-org/GLM-5.3-BF16 | ZhipuAI/GLM-5.3-BF16 | BF16 | 744B-A40B |
Note that the default repo is the FP8 build; the BF16 build carries the suffix, the reverse of GLM-5.2’s naming.
Memory for the weights
Simple arithmetic for the weights alone, before KV cache and runtime overhead: FP8 at 1 byte per parameter is about 744 GB, and BF16 at 2 bytes is about 1,488 GB. Only 40B parameters are active per token, but all experts must be loaded, so this is multi-GPU server territory.
Serving frameworks
Z.ai lists SGLang, vLLM, Transformers, KTransformers and Unsloth for GLM-5.3 and the earlier GLM-5 models, with an SGLang cookbook and a vLLM recipe for GLM-5.3 specifically. On Ascend NPUs, vLLM-Ascend, xLLM and SGLang are supported. Remember the template defaults: reasoning_effort falls back to max, and chat use should pass clear_thinking=true. For fine-tuning, the GLM-5 series supports slime (v0.3.0+) and ms-swift (v4.4.0+). The GLM-5 GitHub repository has the download table and deployment links.
GLM-5.3 FAQ
Is GLM 5.3 free?
The API is paid: $1.40 per million input tokens and $4.40 per million output tokens. You can use GLM-5.3 free in the chat on this site (40 messages a day on paid models, no account) and in Z.ai’s chat app at chat.z.ai. The weights are free to download under the GLM-5.3 License.
What is GLM-5.3’s context window?
1M tokens, with up to 128K output tokens (131,072) per response; the default max_tokens is 65,536. OpenRouter lists a 1,310,720-token context for its route.
Can you turn off thinking in GLM-5.3?
No. GLM-5.3 always reasons. Sending "thinking": {"type": "disabled"} makes the request fail on the API. Use "thinking": {"type": "enabled"} with reasoning_effort: "low" for the fastest answers, or switch to GLM-5.2 if you need no reasoning at all.
Is GLM-5.3 open source?
Its weights are open on Hugging Face and ModelScope under the GLM-5.3 License, which allows commercial use, modification and redistribution. The only extra condition applies to Model-as-a-Service businesses with more than $10 billion in 12-month revenue, which need Z.ai’s security review first.
How much does the GLM 5.3 API cost?
$1.40 per million input tokens, $0.26 per million cached input tokens and $4.40 per million output tokens on the Z.ai API, the same as GLM-5.2 and GLM-5.1. Reasoning tokens are billed as output. On the Coding Plan, GLM-5.3 is included from $18 a month.
When was GLM-5.3 released?
Z.ai announced GLM-5.3 on August 14, 2026. The API release-notes entry is dated August 18, 2026, and the weights were published on Hugging Face on August 25, 2026.
Does GLM-5.3 support images?
No, GLM-5.3 is text-only. For image, video or file input, use GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, which also costs much less.
What is the difference between GLM-5.3 and GLM-5.3-Flash?
GLM-5.3 is the 744B flagship built on GLM-5.2’s base model, text-only, at $1.40 / $4.40. GLM-5.3-Flash is a new 320B / 18B-active multimodal model at $0.15 / $0.50 with MIT weights. GLM-5.3 is stronger on the hardest coding tasks; Flash is the better default for volume and vision. The GLM models comparison and the release timeline show how both fit into the family.
Want to see GLM-5.3 reason through your own problem? Chat with GLM-5.3 now, then compare it with GLM-5.2 or GLM-5.3-Flash using the model chips.