GLM-5.3: Specs, Pricing, API & Free Chat

Z.ai's current flagship: GLM-5.2's base model with far stronger post-training for coding, agents and security work. Specs, benchmarks, prices, code and the license explained simply.

Chat with GLM-5.3 now Use it via API

GLM 5.3 (officially GLM-5.3) is Z.ai’s current flagship model, announced on August 14, 2026 by the company formerly known as Zhipu AI. It keeps GLM-5.2’s base model (744B total parameters, 40B active, Mixture-of-Experts) and gets every improvement from post-training. It reads text only, holds a 1M-token context, writes up to 128K tokens per answer and always reasons, with three effort levels: low, high and max. The Z.ai API charges $1.40 per million input tokens and $4.40 per million output tokens, and the weights are open under the custom GLM-5.3 License.

Z.ai’s headline claims: a 50% gain over GLM-5.2 on its in-house Z.ai Code Bench, the best open-model scores on Terminal Bench 3.0 and Agents’ Last Exam (CLI), and the top result on the CyberGym vulnerability-discovery benchmark. The API release-notes entry followed on August 18, and the weights reached Hugging Face on August 25 after a safety evaluation.

Below you will find the full benchmark table, the one API change that breaks old code (thinking can no longer be disabled), prices, free routes, working code, the Coding Plan, what the GLM-5.3 License actually requires and how to download the weights. To try it right away, Chat with GLM-5.3 now on this site, free and without an account.

GLM 5.3 specs: 1M-token context, 744B total and 40B active parameters, always-on reasoning, $1.40 in and $4.40 out per million tokens
GLM-5.3 key specs at a glance.

What GLM-5.3 is good at

Z.ai describes GLM-5.3 as its latest flagship for complex software engineering and agent work. Because it shares GLM-5.2’s base model, the differences you feel come from how it was trained after pre-training: more environments, more varied tasks and more compute spent on reinforcement learning.

GLM-5.3 specs at a glance

SpecGLM-5.3
DeveloperZ.ai (formerly Zhipu AI)
AnnouncedAugust 14, 2026 (API release notes August 18)
Base modelSame as GLM-5.2; all gains from post-training
Parameters744B total, 40B active (MoE)
Context / max output1M / 128K tokens
Input / outputText / text
ReasoningAlways on; effort low, high or max (default max)
API model IDglm-5.3
API price$1.40 in, $0.26 cached, $4.40 out per 1M
WeightsGLM-5.3 License; FP8 and BF16 on Hugging Face and ModelScope
CapabilitiesStreaming, function calling, context caching, structured output
GLM-5.3 technical specifications from Z.ai’s documentation, blog and model card.

Real engineering work, not coding exercises

For GLM-5.3, Z.ai pushed its training environments toward “real units of expert work”, some representing several days of effort for an experienced engineer. One example from the announcement: the model gets the same working environment as an ML infrastructure engineer (compute clusters, storage, internal documentation, codebases, experiment results) and must diagnose bottlenecks, implement optimizations, run experiments and deliver a measurable speedup without breaking correctness. The aim is a model that owns a substantial task end to end instead of waiting for a human to break it into steps.

Better results with fewer tokens

On Z.ai Code Bench, a private benchmark of realistic coding-agent tasks, GLM-5.3 improves both success and efficiency. At Max effort it completes 34.5% of tasks with roughly 75K output tokens per task, against 23.4% at 96K for GLM-5.2. At High effort it reaches 31.4% with about 50K tokens, ahead of Claude Opus 4.8 at 29.5% with 120K. It still trails Claude Fable 5, which reaches 39.5% at Max effort. Fewer output tokens per task means lower bills and faster turnarounds, not just higher scores.

Terminal and long-horizon agent tasks

The largest public gains are on hard agent benchmarks: Terminal Bench 3.0 from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9 and Agents’ Last Exam (CLI) from 23.8 to 28.5. GLM-5.3 carries over GLM-5.2’s long-horizon RL techniques, including compaction, which helps these gains hold on tasks that run for hours.

General agent and knowledge work

The gains are not limited to code. AutomationBench rises from 26.2 to 48.2, Toolathlon Verified from 59.9 to 73.0, and GDPval-AA v2 (evaluated by Artificial Analysis) from 1508 to 1769, the highest score in Z.ai’s table.

Security research

GLM-5.3 is the strongest model in Z.ai’s table on CyberGym and more than doubles GLM-5.2 on exploitation benchmarks. That makes it a capable code-review and vulnerability-discovery assistant; see the cyber section below for the numbers and the real-world findings.

How Z.ai trained it

GLM-5.3’s story is about training environments more than architecture. Z.ai built pipelines that synthesize long-horizon task environments end to end: research agents collect task patterns from real work and turn them into runnable environments with multi-step dependencies and hidden state, and a judge agent attempts each task to confirm it is solvable. Verifiers are written without access to the reference solution, and solver trajectories are used to find and close reward shortcuts. Only a verifier that passes oracle, no-op and unsolved-state checks is trusted as a reward signal. Z.ai notes the pipeline still needs meaningful human-in-the-loop work.

Under the hood, the same slime RL framework used for GLM-5.2 was extended: Z.ai reports that training-rollout log-probability differences were brought down to the 1e-7 level (a reduction of more than 99.99%), and that system-level optimizations raised end-to-end RL throughput on long-horizon coding tasks by more than 2.3x.

Choosing an effort level

  • low: chat, classification, extraction, quick edits. The fastest option, and the lightest setting GLM-5.3 allows.
  • high: routine agent steps and everyday coding. Z.ai’s Code Bench shows High already beating Claude Opus 4.8 with far fewer tokens.
  • max: the default. Use it for hard debugging, large refactors and long agent runs, as Z.ai recommends for coding.

Where GLM-5.3 is weaker

  • Text only. No image, video or file input. For screenshots and UI work use GLM-5.3-Flash, which is natively multimodal and about a tenth of the price.
  • Always reasons. Every request includes some thinking, so simple lookups cost more latency and output tokens than on a model that can skip reasoning.
  • Not the frontier leader everywhere. In Z.ai’s own table, GPT-5.6 Sol leads Terminal Bench 2.1, Terminal Bench 3.0 and DeepSWE; Claude Opus 4.8 leads NL2Repo and SWE-Marathon; Fable 5 leads ProgramBench, FrontierSWE and PostTrainBench.

GLM 5.3 benchmarks

These are Z.ai’s published results from the GLM-5.3 announcement and model card. A dash means Z.ai did not report that model.

Coding

BenchmarkGLM-5.3GLM-5.2Kimi K3DeepSeek-V4-Pro-0813Qwen3.8-MaxOpus 4.8Fable 5 (w/ fallback)GPT-5.6 Sol
Terminal Bench 2.188.281.088.387.986.685.088.088.8
Terminal Bench 3.028.34.617.4––21.133.734.6
DeepSWE v1.166.946.267.562.756.658.069.772.7
NL2Repo58.048.958.061.155.969.7––
ProgramBench (Almost Solved)19.09.517.5–10.515.533.023.0
FrontierSWE78.167.5–––66.588.2–
SWE-Marathon v1.142.519.448.1––48.833.142.5
PostTrainBench39.831.732.0––32.941.836.2
GLM-5.3 coding benchmarks as reported by Z.ai.

Cyber

BenchmarkGLM-5.3GLM-5.2Kimi K3DeepSeek-V4-Pro-0813Qwen3.8-MaxOpus 4.8Fable 5 (w/ fallback)GPT-5.6 Sol
CyberGym84.577.280.083.378.578.183.883.6
ExploitGym 2h / 6h105 / 13029 / 3936 / 70–14 / 2680 / 120181 / 247216 / 293
ExploitBench54.424.432.2–28.840.078.076.5
GLM-5.3 cybersecurity benchmarks as reported by Z.ai.

Agentic

BenchmarkGLM-5.3GLM-5.2Kimi K3DeepSeek-V4-Pro-0813Qwen3.8-MaxOpus 4.8Fable 5 (w/ fallback)GPT-5.6 Sol
Toolathlon Verified73.059.976.574.172.576.274.774.9
AutomationBench v1.0.648.226.246.743.239.841.046.245.8
Agents’ Last Exam (CLI)28.523.827.625.727.025.723.828.6
HLE w/ Tools62.554.759.860.056.257.963.964.5
GDPval-AA v217691508168215901739158817431730
GLM-5.3 agentic benchmarks as reported by Z.ai.

What stands out

  • Against GLM-5.2, GLM-5.3 wins every row. The biggest relative jumps are Terminal Bench 3.0 (4.6 to 28.3), ExploitGym (29 to 105 tasks in two hours), ExploitBench (24.4 to 54.4) and AutomationBench (26.2 to 48.2).
  • Against other open-weight models, it is level with Kimi K3 on Terminal Bench 2.1 and NL2Repo, ahead on Terminal Bench 3.0 (28.3 vs 17.4) and PostTrainBench, and behind on SWE-Marathon (42.5 vs 48.1) and Toolathlon (73.0 vs 76.5). DeepSeek-V4-Pro-0813 leads it on NL2Repo (61.1 vs 58.0).
  • Against closed models, GLM-5.3 beats Opus 4.8 on 13 of the 16 rows Z.ai published for both, but GPT-5.6 Sol and Fable 5 remain ahead on most coding and exploitation rows.

How Z.ai ran the evaluations

Most agentic suites ran in the Claude Code harness at max effort. Terminal Bench 3.0 used a 400K context and 128K output, averaged over three rollouts with up to 600 agent turns and a 10-hour limit. Agents’ Last Exam ran its 105 tasks in isolated containers with a 1M context. CyberGym ran 1,507 tasks with no web tools and a domain allowlist to prevent cheating. ExploitGym normalized its two-hour and six-hour time limits by each model’s measured throughput. FrontierSWE was run by Proximal at 1M context. These are heavy, max-effort settings; expect different numbers at low or high effort.

GLM-5.3’s cyber capability

Z.ai calls this an “emergent” capability. It added vulnerability-discovery data and environments to post-training expecting modest gains, and found the skill kept growing as training scaled: GLM-5.3 began planning multi-stage exploitation chains, not just spotting isolated flaws.

GLM 5.3 cyber results: CyberGym 84.5, ExploitBench 54.4, 2,436 real-world vulnerabilities across 269 projects
GLM-5.3’s security results in numbers.

Three benchmarks, three stages

  • CyberGym (finding bugs): starting from white-box source code, the model must identify a vulnerability and prove it by triggering a fault. GLM-5.3 scores 84.5%, up from 77.2% for GLM-5.2 and the best result on the benchmark, ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%) in Z.ai’s write-up. (Z.ai’s table lists the 83.8% score under Fable 5 with fallback.)
  • ExploitBench (understanding exploitation): deeper reasoning about real vulnerabilities and how they can be exploited. GLM-5.3 reaches 54.4%, more than double GLM-5.2’s 24.4%, while the top closed models score 78.0% and 76.5%.
  • ExploitGym (completing exploits under time limits): GLM-5.3 completes 105 tasks in two hours and 130 in six, against 29 and 39 for GLM-5.2. The best closed models remain well ahead at 181 / 247 and 216 / 293.

Z.ai sums up the pattern candidly: the further up the exploitation chain a benchmark sits, the larger the gain over GLM-5.2, and the wider the remaining gap to the closed frontier.

Real-world findings

Since GLM-5.2, Z.ai has worked with several security teams to run its models against real codebases. After expert review, screening and deduplication, the model identified 2,436 vulnerabilities across 269 projects, spanning system kernels, operating systems, browser engines, open-source infrastructure, web applications and network protocols. Z.ai tracks them in its public Security Disclosure Ledger, which at announcement showed:

Ledger itemCount
Findings tracked2,436
Publicly disclosed53
Under embargo2,383
Critical severity107
High severity990
Medium severity1,286
Low severity53
Open-source projects269
Z.ai Security Disclosure Ledger figures published with GLM-5.3.

The oldest flaw dated back to code written in 1981, and on average a vulnerability had lived 26.6 years before discovery. The ledger records the affected project, severity, CVE where available and how long each bug had been in the codebase.

What this means for you

For defenders, GLM-5.3 is a strong assistant for reviewing code you own, triaging suspected bugs and writing proof-of-concept tests in an authorized setting. Z.ai released the weights two weeks after launch, once safety evaluation and hardening were complete, and the GLM-5.3 License adds a security-review condition for the very largest Model-as-a-Service providers. Use it on systems you are allowed to test.

We tested it: 5 tasks

Here is how GLM-5.3 handles the five standard tasks every model on GLM Chat runs through. Each task runs three times in the site’s own harness with the public chat’s settings: temperature 0.7, a 2,048-token output cap and the lightest reasoning setting. Because GLM-5.3 cannot turn thinking off, that means reasoning_effort: "low". The harness logs the full answer, latency and token counts, and an editor scores each answer 0 to 2.

  1. Python: stream-deduplicate a large CSV, keep first occurrences, support key columns and preserve the header.
  2. JavaScript: debounce(fn, wait, { leading, trailing }) with cancel() and tests for node --test.
  3. PHP refactor: turn a messy 60-line legacy function into clean PHP 8 without changing behavior, including the VIP discount, escaped output and parameterized SQL.
  4. Explanation: transformer attention for a beginner in about 150 words.
  5. Extraction: a messy product description to strict JSON with fixed keys, centimeter conversions and null for missing facts.

Our test in GLM Chat

Runs: 3 runs per task, median score, September 24, 2026.

GLM-5.3 answering task 1 (Python: stream-safe CSV dedupe) in GLM Chat
GLM-5.3’s best-scoring answer in our battery (best of 3 runs): task 1, Python: stream-safe CSV dedupe.
TaskMedian scoreMedian latencyMedian output tokensWhy (across runs)
Python: stream-safe CSV dedupe2 / 2 (runs: 2, 2, 2)7.0 s474Streams with csv, keeps the header and first occurrence, key columns work in every run.
JavaScript: debounce with options + tests2 / 2 (runs: 2, 1, 2)12.0 s1,190Passes every check in most runs; its own tests fail (1/3 runs).
PHP: refactor a messy 60-line function1 / 2 (runs: 0, 1, 1)12.0 s864Output not escaped (2/3 runs); changes the function’s behaviour (1/3 runs).
Explain attention to a beginner1 / 2 (runs: 1, 2, 1)4.5 s209No queries/keys/values (2/3 runs).
Extract JSON from a messy description1 / 2 (runs: 1, 1, 1)2.2 s140Name padded (3/3 runs).
Total7 / 10
Our test results. Runs: 3 runs per task, median score, September 24, 2026, in GLM Chat with the settings in our editorial policy.

Scores run from 0 (wrong or unusable) through 1 (usable with a small fix) to 2 (correct and complete), for a maximum of 10. For GLM-5.3, the most telling numbers are the latency and token counts next to the scores: even at low effort it reasons before answering, so compare its speed with GLM-5.2, which runs with thinking fully off in the same battery. On quality, watch the PHP refactor (behavior must stay identical across all three order types) and the debounce tests, where subtle timing bugs separate careful models from fast ones.

This battery measures GLM-5.3 as an everyday assistant under a tight output cap. Its long-horizon agent strength shows up in the benchmark tables above, run at max effort with 128K outputs.

GLM-5.3 vs GLM-5.2 vs GLM-5.3-Flash

SpecGLM-5.3GLM-5.2GLM-5.3-FlashGLM-5.1
ReleaseAug 14, 2026Jun 16, 2026Aug 26, 2026Apr 7, 2026
Parameters744B / 40B active744B / 40B active320B / 18B active744B / 40B active
Context / output1M / 128K1M / 128K1M / 128K200K / 128K
InputTextTextVideo, image, text, fileText
ThinkingAlways onOn, can be disabledAlways onOn, can be disabled
Effort levelslow, high, maxhigh, maxlow, high, maxNot supported
Input / output price$1.40 / $4.40$1.40 / $4.40$0.15 / $0.50$1.40 / $4.40
Cached input$0.26$0.26$0.03$0.26
WeightsGLM-5.3 License (FP8, BF16)MIT (BF16, FP8)MIT (FP8, BF16)MIT (BF16, FP8)
Coding PlanIncludedRouted to GLM-5.3Included, 3x quotaRouted to GLM-5.3
GLM-5.3 compared with its closest GLM siblings. Prices per 1M tokens on the Z.ai API.

GLM-5.3 vs GLM-5.2

Same base model, same price, better post-training. GLM-5.3 wins every benchmark Z.ai published for both and uses fewer output tokens per task on Z.ai Code Bench. Stay on GLM-5.2 only if you need thinking fully disabled, prefer a plain MIT license, or pay through OpenRouter, where GLM-5.2 is listed cheaper and has a free variant. The GLM-5.2 page covers those cases in detail.

GLM-5.3 vs GLM-5.3-Flash

GLM-5.3-Flash is a newly trained, smaller model (320B / 18B active) with native image and video input and a hybrid linear-plus-sparse attention design. It costs about a tenth as much and gets 3x the quota on the Coding Plan. Z.ai reports 29.0 for Flash on Z.ai Code Bench v1.0 at max effort, while the GLM-5.3 announcement gives GLM-5.3 34.5% at Max effort. Choose GLM-5.3 for the hardest coding and agent tasks; choose Flash for volume, vision and cost.

GLM-5.3 vs GLM-5.1

GLM-5.1 has a 200K context, no effort levels and much lower scores on long-horizon tasks, at the same price. There is no reason to start new work on it; on the Coding Plan, GLM-5.1 requests already route to GLM-5.3. See the GLM-5.1 page for its history.

Forced thinking: migrating from GLM-5.2

This is the one change that breaks existing code. GLM-5.3 always reasons, and disabling thinking is no longer supported.

ParameterValuesDefaultNotes
thinking.typeenabledenableddisabled is rejected
reasoning_effortlow, high, maxmaxlow = light, high = enhanced, max = deep; max recommended for coding
GLM-5.3 thinking parameters from Z.ai’s documentation.

If your application sends "thinking": {"type": "disabled"} (common for fast chat or extraction on GLM-5.2, GLM-5.1 or GLM-4.x), change it to enabled and set reasoning_effort to low before switching the model ID to glm-5.3. Otherwise the request fails.

Migrating to GLM 5.3: set thinking to enabled, add reasoning_effort low, then change the model ID
Three steps to move a thinking-off app to GLM-5.3.

Before (GLM-5.2 with thinking off):

{
  "model": "glm-5.2",
  "thinking": {"type": "disabled"},
  "messages": [{"role": "user", "content": "Classify: billing, bug or feature?"}]
}

After (GLM-5.3 with the lightest reasoning):

{
  "model": "glm-5.3",
  "thinking": {"type": "enabled"},
  "reasoning_effort": "low",
  "messages": [{"role": "user", "content": "Classify: billing, bug or feature?"}]
}

Effort values on the API vs the Coding Plan

  • Pay-as-you-go API: only low, high and max are accepted for GLM-5.3; any other value returns an error.
  • Coding Plan endpoints: values are mapped for compatibility: none, minimal and low become low; medium and high become high; xhigh and max become max. A tool that sends “thinking disabled” gets low effort instead of an error.
  • Self-hosted weights: in the open chat template, reasoning_effort defaults to max if it is missing or unrecognized. The template’s clear_thinking defaults to false; for chat use, pass clear_thinking=true.

The GLM thinking mode guide compares these controls across all GLM models.

GLM 5.3 pricing and free access

ModelInputCached inputOutput
GLM-5.3$1.40$0.26$4.40
GLM-5.3-FlashX$0.37$0.075$1.25
GLM-5.3-Flash$0.15$0.03$0.50
GLM-5.2$1.40$0.26$4.40
GLM-4.7-FlashFreeFreeFree
Z.ai API prices per 1M tokens. Cached-input storage is listed as free for a limited time.

A worked example: one agent session

Say a coding agent makes 40 requests, each with 60,000 input tokens (54,000 of them repeated context that hits the cache) and 3,000 output tokens:

  • Cached input: 0.054M x $0.26 = $0.01404
  • New input: 0.006M x $1.40 = $0.0084
  • Output: 0.003M x $4.40 = $0.0132
  • Per request: about $0.0356; for 40 requests: about $1.43

Without cache hits the same session would cost 40 x (0.06M x $1.40 + $0.0132) = about $3.89. Agents resend a lot of context, so caching matters more than the headline input price. Reasoning tokens are billed as output, which is why choosing high over max for routine steps also saves money. The pricing page compares every GLM model and the context caching guide explains how to keep cache hit rates high.

Is GLM 5.3 free?

Not on the Z.ai API, but you can use it without paying per token:

  1. GLM Chat (this site): select GLM-5.3 and chat with GLM-5.3 for free. No account; 40 messages a day on paid models, 8,000 characters per message, with the last 20 messages (up to 24,000 characters) sent as context. The chat uses reasoning_effort: "low" for quick replies. If the site’s daily capacity for paid models runs out, answers come from GLM-4.7-Flash for the rest of the UTC day, labeled “GLM-4.7-Flash (daily limit reached)”.
  2. Z.ai’s chat app: GLM-5.3 is available at chat.z.ai.
  3. Free API alternatives: GLM-4.7-Flash is free on the Z.ai API; GLM-5.3-Flash is not free but costs $0.15 / $0.50.

GLM-5.3 on OpenRouter

OpenRouter lists z-ai/glm-5.3 at $1.40 in and $4.40 out per million tokens with a 1,310,720-token context, plus a batch route, z-ai/glm-5.3:batch, at $0.45 / $2.00. There is no free GLM-5.3 variant listed. OpenRouter prices change often, so confirm on the GLM-5.3 page on OpenRouter before you plan costs.

Using GLM 5.3 via API

The model ID is glm-5.3. For pay-as-you-go use, send OpenAI-compatible Chat Completions requests to https://api.z.ai/api/paas/v4/chat/completions with your key as a bearer token. Get a key on the API Keys page at z.ai/manage-apikey/apikey-list; the GLM API quickstart covers account setup.

UseBase URL
Pay-as-you-go API (Chat Completions)https://api.z.ai/api/paas/v4
Coding Plan, OpenAI-compatible toolshttps://api.z.ai/api/coding/paas/v4
OpenAI Responses protocol (e.g. Codex)https://api.z.ai/api/v1
Anthropic Messages protocol (Claude Code, Goose)https://api.z.ai/api/anthropic
Base URLs for GLM-5.3.

Z.ai notes that accounts which have ever subscribed to a Coding Plan, including an expired one, can currently reach the model API only through the OpenAI Chat Completion-compatible protocol.

curl

curl -X POST "https://api.z.ai/api/paas/v4/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $ZAI_API_KEY" \
  -d '{
    "model": "glm-5.3",
    "messages": [
      {"role": "system", "content": "You are a senior software engineer."},
      {"role": "user", "content": "Find the race condition in this code: ..."}
    ],
    "thinking": {"type": "enabled"},
    "reasoning_effort": "max",
    "max_tokens": 16384,
    "temperature": 1.0
  }'

Python with the OpenAI SDK

# pip install --upgrade "openai>=1.0"
import os
from openai import OpenAI
client = OpenAI(
    api_key=os.environ["ZAI_API_KEY"],
    base_url="https://api.z.ai/api/paas/v4/",
)
resp = client.chat.completions.create(
    model="glm-5.3",
    messages=[{"role": "user", "content": "Write a migration plan for splitting this monolith: ..."}],
    max_tokens=16384,
    extra_body={
        "thinking": {"type": "enabled"},   # "disabled" is rejected on glm-5.3
        "reasoning_effort": "high",        # low | high | max
    },
)
msg = resp.choices[0].message
print(getattr(msg, "reasoning_content", None))
print(msg.content)

Streaming with the official SDK

from zai import ZaiClient
client = ZaiClient(api_key="your-api-key")
stream = client.chat.completions.create(
    model="glm-5.3",
    messages=[{"role": "user", "content": "Summarize the risks in this diff: ..."}],
    thinking={"type": "enabled"},
    reasoning_effort="low",
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta
    if delta.reasoning_content:
        print(delta.reasoning_content, end="", flush=True)
    if delta.content:
        print(delta.content, end="", flush=True)

Reasoning streams in delta.reasoning_content and the answer in delta.content; the stream ends with data: [DONE].

Parameters and limits

  • max_tokens: default 65,536, maximum 131,072.
  • temperature: 0.0 to 1.0, default 1.0; top_p default 0.95.
  • Tools: OpenAI-style function calling with tool_choice: "auto", plus tool_stream for streamed tool arguments. See the function calling guide.
  • JSON output: response_format: {"type": "json_object"}.
  • Errors: sending thinking disabled makes the request fail; rate limits return 429 with code 1302. The rate limits and error codes guide lists every code and its fix.

GLM-5.3 on the Coding Plan

GLM-5.3 is included on every GLM Coding Plan tier, alongside GLM-5.3-Flash. Requests for GLM-5.2 and GLM-5.1 are routed to GLM-5.3 automatically, and GLM-4.7 requests go to GLM-5.3-Flash.

TierPriceCredits per 5 hoursCredits per week
Lite$18 / month2,00010,000
Pro$80 / month12,00060,000
Max$168 / month28,000140,000
GLM Coding Plan tiers (monthly billing; quarterly saves 20%, yearly 30%).

How GLM-5.3 uses credits

Credits = (input tokens x 6.9 + cached input x 1.7 + output x 24) / 10,000 for GLM-5.3. Off-peak calls cost 50% of that. Peak hours are Monday to Friday, 14:00 to 18:00 UTC+8; everything else, including weekends, is off-peak. From September 25 to October 7, 2026, Z.ai charges all hours at the off-peak rate.

Take a request with 5,000 new input tokens, 45,000 cached tokens and 2,000 output tokens: (5,000 x 6.9 + 45,000 x 1.7 + 2,000 x 24) / 10,000 = (34,500 + 76,500 + 48,000) / 10,000 = 15.9 credits at peak, or about 7.95 off-peak. A Lite plan’s 2,000 credits per five hours covers roughly 125 such requests at peak or about 251 off-peak. The same request on pay-as-you-go would cost about $0.0275.

Z.ai estimates weekly GLM-5.3 allowances at a 95% cache hit rate of 48 to 97 million tokens on Lite, 290 to 580 million on Pro and 676 to 1,352 million on Max (the range runs from all-peak to all-off-peak use). The Coding Plan guide covers the break-even against pay-as-you-go.

Claude Code setup

Point Claude Code at the Anthropic-compatible endpoint and map its model slots in ~/.claude/settings.json:

{
  "env": {
    "ANTHROPIC_AUTH_TOKEN": "your-api-key",
    "ANTHROPIC_BASE_URL": "https://api.z.ai/api/anthropic",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "glm-5.3-flash[1m]",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "glm-5.3[1m]",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "glm-5.3[1m]",
    "CLAUDE_CODE_AUTO_COMPACT_WINDOW": "1000000",
    "API_TIMEOUT_MS": "3000000"
  }
}

The [1m] suffix enables the 1M context. Inside a session, /effort switches thinking intensity (default max) and /status confirms the active model. In OpenAI-compatible tools such as Cline, use the Coding Plan base URL, the custom model glm-5.3, a 1,000,000-token context window and leave image support off. Full walkthrough: How to use GLM in Claude Code.

Z.ai’s own desktop agent, ZCode, runs GLM-5.3 with a Goal mode that plans, codes, tests and verifies until a target is met; Z.ai reports a cache hit rate above 98% there.

The GLM-5.3 License explained

GLM-5.3 is the first GLM-5 flagship not released under plain MIT. Its custom GLM-5.3 License keeps MIT-style freedoms and adds one condition that affects only very large companies.

What you can do

The license grants, free of charge, the right to use, copy, modify, merge, publish, distribute, sublicense and sell the software, including the model weights, configuration files, inference and training code, and to run, deploy and fine-tune it and create derivative works. You must keep the copyright and permission notice in copies and follow applicable laws. The software comes “as is”, without warranty.

The one extra condition

“Model as a Service” means giving a third party access to language model inference or fine-tuning (e.g., via API) in a manner that allows such third party to exercise meaningful control over the inputs, parameters, or training data.

GLM-5.3 License, clause 2

If you (together with your affiliates) operate such a Model-as-a-Service business and your combined revenue exceeds $10 billion over any consecutive 12 months, you must pass Z.ai’s security review before using GLM-5.3 or its derivatives for any commercial purpose. Two things are explicitly not Model as a Service: end-user products where the model is embedded in specific features or harnesses, and simply relaying requests to models hosted by others.

What it means in practice

  • Startups, most companies, researchers and hobbyists: effectively MIT terms. Self-host, fine-tune, sell products built on it.
  • A SaaS app with GLM-5.3 inside a feature: not Model as a Service, so no review, whatever your size.
  • A giant cloud or API provider selling raw GLM-5.3 inference: needs Z.ai’s security review first. Questions go to glmlicense@z.ai.

This is a summary, not legal advice. The Is GLM open source? guide compares the GLM-5.3 License with the MIT license used by every other open GLM model.

Downloading the weights

Z.ai published GLM-5.3’s weights on August 25, 2026, after the two-week safety evaluation it promised at launch.

VariantHugging FaceModelScopePrecisionSize
GLM-5.3zai-org/GLM-5.3ZhipuAI/GLM-5.3FP8744B-A40B
GLM-5.3-BF16zai-org/GLM-5.3-BF16ZhipuAI/GLM-5.3-BF16BF16744B-A40B
Official GLM-5.3 weight repositories.

Note that the default repo is the FP8 build; the BF16 build carries the suffix, the reverse of GLM-5.2’s naming.

Memory for the weights

Simple arithmetic for the weights alone, before KV cache and runtime overhead: FP8 at 1 byte per parameter is about 744 GB, and BF16 at 2 bytes is about 1,488 GB. Only 40B parameters are active per token, but all experts must be loaded, so this is multi-GPU server territory.

Serving frameworks

Z.ai lists SGLang, vLLM, Transformers, KTransformers and Unsloth for GLM-5.3 and the earlier GLM-5 models, with an SGLang cookbook and a vLLM recipe for GLM-5.3 specifically. On Ascend NPUs, vLLM-Ascend, xLLM and SGLang are supported. Remember the template defaults: reasoning_effort falls back to max, and chat use should pass clear_thinking=true. For fine-tuning, the GLM-5 series supports slime (v0.3.0+) and ms-swift (v4.4.0+). The GLM-5 GitHub repository has the download table and deployment links.

GLM-5.3 FAQ

Is GLM 5.3 free?

The API is paid: $1.40 per million input tokens and $4.40 per million output tokens. You can use GLM-5.3 free in the chat on this site (40 messages a day on paid models, no account) and in Z.ai’s chat app at chat.z.ai. The weights are free to download under the GLM-5.3 License.

What is GLM-5.3’s context window?

1M tokens, with up to 128K output tokens (131,072) per response; the default max_tokens is 65,536. OpenRouter lists a 1,310,720-token context for its route.

Can you turn off thinking in GLM-5.3?

No. GLM-5.3 always reasons. Sending "thinking": {"type": "disabled"} makes the request fail on the API. Use "thinking": {"type": "enabled"} with reasoning_effort: "low" for the fastest answers, or switch to GLM-5.2 if you need no reasoning at all.

Is GLM-5.3 open source?

Its weights are open on Hugging Face and ModelScope under the GLM-5.3 License, which allows commercial use, modification and redistribution. The only extra condition applies to Model-as-a-Service businesses with more than $10 billion in 12-month revenue, which need Z.ai’s security review first.

How much does the GLM 5.3 API cost?

$1.40 per million input tokens, $0.26 per million cached input tokens and $4.40 per million output tokens on the Z.ai API, the same as GLM-5.2 and GLM-5.1. Reasoning tokens are billed as output. On the Coding Plan, GLM-5.3 is included from $18 a month.

When was GLM-5.3 released?

Z.ai announced GLM-5.3 on August 14, 2026. The API release-notes entry is dated August 18, 2026, and the weights were published on Hugging Face on August 25, 2026.

Does GLM-5.3 support images?

No, GLM-5.3 is text-only. For image, video or file input, use GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, which also costs much less.

What is the difference between GLM-5.3 and GLM-5.3-Flash?

GLM-5.3 is the 744B flagship built on GLM-5.2’s base model, text-only, at $1.40 / $4.40. GLM-5.3-Flash is a new 320B / 18B-active multimodal model at $0.15 / $0.50 with MIT weights. GLM-5.3 is stronger on the hardest coding tasks; Flash is the better default for volume and vision. The GLM models comparison and the release timeline show how both fit into the family.

Want to see GLM-5.3 reason through your own problem? Chat with GLM-5.3 now, then compare it with GLM-5.2 or GLM-5.3-Flash using the model chips.

Try GLM-5.3 on your own prompt

Free, no sign-up. Your conversation stays in your browser.

Open chat