GLM vs Claude, short answer: Claude still holds the top coding scores, but GLM gets surprisingly close for a fraction of the price. In Z.ai’s published table, Claude Fable 5 leads GLM-5.3 on the hardest coding benchmarks, while GLM-5.3 beats Claude Opus 4.8 on 10 of 13 shared coding and agent tests. GLM-5.3 costs $1.40 input and $4.40 output per million tokens. Claude Opus 5.5 costs $4 / $20 and Claude Fable 5.1 costs $10 / $50. GLM also has open weights. Claude does not.
There is also a middle path: keep Claude Code as the agent and plug GLM in behind it through Z.ai’s GLM Coding Plan. This guide compares the two families on price, benchmarks, features and weights, shows that setup, and ends with a clear “pick GLM if… / pick Claude if…” verdict.

GLM vs Claude at a glance
GLM is made by Z.ai (formerly Zhipu AI). Its current models are GLM-5.3, the flagship, and GLM-5.3-Flash, the cheap multimodal model. Claude is made by Anthropic. Anthropic recommends Claude Opus 5.5 for most workloads and Claude Fable 5.1 for the most demanding reasoning and long-horizon agent work, with Sonnet 5 and Haiku 4.5 as faster, cheaper tiers.
| Compared | GLM (Z.ai) | Claude (Anthropic) |
|---|---|---|
| Top model | GLM-5.3 | Claude Fable 5.1 |
| Default pick | GLM-5.3 | Claude Opus 5.5 |
| Budget model | GLM-5.3-Flash ($0.15 / $0.50) | Claude Haiku 4.5 ($1 / $5) |
| Flagship price (in / out per 1M) | $1.40 / $4.40 | $4 / $20 (Opus 5.5), $10 / $50 (Fable 5.1) |
| Free API model | GLM-4.7-Flash, GLM-4.5-Flash, GLM-4.6V-Flash | Small free credits for new users |
| Context / max output | 1M / 128K | 1M / 128K (Haiku 4.5: 200K / 64K) |
| Image input | GLM-5.3-Flash (also video and files) | All current models |
| Weights | Open (MIT; GLM-5.3 License for GLM-5.3) | Closed |
| Coding subscription | GLM Coding Plan from $18/month | See claude.com/pricing |
| Works in Claude Code | Yes, via the Coding Plan endpoint | Yes, natively |

Price: GLM vs Claude API costs
Price is where the gap is widest. Both providers bill per million tokens and charge much less for input read from the cache. Claude also charges for writing to the cache: 1.25 times the input price for the 5-minute cache and 2 times for the 1-hour cache. Z.ai’s cache storage is free for a limited time.
| Model | Input | Cache hit | Output |
|---|---|---|---|
| GLM-5.3 | $1.40 | $0.26 | $4.40 |
| GLM-5.2 | $1.40 | $0.26 | $4.40 |
| GLM-5.3-Flash | $0.15 | $0.03 | $0.50 |
| GLM-4.7-Flash | Free | Free | Free |
| Claude Fable 5.1 | $10 | $0.25 | $50 |
| Claude Opus 5.5 | $4 | $0.20 | $20 |
| Claude Sonnet 5 | $2 | $0.20 | $10 |
| Claude Haiku 4.5 | $1 | $0.10 | $5 |
| Claude Fable 5 (legacy) | $10 | $1 | $50 |
| Claude Opus 4.8 (legacy) | $5 | $0.50 | $25 |
Worked example: 10M input and 1M output tokens
Take a month of steady API use: 10 million input tokens and 1 million output tokens. Scenario A has no caching. Scenario B has 8M of the input tokens read from cache. Claude’s cache-write charges are left out of B, so its real figure would be somewhat higher.
| Model | A: no cache | B: 80% cache hits |
|---|---|---|
| GLM-5.3 | $18.40 | $9.28 |
| GLM-5.3-Flash | $2.00 | $1.04 |
| Claude Haiku 4.5 | $15.00 | $7.80 |
| Claude Sonnet 5 | $30.00 | $15.60 |
| Claude Opus 5.5 | $60.00 | $29.60 |
| Claude Opus 4.8 | $75.00 | $39.00 |
| Claude Fable 5.1 | $150.00 | $72.00 |
The formula is simple. GLM-5.3 in scenario A is 10 × $1.40 + 1 × $4.40 = $18.40. Claude Opus 5.5 is 10 × $4 + 1 × $20 = $60, about 3.3 times as much. In scenario B, Opus 5.5 is 2 × $4 + 8 × $0.20 + 1 × $20 = $29.60, against GLM-5.3’s 2 × $1.40 + 8 × $0.26 + 1 × $4.40 = $9.28. Fable 5.1 costs about 8 times GLM-5.3 in both scenarios.
Output tokens decide most agent bills, and here Z.ai’s own measurements add a twist. On Z.ai Code Bench, Z.ai reports GLM-5.3 at High effort using about 50K output tokens per task, while Claude Opus 4.8 used about 120K. At list prices, that works out to roughly 50,000 × $4.40 / 1M = $0.22 of output per task for GLM-5.3, against 120,000 × $25 / 1M = $3.00 for Opus 4.8. Treat this as an illustration built on Z.ai’s figures, not a guarantee for your tasks.
Three price caveats apply. Anthropic’s Batch API halves Claude prices for jobs that can wait; Opus 5.5 drops to $2 / $10. OpenRouter lists a batch variant of GLM-5.3 (z-ai/glm-5.3:batch) at $0.45 / $2.00; confirm that on its OpenRouter page before relying on it. Tokenizers differ, too. Anthropic says the tokenizer used by Claude 4.7 and later produces about 30% more tokens for the same text than its previous tokenizer, so identical prompts do not cost identical token counts across providers. And Claude includes the full 1M context at standard prices, so long prompts cost the same per token as short ones on both sides. Every GLM price is on the GLM pricing page.
Coding benchmarks: GLM-5.3 vs Opus 4.8 and Fable 5
Every number in this section is Z.ai’s reported result from its GLM-5.3 and GLM-5.2 announcements. Z.ai ran the Claude models itself, and many of the agentic tests used Claude Code as the harness. The Fable 5 column is labelled “with fallback” in Z.ai’s table. Z.ai’s tables do not include Claude Fable 5.1 or Claude Opus 5.5, Anthropic’s current recommended models.
| Benchmark | GLM-5.3 | GLM-5.2 | Opus 4.8 | Fable 5 |
|---|---|---|---|---|
| Terminal Bench 2.1 | 88.2 | 81.0 | 85.0 | 88.0 |
| Terminal Bench 3.0 | 28.3 | 4.6 | 21.1 | 33.7 |
| DeepSWE v1.1 | 66.9 | 46.2 | 58.0 | 69.7 |
| NL2Repo | 58.0 | 48.9 | 69.7 | — |
| ProgramBench (Almost Solved) | 19.0 | 9.5 | 15.5 | 33.0 |
| FrontierSWE | 78.1 | 67.5 | 66.5 | 88.2 |
| SWE-Marathon v1.1 | 42.5 | 19.4 | 48.8 | 33.1 |
| PostTrainBench | 39.8 | 31.7 | 32.9 | 41.8 |
| Toolathlon Verified | 73.0 | 59.9 | 76.2 | 74.7 |
| AutomationBench v1.0.6 | 48.2 | 26.2 | 41.0 | 46.2 |
| Agents’ Last Exam (ALE-CLI) | 28.5 | 23.8 | 25.7 | 23.8 |
| HLE with tools | 62.5 | 54.7 | 57.9 | 63.9 |
| GDPval-AA v2 | 1769 | 1508 | 1588 | 1743 |
Here is what the table says. GLM-5.3 beats Opus 4.8 on 10 of these 13 rows. Opus 4.8 keeps NL2Repo (69.7 vs 58.0), SWE-Marathon (48.8 vs 42.5) and Toolathlon (76.2 vs 73.0). Against Fable 5, GLM-5.3 wins 5 of the 12 rows where Fable has a score: Terminal Bench 2.1, SWE-Marathon, AutomationBench, Agents’ Last Exam and GDPval-AA. Fable 5 leads clearly on the hardest long-horizon engineering tests, including ProgramBench (33.0 vs 19.0), FrontierSWE (88.2 vs 78.1) and Terminal Bench 3.0 (33.7 vs 28.3).
Z.ai’s in-house Z.ai Code Bench tells the same story. GLM-5.3 reaches 31.4% at High effort, ahead of Claude Opus 4.8 at 29.5%, and 34.5% at Max effort. Claude Fable 5 stays ahead at 39.5% at Max effort. The cheap GLM-5.3-Flash scores 29.0% at max effort, against Opus 4.8’s 29.5%.
On security benchmarks, Z.ai’s announcement text compares GLM-5.3 with Claude Mythos 5, Anthropic’s limited-availability model. GLM-5.3 posts 84.5% on CyberGym against Mythos 5’s 83.8%, the best result Z.ai reports on that benchmark. Mythos 5 stays well ahead further up the exploitation chain: 78.0% vs 54.4% on ExploitBench, and 181 vs 105 tasks within two hours on ExploitGym.
GLM-5.2 vs Opus 4.8
For “GLM 5.2 vs Opus” searches, here is Z.ai’s June 2026 table. GLM-5.2 was clearly behind Opus 4.8 on coding at launch. Z.ai’s own framing places it between Opus 4.7 and Opus 4.8 at similar token use.
| Benchmark | GLM-5.2 | Claude Opus 4.8 |
|---|---|---|
| HLE | 40.5 | 49.8 (full set) |
| AIME 2026 | 99.2 | 95.7 |
| IMOAnswerBench | 91.0 | 83.5 |
| GPQA-Diamond | 91.2 | 93.6 |
| SWE-bench Pro | 62.1 | 69.2 |
| NL2Repo | 48.9 | 69.7 |
| DeepSWE | 46.2 | 58.0 |
| ProgramBench | 63.7 | 71.9 |
| Terminal Bench 2.1 (Terminus-2) | 81.0 | 85.0 |
| Terminal Bench 2.1 (in Claude Code) | 82.7 | 78.9 |
| FrontierSWE | 74.4 | 75.1 |
| PostTrainBench | 34.3 | 37.2 |
| SWE-Marathon | 13.0 | 26.0 |
| MCP-Atlas (public set) | 76.8 | 77.8 |
| Tool-Decathlon | 48.2 | 59.9 |
GLM-5.2 wins AIME 2026 and IMOAnswerBench. It also wins Terminal Bench 2.1 when both models run inside Claude Code, 82.7 against 78.9. Opus 4.8 wins nearly everything else, most sharply NL2Repo, SWE-Marathon and Tool-Decathlon. If you are choosing today, GLM-5.3 is the better GLM to put against Claude. It costs the same as GLM-5.2 on Z.ai’s API and closes most of those gaps.
Z.ai has benchmarked each GLM release against the Claude generation of its day. It described GLM-5.1 as overall aligned with Claude Opus 4.6, positioned GLM-5 against Opus 4.5, and called GLM-4.7’s coding aligned with Claude Sonnet 4.5.
Specs and features compared
| Feature | GLM-5.3 | GLM-5.3-Flash | Claude Opus 5.5 | Claude Fable 5.1 |
|---|---|---|---|---|
| Context | 1M | 1M | 1M | 1M |
| Max output | 128K | 128K | 128K | 128K |
| Input | Text | Video, image, text, file | Text, image | Text, image |
| Thinking | Always on; effort low, high, max | Always on; effort low, high, max | Adaptive, always on; default effort medium | Adaptive, always on; default effort high |
| Parameters | 744B total, 40B active | 320B total, 18B active | — | — |
| Weights | Open (GLM-5.3 License) | Open (MIT) | Closed | Closed |
| API format | OpenAI-compatible; Anthropic-compatible on the Coding Plan | Same | Anthropic Messages API | Anthropic Messages API |
| Where to run | Z.ai API, OpenRouter, self-host | Z.ai API, OpenRouter, self-host | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry | Same |
| Reliable knowledge cutoff | — | — | Jun 2026 | Jun 2026 |
| Batch discount | — | — | 50% | 50% |
Three differences matter most in practice:
- Images. Every current Claude model reads images. On the GLM side only GLM-5.3-Flash does, and it adds video and file input. GLM-5.3 is text-only, so screenshot-heavy work belongs on GLM-5.3-Flash or Claude.
- Thinking control. Both families now keep reasoning on at the top end and let you steer how much through an effort setting. GLM-5.3 accepts
reasoning_effortlow,highormaxand rejects attempts to disable thinking. The GLM thinking mode guide covers every model. - Control over the model. GLM weights can be downloaded, fine-tuned and served on your own hardware. Claude runs only on Anthropic’s API and the cloud platforms that resell it.
The middle path: GLM inside Claude Code
If you like Claude Code as a tool but not the Claude bill, you can keep Claude Code and swap the model. Z.ai’s GLM Coding Plan exposes an Anthropic-compatible endpoint, https://api.z.ai/api/anthropic. Claude Code talks to it exactly as it talks to Anthropic, and Z.ai maps Claude Code’s model slots to GLM models. Z.ai also ran many of its GLM-5.3 agentic benchmarks with Claude Code as the harness, so this is a setup Z.ai actively benchmarks.
Once you have a Coding Plan subscription and an API key from z.ai/manage-apikey/apikey-list, add these keys to ~/.claude/settings.json. Add them to the existing file rather than replacing it:
{
"env": {
"ANTHROPIC_AUTH_TOKEN": "your-api-key",
"ANTHROPIC_BASE_URL": "https://api.z.ai/api/anthropic",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "glm-5.3-flash[1m]",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "glm-5.3[1m]",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "glm-5.3[1m]",
"CLAUDE_CODE_AUTO_COMPACT_WINDOW": "1000000",
"API_TIMEOUT_MS": "3000000"
}
}
The [1m] suffix enables GLM’s 1M-token context in Claude Code. Open a new terminal, run claude, then type /status to confirm the model shows as glm-5.3. Use /effort to change thinking intensity; the default is max. Z.ai also ships a helper, npx @z_ai/coding-helper, that writes the configuration for you. The full walkthrough, with fixes for errors such as 1113, is in how to use GLM in Claude Code.
What the plan costs: Lite is $18 a month, Pro $80 and Max $168, with 20% off quarterly and 30% off yearly billing. Usage is measured in credits that refresh every 5 hours and every week. Off-peak hours cost half the credits (peak is 14:00 to 18:00 UTC+8 on weekdays, and from September 25 to October 7, 2026 every hour is billed at the off-peak rate), and GLM-5.3-Flash gets three times the quota of GLM-5.3. Z.ai estimates the Lite tier at 48 to 97 million GLM-5.3 tokens a week at a 95% cache-hit rate. The GLM Coding Plan guide covers tiers, credits and break-even maths.
To switch back to Claude, remove the env entries and start a new session. A common split is to keep Claude for the hardest tasks, such as large repository builds where Opus and Fable lead, and run GLM for the long tail of everyday edits.
Our test results
Here is how the GLM models handle GLM Chat’s five standard tasks. Each model runs three times through the same five fixed prompts, with the public chat’s settings: temperature 0.7, at most 2,048 output tokens and the lightest reasoning setting available. An editor scores each answer 0, 1 or 2, for a maximum of 10.
- Python CSV deduplication: stream a large CSV row by row, keep the header and the first occurrence, and support key columns. This shows whether the model respects a memory constraint instead of loading everything.
- JavaScript debounce:
debounce(fn, wait, { leading, trailing })withcancel()and tests that pass undernode --test. This probes timing edge cases, especially leading and trailing together. - PHP refactor: turn a messy 60-line
proc()function into clean PHP 8 without changing behaviour, including the VIP discount, with escaped output and parameterised SQL. This is the closest proxy for real maintenance work. - Beginner explanation: transformer attention in about 150 words, accurate on queries, keys, values and weights, and between 120 and 180 words long.
- JSON extraction: return valid JSON with exact keys, dimensions converted to centimetres and
nullwhere the text is silent. This catches models that invent data.
A 2 means correct and complete, a 1 means usable after a small fix, and a 0 means wrong, broken or invented. The table shows each model’s scores and the date of the run.
Our test results
| Model | Task 1 | Task 2 | Task 3 | Task 4 | Task 5 | Total /10 | Avg latency | Why (across runs) |
|---|---|---|---|---|---|---|---|---|
| GLM-5.3 | 2 | 2 | 1 | 1 | 1 | 7 | 7.6 s | Task 3: output not escaped (2/3 runs) · Task 4: no queries/keys/values (2/3 runs) · +1 more |
| GLM-5.3-Flash | 2 | 0 | 1 | 2 | 2 | 7 | 12.9 s | Task 2: fails our debounce behaviour checks (2/3 runs) · Task 3: output not escaped (3/3 runs) |
| GLM-5.2 | 2 | 1 | 1 | 1 | 2 | 7 | 6.6 s | Task 2: its own tests fail (2/3 runs) · Task 3: output not escaped (3/3 runs) · +1 more |
| GLM-5.1 | 2 | 2 | 1 | 1 | 2 | 8 | 14.3 s | Task 3: output not escaped (3/3 runs) · Task 4: no queries/keys/values (2/3 runs) |
| GLM-4.7 | 2 | 1 | 1 | 2 | 2 | 8 | 14.4 s | Task 2: fails our debounce behaviour checks (1/3 runs) · Task 3: output not escaped (3/3 runs) |
| GLM-4.7-Flash | 2 | 0 | 0 | 1 | 1 | 4 | 38.5 s | Task 2: fails our debounce behaviour checks (2/3 runs) · Task 3: changes the function’s behaviour (3/3 runs) · +2 more |
To place Claude on the same scale, send the same five prompts, three times each, to a Claude model with temperature 0.7 and a 2,048-token output cap, then score them with the rubric in our editorial policy. Look hardest at task 3: a behaviour-preserving refactor is where coding models most often slip, whatever their benchmark scores. You can run the GLM side yourself right now in the free GLM-5.3 chat.
Open weights, privacy and control
For some teams this section decides everything. GLM weights are published on Hugging Face under zai-org. GLM-5.2, GLM-5.1, GLM-5, GLM-5.3-Flash and GLM-4.7-Flash use the MIT license. GLM-5.3 uses the GLM-5.3 License, which gives MIT-style rights with one condition: a “Model as a Service” business with more than $10 billion in revenue over any 12 months must pass Z.ai’s security review before commercial use. Almost every company falls below that bar. The details are in Is GLM open source?
Open weights mean you can run GLM on your own servers, keep every prompt inside your network, fine-tune on private data and pin a version for as long as you like. The flagships are big, though. GLM-5.3 has 744B total parameters, so the FP8 weights alone need roughly 744 GB of GPU memory (one byte per parameter) before any KV cache. Plan on multi-GPU servers with vLLM or SGLang. GLM-4.7-Flash, a 30B-A3B model, is the practical choice for a single workstation.
Claude is closed-weights. You use it through Anthropic’s Claude API, the claude.ai apps, or cloud platforms such as Amazon Bedrock, Google Cloud and Microsoft Foundry. Anthropic publishes retirement commitments for each model: Opus 5.5 will not be retired before September 22, 2027, and Fable 5.1 not before September 1, 2027.
GLM vs Claude FAQ
Is GLM as good as Claude for coding?
Close, but not at the very top. In Z.ai’s reported results, GLM-5.3 beats Claude Opus 4.8 on 10 of 13 coding and agent benchmarks and on Z.ai Code Bench. Claude Fable 5 still leads on the hardest long-horizon tests, such as ProgramBench, FrontierSWE and Terminal Bench 3.0.
Is GLM-5.2 better than Claude Opus?
Not overall. In Z.ai’s June 2026 table, Opus 4.8 leads GLM-5.2 on most coding rows. GLM-5.2 wins AIME 2026, IMOAnswerBench and Terminal Bench 2.1 when run in Claude Code. Z.ai places GLM-5.2 between Opus 4.7 and Opus 4.8 at similar token use.
How much cheaper is GLM than Claude?
GLM-5.3 ($1.40 / $4.40 per 1M tokens) costs about a third of Claude Opus 5.5 ($4 / $20) for input-heavy use and about a fifth for output, and roughly an eighth of Claude Fable 5.1 overall. GLM-5.3-Flash ($0.15 / $0.50) is far cheaper than Claude Haiku 4.5 ($1 / $5).
Can I use GLM in Claude Code?
Yes. With a GLM Coding Plan, set ANTHROPIC_BASE_URL to https://api.z.ai/api/anthropic, add your Z.ai key as ANTHROPIC_AUTH_TOKEN, and map the model slots to glm-5.3 and glm-5.3-flash. Plan usage comes from your credits, not your account balance.
Is Claude open source like GLM?
No. Claude’s weights are not published. GLM’s weights are, under MIT for most models and the GLM-5.3 License for GLM-5.3.
Does GLM read images like Claude?
GLM-5.3-Flash does, and it also accepts video and files. The flagship GLM-5.3 is text-only. Every current Claude model accepts images.
Which is better for reasoning and knowledge work?
The results are mixed. In Z.ai’s GLM-5.3 table, GLM-5.3 leads on GDPval-AA, a knowledge-work benchmark (1769 against Opus 4.8’s 1588 and Fable 5’s 1743). It sits between them on HLE with tools (62.5 against 57.9 and 63.9). In the earlier GLM-5.2 table, Opus 4.8 led on GPQA-Diamond and HLE, while GLM-5.2 led on AIME 2026.
Is there a free way to try GLM and Claude?
GLM-4.7-Flash is free on the Z.ai API, chat.z.ai is free, and the GLM chat on this site needs no sign-up. Anthropic gives new API users a small amount of free credits to test the Claude API.
Pick GLM if… / Pick Claude if…
Pick GLM if…
- Cost matters. GLM-5.3 is about a third to a fifth of Opus 5.5’s price per token, and GLM-5.3-Flash costs pennies.
- You want Opus-4.8-class agentic coding. In Z.ai’s table, GLM-5.3 beats it on 10 of 13 rows.
- You live in Claude Code and want a flat monthly bill. The GLM Coding Plan starts at $18.
- You need open weights to self-host, fine-tune or keep data on your own infrastructure.
- You need video or file input at a low price. GLM-5.3-Flash handles both.
- You want a free API model for prototypes (GLM-4.7-Flash).
Pick Claude if…
- You need the strongest results on the hardest long-horizon engineering. Fable 5 leads GLM-5.3 on ProgramBench, FrontierSWE and Terminal Bench 3.0 in Z.ai’s own table, and Fable 5.1 is Anthropic’s newer top model.
- Repository generation from a spec is core to your work. Opus 4.8 leads NL2Repo 69.7 to 58.0.
- You want image input on every model tier, from Haiku to Fable.
- You buy through Amazon Bedrock, Google Cloud or Microsoft Foundry, or you need published model retirement dates.
- You run large offline jobs that fit Anthropic’s 50% Batch API discount.
Verdict
For most developers, GLM is the better value and Claude is the higher ceiling. Z.ai vs Claude is no longer a question of whether GLM can code. In Z.ai’s reported results, GLM-5.3 matches or beats Claude Opus 4.8 on most agentic benchmarks at roughly a third of Opus 5.5’s input price. Claude Fable 5 and its successor remain the choice when a task is long, hard and worth paying for.
The smartest setup is often both. Run GLM-5.3 inside Claude Code on the Coding Plan for daily work, and keep a Claude key for the few jobs where the top model earns its price. Try GLM first in the free GLM chat, see how GLM fares against another open-weights rival in GLM vs DeepSeek, or start building with the GLM API quickstart.