GLM Pricing 2026: API Prices, Coding Plan & Free Options

Every Z.ai API price for GLM models, what the Coding Plan really buys, and the free ways to use GLM, with the arithmetic shown.

GLM pricing on the Z.ai API runs from free (GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash cost nothing per token) up to $1.40 per 1M input tokens and $4.40 per 1M output tokens for the flagship GLM-5.3, GLM-5.2 and GLM-5.1. The cheapest current-generation model, GLM-5.3-Flash, costs $0.15 in and $0.50 out. Cached input is billed at roughly one fifth of the normal input price, and cache storage is free for a limited time.

If you code with agents all day, the flat-rate GLM Coding Plan costs $18, $80 or $168 per month (less with quarterly or yearly billing) and usually beats pay-as-you-go. If you only want to talk to the models, you do not need to pay anything: the free GLM chat on this site runs every current model without sign-up. This page lists every price Z.ai (formerly Zhipu AI) publishes, then shows the arithmetic so you can estimate your own bill.

GLM pricing overview: Z.ai API prices per 1M tokens, Coding Plan tiers and free GLM models
GLM pricing in one view: pay-as-you-go API, Coding Plan subscriptions and free options.

GLM pricing at a glance

Z.ai sells GLM access in three ways. Pick the one that matches how you work, because the same workload can cost ten times more on the wrong route.

  • Pay-as-you-go API. You top up a balance and pay per token (or per image, video or tool call). Best for apps, scripts, batch jobs and anything that runs outside a coding tool. Prices are per 1M tokens and split into input, cached input and output.
  • GLM Coding Plan. A monthly subscription (Lite, Pro, Max, plus Team seats) that gives you a credit allowance every 5 hours and every week. It works only inside supported coding tools such as Claude Code, Cline, OpenCode, Cursor, ZCode and AutoClaw. Plan quota never draws from your API balance, and it cannot be used for general API calls.
  • Free routes. Three API models are free per token, the chat.z.ai web app is free, OpenRouter lists a free rate-limited GLM-5.2 variant, and this site’s chat gives you daily messages on the paid models plus unlimited GLM-4.7-Flash.
GLM pricing summary card with free models, flagship prices and Coding Plan tiers
The numbers most people need, on one card.

Z.ai API pricing for GLM text models

These are the pay-as-you-go prices from Z.ai’s official pricing page (docs.z.ai/guides/overview/pricing). All prices are in dollars per 1M tokens. “Cached input” is what you pay for input tokens that hit the context cache; the storage of cached input is listed as limited-time free for every paid model below except GLM-4-32B-0414-128K, which has no cached price.

ModelInputCached inputOutputModel ID
GLM-5.3$1.40$0.26$4.40glm-5.3
GLM-5.2$1.40$0.26$4.40glm-5.2
GLM-5.1$1.40$0.26$4.40glm-5.1
GLM-5$1.00$0.20$3.20glm-5
GLM-5.3-FlashX$0.37$0.075$1.25glm-5.3-flashx
GLM-5.3-Flash$0.15$0.03$0.50glm-5.3-flash
GLM-4.7$0.60$0.11$2.20glm-4.7
GLM-4.7-FlashX$0.07$0.01$0.40glm-4.7-flashx
GLM-4.7-FlashFreeFreeFreeglm-4.7-flash
GLM-4.6$0.60$0.11$2.20glm-4.6
GLM-4.5$0.60$0.11$2.20glm-4.5
GLM-4.5-X$2.20$0.45$8.90glm-4.5-x
GLM-4.5-Air$0.20$0.03$1.10glm-4.5-air
GLM-4.5-AirX$1.10$0.22$4.50glm-4.5-airx
GLM-4.5-FlashFreeFreeFreeglm-4.5-flash
GLM-4-32B-0414-128K$0.10–$0.10glm-4-32b-0414-128k
Z.ai pay-as-you-go prices per 1M tokens for text models.

Three things stand out. First, the whole flagship line (GLM-5.3, GLM-5.2 and GLM-5.1) shares one price, so there is no cost reason to stay on an older flagship. Second, the “X” variants (GLM-5.3-FlashX, GLM-4.7-FlashX, GLM-4.5-X, GLM-4.5-AirX) are faster, pricier versions of their base models; GLM-5.3-FlashX, for example, runs at around 200 tokens per second according to Z.ai. Third, output tokens cost roughly three to six times more than input on nearly every paid model, so long answers and heavy reasoning drive most bills.

Bar chart comparing GLM API output price per 1M tokens across models
Output price per 1M tokens by model. Output is where most of a GLM bill comes from.

GLM-5-Turbo pricing

GLM-5-Turbo, the GLM-5 variant tuned for OpenClaw agent workflows, does not appear on Z.ai’s public pricing page. OpenRouter lists it as z-ai/glm-5-turbo at $1.20 input and $4.00 output per 1M tokens. On Z.ai itself, the Coding Plan subscription page names GLM-5-Turbo for agents, so the plan is the main way most people reach it. See GLM with OpenClaw for when Turbo is the right pick.

Vision, image, video, audio and tool pricing

Z.ai’s multimodal and utility models use the same per-token structure where they read text or images, and per-item prices where they generate media. GLM-5.3-Flash is itself natively multimodal (video, image, text and file input), so for most image-understanding tasks it is the model to price first.

Vision modelInputCached inputOutput
GLM-4.6V$0.30$0.05$0.90
GLM-4.6V-FlashX$0.04$0.004$0.40
GLM-4.6V-FlashFreeFreeFree
GLM-4.5V$0.60$0.11$1.80
GLM-OCR$0.03–$0.03
Vision model prices per 1M tokens on the Z.ai API.
ProductTypePrice
GLM-ImageImage generation$0.015 per image
CogView-4Image generation$0.01 per image
CogVideoX-3Video generation$0.20 per video
GLM-ASR-2512Speech recognition$0.03 per 1M tokens (about $0.0024 per minute)
Web Search (built-in tool)Tool call$0.01 per use
GLM Slide/Poster Agent (beta)Agent$0.70 per 1M tokens
General-Purpose TranslationAgent$3 per 1M tokens
Special Effects Video TemplatesAgent$0.20 per video
Per-item and per-use prices from Z.ai’s pricing page.

Two details matter here. GLM-Image returns a URL to the generated picture, and that link expires after 30 days, so download anything you want to keep. And the built-in Web Search tool adds $0.01 per search on top of the tokens the model reads, which adds up fast in agent loops that search on every turn. The GLM web search API guide covers how to limit that.

GLM-5.2 pricing and cost

GLM-5.2 costs $1.40 per 1M input tokens, $0.26 per 1M cached input tokens and $4.40 per 1M output tokens on the Z.ai API, exactly the same as GLM-5.3 and GLM-5.1. Its 1M-token context window does not carry a long-context surcharge on the published price list: a request is billed by the tokens it uses.

What does that mean in practice? A single request that fills 800,000 tokens of context costs 0.8 × $1.40 = $1.12 in input alone if nothing is cached. The same request with 90% of those tokens served from cache costs (0.08 × $1.40) + (0.72 × $0.26) = $0.112 + $0.187 = about $0.30. That gap is why long-context agents only make economic sense with caching.

GLM-5.2 is also cheaper on some third-party routers. OpenRouter lists z-ai/glm-5.2 at roughly $0.65 input and $2.04 output per 1M tokens (it routes across several providers, so the figure moves), plus a free variant, z-ai/glm-5.2:free, limited to a 32,768-token context and rate-limited. On the GLM Coding Plan, GLM-5.2 requests are now routed automatically to GLM-5.3. Full specs are on the GLM-5.2 model page.

Cached input pricing explained

Context caching is automatic on the Z.ai API. When a request repeats content the platform has recently processed (a long system prompt, a document you keep asking about, or earlier turns of a conversation), the repeated tokens are served from cache and billed at the cached-input rate instead of the full input rate. You do not set a flag or create a cache object.

  • How much you save: Z.ai’s FAQ says cache hits are charged at approximately one fifth of the original price. The table bears that out: GLM-5.3 input is $1.40 and cached input is $0.26 (about 19%).
  • Storage: cached-input storage is listed as “Limited-time Free” on the pricing page.
  • Where to see it: each response reports cached tokens in usage.prompt_tokens_details.cached_tokens. Subtract that from prompt_tokens to get the tokens billed at the full rate.
  • What breaks it: caching keys on identical or highly similar content. Z.ai notes that minor formatting differences can reduce cache hits, and that cached content expires after a time limit.

A response with cache usage looks like this:

{
  "usage": {
    "prompt_tokens": 1200,
    "completion_tokens": 300,
    "total_tokens": 1500,
    "prompt_tokens_details": {
      "cached_tokens": 800
    }
  }
}

On GLM-5.3 that request costs (400 × $1.40 + 800 × $0.26 + 300 × $4.40) / 1,000,000 = (560 + 208 + 1,320) / 1,000,000 ≈ $0.0021. Without the cache hit it would be (1,200 × $1.40 + 300 × $4.40) / 1,000,000 = $0.0030. The step-by-step guide to structuring prompts for cache hits is in GLM context caching.

Worked cost examples

The formula for any token-priced GLM model is simple:

cost = (uncached_input × input_price
      + cached_input × cached_price
      + output × output_price) / 1,000,000

Example 1: 1,000 requests of 2K in and 500 out

A typical chatbot or classification job: 1,000 requests, each with 2,000 input tokens and 500 output tokens, no caching. That is 2,000,000 input tokens (2M) and 500,000 output tokens (0.5M) in total.

ModelInput costOutput costTotalPer request
GLM-5.2 or GLM-5.32 × $1.40 = $2.800.5 × $4.40 = $2.20$5.00$0.005
GLM-52 × $1.00 = $2.000.5 × $3.20 = $1.60$3.60$0.0036
GLM-4.72 × $0.60 = $1.200.5 × $2.20 = $1.10$2.30$0.0023
GLM-5.3-FlashX2 × $0.37 = $0.740.5 × $1.25 = $0.625$1.365$0.0014
GLM-4.5-Air2 × $0.20 = $0.400.5 × $1.10 = $0.55$0.95$0.00095
GLM-5.3-Flash2 × $0.15 = $0.300.5 × $0.50 = $0.25$0.55$0.00055
GLM-4.7-Flash$0$0$0$0
1,000 requests × (2,000 input + 500 output tokens), Z.ai list prices, no cache hits.

The spread is roughly 9x between the flagship and GLM-5.3-Flash, and infinite against GLM-4.7-Flash. Run your prompts through the cheaper model first; move up only for the requests where quality clearly improves. You can compare answers side by side for free in GLM-5.3-Flash in the chat and GLM-5.2 in the chat before you commit.

Example 2: a coding agent with a large, repeated context

An agent makes 200 calls to GLM-5.3. Each call sends 100,000 input tokens (repo files, tool results, history), of which 90,000 repeat from the previous call and hit the cache, and returns 2,000 output tokens.

  • Without caching: input 200 × 100,000 = 20M tokens × $1.40 = $28.00. Output 200 × 2,000 = 0.4M × $4.40 = $1.76. Total $29.76.
  • With 90% cache hits: cached 18M × $0.26 = $4.68. Uncached 2M × $1.40 = $2.80. Output $1.76. Total $9.24, a 69% saving.
  • Same run on GLM-5.3-Flash with caching: 18M × $0.03 = $0.54, plus 2M × $0.15 = $0.30, plus 0.4M × $0.50 = $0.20. Total $1.04.

Example 3: reasoning tokens on a hard question

GLM-5.3 and GLM-5.3-Flash always think before they answer, and GLM-5.2, GLM-5.1, GLM-5 and GLM-4.7 think by default. The reasoning text is generated output, so it is counted in completion_tokens and billed at the output price. If GLM-5.3 writes 6,000 reasoning tokens and a 1,000-token answer, you pay for 7,000 output tokens: 7,000 × $4.40 / 1,000,000 = $0.031. Setting reasoning_effort to "low" on GLM-5.3 (the default is "max") is the single biggest lever on that number. The GLM thinking mode guide lists the settings per model.

A marketing workflow generates 500 GLM-Image posters: 500 × $0.015 = $7.50. A research agent that runs 3 web searches per question over 1,000 questions pays 3,000 × $0.01 = $30 for search alone, before any tokens. Tool calls are often the hidden cost in agent budgets.

GLM Coding Plan pricing

The GLM Coding Plan is Z.ai’s subscription for AI coding tools. Since July 30, 2026, individual plans are credits-based: every plan has a 5-hour credit limit and a weekly credit limit, and each request burns credits according to the tokens it uses. All plans include GLM-5.3 and GLM-5.3-Flash; requests for GLM-5.2 or GLM-5.1 are routed to GLM-5.3, and requests for GLM-4.7 are routed to GLM-5.3-Flash.

PlanMonthly billingQuarterly (20% off)Yearly (30% off)5-hour creditsWeekly credits
Lite$18/month$14.40/month$12.60/month2,00010,000
Pro$80/month$64/month$56/month12,00060,000
Max$168/month$134.40/month$117.60/month28,000140,000
Individual GLM Coding Plan tiers. Quarterly prices are the monthly price minus 20%.

Pro gives 6x the Lite usage and Max gives 14x. Beyond credits, the subscribe page lists these tier benefits: Lite is “built for lightweight iteration on small repo” with 20+ agent tools and default data privacy; Pro adds a curated selection of MCP tools, faster generation and priority access to the latest flagship models; Max adds first access to new models and dedicated resources during peak times. Every plan includes the Vision Understanding, Web Search, Web Reader and Zread MCP servers.

Team plan prices

SeatMonthlyAnnual (10% off)5-hour creditsWeekly credits
Standard Seat$88/seatfrom $79.20/seat15,00066,000
Premium Seat$188/seatfrom $169.20/seat35,000155,000
GLM Coding Team Plan seats, per seat per month.

Team seats add centralized seat and permission management, usage dashboards, centralized billing and invoicing, and optional on-demand overage (billed, as a limited-time offer, at 10% off the API list price). Premium seats also get early access to new models and priority during peak hours.

How Coding Plan credits are calculated

credits = (input × input_multiplier
         + cached_input × cached_multiplier
         + output × output_multiplier) / 10,000

MCP tool call credits = number of calls × 1.2
ProductInput multiplierCached input multiplierOutput multiplier
GLM-5.36.91.724
GLM-5.3-Flash (incl. vision MCP)2.30.568
Web Search, Web Reader, Zread MCP––1.2 per call
Credit multipliers from Z.ai’s Coding Plan documentation.

Peak and off-peak. Peak hours are Monday to Friday, 14:00–18:00 UTC+8. Outside those hours, and all day on weekends, model usage is charged at 50% of the standard credit rate.

GLM Coding Plan credit formula, model multipliers and peak hours
Coding Plan credit math in six lines.

Credits math: what a Lite plan buys

Take one agent call on GLM-5.3 with 100,000 input tokens, 95,000 of them cached (so 5,000 uncached), and 2,000 output tokens:

  1. Uncached input: 5,000 × 6.9 = 34,500
  2. Cached input: 95,000 × 1.7 = 161,500
  3. Output: 2,000 × 24 = 48,000
  4. Sum 244,000 / 10,000 = 24.4 credits at peak, or 12.2 credits off-peak.

Lite’s 10,000 weekly credits cover about 409 of those calls if they all land in peak hours (10,000 / 24.4), or about 819 off-peak. The 5-hour cap of 2,000 credits allows about 81 peak calls per window. The same call on GLM-5.3-Flash costs (5,000 × 2.3 + 95,000 × 0.56 + 2,000 × 8) / 10,000 = (11,500 + 53,200 + 16,000) / 10,000 = 8.07 credits, so Lite covers roughly 1,239 of them per week at peak.

Compare the API. That GLM-5.3 call costs (5,000 × $1.40 + 95,000 × $0.26 + 2,000 × $4.40) / 1,000,000 = $0.007 + $0.0247 + $0.0088 = $0.0405 pay-as-you-go. 409 calls would be about $16.56 on the API, in a single week, while Lite costs $18 for the whole month. If you actually use the quota, the plan wins by a wide margin; if you code a few hours a month, the API can be cheaper. Z.ai claims plan users can save up to 92% against pay-as-you-go GLM-5.3 by making full use of off-peak discounts.

Z.ai’s estimated weekly token allowance

Cache hit rateModelLite (M tokens/week)Pro (M tokens/week)Max (M tokens/week)
95%GLM-5.348–97290–580676–1,352
95%GLM-5.3-Flash146–292877–1,7552,047–4,095
98%GLM-5.352–104313–627731–1,463
98%GLM-5.3-Flash158–317950–1,9002,217–4,433
Low end = all usage at peak rate; high end = all usage off-peak. Source: Z.ai Coding Plan overview.

Other plan rules worth knowing before you pay: subscriptions are non-refundable, so cancel auto-renewal at least 24 hours before the renewal date; upgrading to a higher tier applies immediately with prorated credit; the plan only works in supported tools; and an “1113 Insufficient Balance” error while subscribed usually means the tool is pointed at the wrong base URL (the plan uses https://api.z.ai/api/coding/paas/v4 for OpenAI-compatible tools and https://api.z.ai/api/anthropic for Claude Code). Legacy prompt-based plans sold before July 30, 2026 keep their terms until the current billing cycle ends. Setup steps are in GLM in Claude Code.

Free ways to use GLM

You can get a lot done with GLM without spending anything. Here are the real free routes and their limits.

RouteModelsLimits
GLM Chat (this site)GLM-5.3-Flash, GLM-5.3, GLM-5.2, GLM-5.1, GLM-4.7, GLM-4.6, GLM-4.5, GLM-4.7-Flash, GLM-Image40 messages/day on paid models; GLM-4.7-Flash unlimited (one request every 3 s); 3 images/day; no sign-up
chat.z.aiZ.ai’s current GLM modelsFree web chat app from Z.ai
Z.ai API, free modelsGLM-4.7-Flash, GLM-4.5-Flash, GLM-4.6V-FlashFree input, cached input and output; your account’s rate limits apply
OpenRouter free variantz-ai/glm-5.2:free$0, 32,768-token context, rate-limited
Free GLM access routes and their limits.

Be precise about “free” when you read other sites. GLM-5.3-Flash is free to use in this site’s chat, where it is the default model, but on the Z.ai API it costs $0.15 in and $0.50 out. GLM-4.7-Flash is the one current model that is genuinely free per token on the API, and it is a capable 30B-A3B model with a 200K context. The GLM-4.7-Flash page shows how to call it at zero cost.

To call the free model, use the standard endpoint with its model ID:

curl https://api.z.ai/api/paas/v4/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $ZAI_API_KEY" \
  -d '{
    "model": "glm-4.7-flash",
    "messages": [{"role": "user", "content": "Summarize the GLM pricing tiers in 3 bullets."}]
  }'

OpenRouter prices for GLM models

OpenRouter is a third-party router that resells GLM access alongside other providers. Its listed prices differ from Z.ai’s, sometimes lower, because several hosts serve the open weights. Prices change often: treat the table as OpenRouter’s listed price and confirm on the model’s OpenRouter page before you budget.

OpenRouter slugInput / 1MOutput / 1MContext
z-ai/glm-5.3$1.40$4.401,310,720
z-ai/glm-5.3:batch$0.45$2.00–
z-ai/glm-5.3-flash$0.15$0.50–
z-ai/glm-5.3-flashx$0.37$1.25–
z-ai/glm-5.2about $0.65about $2.041,048,576
z-ai/glm-5.2:free$0$032,768
z-ai/glm-5.1$0.97$3.04204,800
z-ai/glm-5-turbo$1.20$4.00–
z-ai/glm-5$0.60$1.92–
z-ai/glm-4.7$0.40$1.75–
z-ai/glm-4.7-flash$0.06$0.40–
z-ai/glm-4.6$0.43$1.75–
z-ai/glm-4.5$0.60$2.20–
z-ai/glm-4.5-air$0.13$0.85–
OpenRouter’s listed prices per 1M tokens. A dash means no context figure in our source.

Note one reversal: GLM-4.7-Flash is free on Z.ai’s own API but costs money on OpenRouter. For older open models such as GLM-5.1, GLM-5 and GLM-4.7, OpenRouter’s listed rates run below Z.ai’s; for GLM-5.3 and GLM-5.3-Flash they match. OpenRouter also sells a batch variant of GLM-5.3 at $0.45 in and $2.00 out, which suits offline jobs that can wait.

How to cut your GLM API bill

  1. Start on the cheapest model that passes. Test GLM-4.7-Flash (free) and GLM-5.3-Flash ($0.15/$0.50) before reaching for the $1.40/$4.40 flagship. The GLM models comparison shows what each tier is good at.
  2. Turn reasoning down. Reasoning tokens are output tokens. On GLM-5.3 and GLM-5.3-Flash pass "reasoning_effort": "low" (thinking cannot be disabled there). On GLM-5.2, "none" skips thinking. On GLM-5.1, GLM-5 and GLM-4.7 send "thinking": {"type": "disabled"} for simple tasks.
  3. Cap output. Set max_tokens to what you really need. Output costs several times more than input on almost every paid model.
  4. Design for cache hits. Keep system prompts and long documents byte-identical and at the start of the message list; append new turns at the end. Cached input costs about a fifth of fresh input.
  5. Trim conversation history. Every past turn you resend is billed again as input (cached or not). Summarize old turns in long sessions.
  6. Control tool calls. Each built-in web search costs $0.01. Let the model search only when the question needs fresh facts.
  7. Move heavy coding to the Coding Plan, and schedule long agent runs off-peak (outside Monday to Friday, 14:00–18:00 UTC+8) to halve credit use.
  8. Watch the bill with a delay in mind. Z.ai’s billing history shows the previous day’s consumption, so today’s usage is not visible immediately. Check the billing page at z.ai/manage-apikey/billing daily during a new rollout.

If requests start failing rather than costing more, see GLM API rate limits and error codes: a 429 with code 1113 means your balance is empty (or, on the plan, a wrong base URL), while 1302 and 1305 are rate and capacity limits. New to the API? The GLM API quickstart gets you from key to first request in minutes.

GLM pricing FAQ

How much does the GLM API cost?

For text models, between $0 and $2.20 per 1M input tokens and between $0 and $8.90 per 1M output tokens, depending on the model. The current flagships (GLM-5.3, GLM-5.2, GLM-5.1) cost $1.40 in and $4.40 out; GLM-5.3-Flash costs $0.15 in and $0.50 out; GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash are free.

How much does GLM-5.2 cost per request?

It depends on tokens. A request with 2,000 input and 500 output tokens costs (2,000 × $1.40 + 500 × $4.40) / 1,000,000 = $0.005 on the Z.ai API. Reasoning tokens count as output, so thinking-heavy requests cost more.

Is there a free GLM API?

Yes. GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash are free on the Z.ai API for input, cached input and output. You still need an API key, and your account’s rate limits apply.

Is GLM-5.3-Flash free?

It is free to chat with on this site (it is the default model), but not on the Z.ai API, where it costs $0.15 per 1M input tokens and $0.50 per 1M output tokens. That is still about one ninth of the flagship price.

How much is the GLM Coding Plan?

Lite is $18 per month, Pro $80 and Max $168 on monthly billing. Quarterly billing takes 20% off and yearly billing 30% off, which brings Lite to $12.60, Pro to $56 and Max to $117.60 per month. Team seats cost $88 (Standard) or $188 (Premium) per seat per month.

Can I use Coding Plan credits for normal API calls?

No. The plan works only inside supported coding tools through the Coding Plan base URLs, and its quota never draws from or adds to your API balance. For apps and scripts, pay per token.

What is cached input on Z.ai?

Input tokens that repeat content the platform has recently processed, such as a fixed system prompt or earlier conversation turns. They are billed at the cached rate (for GLM-5.3, $0.26 instead of $1.40 per 1M), and caching happens automatically.

Is GLM cheaper on OpenRouter?

For some models, yes. OpenRouter’s listed prices for GLM-5.2, GLM-5.1, GLM-5 and GLM-4.7 are below Z.ai’s, and it offers a free, rate-limited z-ai/glm-5.2:free. For GLM-5.3 and GLM-5.3-Flash the listed prices match Z.ai’s, and GLM-4.7-Flash is cheaper on Z.ai (free). Always confirm on OpenRouter before budgeting.

Want to see which model is worth paying for? Try the same prompt on a few of them in the free GLM chat, then check the GLM release timeline to make sure you are pricing the latest version.