GLM thinking mode is the chain-of-thought switch in Z.ai’s Chat Completions API. You control it with two request fields: thinking ({"type": "enabled"} or {"type": "disabled"}) and, on GLM-5.2 and newer, reasoning_effort, which sets how hard the model thinks. Thinking is on by default for every GLM-5 model and for the GLM-4.7 series. The model returns its reasoning in a separate reasoning_content field, and those tokens are generated output, so you pay for them at the output rate.
The one rule that trips people up: GLM-5.3 and GLM-5.3-Flash cannot turn thinking off. If you send "disabled" to either model, the request fails. The lightest setting is thinking: {"type": "enabled"} plus reasoning_effort: "low". This guide covers the per-model defaults, the exact values each model accepts, working code, the three multi-turn modes (interleaved, preserved and turn-level), what thinking costs, and when to switch it off or dial it down.

How GLM thinking mode works
Z.ai (formerly Zhipu AI) calls this feature Deep Thinking. With thinking enabled, the model writes out a chain of reasoning before it answers. That helps most with multi-step problems, logic, code and planning, and it also lets you see why the model reached its answer. Three request parameters control it:
thinking.type:enabled(the default) ordisabled. Supported on GLM-4.5 and newer.reasoning_effort: how much reasoning to do once thinking is enabled. Only GLM-5.2 and newer support it. The documented levels arelow(light thinking),high(enhanced thinking) andmax(deep thinking, the default).thinking.clear_thinking: whether reasoning from earlier assistant turns is dropped from the context (true, the default on the standard API) or kept (false, which Z.ai calls Preserved Thinking).
“Enabled” doesn’t mean the same thing on every model. Z.ai’s parameter reference says that with thinking enabled, GLM-5.2, GLM-5.1, GLM-5, GLM-4.6 and GLM-4.5 decide for themselves whether a request needs thinking. That’s the dynamic, or hybrid, behaviour. GLM-5.3, GLM-5.3-Flash, GLM-4.7 and GLM-4.5V use forced thinking: when enabled, they always think. The difference matters for cost. A dynamic model may answer “What is 2 + 2?” with no reasoning at all, while a forced-thinking model will always spend some reasoning tokens first.
Where the reasoning shows up
In a normal response, the answer is in choices[0].message.content and the reasoning is in choices[0].message.reasoning_content. When streaming, reasoning arrives as choices[0].delta.reasoning_content chunks, followed by the answer as delta.content chunks. The usage object and the finish_reason come in the last chunk, and the stream ends with data: [DONE]. Show the reasoning to users, log it or throw it away. Just don’t mix it into the answer text. Z.ai’s reference pages are the Thinking Mode guide and the Deep Thinking guide.
GLM reasoning defaults and reasoning_effort values by model
This table covers every current text model, based on Z.ai’s thinking-mode guide, the Deep Thinking guide and the Chat Completions API reference. Use it as your quick reference before you change a model ID.
| Model | Thinking default | Can you disable it? | reasoning_effort | When enabled |
|---|---|---|---|---|
| GLM-5.3 | On (forced) | No, “disabled” fails | low / high / max (default max) | Always thinks |
| GLM-5.3-Flash / FlashX | On (forced) | No | low / high / max (default max) | Always thinks |
| GLM-5.2 | On | Yes: “disabled”, or effort none/minimal | high / max (default max) | Decides per request |
| GLM-5.1 | On | Yes | Not supported | Decides per request |
| GLM-5 | On | Yes | Not supported | Decides per request |
| GLM-4.7 series | On | Yes, per turn | Not supported | Always thinks |
| GLM-4.6 | Hybrid (dynamic) | Yes | Not supported | Decides per request |
| GLM-4.5 series | Hybrid (dynamic) | Yes | Not supported | Decides per request |
How each reasoning_effort value is mapped
Z.ai accepts extra effort names so that clients written for other APIs keep working, but the mapping depends on the model and on the endpoint. On the standard API, GLM-5.3 and GLM-5.3-Flash are strict: anything other than low, high or max returns an error. The Coding Plan endpoint is more forgiving and folds every name into one of the three levels.
| Value sent | GLM-5.3 / 5.3-Flash (API) | GLM-5.3 / 5.3-Flash (Coding Plan) | GLM-5.2 |
|---|---|---|---|
| none | Error | low | Skips thinking |
| minimal | Error | low | Skips thinking |
| low | low | low | high |
| medium | Error | high | high |
| high | high | high | high |
| xhigh | Error | max | max |
| max | max | max | max |
Two practical consequences. First, low on GLM-5.2 is really high, so GLM-5.2 has only two effort levels. If you want less reasoning than that, turn thinking off. Second, on the GLM Coding Plan, requests for GLM-5.2 and GLM-5.1 are now routed to GLM-5.3, so in agents like Claude Code the GLM-5.3 column is the one that applies.
Self-hosted weights behave differently
If you serve the open weights yourself with vLLM or SGLang, the chat template sets the rules, not Z.ai’s API. The GLM-5 GitHub README says GLM-5.3 and GLM-5.3-Flash fall back to max when reasoning_effort is missing or set to any other value, so you must pass low or high explicitly. GLM-5.2’s template accepts only high and max, with the same fallback. The template also defaults clear_thinking to false. For plain chat, pass clear_thinking=true explicitly.

Code: enable, tune and disable thinking
All examples call the standard endpoint https://api.z.ai/api/paas/v4/chat/completions. Replace your-api-key or set ZAI_API_KEY in your environment. If you haven’t made a key yet, the GLM API quickstart walks through it.
GLM-5.3 with light reasoning (curl)
This is the cheapest and fastest way to call GLM-5.3. It’s also the migration target for any code that used to send "disabled".
curl -X POST "https://api.z.ai/api/paas/v4/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer your-api-key" \
-d '{
"model": "glm-5.3",
"messages": [
{"role": "user", "content": "Summarise the difference between TCP and UDP in three bullet points."}
],
"thinking": {"type": "enabled"},
"reasoning_effort": "low",
"max_tokens": 4096
}'
Streaming reasoning and answer separately (Python, OpenAI SDK)
The OpenAI Python SDK works with Z.ai’s base URL. thinking isn’t an OpenAI parameter, so send it through extra_body. reasoning_effort goes there too, so the request body comes out exactly as Z.ai documents it. reasoning_content isn’t a typed field in the SDK, so read it with getattr.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["ZAI_API_KEY"],
base_url="https://api.z.ai/api/paas/v4/",
)
stream = client.chat.completions.create(
model="glm-5.3",
messages=[{"role": "user", "content": "Design a rate limiter for a public REST API."}],
stream=True,
max_tokens=8192,
extra_body={
"thinking": {"type": "enabled"},
"reasoning_effort": "high",
},
)
reasoning = []
usage = None
for chunk in stream:
if getattr(chunk, "usage", None):
usage = chunk.usage
if not chunk.choices:
continue
delta = chunk.choices[0].delta
thought = getattr(delta, "reasoning_content", None)
if thought:
reasoning.append(thought)
if delta.content:
print(delta.content, end="", flush=True)
print("\n\nReasoning characters:", len("".join(reasoning)))
if usage:
print("Prompt tokens:", usage.prompt_tokens)
print("Output tokens (reasoning + answer):", usage.completion_tokens)
Turning thinking off on GLM-5.2 (curl)
GLM-5.2 is the newest model that still lets you disable thinking. Either of these requests produces a direct answer with no reasoning phase:
curl -X POST "https://api.z.ai/api/paas/v4/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer your-api-key" \
-d '{
"model": "glm-5.2",
"messages": [{"role": "user", "content": "Convert 72 degrees Fahrenheit to Celsius."}],
"thinking": {"type": "disabled"}
}'
curl -X POST "https://api.z.ai/api/paas/v4/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer your-api-key" \
-d '{
"model": "glm-5.2",
"messages": [{"role": "user", "content": "Convert 72 degrees Fahrenheit to Celsius."}],
"thinking": {"type": "enabled"},
"reasoning_effort": "none"
}'
Official Python SDK (zai-sdk)
With pip install zai-sdk, thinking and reasoning_effort are ordinary keyword arguments:
from zai import ZaiClient
client = ZaiClient(api_key="your-api-key")
response = client.chat.completions.create(
model="glm-5.3",
messages=[{"role": "user", "content": "Find the bug: def avg(xs): return sum(xs) / len(xs) - 1"}],
thinking={"type": "enabled"},
reasoning_effort="max",
)
print(response.choices[0].message.content)
print("---")
print(response.choices[0].message.reasoning_content)
Migrating a “thinking disabled” app to GLM-5.3
Z.ai’s migration notice is short: if your app sends thinking.type: "disabled", change it to enabled and add reasoning_effort: "low" before you switch the model ID to glm-5.3. Otherwise every request fails with a 400 invalid-parameter error. The GLM API error code guide covers what that error looks like and how to find the cause. A small helper keeps one code path for every model:
FORCED_THINKING = {"glm-5.3", "glm-5.3-flash", "glm-5.3-flashx"}
def light_thinking(model):
"""Return the lightest valid reasoning settings for a GLM model."""
if model in FORCED_THINKING:
return {"thinking": {"type": "enabled"}, "reasoning_effort": "low"}
return {"thinking": {"type": "disabled"}}
# usage with the OpenAI SDK:
# client.chat.completions.create(model=m, messages=msgs, extra_body=light_thinking(m))
Interleaved, preserved and turn-level thinking
Once a conversation runs to several turns or several tool calls, you have to decide what happens to the reasoning from earlier steps. GLM supports three modes, each added in a different model generation.
Interleaved thinking (since GLM-4.5)
With interleaved thinking, the model reasons between tool calls: it reads each tool result, thinks about it, then decides the next step. It’s on by default. When you run a tool loop, Z.ai’s guide says the thinking blocks should be kept and sent back along with the tool results. In practice, that means putting the assistant message back into messages with its reasoning_content, its tool_calls and its content.
Preserved thinking (clear_thinking: false)
By default, the standard API drops reasoning_content from earlier turns (clear_thinking: true). It keeps only the visible text, the tool calls and the tool results. That’s the right choice for ordinary chat because it keeps the context short and cheap. For coding agents, Z.ai recommends Preserved Thinking: set "clear_thinking": false inside the thinking object and send back the full, unmodified reasoning from every earlier assistant turn. The model keeps its train of thought across turns, and Z.ai says this also raises cache hit rates. Preserved Thinking is on by default on the Coding Plan endpoint and off by default on the standard API.
The rule is strict. Consecutive reasoning_content blocks must exactly match what the model generated, in the same order. If you truncate, rewrite or reorder them, performance may drop and the feature may stop working. Here’s a minimal non-streaming tool loop with Preserved Thinking on GLM-4.7:
import json
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["ZAI_API_KEY"], base_url="https://api.z.ai/api/paas/v4/")
tools = [{
"type": "function",
"function": {
"name": "get_order_status",
"description": "Look up the shipping status of an order by its ID",
"parameters": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
},
},
}]
messages = [{"role": "user", "content": "Where is order A1234?"}]
extra = {"thinking": {"type": "enabled", "clear_thinking": False}}
first = client.chat.completions.create(
model="glm-4.7", messages=messages, tools=tools, extra_body=extra
)
msg = first.choices[0].message
if not msg.tool_calls:
print(msg.content)
else:
# Send the reasoning back unchanged, together with the tool calls
messages.append({
"role": "assistant",
"content": msg.content or "",
"reasoning_content": getattr(msg, "reasoning_content", None) or "",
"tool_calls": [
{"id": tc.id, "type": "function",
"function": {"name": tc.function.name, "arguments": tc.function.arguments}}
for tc in msg.tool_calls
],
})
for tc in msg.tool_calls:
args = json.loads(tc.function.arguments)
result = {"order_id": args["order_id"], "status": "shipped"} # your real lookup here
messages.append({"role": "tool", "tool_call_id": tc.id, "content": json.dumps(result)})
second = client.chat.completions.create(
model="glm-4.7", messages=messages, tools=tools, extra_body=extra
)
print(second.choices[0].message.content)
For the full tool schema, parallel calls and streaming tool arguments, see the GLM function calling guide.
Turn-level thinking (since GLM-4.7)
Turn-level thinking means each request in a session can switch thinking on or off by itself, and the model stays coherent across the switches. Use it to spend reasoning only where it pays off. Turn it off for “rephrase this” or “what’s the capital of…” turns, and on for debugging, planning or multi-constraint turns. In agent loops, Z.ai suggests turning it down on turns that just execute a tool and turning it up on turns that decide what to do with a tool’s result. On models that don’t allow disabling, such as GLM-5.3, do the same thing with reasoning_effort instead: low for easy turns, max for hard ones.
What thinking costs: reasoning tokens and your bill
Z.ai’s pricing has three lines per model: input, cached input and output. There’s no separate price for reasoning. The chain of thought is text the model generates, so it’s billed at the output rate, the most expensive of the three. Z.ai’s guide says it plainly: the thinking process consumes extra tokens, and higher effort levels increase response time. On GLM-5.3, GLM-5.2 and GLM-5.1, output costs $4.40 per 1M tokens. On GLM-5.3-Flash it’s $0.50, and on GLM-4.7 it’s $2.20. Full tables are on the GLM pricing page.
Z.ai’s own coding benchmark shows how much reasoning a model can use. In Z.ai’s published results for its in-house Z.ai Code Bench, GLM-5.3 at Max effort averaged about 75K output tokens per task and scored 34.5%. At High effort it used about 50K tokens and scored 31.4%. GLM-5.2 at Max used 96K tokens and scored 23.4%. Here’s what that works out to per task at API list prices:
| Setup | Output tokens per task | Output cost per task | Z.ai Code Bench |
|---|---|---|---|
| GLM-5.3, max | ~75,000 | 75,000 × $4.40 / 1M = $0.33 | 34.5% |
| GLM-5.3, high | ~50,000 | 50,000 × $4.40 / 1M = $0.22 | 31.4% |
| GLM-5.2, max | ~96,000 | 96,000 × $4.40 / 1M ≈ $0.42 | 23.4% |
Moving from High to Max on GLM-5.3 costs about 50% more output for about 3 points more on that benchmark. For hard agentic coding that trade is often worth it. For a chatbot it usually isn’t. For short tasks the ratio matters more than the totals. Take an illustrative case: a 300-token answer that comes with 1,200 tokens of reasoning is billed as 1,500 output tokens, five times the answer alone. At 10,000 such requests a day on GLM-5.3, that’s 15M output tokens × $4.40 = $66 a day instead of $13.20. The same traffic on GLM-5.3-Flash costs 15M × $0.50 = $7.50.
Thinking on the Coding Plan
Coding Plan usage is counted in credits, and output tokens carry the heaviest weight: GLM-5.3’s output multiplier is 24, against 6.9 for input. A task that generates 75,000 output tokens uses 75,000 × 24 / 10,000 = 180 credits for output alone at peak rates, or half that off-peak. The Lite plan’s 5-hour window holds 2,000 credits, so the effort level you choose in your coding agent directly decides how many tasks fit in a window. GLM-5.3-Flash’s output multiplier is 8, a third of GLM-5.3’s.
Give the reasoning room
GLM-5.x models accept max_tokens up to 131,072, and GLM-4.5 up to 96K. If you set a small max_tokens with a high effort level, check finish_reason. A value of length means generation stopped at your cap before the reply was complete. Raise the cap or lower the effort. Don’t just retry.
When to turn GLM reasoning off (or down)
Z.ai’s own best-practice list is a good starting point. Keep thinking on for complex problem analysis, multi-step reasoning, technical design, strategy and planning, research and analysis, and creative writing. Turn it off for simple fact lookups, basic translation, simple classification and quick Q&A. Here’s how that works out in practice:
| Workload | GLM-5.3 / 5.3-Flash | GLM-5.2 | GLM-5.1, GLM-5, GLM-4.x |
|---|---|---|---|
| Classification, tagging, routing | low | disabled | disabled |
| JSON extraction from clean text | low | disabled | disabled |
| Customer-support chat | low | disabled or high | disabled or enabled |
| Summaries of long documents | low or high | high | enabled |
| Maths, logic, tricky reasoning | max | max | enabled |
| Agentic coding, long tool chains | max (high to save tokens) | max | enabled + preserved |
Three signs you should dial reasoning down: users complain about the wait before the first visible token, your output-token bill is several times your answer length, or the task has one correct format and no real decisions to make. Three signs you should dial it up: answers are fluent but wrong on multi-step problems, the model skips constraints in long instructions, or an agent makes a poor choice right after reading a tool result.
One more lever: model choice. Forced thinking at low on GLM-5.3-Flash ($0.15 in / $0.50 out per 1M) often costs less than thinking-off on a bigger model. If you want a model that’s completely free, GLM-4.7-Flash costs nothing on the API and still supports thinking. It’s covered in the GLM-4.7-Flash guide.
How GLM Chat uses thinking
The free GLM chat on this site always uses the lightest reasoning setting each model allows. Thinking is disabled where it can be, and reasoning_effort is low on GLM-5.3 and GLM-5.3-Flash, where it can’t be turned off. That keeps replies quick. Want to see how a model handles everyday prompts at that setting? Try GLM-5.3 in the chat or compare it with GLM-5.2 in the chat. For model-level details, see the GLM-5.3 page, the GLM-5.2 page and the GLM-4.7 page.
Next step: pair thinking with GLM context caching to cut the input side of your bill, or add live results with the GLM web search API.
GLM thinking mode FAQ
Can I turn off thinking on GLM-5.3?
No. GLM-5.3 and GLM-5.3-Flash use forced thinking, and a request with thinking.type: "disabled" fails. Send thinking: {"type": "enabled"} with reasoning_effort: "low" for the lightest reasoning. If you need no reasoning at all, use GLM-5.2 with thinking disabled, or an older model.
What is the default reasoning_effort?
max on every model that supports the parameter (GLM-5.2, GLM-5.3 and GLM-5.3-Flash). If you leave it out, you get the deepest and most expensive reasoning, so set it explicitly in production.
Does GLM-4.7 or GLM-5.1 support reasoning_effort?
No. Z.ai documents reasoning_effort only for GLM-5.2 and newer. On GLM-5.1, GLM-5, GLM-4.7, GLM-4.6 and GLM-4.5, the only control is thinking.type: enabled or disabled.
Are reasoning tokens billed?
Yes. Reasoning is generated text, and Z.ai prices only input, cached input and output, so the reasoning is billed as output at the model’s output rate. Check usage.completion_tokens to see the combined count for reasoning and answer.
What does clear_thinking do?
It controls whether reasoning from earlier turns stays in the context. true, the API default, strips it. false keeps it (Preserved Thinking), but only if you send back the exact, complete reasoning_content of each earlier assistant turn. It doesn’t change whether the model thinks in the current turn.
Is thinking on when I use GLM in Claude Code?
Yes. Coding Plan requests go to GLM-5.3 or GLM-5.3-Flash, which always think, and Preserved Thinking is on by default on the Coding Plan endpoint. Effort names from the client are folded into low, high or max. Setup is in the GLM in Claude Code guide.
Why is reasoning_content empty on some GLM-5.2 replies?
GLM-5.2, GLM-5.1, GLM-5, GLM-4.6 and GLM-4.5 decide for themselves whether a request needs thinking. For simple prompts they may answer directly, even with thinking enabled. That’s expected, and it saves you output tokens.