GLM API & Docs

GLM Function Calling & Structured Output: Complete Guide

Define tools, run the full call-and-respond loop, stream tool arguments, keep reasoning intact, and get validated JSON from GLM models.

GLM function calling follows the OpenAI tool-calling pattern. You send a tools array of functions described in JSON Schema, the model replies with tool_calls (a function name plus JSON arguments), your code runs the function, and you send the result back as a role: "tool" message with the matching tool_call_id. The model then writes its final answer. Structured output works through JSON mode: set response_format to {"type": "json_object"} and describe the shape you want in the prompt.

This guide covers both, with code that runs as written: the tool schema, the one tool_choice value Z.ai supports, a complete multi-round loop in Python, several tool calls in one response, streaming tool arguments with tool_stream, how thinking interacts with tools, and a validation-and-retry strategy that makes JSON output dependable. Z.ai (formerly Zhipu AI) documents these features on docs.z.ai; everything below matches its request formats. New to the API? Start with the GLM API quickstart to get a key and a first request working.

GLM function calling loop: tools schema, tool_calls, tool results and a final answer, plus JSON mode
Function calling is a loop: the model asks, your code runs the function, the model answers.

Which GLM models support function calling

All current GLM text models accept the tools parameter on the chat completions endpoint, including glm-5.3, glm-5.2, glm-5.1, glm-5, glm-4.7, glm-4.7-flash, glm-4.6 and the GLM-4.5 series. On the vision side, Z.ai’s API reference limits tools to the GLM-5.3-Flash series, the GLM-4.6V series and autoglm-phone-multilingual. Two related features have narrower support:

FeatureParameterSupported on
Function callingtools, tool_choiceGLM text models; GLM-5.3-Flash series and GLM-4.6V series for vision
Streaming tool argumentstool_stream: trueGLM-5.3, GLM-5.2, GLM-5.1, GLM-5, GLM-4.7, GLM-4.6
JSON moderesponse_format: {"type": "json_object"}Text models only
Built-in web searchtools: [{"type": "web_search", ...}]Chat completions; $0.01 per use
Tool and structured-output support, from Z.ai’s chat completions reference.

Which model should run your agent? For demanding multi-step tool use, GLM-5.3 is the flagship. For high-volume tool calls where cost matters, GLM-5.3-Flash costs $0.15 input and $0.50 output per 1M tokens and also accepts images. To prototype for free, GLM-4.7-Flash supports tools at zero cost per token.

Defining tools: the schema

Each entry in tools has type: "function" and a function object with three fields. Z.ai’s reference marks all three as required:

  • name: letters, digits, underscores and dashes only (pattern ^[a-zA-Z0-9_-]+$), 1 to 64 characters.
  • description: what the function does and when to use it. The model reads this to decide whether to call it, so write it for the model, not for a colleague.
  • parameters: a JSON Schema object describing the arguments. Use type, properties, required, enum and per-field description.

You can pass up to 128 functions in one request. Here are two well-specified tools used throughout this guide:

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_order_status",
            "description": "Look up the current status of a customer order by its order ID. Use it whenever the user asks where an order is.",
            "parameters": {
                "type": "object",
                "properties": {
                    "order_id": {"type": "string", "description": "Order ID, for example A1001"}
                },
                "required": ["order_id"],
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "calculate_shipping",
            "description": "Calculate the shipping cost for a parcel from its weight.",
            "parameters": {
                "type": "object",
                "properties": {
                    "weight_kg": {"type": "number", "description": "Parcel weight in kilograms"},
                    "express": {"type": "boolean", "description": "True for express delivery", "default": False},
                },
                "required": ["weight_kg"],
            },
        },
    },
]

The same request in raw HTTP, if you are not using Python. Note the tool_choice value and that the body is plain Chat Completions JSON:

curl -X POST "https://api.z.ai/api/paas/v4/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $ZAI_API_KEY" \
  -d '{
    "model": "glm-5.3",
    "messages": [{"role": "user", "content": "Where is order A1001?"}],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_order_status",
        "description": "Look up the current status of a customer order by its order ID.",
        "parameters": {
          "type": "object",
          "properties": {"order_id": {"type": "string", "description": "Order ID, for example A1001"}},
          "required": ["order_id"]
        }
      }
    }],
    "tool_choice": "auto",
    "reasoning_effort": "low"
  }'

If the model decides to call the tool, choices[0].message.tool_calls[0] holds an id, type: "function" and a function object with name and arguments, and finish_reason is tool_calls. Your next request repeats the conversation, adds that assistant message, and adds a tool message such as {"role": "tool", "tool_call_id": "<id from the call>", "content": "{\"status\": \"shipped\"}"}.

Schema design rules that improve tool choice

Z.ai’s best-practice list for tools comes down to three principles: single responsibility (each function does one thing), clear naming, and complete descriptions. In practice that means:

  • Use enum for any argument with a fixed set of values, such as units or operation types. It removes a whole class of invalid arguments.
  • Put formats and examples in the parameter description (“Date in YYYY-MM-DD format”, “Order ID, for example A1001”).
  • Mark only truly mandatory fields as required and document defaults for the rest.
  • Prefer several narrow tools over one tool with a mode switch. A get_order_status and a cancel_order are easier for the model to pick correctly than a single manage_order.
  • Keep the tool list relevant to the conversation. Every definition is part of the prompt, so 100 unused tools cost input tokens on every call.

Besides function, the chat completions reference defines two other tool types: web_search, which lets the model search the web during the call, and retrieval, which answers from a knowledge base by knowledge_id. Web search is covered in the GLM web search API guide.

tool_choice: only auto is supported

Z.ai’s reference is explicit: tool_choice controls how the model selects a function, the default is auto, and only auto is supported. You cannot force a specific function, force “some tool”, or pass none the way some other APIs allow. Code ported from elsewhere that sends "required" or {"type": "function", "function": {"name": ...}} needs changing.

The workarounds are simple:

  • To prevent tool calls on a turn, leave tools out of that request.
  • To make one tool very likely, send only that tool and state in the system prompt when it must be used.
  • To guarantee structured data without a tool, use JSON mode instead (covered below).
  • To catch a missed call, check finish_reason: tool_calls means the model wants a function run, stop means it answered directly.

The full GLM function calling loop in Python

A real agent does not stop after one tool call. The model may call a tool, read the result, call another, and only then answer. The loop below handles any number of rounds, several calls per round, malformed arguments and unknown function names. It uses the OpenAI Python SDK pointed at Z.ai; install it with pip install --upgrade 'openai>=1.0' and set ZAI_API_KEY.

import json
import os
from openai import OpenAI
client = OpenAI(
    api_key=os.environ["ZAI_API_KEY"],
    base_url="https://api.z.ai/api/paas/v4/",
)
MODEL = "glm-5.3"
# Stand-in data so the example runs without a database
ORDERS = {
    "A1001": {"status": "shipped", "eta_days": 2},
    "A1002": {"status": "processing", "eta_days": 5},
}
def get_order_status(order_id: str) -> dict:
    order = ORDERS.get(order_id.strip().upper())
    if order is None:
        return {"error": f"No order with id {order_id}"}
    return {"order_id": order_id.strip().upper(), **order}
def calculate_shipping(weight_kg: float, express: bool = False) -> dict:
    cost = (4.0 + 1.5 * weight_kg) * (2 if express else 1)
    return {"weight_kg": weight_kg, "express": express, "cost": round(cost, 2)}
FUNCTIONS = {"get_order_status": get_order_status, "calculate_shipping": calculate_shipping}
tools = [
    {"type": "function", "function": {
        "name": "get_order_status",
        "description": "Look up the current status of a customer order by its order ID.",
        "parameters": {"type": "object",
                       "properties": {"order_id": {"type": "string", "description": "Order ID, for example A1001"}},
                       "required": ["order_id"]}}},
    {"type": "function", "function": {
        "name": "calculate_shipping",
        "description": "Calculate the shipping cost for a parcel from its weight.",
        "parameters": {"type": "object",
                       "properties": {"weight_kg": {"type": "number", "description": "Parcel weight in kilograms"},
                                      "express": {"type": "boolean", "description": "True for express delivery"}},
                       "required": ["weight_kg"]}}},
]
def as_text(arguments) -> str:
    # The docs describe arguments as a JSON string; accept an object too
    return arguments if isinstance(arguments, str) else json.dumps(arguments)
def run_tool(name: str, arguments) -> dict:
    fn = FUNCTIONS.get(name)
    if fn is None:
        return {"error": f"Unknown function {name}"}
    try:
        args = json.loads(as_text(arguments) or "{}")
        return fn(**args)
    except (json.JSONDecodeError, TypeError) as exc:
        return {"error": f"Bad arguments for {name}: {exc}"}
def run(user_text: str, max_rounds: int = 6) -> str:
    messages = [
        {"role": "system", "content": "You are a support assistant. Use the tools for order and shipping questions."},
        {"role": "user", "content": user_text},
    ]
    for _ in range(max_rounds):
        response = client.chat.completions.create(
            model=MODEL,
            messages=messages,
            tools=tools,
            tool_choice="auto",
            extra_body={"reasoning_effort": "low"},
        )
        msg = response.choices[0].message
        if not msg.tool_calls:
            return msg.content
        # 1) Keep the assistant turn, with its reasoning, before the tool results
        messages.append({
            "role": "assistant",
            "content": msg.content or "",
            "reasoning_content": getattr(msg, "reasoning_content", None) or "",
            "tool_calls": [
                {"id": tc.id, "type": "function",
                 "function": {"name": tc.function.name, "arguments": as_text(tc.function.arguments)}}
                for tc in msg.tool_calls
            ],
        })
        # 2) Answer every tool call, matched by id
        for tc in msg.tool_calls:
            result = run_tool(tc.function.name, tc.function.arguments)
            messages.append({
                "role": "tool",
                "tool_call_id": tc.id,
                "content": json.dumps(result, ensure_ascii=False),
            })
    raise RuntimeError("Stopped after too many tool rounds")
print(run("Where is order A1001, and what would express shipping cost for a 3 kg parcel?"))
The GLM function calling loop in five steps: tools, tool_calls, run the function, tool message, final answer

What each part of the loop does

  1. First call. The model sees the tools and the question. If it needs data, it returns tool_calls and finish_reason: "tool_calls" instead of an answer.
  2. Assistant turn. You append the model’s message exactly as a history entry: its content, its tool_calls with their ids, and its reasoning_content. Without this entry, the tool results that follow have nothing to attach to.
  3. Tool results. For every call you append a role: "tool" message whose tool_call_id matches the call’s id, with the result serialised as a string. JSON is the clearest format for the model to read.
  4. Next call. Send the whole history again with the same tools. The model either calls more functions or writes the final answer.
  5. Safety valve. max_rounds stops a confused model from looping forever and running up a bill.

Errors go back to the model as results, not exceptions. When the arguments are malformed or a lookup fails, returning {"error": "..."} lets the model correct itself or explain the problem to the user, which is usually better than crashing the request.

Several tool calls in one response

tool_calls is an array, and a single assistant message can contain more than one entry. The question in the example above (“where is my order, and what would shipping cost?”) can produce both a get_order_status and a calculate_shipping call in the same response. Z.ai’s examples always loop over the whole array, and so should you:

  • Run every call in the list, not just the first one.
  • Send one role: "tool" message per call, each with its own tool_call_id.
  • Independent calls can run concurrently in your own code, for example with a thread pool. Collect the results, then append one tool message per call with its ID.

Z.ai’s documentation does not describe a switch to turn multiple calls per response on or off, so write your handler for a list from day one. In streaming mode, each call in the list is identified by its index, as shown in the next section.

Streaming tool calls with tool_stream

Normally, when you stream a response that ends in a tool call, the arguments arrive only once they are complete. With tool_stream: true (together with stream: true), GLM streams the function name and argument fragments as they are generated, alongside reasoning_content and content. Z.ai supports it on GLM-5.3, GLM-5.2, GLM-5.1, GLM-5, GLM-4.7 and GLM-4.6. It is off by default.

Use it when you want to show progress in a UI (“looking up order A1002…”) or start preparing work before the arguments finish. The rule for handling it: collect fragments per index and concatenate function.arguments until the stream ends. This snippet reuses client and tools from the loop above:

stream = client.chat.completions.create(
    model="glm-5.3",
    messages=[{"role": "user", "content": "Where is order A1002?"}],
    tools=tools,
    stream=True,
    extra_body={"tool_stream": True, "reasoning_effort": "low"},
)
reasoning, content, calls = "", "", {}
for chunk in stream:
    if not chunk.choices:
        continue
    delta = chunk.choices[0].delta
    if getattr(delta, "reasoning_content", None):
        reasoning += delta.reasoning_content
    if delta.content:
        content += delta.content
    for tc in delta.tool_calls or []:
        slot = calls.setdefault(tc.index, {"id": None, "name": "", "arguments": ""})
        if tc.id:
            slot["id"] = tc.id
        if tc.function and tc.function.name:
            slot["name"] = tc.function.name
        if tc.function and tc.function.arguments:
            slot["arguments"] += tc.function.arguments
            print(f"[{slot['name']}] {slot['arguments']}", flush=True)
for index, call in sorted(calls.items()):
    print(index, call["id"], call["name"], json.loads(call["arguments"] or "{}"))

Do not parse the argument string until the stream has finished; a partial fragment is not valid JSON. Once complete, the collected calls feed into the same tool-execution step as the non-streaming loop. Z.ai’s migration guide recommends testing specifically for “parameter completeness in tool streams” after you switch it on.

Thinking and tools: preserve reasoning_content

GLM models think between tool calls. Z.ai calls this interleaved thinking, and it has been on by default since GLM-4.5: the model reasons, calls a tool, reads the result, reasons again and decides what to do next. The docs give one instruction for it: when using interleaved thinking with tools, thinking blocks should be explicitly preserved and returned together with the tool results. That is why the loop above stores reasoning_content in the assistant message before appending tool results.

Two related settings matter for agents:

  • Preserved thinking keeps reasoning from earlier assistant turns in context. It is on by default on the Coding Plan endpoint and off by default on the standard API. Turn it on by adding "clear_thinking": false inside the thinking object. You must then send every historical reasoning_content back complete, unmodified and in the original order; Z.ai warns that edited or reordered blocks can degrade performance and cache hit rates.
  • Turn-level thinking (since GLM-4.7) lets each request in a session switch thinking on or off. Use it to skip reasoning on quick tool-execution turns and keep it for turns that must decide what to do with tool results. On GLM-5.3 and GLM-5.3-Flash thinking cannot be disabled, so lower reasoning_effort to low instead.

To enable preserved thinking with the OpenAI SDK, pass the thinking object through extra_body. In an agent, add tools=tools and keep appending each assistant turn with its reasoning_content, exactly as the loop above does:

import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["ZAI_API_KEY"], base_url="https://api.z.ai/api/paas/v4/")
messages = [{"role": "user", "content": "Plan the steps to move a small web app to a new server."}]
response = client.chat.completions.create(
    model="glm-4.7",
    messages=messages,
    extra_body={"thinking": {"type": "enabled", "clear_thinking": False}},
)
msg = response.choices[0].message
messages.append({
    "role": "assistant",
    "content": msg.content,
    "reasoning_content": getattr(msg, "reasoning_content", None) or "",
})
print(msg.content)

For plain chat without tools, leave clear_thinking at its default of true: old reasoning is dropped automatically, which keeps context short and cheap. The GLM thinking mode guide covers each mode and the per-model reasoning_effort values.

GLM structured output with JSON mode

When you need data rather than an action, skip tools and use GLM JSON mode. Set response_format to {"type": "json_object"} and the model returns a JSON object in message.content. The field accepts two values, text (the default) and json_object, and only text models support it.

JSON mode guarantees JSON, not your JSON. The model decides the keys unless you tell it, so Z.ai’s own examples always describe the expected structure in the system message. A minimal example:

import json
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["ZAI_API_KEY"], base_url="https://api.z.ai/api/paas/v4/")
response = client.chat.completions.create(
    model="glm-4.7-flash",
    messages=[
        {"role": "system", "content": (
            "You are a sentiment analysis expert. Return only a JSON object in this format: "
            '{"sentiment": "positive|negative|neutral", "confidence": 0.0, "keywords": ["..."]}'
        )},
        {"role": "user", "content": "The weather is really nice today, I am feeling very happy!"},
    ],
    response_format={"type": "json_object"},
    extra_body={"thinking": {"type": "disabled"}},
)
result = json.loads(response.choices[0].message.content)
print(result["sentiment"], result["confidence"], result["keywords"])

The official list of models that support structured output includes glm-5, glm-4.7, glm-4.6 and glm-4.5, and the reference applies response_format to all text models. Note what is missing: there is no json_schema type, so you cannot hand the API a schema and have it enforced server-side. Enforcement is your job, and it is not hard.

GLM JSON mode checklist: response_format json_object, schema in the prompt, validate and retry

Schema validation strategy for reliable JSON

Z.ai recommends multi-layer validation (schema checks plus business-logic checks) and a fallback plan. A pattern that works well in production:

  1. Put the JSON Schema itself in the system prompt, along with rules such as “use null for anything not stated”.
  2. Call with response_format: {"type": "json_object"}.
  3. Parse with json.loads and validate with the jsonschema package.
  4. On failure, send the invalid output back with the validation error and ask for a corrected object. Two or three attempts are plenty.
  5. After schema validation, run business checks the schema cannot express (a price that must be positive, a date that must be in the future).

Install the validator with pip install jsonschema. The extractor below uses the free glm-4.7-flash with thinking off, which suits simple extraction well:

import json
import os
from jsonschema import ValidationError, validate
from openai import OpenAI
client = OpenAI(api_key=os.environ["ZAI_API_KEY"], base_url="https://api.z.ai/api/paas/v4/")
SCHEMA = {
    "type": "object",
    "properties": {
        "name": {"type": "string"},
        "brand": {"type": ["string", "null"]},
        "price": {"type": ["number", "null"]},
        "colors": {"type": "array", "items": {"type": "string"}},
        "battery_hours": {"type": ["number", "null"]},
        "warranty_years": {"type": ["number", "null"]},
    },
    "required": ["name", "brand", "price", "colors", "battery_hours", "warranty_years"],
    "additionalProperties": False,
}
SYSTEM = (
    "Extract product data. Return only a JSON object that matches this JSON Schema. "
    "Use null for anything the text does not state. Do not invent values.\n"
    + json.dumps(SCHEMA)
)
def clean(raw: str) -> str:
    # Drop anything outside the outermost braces, such as code fences
    start, end = raw.find("{"), raw.rfind("}")
    return raw[start:end + 1] if start != -1 and end > start else raw
def extract(text: str, attempts: int = 3) -> dict:
    messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": text}]
    for _ in range(attempts):
        response = client.chat.completions.create(
            model="glm-4.7-flash",
            messages=messages,
            response_format={"type": "json_object"},
            extra_body={"thinking": {"type": "disabled"}},
        )
        raw = response.choices[0].message.content or ""
        try:
            data = json.loads(clean(raw))
            validate(instance=data, schema=SCHEMA)
            return data
        except (json.JSONDecodeError, ValidationError) as exc:
            problem = exc.message if isinstance(exc, ValidationError) else str(exc)
            messages.append({"role": "assistant", "content": raw})
            messages.append({"role": "user", "content": f"That output was invalid: {problem}. Return the corrected JSON object only."})
    raise ValueError("No valid JSON after retries")
print(extract("The Aurora X2 speaker by Lumen Audio sells for 129.99, comes in black or sand, and runs 18 hours per charge."))

Three details make this robust. The schema lists every key as required but allows null, so the model must decide explicitly whether a value is present rather than silently omitting it. additionalProperties: false catches invented keys. And the retry message quotes the exact validation error, which gives the model something concrete to fix. If you use Pydantic, the same pattern works with Model.model_validate_json(raw) in place of jsonschema.

Z.ai also warns that strict JSON output can make responses feel less natural in complex scenarios. Keep JSON mode for machine-read output, and let user-facing replies stay free text.

Function calling or JSON mode: which to use

You need to…UseWhy
Fetch live data or take an actionFunction callingThe model chooses when to call and with what arguments
Extract fields from text into a fixed shapeJSON modeOne call, no loop, validated in your code
Classify or score inputJSON modeSmall object with an enum field
Let the model pick among several actionsFunction callingEach action is a separate tool
Search the web during a replyBuilt-in web_search toolNo function of your own needed
Force a specific function every timeJSON mode with that function’s schematool_choice only supports auto

The last row is a useful trick. Because you cannot force a tool, a JSON-mode call whose system prompt contains the function’s parameter schema gives you guaranteed arguments; you then call the function yourself.

Keep tool calls safe

Function calling lets model output trigger real code, so treat every argument as untrusted input. Z.ai’s docs flag this twice: they ask for security validation and permission control around any external API or database a tool touches, and they point out that function calling involves code execution. A short checklist:

  • Validate arguments before acting. Check types, lengths and allowed values in your function, even when the schema already declares them. The model can still send an order ID with stray characters or a negative weight.
  • Never pass raw SQL or shell text through a tool. Expose narrow operations (get_order_status) rather than general ones (run_query), and use parameterised queries inside them.
  • Check permissions per user. The model does not know who is allowed to see which order. Resolve the current user in your code and enforce access there.
  • Confirm destructive actions. For refunds, deletions or payments, have the tool return a confirmation request and let the user approve before the action runs.
  • Log every call. Record the function name, arguments, result and the request that triggered it. It is the only way to debug an agent after the fact.
  • Bound the loop. Cap rounds per request and total tool calls per session, as the max_rounds guard in the loop above does.

A tool that returns a structured error ({"error": "Order not found"}) is safer than one that raises, because the model can explain the problem instead of retrying blindly. Keep error messages factual and free of internal details such as stack traces or connection strings, since the model may repeat them to the user.

Common function-calling errors and fixes

  • HTTP 400, code 1210 or 1214: a parameter is invalid. Common causes are a tool_choice other than auto, a function name with spaces or dots, more than 128 tools, or a response_format type other than text or json_object. The message names the field.
  • Code 1213: a required parameter is missing. For tool messages, check that tool_call_id and content are both present.
  • The model answers without calling the tool: sharpen the tool description, say in the system prompt when the tool must be used, and send fewer, more relevant tools.
  • Arguments are not valid JSON: catch the parse error and return it to the model as the tool result; it will usually retry with fixed arguments.
  • finish_reason: "length": the output hit max_tokens, possibly mid-argument. Raise the limit or reduce reasoning effort.
  • Tool result ignored: make sure the tool_call_id matches the call’s id exactly and that the assistant message with tool_calls comes before the tool messages.
  • Thinking disabled on GLM-5.3: GLM-5.3 and GLM-5.3-Flash reject thinking: {"type": "disabled"}. Use reasoning_effort: "low" instead.

Rate limits and all other error codes are covered in GLM API rate limits and error codes. For tool-heavy agent frameworks built on GLM, see GLM with OpenClaw, and compare per-call costs on the GLM pricing page.

GLM function calling FAQ

Does GLM support function calling?

Yes. Pass a tools array of JSON Schema function definitions to the chat completions endpoint. The model returns tool_calls with a function name and arguments; you run the function and send the result back as a role: "tool" message. The format matches OpenAI-style tool calling.

Which GLM models support tool calling?

All current text models, including GLM-5.3, GLM-5.2, GLM-5.1, GLM-5, GLM-4.7, GLM-4.7-Flash, GLM-4.6 and the GLM-4.5 series. For image input with tools, use the GLM-5.3-Flash series or the GLM-4.6V series. Streaming tool arguments (tool_stream) work on GLM-5.3, 5.2, 5.1, 5, 4.7 and 4.6.

Does GLM support tool_choice required or a named function?

No. Only tool_choice: "auto" is supported. To make a call very likely, send only the tool you want and instruct the model in the system prompt. To guarantee structured arguments, use JSON mode with that function’s schema and call the function yourself.

Does GLM have a JSON mode?

Yes. Set response_format to {"type": "json_object"} on a text model and describe the expected structure in the system message. The response content is then a JSON object you can parse with json.loads.

Does GLM support JSON Schema structured outputs?

Not as a server-side feature: response_format accepts only text and json_object. Put your schema in the prompt, validate the result with a library such as jsonschema or Pydantic, and retry with the validation error when it fails.

Can GLM call several functions in one response?

Yes, tool_calls is a list and can hold more than one call. Run each one and send a separate tool message per call with its own tool_call_id. When streaming with tool_stream, group argument fragments by each call’s index.

How many tools can I pass to GLM?

Up to 128 functions per request. Function names may use letters, digits, underscores and dashes, up to 64 characters. Fewer, well-described tools give better tool selection and cost fewer input tokens.

Do I need to send reasoning_content back with tool results?

For agent loops, yes. Z.ai’s docs say thinking blocks should be preserved and returned together with tool results, so include the assistant message’s reasoning_content in history before the tool messages. With preserved thinking (clear_thinking: false), send all earlier reasoning back complete and in order. You can try tool-free prompts first in the free GLM chat.

More in GLM API & Docs