Editorial Policy

The sources we trust, how we test GLM models in our own chat, how pages change when Z.ai updates a model, and how corrections work.

This editorial policy explains how GLM Chat researches, writes, tests and corrects its pages about GLM models: which sources we trust, how we test GLM models in our own chat, how pages change when Z.ai changes a model, and what happens when you report a mistake. It applies to every model page, guide, comparison and help article on the site.

GLM Chat editorial policy: official sources, a five-task model test and open corrections

What this editorial policy covers

GLM Chat publishes two kinds of content. The first is factual reference material: specs, prices, model IDs, API parameters, licenses and release dates for the GLM models made by Z.ai (formerly Zhipu AI). The second is practical guidance: which model to pick, how to call it, how to set it up in a coding tool and how to fix common errors. Both follow the same rule. Every fact must trace back to an official source, and every recommendation must be something you can check yourself.

Sources we use

We build our pages from primary sources only:

  • Z.ai documentation at docs.z.ai: model guides, the API reference, the pricing page, error codes, release notes and the Coding Plan documentation.
  • The Z.ai blog, where each flagship launch is announced with its benchmark tables and architecture notes.
  • GitHub: the zai-org repositories, including READMEs with download tables and serving instructions.
  • Hugging Face: the zai-org model cards and the license files that ship with the weights.
  • Official competitor pricing pages, such as DeepSeek’s API docs and Anthropic’s Claude pricing page, for comparisons.
  • OpenRouter’s public model listings, for third-party prices, which we label as OpenRouter’s listed price.

A few rules follow from this. Benchmark numbers are attributed to whoever published them (“Z.ai reports…”), because they come from the vendor, not from us. We do not publish rumored specs or predicted release dates; if Z.ai has not announced something, it is not on our pages. Prices are always given per 1M tokens and tied to the provider that charges them.

We do not use forum threads, social media posts, leaked screenshots or other sites’ summaries as sources for facts. They can alert us to a change worth looking into, but a fact only goes on the page once an official source confirms it.

How we write

Every page opens with the direct answer, then adds the detail: exact model IDs, prices, limits, tables and working code. Code samples use the real endpoints and parameters from Z.ai’s documentation, and they never show a real API key; placeholders such as $ZAI_API_KEY or your-api-key stand in for it. We avoid hype and say plainly when an older model is no longer the best choice.

How pages are updated when Z.ai changes a model

Z.ai releases often. In 2026 it shipped a new flagship roughly every two months, from GLM-5 in February to GLM-5.3 and GLM-5.3-Flash in August. We update our pages when any of these happen:

  • A new model is released. It gets its own page, and the GLM models comparison, the release timeline and the pricing page are updated. The previous model’s page stays online and points to its successor.
  • A price changes. Every table and worked cost example that uses the old price is corrected.
  • Routing changes. For example, on the Coding Plan, requests for GLM-5.2 and GLM-5.1 now go to GLM-5.3, and GLM-4.7 requests go to GLM-5.3-Flash. Every page that describes the plan reflects this.
  • An API parameter changes. For example, GLM-5.3 uses forced thinking, so a request that sends "disabled" fails; code samples use "enabled" with a low effort level instead. Our thinking mode guide tracks these settings per model.

When a change reaches us through a reader instead of an official announcement, we confirm it against the official source before we change the page.

How we test GLM models

Official benchmarks tell you how Z.ai measures its models. Our own test shows how they behave in the chat you can use on this site. Every model is run through five fixed prompts in GLM Chat’s own test harness, three times per prompt, with the same settings as the public chat: temperature 0.7, a maximum of 2,048 output tokens and the lightest reasoning setting the model allows. The models in the battery are glm-5.3, glm-5.3-flash, glm-5.2, glm-5.1, glm-4.7 and glm-4.7-flash.

For each answer, the harness records the full response, the latency and the token counts. Each of the three answers to every task is scored from 0 to 2 with the same checks, and the published score for the task is the median of the three. That is 15 answers per model and 90 for the whole battery.

How we test GLM models: five fixed tasks, public chat settings, 0 to 2 scoring per task

The five tasks

  1. Python. “Write a Python function that removes duplicate rows from a large CSV file, keeping the first occurrence of each row. It must stream the file (constant memory apart from the set of seen keys), support choosing key columns, and preserve the header.”
  2. JavaScript. “Write a debounce(fn, wait, { leading, trailing }) function in JavaScript with leading and trailing options and a cancel() method, plus unit tests using Node’s built-in test runner.”
  3. PHP refactor. The model receives a deliberately messy 60-line PHP function, proc(), with three copy-pasted loops for order types, magic numbers for discounts, string-built HTML without escaping, a SQL insert built by string concatenation, and echo and return mixed together. The prompt: “Refactor this into clean, modern PHP 8 without changing behaviour. Explain the changes briefly.”
  4. Explanation. “Explain transformer attention to a complete beginner in about 150 words.”
  5. Extraction. The model receives a messy product description containing a name, brand, price with currency, dimensions in mixed units, colour options, battery life, a warranty mention and one missing field. The prompt: “Return only valid JSON with the keys name, brand, price, currency, dimensions_cm, colors, battery_hours, warranty_years. Use null for anything not stated.”

Scoring rubric

Each task is scored 0 to 2, for a maximum of 10 per model:

  • 2 = correct and complete: runs or reads as intended, meets every stated constraint, no factual errors.
  • 1 = usable with fixes: one minor bug, a missed constraint, or small inaccuracy that a reader could fix in minutes.
  • 0 = wrong or unusable: does not run, misses the task, invents data, or breaks the output format.

How we apply it: a 2 needs every task-specific check below. If the core of the task is right (the code runs and does the job, the explanation is accurate, the JSON parses with the right values) but a secondary check is missed, the answer gets a 1. A 0 means the core of the task is wrong. We run the code answers ourselves, automatically and for every run: the Python against a test CSV, the JavaScript with its own tests plus our own five behaviour checks under node --test, and the PHP refactor side by side with the original function on every order type. The JSON is parsed and compared field by field, and an editor reads every explanation for accuracy.

Task-specific checks

  1. Python: streams row by row with the csv module, header preserved, key columns supported, first occurrence kept.
  2. JavaScript: leading/trailing both honoured including leading+trailing together, cancel works, tests run with node --test and pass.
  3. PHP: behaviour identical for all three order types incl. VIP discount, output escaped, parameterised SQL, duplication removed.
  4. Explanation: technically accurate (queries, keys, values, weights), beginner-friendly, 120–180 words.
  5. Extraction: valid JSON, exact keys, values correct, units converted to cm, null where the text is silent, nothing invented.

What each task reveals

The Python and JavaScript tasks test whether a model respects every constraint in a precise spec, not just the obvious one. The PHP refactor tests whether it can improve code without changing what it does, which is the everyday reality of maintenance work. The explanation tests accuracy under a tight word limit. The extraction tests discipline: valid JSON, no invented values and correct unit conversion.

Why three runs and a median

At temperature 0.7 the same model can give a strong answer to a prompt one time and a flawed one the next. A single run can flatter a model or punish it for one unlucky sample. Three runs with the median score show what you will usually get: one bad or unusually good answer out of three does not move the result.

  • Median, not average. Each task’s published score is the middle of its three run scores, so it is always a whole number (0, 1 or 2) and a model total stays out of 10. The individual run scores are shown next to each median.
  • Latency and tokens are medians too. The latency and output-token figures in each results table are the median of the three runs for that task.
  • “Why” describes the pattern. The Why column explains the failure that shows up across the runs, not the quirk of a single answer.
  • The result card shows the best answer. The chat-style card on each model page shows that model’s best-scoring answer across all runs, so you can see what it does at its best, while the table shows what it does typically.
  • Every run is kept. All 90 answers, with their latency, token counts, scores and notes, are stored with the published results, so any score can be traced back to the answers behind it.

How results are published

Results appear on each model page and side by side on comparison pages such as GLM vs Claude, always with the number of runs, the median method and the date the battery ran. GLM-4.6, GLM-4.5, GLM-4.5-Air and GLM-5 are outside the six-model battery, so their pages show the full results table for the models we test. Because the test uses the public chat settings, the results show what you get when you ask the same model here. They are a practical snapshot from three runs per task, and they sit next to Z.ai’s published benchmarks rather than replacing them.

Corrections

If you find an error, email info@glm-ai.chat with the URL of the page and a link to the official source that shows the correct information. We compare the report with that source. If the page is wrong, we fix it on every page where the same fact appears, not only the one you reported. Corrections to prices, limits, model IDs and code get priority, because those are the facts readers act on. If the official source itself is unclear, we say so on the page instead of guessing.

More ways to reach us are on the contact page.

Advertising and editorial separation

GLM Chat is free and funded by ads served through Google AdSense. Ads are placed by the ad network and are kept separate from the writing. No company pays for coverage, rankings, test scores or recommendations, and advertisers do not see or approve pages before they are published. When we recommend a model, a free option or a setup, it is because the facts on the page support it.

Use of AI

Our pages are drafted with AI assistance and edited by the GLM Chat Team. AI helps us turn long official documents into readable guides and working code samples. Editors match every number, model ID and parameter to its official source before a page goes live, and test scores are assigned by an editor using the rubric above, never by a model.

Who writes GLM Chat

The GLM Chat Team writes, tests and maintains the site. You can read more about GLM Chat, see how we handle data in the privacy policy, or try the models yourself in the free GLM chat.