Ai News
Ai News

Z.ai Releases GLM-5.3 Open Weights After Safety Hardening

Published Aug 28, 2026 Sources checked Aug 28, 2026

Z.ai has published GLM-5.3 FP8 and BF16 weights after the two-week safety-review window announced at launch, enabling local deployment of its 744B-A40B model.

GLM-5.3's model launch and its weight release are separate milestones

Z.ai launched GLM-5.3 on August 14, 2026 through its hosted services and GLM Coding Plan, but deliberately did not publish the model weights at the same time. The company said it would hold the weights for roughly two weeks while it completed additional safety evaluation and hardening because cybersecurity capability had advanced faster than expected during post-training.

That second milestone has now happened. Z.ai's official GLM-5 repository lists downloadable GLM-5.3 FP8 and GLM-5.3-BF16 checkpoints on Hugging Face and ModelScope, and the official Hugging Face repositories were updated on August 28.

This means the main GLM-5.3 model is now available for self-hosted deployment and further work under its model license. It should not be confused with GLM-5.3-Flash, a separate 320B-A18B natively multimodal model whose weights were released on August 26.

The published checkpoint is a 744B-A40B mixture-of-experts model

Z.ai's official repository describes GLM-5.3 as 744 billion total parameters with 40 billion active parameters. The FP8 and BF16 weight variants are both listed for download.

The company says GLM-5.3 uses the same base model as GLM-5.2 and that the capability gains come from scaled post-training rather than a new pretraining run. That distinction matters: GLM-5.3 is primarily a post-training upgrade focused on coding, long-horizon agentic work and cybersecurity rather than a new underlying base architecture.

For local serving, Z.ai currently documents support paths through SGLang, vLLM, Transformers, KTransformers and Unsloth, with additional deployment guidance for Ascend NPU environments. The model also exposes reasoning_effort levels of low, high and max, with max as the documented default.

A checkpoint this large remains data-center-class infrastructure at full precision. The open-weight release expands deployment choice, but it does not turn GLM-5.3 into a typical single-GPU local model.

Z.ai reports large coding gains from post-training

At the August 14 launch, Z.ai reported a 50% improvement over GLM-5.2 on its private Z.ai Code Bench. On public evaluations in the company's published configuration, GLM-5.3 scored 28.3 on Terminal-Bench 3.0 versus 4.6 for GLM-5.2, and 66.9 on DeepSWE v1.1 versus 46.2.

Those numbers are vendor-reported and depend on Z.ai's stated harnesses, context limits, sampling settings and tool configurations. The open-weight release makes independent testing easier, but it does not by itself independently validate the launch benchmarks.

Cyber capability explains the unusual release sequence

The notable part of the GLM-5.3 rollout is that Z.ai explicitly tied the delayed weight publication to dual-use cybersecurity capability.

The company says vulnerability-discovery training produced broader exploitation-chain reasoning than expected. In its published evaluations, GLM-5.3 reached 84.5% on CyberGym and 54.4% on ExploitBench, while completing 105 ExploitGym tasks within two hours and 130 within six hours under the company's normalized evaluation setup.

Z.ai also says that, after expert review, screening and deduplication, its security work identified 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high severity findings. These are company-reported results and should not be treated as an independently audited vulnerability census.

The extra safety window does not mean the open checkpoint is risk-free. It means Z.ai chose a staged release: hosted access first, followed by model weights after additional evaluation and hardening.

What is available now

Released now: the GLM-5.3 FP8 and BF16 weights are listed in Z.ai's official repository with Hugging Face and ModelScope download links, and the Hugging Face model pages are live.

Already available before the weight release: hosted GLM-5.3 access and GLM Coding Plan integration date back to August 14.

Separate product: GLM-5.3-Flash is a different, smaller natively multimodal model and was already open-weight before this main GLM-5.3 checkpoint release.

The practical change on August 28 is therefore not a new GLM-5.3 model announcement. It is the transition of the flagship from hosted-only availability to an openly downloadable weight release, enabling independent deployment, evaluation and adaptation at a scale appropriate for very large mixture-of-experts models.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books