AI

Z.ai Releases GLM-5.3 with Major Coding Gains Without Retraining Base Model

Z.ai releases GLM-5.3, achieving significant coding and cybersecurity gains through scaled post-training without retraining the 743B-parameter base model.

Z.ai has released GLM-5.3, achieving significant coding and cybersecurity gains through scaled post-training without retraining the base model. Terminal-Bench 3.0 improved from 4.6 to 28.3, and CyberGym reached 84.5%. Weights will be released in roughly two weeks.

Z.ai has released GLM-5.3, a new version of its flagship model that achieves significant performance gains without retraining the underlying 743 billion-parameter base model. Every reported improvement comes from scaled post-training: more task environments, more environment types, and longer training.

Coding Benchmarks Show Dramatic Improvements

The results land in two places. Coding jumps most on the longest-horizon benchmarks. Terminal-Bench 3.0 moves from 4.6 to 28.3 against GLM-5.2. DeepSWE v1.1 moves from 46.2 to 66.9. Agents' Last Exam (CLI) moves from 23.8 to 28.5. On GDPval-AA v2, which spans 44 occupations, GLM-5.3 scores 1,769.

On Z.ai Code Bench, an internal evaluation, the company reports a 50% improvement over GLM-5.2. It reports 31.4% at roughly 50,000 output tokens per task. Claude Opus 4.8 scores 29.5% at 120,000 tokens. Claude Fable 5 still leads at 39.5% at maximum effort. Z.ai argues a private benchmark reduces contamination risk. On public suites, GLM-5.3 trails GPT-5.6 Sol and Fable 5 on several harder coding evaluations.

Cybersecurity Capabilities Compound Unexpectedly

Z.ai flags this one as unplanned. It added vulnerability-discovery data expecting better single-bug reasoning. Instead, capability kept compounding as training scaled. The model began forming coherent plans across complete exploitation chains.

CyberGym, which tests discovery and validation from white-box source, moves from 77.2% to 84.5%. That edges past Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. ExploitBench, which requires root-cause reasoning and a working exploit, moves from 24.4% to 54.4%. Mythos 5 sits at 78.0%. On ExploitGym, GLM-5.3 completes 105 tasks in two hours and 130 in six. GLM-5.2 completes 29 and 39. Mythos 5 completes 181 and 247.

Availability and Timeline

GLM-5.3 is live through the Z.ai API, the GLM Coding Plan, and ZCode. Weights are not yet public. Z.ai says it will publish them roughly two weeks after launch, once safety evaluation and hardening finish. The model is available today for startups and mid-market engineering orgs via the Coding Plan or API. Enterprises with data-residency or vendor-review rules should wait for weights. Security vendors and MSSPs get the most signal, and the most policy exposure.