Every fact in this article comes from the official Z.ai announcement: "GLM-5.3: Frontier Coding with Emergent Cyber Capabilities". No rumors, no leaks โ just what Z.ai published, organized for developers.
Z.ai shipped GLM-5.3, and their one-line summary is unusually blunt: "Scaling post-training is all we did for GLM-5.3." It's the same base model as GLM-5.2 โ every gain comes from post-training. The result is what Z.ai calls their most capable open-weights model for coding, plus an emergent cyber capability for vulnerability discovery that they themselves describe as developing "faster than we expected."
GLM-5.3 keeps the GLM-5.2 stack: IndexShare for efficient long-context processing and SAO for reinforcement learning on long-horizon tasks. What changed is the post-training recipe. Z.ai scaled environments toward real units of expert work, on the premise that "scaling toward real units of expert work translates to generalizable performance gains" more reliably than increases in pretraining scale.
On the infrastructure side, they built synthesized environment pipelines to generate larger training corpora, and evolved SAO with context compression to reduce the amount of context that needs to be stored and processed. They also re-implemented their RL codebase in slime โ a minimalist, open-source RL framework โ doubling throughput to 2.3ร the RL training throughput of the previous stack while keeping results numerically equivalent (logprob diff within 1e-7).
Z.ai's stated problem with naive agentic-RL scaling: without deliberate environment design, much of the learning signal is spent memorizing surface-level patterns that don't generalize. So they focused on environments that mirror expert workflows and synthesize larger corpora instead of scaling generic tasks. The Coding subset of their benchmark table:
| Benchmark | GLM-5.3 | GLM-5.2 |
|---|---|---|
| Terminal Bench 2.1 (Overall) | 88.2 | 81.0 |
| Terminal Bench 3.0 | 28.3 | 4.6 |
| DeepSWE v1.1 (K = 4) | 66.9 | 46.2 |
| NL2Repo Real-World (K = 4) | 58.0 | 48.9 |
| FrontierSWE v1.0 (K = 4) | 78.1 | 67.5 |
| SWE-Marathon | 42.5 | 19.4 |
| PostTrainBench (Coding) | 39.8 | 31.7 |
Token-efficiency, measured as the fraction of expert tokens the model needs to match its final score on PostTrainBench Coding: GLM-5.3 uses 34.5% of expert tokens at ~75K average tokens, versus 23.4% at 96K for GLM-5.2. For comparison, Z.ai reports Claude Opus 4.8 at 31.4% at High thinking effort, and GLM-5.3 again at 29.5% with 120K average tokens โ while noting they still trail Claude Fable 5 (39.5%).
The announcement's most interesting section. As post-training scaled, cyber capability developed faster than Z.ai expected. Their framing: this capability is vital for "finding bugs before bad actors can leverage them." They released the full evaluation, including cases where other models score higher, "because it's the right thing to do."
| Benchmark | GLM-5.3 | GLM-5.2 |
|---|---|---|
| CyberGym (Vulnerability Discovery) | 84.5 | 77.2 |
| ExploitBench (Exploit Discovery) | 54.4 | 24.4 |
| ExploitGym (Num. Pwned) | 105 / 130 | 29 / 39 |
On CyberGym, that 84.5% is the best result Z.ai reports โ ahead of Mythos 5 (83.8%) and GPT-5.6 Sol (83.6%). On ExploitGym, Mythos 5 remains ahead at 181/247.
The real-world test is the stronger signal: working with open-source projects and security teams (1NG, Trail of Bits), GLM-5.3 scanned 269 projects and identified 2,436 vulnerabilities โ 1,097 of them medium-to-high severity. The flaws span roughly 45 years of software history, the oldest dating back to 1981, with an average age of 26.6 years. By severity: 107 Critical, 990 High, 1,286 Medium, 53 Low. Affected projects were notified privately and given time to patch before publication; reports are tracked in the new Z.ai Security Disclosure Ledger.
| Benchmark | GLM-5.3 | GLM-5.2 |
|---|---|---|
| Toolathlon | 73.0 | 59.9 |
| AutomationBench (Soft.) | 48.2 | 26.2 |
| Agents' Last Exam (CLI) | 28.5 | 23.8 |
| HLE w/ Tools | 62.5 | 54.7 |
| GDPval-AA (Avg Reward) | 1,769 | 1,508 |
General reasoning (PostTrainBench Reasoning): 63.3 vs 61.4. And one honest step back, as disclosed by Z.ai: MMStar dropped to 60.2 from 66.3 โ the trade-off of the post-training focus.
low, high, and max; the disabled option is no longer supporteddisabled will error โ check your integrations before migratingGLM-5.3 has been rolled out to all GLM Coding Plan subscribers and works in ZCode, Claude Code, OpenCode, and more. The GLM Coding Plan is now points-based: subscribers receive a monthly quota that can be allocated to any GLM coding model โ at half cost during off-peak hours. ZCode subscription benefits include early access to new GLM releases and priority access during peak hours.
Same base model as GLM-5.2 โ but the post-training jump makes GLM-5.3 Z.ai's strongest coding release yet, with an emergent cyber capability for vulnerability discovery that's already finding real bugs.
๐ป 50% improvement over GLM-5.2 on Z.ai Code Bench; open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam
๐ก๏ธ CyberGym 84.5% โ state of the art for vulnerability discovery
๐ 2,436 vulnerabilities found across 269 real projects, tracked in the Z.ai Security Disclosure Ledger
๐ Open source โ weights in two weeks, after safety evaluation and hardening