Skip to content
← All news
5 min read

Z.ai shipped GLM-5.3 and said its security skills outran the plan

The new flagship doubled its exploit benchmark score, found 2,436 vulnerabilities across 269 projects, and ships open weights in about two weeks.

Z.ai says GLM-5.3's exploit score more than doubled in post-training. It found bugs up to 40 years old.

Z.ai released GLM-5.3 on August 14 with an unusual admission in its own launch post: the model's cybersecurity ability grew faster than the lab planned for. The base model is unchanged from GLM-5.2. Every gain came from scaled-up post-training, and the security numbers moved the most.

The numbers Z.ai is reporting

  • ExploitBench: 54.4%, more than double GLM-5.2's score.
  • CyberGym: 84.5%, up from 77.2%, ahead of every model in Z.ai's comparison set.
  • 2,436 vulnerabilities found across 269 open source projects during testing with security teams, some in code roughly 40 years old.
  • Every finding tracked in a public Z.ai Security Disclosure Ledger.

VentureBeat reports that one flagged finding was a serious vulnerability in Cursor, the AI code editor. Z.ai says it is holding the open weights back for about two weeks of safety evaluation and hardening. The model is live now through the API and the GLM Coding Plan.

Emergent here means unplanned: the lab scaled post-training for coding, and exploit-finding ability came along at a steeper curve than the coding metrics.

The dual-use tension, stated plainly

A model that finds 40-year-old bugs in open source is a gift for maintainers and a tool for attackers, and the same benchmark measures both. Z.ai's answer is the disclosure ledger plus the delayed weights release. Whether two weeks of hardening changes what an open-weight model can later be fine-tuned to do is an open question, and the field's honest answer so far is: not much. What the delay does buy is patch time for the disclosed findings.

One more caveat. Until the weights ship, the ExploitBench and CyberGym numbers are Z.ai's own runs, not independent ones. Treat them the way you treat any vendor benchmark: a claim with a release date attached.

Why a build studio cares

Software audits are our flagship service, and models that find 40-year-old bugs change what an audit can cover in a week. Vulnerability-hunting models are entering the attacker toolkit and the defender toolkit in the same quarter. If you maintain anything public, assume a model at this level scans it within months. The cheap move is to run that scan on yourself first.

Next step: pick your oldest, least-touched repository and point a current coding model at it with one prompt: find memory-safety and injection issues, rank by exploitability. Reading the output costs an afternoon and tells you what a stranger with the same model already knows. Want that pass run properly, with findings verified by senior engineers? Write to us at hello@gattyworks.com.

GLM-5.3Z.aiOpen WeightsAI SecurityGLM53ZAIOpenWeightsAISecurityCyberSecurityFrontierAICodingModelsVulnerabilityResearchOpenSourceAIAINews

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.