Z.ai released GLM-5.3 as its latest flagship model, saying the update improves complex software engineering and agent tasks through post-training on the same base as GLM-5.2. In company documentation, GLM-5.3 reaches state-of-the-art results among open-source models on several coding and agent benchmarks and shows sharp gains on cybersecurity evaluations.

According to Z.ai’s model guide, GLM-5.3 scores 84.5% on CyberGym, which tests whether a model can identify and validate vulnerabilities from white-box source code by triggering faults. That edges Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. ExploitBench rose from 24.4% on GLM-5.2 to 54.4% on GLM-5.3, though Mythos 5 and GPT-5.6 Sol remain ahead at 78.0% and 76.5%.

Coding gains and a clear exploit gap

Z.ai said GLM-5.3 improves about 50% over GLM-5.2 on Z.ai Code Bench and posts open-source state-of-the-art results on Terminal-Bench 3.0 and Agents’ Last Exam. The company also reported stronger professional-task performance on GDPval-AA v2. On ExploitGym, which measures exploitation tasks under time-normalized budgets, GLM-5.3 completed 105 tasks in two hours and 130 in six hours, up from 29 and 39 for GLM-5.2, while Mythos 5 remained ahead at 181 and 247.

Z.ai’s own framing is careful: capability is growing fastest where the model is still furthest behind closed systems, especially deeper in the exploitation chain. The API for GLM-5.3 is described as coming soon, with the model already available to GLM Coding Plan users.

Decoded Take

GLM-5.3 is another marker that open-source Chinese models are competitive on discovery and coding loops, while closed frontier systems still dominate full exploit workflows. That split matters for policy and procurement: vulnerability finding and vulnerability weaponization are not the same capability. Watch how Z.ai handles public release safeguards, and whether Western labs answer with more transparent cyber evals of their own.