A new coding AI from Zhipu claims the top spot in vulnerability detection, but falls short in exploitation tasks. The company plans to release open weights after safety checks, raising questions about global access and security.
Zhipu’s latest coding model, GLM-5.3, is making headlines for its performance in cybersecurity benchmarks, but the story is more nuanced than the top-line numbers suggest. The Beijing-based company, also known as Z.ai, launched GLM-5.3 on August 14 and released a technical note comparing its results to leading American models. While much of the coverage focused on GLM-5.3’s strong showing in vulnerability discovery, the full data reveals a mixed picture across different security tasks.
On the CyberGym benchmark, GLM-5.3 scored 84.5%, edging out Anthropic’s Mythos 5 at 83.8% and OpenAI’s GPT-5.6 Sol at 83.6%. This result fueled reports that a Chinese model now leads in bug hunting. However, Zhipu’s own release notes that this is just one of three cybersecurity benchmarks, and the other two show GLM-5.3 trailing its American competitors. On ExploitBench, which tests a model’s ability to reason about and exploit real vulnerabilities, GLM-5.3 achieved 54.4%, a significant jump from its predecessor’s 24.4%, but well behind Mythos 5 at 78.0% and GPT-5.6 Sol at 76.5%. In ExploitGym, which measures how many exploitation tasks a model can complete within a set time, GLM-5.3 finished 105 tasks in two hours and 130 in six, compared to Mythos 5’s 181 and 247, respectively.
These results highlight a key distinction: finding vulnerabilities and building working exploits are separate challenges. Zhipu acknowledges that its model’s capabilities diminish the further along the exploitation chain the task sits. The company’s transparency stands in contrast to some American labs, which restrict access to similar research due to security concerns. Notably, Zhipu intends to release GLM-5.3’s weights for public download after completing safety evaluations, a move that could broaden access to advanced vulnerability discovery tools worldwide.
Some confusion in media coverage stems from Zhipu’s use of different Anthropic models for various comparisons. The main benchmark table pits GLM-5.3 against Opus 4.8, performance charts use Fable 5, and the cybersecurity section references Mythos 5. This patchwork of comparisons makes it difficult to draw a single, clear conclusion about GLM-5.3’s standing. On coding tasks, the results are similarly mixed: GLM-5.3 leads Opus 4.8 in some tests but trails Claude Fable 5 on Zhipu’s internal benchmarks.
Methodology details in Zhipu’s release also deserve attention. The company evaluated GLM-5.3 using Anthropic’s Claude Code 2.1.207 agent, ensuring a consistent testing environment across models. However, the CyberGym score is based on a single run across 1,507 tasks, with no variance reported, and the ExploitGym time budgets were normalized using throughput rates from Artificial Analysis, with some rescaling factors missing for Mythos 5. These factors suggest caution in interpreting small differences as definitive leads.
Beyond benchmarks, Zhipu collaborated with Chinese security teams to test GLM-5.3 on real-world codebases, identifying 2,436 vulnerabilities across 269 open-source projects. Of these, 107 were labeled critical, 990 high, 1,286 medium, and 53 low severity. The oldest flaw dated back to 1981, with the average vulnerability remaining undiscovered for 26.6 years. However, only 53 findings have been publicly disclosed, while 2,383 remain under embargo. The release does not specify how many vulnerabilities were previously unknown or independently reproduced, leaving open questions about the model’s real-world impact.
Efficiency and distribution may prove more consequential than benchmark margins. Zhipu reports that GLM-5.3 achieves 31.4% on its internal coding benchmark using about 50,000 output tokens per task, compared to Opus 4.8’s 29.5% at 120,000 tokens. This suggests better performance at lower computational cost, which could make advanced security tools more accessible to organizations with limited resources. The company’s commitment to open weights, pending safety review, could further democratize access-especially in regions where American models are export-controlled. The weights are expected to be released by the end of August.
As AI models increasingly demonstrate capabilities in sensitive domains, questions about responsible release and potential misuse are growing. Recent research has shown that generative AI can outperform humans in trust-building for scams, as highlighted in a study on AI chatbots and romance scams. The debate over open access versus security restrictions is likely to intensify as models like GLM-5.3 approach or surpass the frontier in critical tasks.
Zhipu, founded in 2019 and headquartered in Beijing, has positioned itself as one of China’s leading AI labs competing with American giants. The company has raised over $400 million in funding and employs more than 500 staff, focusing on large language models and AI infrastructure. Its previous models have been adopted by major Chinese tech firms and research institutions, signaling growing domestic demand for homegrown AI solutions.