NIST’s Assessing AI for Cybersecurity Information (CAISI) evaluated Z.ai’s GLM-5.3, released on August 14, 2026, with its weights made public two weeks later. CAISI concluded that GLM-5.3 was the most cyber-capable open-weight model it had evaluated, but that its capabilities remained significantly below those of current U.S. frontier models. On CAISI’s composite cyber-capability measure, GLM-5.3 trailed the current U.S. frontier by about four months; the comparison covers released models and excludes developed but unreleased systems. The evaluation used four benchmarks spanning vulnerability discovery and exploit development: SEC-Bench Pro, ExploitBench, ExploitGym (Userspace), and CAISI OSS-Fuzz. GLM-5.3 solved 74 of 183 SEC-Bench Pro tasks, scored 9.8 of 16 on ExploitBench, succeeded on 47 of 498 ExploitGym tasks, and completed 23 of 297 OSS-Fuzz tasks. The corresponding U.S. frontier best results were 165 of 183, 16 of 16, 223 of 502, and 69 of 297, while the PRC frontier best results were 50 of 183, 5.1 of 16, 13 of 502, and 7 of 297. CAISI tested the models as agents in a ReAct harness with bash and Python access, maximum reasoning settings, benchmark-specific turn limits, and disabled cyber safeguards for applicable U.S. models. To combine results across tasks of different difficulty, CAISI fitted a one-parameter logistic item-response model that estimates latent model capability and task difficulty. Its index is calibrated so that a 400-point increase corresponds to tenfold higher statistical odds of solving a task. CAISI placed GLM-5.3 above the previously evaluated Kimi K3 and below the current U.S. frontier, with estimated differences outside the 95% confidence intervals.
