China’s Z.AI Unveils GLM-5.3, Says It Leads Open-Source AI Models in Key Benchmarks
Xu Wei
DATE:  6 hours ago
/ SOURCE:  Yicai
China’s Z.AI Unveils GLM-5.3, Says It Leads Open-Source AI Models in Key Benchmarks China’s Z.AI Unveils GLM-5.3, Says It Leads Open-Source AI Models in Key Benchmarks

(Yicai) Aug. 14 -- Chinese large model developer Z.AI, formerly known as Zhipu AI, today released GLM-5.3, claiming it is the highest-ranked open-source model across multiple mainstream benchmarks, with coding and artificial intelligence agent capabilities approaching those of Claude Fable 5, one of Anthropic’s most advanced AI models.

GLM-5.3 retains the same base model as its predecessor GLM-5.2, but extensive post-training scaling has significantly raised its intelligence ceiling, Beijing-based Z.AI said. The open-source model has stronger coding and cybersecurity capabilities and delivers a better vibe-coding experience than other Chinese-developed models, it added.

Z.AI's high benchmark rankings are notable because its open-source model is narrowing the gap with leading proprietary models from Anthropic and OpenAI, including on advanced coding and AI agent tasks. Unlike proprietary models, open models give developers greater freedom to customize and deploy them independently. Other Chinese open models include Moonshot AI’s Kimi K3 and DeepSeek’s V4 Pro.

GLM-5.3 scored 28.3 percent on Terminal-Bench 3.0, which measures a model’s ability to complete complex tasks in a real terminal environment, up from 4.6 percent for its predecessor, according to Z.AI. That compares with the 42.7 percent scored by Anthropic’s closed-source Claude Opus 5, which ranks first on the benchmark. GLM-5.3’s score on DeepSWE v1.1, which focuses on long-horizon software engineering and sustained code modification, jumped to 66.9 percent from 46.2 percent, versus the leading 74 percent achieved by Claude Opus 5, a new model launched less than a month ago.

On Agents’ Last Exam, which covers multiple real-world professional scenarios and emphasizes cross-tool collaboration and long-horizon tasks, GLM-5.3’s score rose to 28.5 from 23.8.

GLM-5.3 also scored 1,769 on GDPval-AA v2, which covers 44 occupations and evaluates real-world, high-value knowledge work, about 4 percent below Claude Opus 5 Max’s score of 1,849. The result demonstrates GLM-5.3’s ability to perform professional tasks using capabilities emerging from its coding skills, Z.AI said.

AI Model Race Heats Up

Competition among large model developers has intensified recently, with a number of new models being launched.

DeepSeek released V4 Pro early yesterday and made it available for developers through its application programming interface. The new model, DeepSeek-V4-Pro-0813, has enhanced AI agent capabilities.

Earlier this week, SpaceXAI, the AI division of Elon Musk’s SpaceX, released Grok 4.6, an upgrade from Grok 4.5 that focuses on improving performance in long-horizon tasks and the ability to handle more complex interactive, visual, and knowledge-work tasks, the American company said.

Despite the new launch, Z.AI's shares [HK: 2513] closed 3.6 percent lower at HKD1,270 (USD162) today, though they remain almost 11 times above their January initial public offering price.

Editor: Emmi Laine

Follow Yicai Global on
Keywords:   AI Large Model,Z.AI Co.,Open-source model,Claude,Anthropic,China,Zhipu AI,LLM,benchmarks,GLM-5.3