GLM-5.3 Beats Kimi K3 on 5 of 6 Tests — But GPT Still Sets the Pace
Z.ai’s new coding model looks stronger than Kimi K3 across most shown benchmarks, but the chart also reveals a ceiling it hasn’t cracked.
GLM-5.3 has a strong answer for Kimi K3 — just not the whole leaderboard.
Z.ai has released GLM-5.3, a 743-billion-parameter coding model the Chinese lab is presenting as the strongest open-weights coder on the market.
The sharper claim comes from a benchmark chart labeled “LLM Performance Evaluation”: GLM-5.3 beats Kimi K3 on five of the six listed tests. The catch is that GPT-5.6 Sol still posts the top score in four of them.
So can GLM-5.3 beat Kimi K3 in the code arena? On this chart, mostly yes. But whether that makes it the open model to beat depends on which coding test matters most.
The Kimi question has a clear-but-not-clean answer
Across the six visible benchmarks — Terminal Bench 3.0, DeepSWE, Agents’ Last Exam, AutomationBench, HLE with Tools, and GDPVal-AA v2 — GLM-5.3 comes out ahead of Kimi K3 in most direct comparisons.
- Terminal Bench 3.0: GLM-5.3 scores 28.3, ahead of Kimi K3’s 17.4.
- Agents’ Last Exam: GLM-5.3 scores 28.5, just above Kimi K3’s 27.6.
- AutomationBench: GLM-5.3 scores 48.2, beating Kimi K3’s 46.7.
- HLE with Tools: GLM-5.3 scores 62.5, above Kimi K3’s 59.8.
- GDPVal-AA v2: GLM-5.3 scores 1769, ahead of Kimi K3’s 1682.
The one loss is important: on DeepSWE, a benchmark for fixing real GitHub issues end to end, Kimi K3 scores 67.5 while GLM-5.3 scores 66.9.
That narrow miss matters because DeepSWE is one of the more directly coding-heavy tests in the set.
Where GLM-5.3 actually leads the field
GLM-5.3 is not just beating Kimi K3 in selected rows. It leads the entire chart on two benchmarks: AutomationBench at 48.2 and GDPVal-AA v2 at 1769.
The jump over GLM-5.2 is also dramatic. On Terminal Bench 3.0, GLM-5.2 scored 4.6; GLM-5.3 scores 28.3. On AutomationBench, it rises from 26.2 to 48.2.
Z.ai says the improvement came from scaling what happened after the base model was built. “Scaling post-training is all we did for GLM-5.3,” the company wrote in its launch post. “Over the past month we kept scaling on this stack: more environments, more diverse tasks, and more compute spent training on them.”
Post-training is the phase after a model’s broad initial training, when it is tuned for specific behaviors such as coding, tool use, or following instructions.
That helps explain the benchmark shape: GLM-5.3 looks especially strong where agentic work and tool-heavy automation are being tested.
The token story may be the bigger product claim
Z.ai is also arguing that GLM-5.3 is more efficient, not simply higher-scoring. The lab says the model reaches 34.5% on Z.ai Code Bench at Max effort while using about 75,000 output tokens per task.
GLM-5.2, by comparison, is listed at 23.4% while using 96,000 output tokens. Tokens are the small units of text a model reads or writes; fewer output tokens can mean less time and lower cost per task.
The model is large: 743 billion parameters. Parameters are the internal settings a model uses to process information.
But Z.ai’s own comparison still leaves room at the top. The company says GLM-5.3 beats Claude Opus 4.8 on token economy, while also noting that it “remains behind Claude Fable 5, which reaches 39.5% at Max effort.”
In other words, the pitch is not pure domination. It is performance plus efficiency.
The ceiling is still set by closed models
The same chart that makes GLM-5.3 look strong against Kimi K3 also shows why the broader race is unresolved. GPT-5.6 Sol leads four of the six benchmarks: Terminal Bench 3.0, DeepSWE, Agents’ Last Exam, and HLE with Tools.
Fable 5 is also ahead of GLM-5.3 on Terminal Bench 3.0, DeepSWE, and HLE with Tools. On DeepSWE, GPT-5.6 Sol posts 72.7, Fable 5 reaches 69.7, Kimi K3 gets 67.5, and GLM-5.3 lands at 66.9.
GLM-5.3 is live through the GLM Coding Plan subscription and ZCode, with API access and downloadable weights expected after a safety review, according to the launch details.
So the answer is: yes, GLM-5.3 appears to beat Kimi K3 on this six-test comparison. But the chart’s harder message is that beating Kimi is no longer the finish line — it is the qualifying round.
