Skip to content

GLM-5.2

MIT

Zhipu AI · 753B · transformer-moe

2026-06-161.0M context753B params

Use Cases

chatcodereasoningmultilingualtoolsmath

Quantization Options

QuantBitsVRAMQualityStatus
Q2_Krec2274.0 GBModerate
Q4_K_M4486.0 GBGood
Q8_08821.0 GBExcellent

About this model

GLM-5.2 is Zhipu AI's June 2026 flagship for long-horizon agentic work, released under MIT license. A 753B parameter Mixture-of-Experts with ~40B active per token, it pairs a usable 1M-token context window with an IndexShare sparse-attention design that cuts per-token compute roughly 2.9x at full context. It scores 62.1 on SWE-bench Pro and 91.2 on GPQA Diamond — frontier-class results from open weights. The official Ollama library carries it only as the `glm-5.2:cloud` remote tag; local runs use community GGUFs (Unsloth). Even the 2-bit quant is a ~254 GB download needing ~274 GB of memory — a 512 GB Mac Studio or multi-GPU server proposition. Everyone else should treat it as a cloud model.

Benchmarks

62.1
swe-bench-pro
91.2
gpqa-diamond