Voight we run GLM in production across our agent fleet. When
@Zai_org ships a new GLM, that is infrastructure news for us. We went through GLM-5.3 and this is what actually matters:
The release is unusual: same 743B base as GLM-5.2, zero new pretraining. The entire jump came from post-training.
· DeepSWE v1.1: 46.2 → 66.9
· Terminal-Bench 3.0: 4.6 → 28.3 (6x)
· CyberGym: 84.5%, above GPT-5.6 Sol
· ExploitBench: 24.4% → 54.4%
For teams running agents in production, the key stat is efficiency: 31.4% on Code Bench High, beating Claude Opus 4.8 (29.5%), at ~50K output tokens per task instead of ~120K. Tokens per task is the unit of agent economics.
The frontier still leads (Fable 5: 39.5% at Max effort). But the open-weight gap has never been this thin.
1M context. 128K output. MIT weights expected August 28.
When the weights drop, this class of model runs on your own GPUs. For our fleet that is the roadmap: smarter post-training, cheaper tokens, inference on hardware we control.