We just launched Wally, our inference stack for open frontier models.
We're not just fast. We're the fastest.
For GLM-5.3 flash, wally has 43% more throughput than Nebius, 3.1x Fireworks, 4.5x Z ai.
The craziest part...
This entire launch video was made by GLM-5.3 Flash running on Wally + Blender.
Meet Wally.
Our inference stack for open frontier models, built to be the fastest place to run them.
Performance snapshot:
GLM-5.3 Flash: 380 tok/s.
GLM-5.3 Max: 790 tok/s
Qwen3.8-27B: 485 tok/s
DeepSeek-V4.1 Flash: 615 tok/s
1/7