There are 2 open weight models that are dominating the (open weight) Pareto frontier right now, for remote inference: DeepSeek V4.1 Flash and GLM 5.3 Flash
glm53f overall *feels* like a better model. It does not suffer from Claudish as bad as ds41f and feels easier to control. It feels more straight-shooting than ds41f as well, but I still need to benchmark it to be sure
(Notice how ds41f is dominating the agent arena frontier, whereas glm53f is dominating the text arena pareto frontier, see below images)
But for a fair comparison, one must take price and throughput into account
According to the most recent profiling run I ran on Hugging Face inference providers, ds41f on
@baseten is currently ripping at above 200 median tok/s, whereas glm53f is around 50-60% that, depending on which provider you are looking at (OpenRouter UI is reporting different speed for some reason. If you can profile on OR, I would appreciate that in the replies!)
Cached input token prices vary. But in general, ds41f can do this because it has drastically more efficient KV cache characteristics
These two models, glm53f and ds41f currently absolutely mog the "consumer LLM" side of the business, including GPT 5.6 Luna (fast), both in output quality AND speed
This trend in 2026 keeps surprising me. I am in this sector since code-davinci-002 was a thing, and OpenAI has always been at the Pareto frontier on the low end
The king will probably return next week with a Luna that mogs glm53f and ds41f in terms of output quality
We also still haven't seen the full might of cerebras in action. gpt-5.3-codex-spark was ok but it was never a daily driver. are we going to see gpt-6-luna on cerebras ripping through with 1000 tok/s next week? If something like this happens, is it going to be cheaper than ds41f, or more expensive?
Depends on how efficient OAI's next architecture is going to be. Can they beat 890 bytes per KV cache token?
My gut feeling says this: if OAI cannot come up with a model that is both cheaper, faster AND higher quality than these 2 models on the low end, then the implications will be huge
If OAI cannot beat the frontier on the low end with 2.5 months to prepare since Luna's release, then what do you think this will imply? Let me know in the replies below
Excited for next week!