GMT+8. Personal rants. Priors and opinions may change overtime. Powered by coffee.

DeepSeek v4.1 Flash!
18
If only Z.ai can make GLM 5.3 & -Flash cheaper, undercutting Deepseek, then I think most people would switch to it rather than Deepseek. Right now, they're more expensive in real-world use (bad cache hit rate) compared to Deepseek. GLM models are really good!
37
After spending more time with it, same conclusion, Space Bunny Alpha is a useless time-wasting money-burning model. Don't care what's the actual model behind it, won't use it ever again.
82
What are the chances that these providers are just reselling from the same provider? Like they have a deal for a fixed 5% margin. If they don't, I would assume at least one of them would lower their price to attact volume, esp bcs it's a large model which needs scale.
1
12
If I were to enter the market, I'd definitely undercut the price, if I can. If not, then I'd like offer significantly better speed, at least 2x to offer competitive value. Or, I could also happily resell from Tencent cloud, for a fixed margin, slapping my brand/name on it.
1
8
From the first party provider's perspective, I can see that I'd be happy to share 5% margin to someone to bring me volume. Inference overhead is huge. More volume helps me, plus I can capture US/EU market that are averse of using a Chinese (Tencent) inference service.
7
Space Bunny Alpha is a very wasteful model. It does not care about time/turn spent at all. Just keeps on going, doing nothing meaningful. If this is priced 0.30/0.03/1.20, no one would use this model. Even at 0.10/0.01/0.50, I don't think anyone would find this model competitive.
1
139
Asked it to investigate a bug and how to replicate the bug. Assuming 0.30/0.03/1.20, it burned $1.4 in tokens. Same prompt and repo, GLM 5.3 Flash costs just shy of $0.20.
1
61
But, GLM 5.3 Flash finishes < 20 minutes. Space Bunny Alpha took 3.5 HOURS! It keeps on probing and probing non-stop.
107
Today, I got burned in a new way using OpenRouter!
2
1
26
... care about which provider you have enabled/disabled. It does not show what your actual cost would be. In the past, I got burned because request would get routed to shitty providers that just burns money very very quickly and achieves nothing. Waste of time and money.
1
5
Worked around that, set strict routing guardrails. Boy, got burned again, due to their API. So I was paying $0.30 in and $0.60 out, but Pi calculated it using the $0.02 price. Shitty provider never fails to impress!
4
I love Deepseek, they're the only company truly pushing the cost-performance frontier. But I really don't get why people use V4.1-Flash that much. I find it bad at instruction following. It writes unmaintainable code and speaks weirdly (similar to Luna). It's too RL-ed.
1
31
We have GLM 5.3 Flash and Mimo 2.6 Flash at roughly the same price. It's still one of the top models in OpenRouter and OpenCode. I tried official and 3rd party US hosted providers. Had basically the same experience. Hence me wondering why people hold it at such high regard.
32
GLM 5.3 Flash (max) still beats GPT-6 Luna (max) at roughly the same price!
30
Basically, the safety enabled checker cannot be overridden. There is a bool flag, but it does not work. Their safety checker is not great. Had a client barking up to me due to their customers facing errors. Solution to this problem, open models + @modal.
Companies that force their morality onto others will eventually fail. No one should force what they believe onto others. Comply with local regulations, then stop. Fal used to be great. Horrible company now.
1
27
This is the importance of truly open permissively licensed models. We must not want to live in a world where a handful of companies decide who can and cannot do certain things. Companies should never be the final arbiter of who gets what.
1
14
There's a reason democracy won. There's a reason there's an elected government that forms and upholds the law. Not companies, esp not some founder having a superiority complex.
16