GLM-5.3-Flash, 250 tok/s interactive, $25,000. Who would buy the "tinybox red flash"?

Sep 25, 2026 · 6:17 AM UTC

83
19
952
74,732
Sort replies: Relevant Recent Liked
Replying to @__tinygrad__
Not really worth it tbh. It like having a 90 iq assistant when everyone has a 140 iq one. Economics dont make sense either. The genius wants just 200$ a month, while the retard wants 25k upfront+electricity+maintenance and more 😂
8
66
5,665
It's not a replacement for your Codex subscription. It's a fast, private, and aligned model you can go to for 90% of tasks that both models would do perfectly.
5
117
5,622
Replying to @__tinygrad__
Why everything on openrouter is below 125tok/s and usually like 20 😭😭
2
13
3,520
And often completely unusable time to first token.
31
3,317
Replying to @__tinygrad__
I don't need 250t/s Give me better quants at lower speeds I'm good at 50 bruv Let me get that 8 bit Speed is not the answer. I'm not trying to break things faster.
1
9
2,915
Does the 4-bit quant really impact accuracy much? What's the best study on this? I'm not sure exactly how the Unsloth Top-1% accuracy numbers translate to agentic performance.
1
10
2,479
Replying to @__tinygrad__
Is this real? I'd prob sell some shit and buy this if those numbers are real. If these numbers are real let me crunch. Do you have a pre-order form?
2
4
3,903
We can build it if there's demand. Yea, 4-bit quant, speculative exec. If this tweet gets good traction we'll put up a $1,000 down preorder, shipping by the end of the year.
1
29
4,073
Replying to @__tinygrad__
get GLM 5.3 (not flash) for 20K and maybe
1
70
Replying to @__tinygrad__
That speed sounds incredible for local inference, but the price tag is a tough pill to swallow.
260
How? That’s the price of 4x DGX Sparks but significantly faster.
3
1
844
Replying to @__tinygrad__
I might!
62
Replying to @__tinygrad__
25k is a lot for a box that mostly buys you convenience. is the target home lab people or small teams?
1,017
Replying to @__tinygrad__
Flash full weight ? If so it might be interesting
345
Replying to @__tinygrad__
What if you could do it for $10,000?
3
21
1,574
Replying to @__tinygrad__
People balk at this but will buy an M5 ultra Mac Studio which has 1/5 of the performance and will probably be at a similar price point. I’m probably not the regular user but I would happily pay double this price for something that could run a ~1T model at this speed.
4
445
Replying to @__tinygrad__
tinygrad - what’s the power draw at 250 tok/s?
4
974
Replying to @__tinygrad__
@__tinygrad__ Can we purchase with bitcoin?
1
7
Replying to @__tinygrad__
could you stream on twitch running local models in tinybox?
2
456
Replying to @__tinygrad__
Make a teenybox, I'd buy that
1
841
Replying to @__tinygrad__
This is a better deal than 4x Sparks that go right at 1/2 that speed for what ends up being the same price today. Very interesting. GLM5.3 Flash isn't smart enough to fully replace Opus/GPT6, but it also doesn't say no.
2
223
Replying to @__tinygrad__
How many concurrent sessions?
1
329
Replying to @__tinygrad__
definitely interested...
2
636
Replying to @__tinygrad__
can it run MiMo-V2.6-Flash?
1
65
Replying to @__tinygrad__
faaaakk whats the stats for multistream?
2
838
Replying to @__tinygrad__
Whats parallel request stream capability?
1
160
Replying to @__tinygrad__
its too expensive for the value tbh. you gotta bring the price down somehow george. design the custom hardware and manufacture it at scale idk pull the rabbit out
1
1,007
Replying to @__tinygrad__
I would do this only if i had a bunker and would know that the world ends tmrw, otherwise, it's only going to get better and cheaper. Some people commited to 400k$ a year ago to host some open model and get their money back in 5-10 years, i think they have a bunch of hardware to rent today and better models.
593
Replying to @__tinygrad__
25000? bro that is like 500 years worth of inference in a thing like sail research lol
149