enjoyed going through this back and forth between goodalexander and beni on
$CBRS
prefer more of this kind of content and discussion re markets
had jippity make a ld style flow for the graphically inclined
No problem, I have no problem with people who don't agree with me, especially when it's a fair thread like yours right here
Some of these are very legit corrections and I am a little embarrassed about them kek, but some I think you're overreaching on, so going point by point:
1) Jane Street / $200M per MW
Small distinction: Dylan has talked about Jane Street extracting roughly $300–500M/MW of value, but the source for the reported $200M/MW payment is Gavin Baker, who also wrote that Jane Street is “reportedly paying $200 million per MW.”
So yes, I should say reportedly paying ~$200M/MW, not present it as a disclosed contract that's my bad.
Dylan also replied directly to my post about the $200M/MW economics and challenged the assumption that all the capacity was delivered and fully utilized, which sure I guess but he did not challenge the reported price itself.
2) I don't really agree with this one.
Astra Ultrafast proves NVIDIA can serve frontier intelligence quickly sure but It does not prove NVIDIA has closed the Cerebras speed gap whatsoever...
Cerebras can still run any model multiples faster than Nvidia on decode alone... But when you use end to end latency the results are even more telling.
I'm also not claiming Sol is smarter than Astra. That's a separate question. My point is that Cerebras demonstrated that you can take a frontier model and run it at a speed that is still in a different class from conventional GPU serving.
So Astra Ultrafast is absolutely competition.
But “NVIDIA can also offer a fast frontier model” is very different from “NVIDIA has erased Cerebras' latency advantage.”
It hasn't.
3) OpenAI does not “buy the chips”
Correct. Bad wording on my end.
The disclosed arrangement is a 750MW inference-capacity/services agreement. Cerebras owns and deploys the infrastructure and carries the associated capex. The disclosed service tranches run for three or four years, extendable by OpenAI to five, and OpenAI also provided the ~$1B secured working-capital loan.
So yes, describing that as OpenAI simply “buying the chips” was wrong.
That actually makes Cerebras' return on deployed capital more important, not less.
4) 7.5×
Yeah I fucked up here lol.
$200M / $7.5M = 26.7×, not 7.5×.
I mixed annual and lifetime economics in one sentence.
Impressive stuff
5) Was Sol Ultrafast publicly available and then pulled?
Also fair.
OpenAI launched Sol Ultrafast as a limited preview to selected customers, not broadly available public capacity that was later withdrawn.
OpenAI literally said access would expand “as capacity grows.” So my wording implying a generally available Cerebras endpoint that had been pulled because they ran out of capacity is wrong. Was not even what I was trying to say but I guess when you speedrun through 8.5k words you fuck up along the way that's my bad.
The capacity point can stand without saying that
6) NVIDIA- “not the gotcha you think”
This is what I should have explained in the article instead of teasing it and forgetting lmao
What I meant is pretty simple: Astra Ultrafast proves NVIDIA can serve frontier models quickly, but it does not show that Blackwell has closed Cerebras’ latency advantage. Same-model comparisons still put Cerebras materially ahead on single-user decode, and on some long reasoning workloads the gap gets even larger end-to-end. Cerebras’ Llama 3 70B comparison, for example, showed >21× lower end-to-end latency than B200 Blackwell on a 1,024-token input / 4,096-token output workload. Even with very low batching the numbers won't be close.
I heard rumors on later models, but I have not been able to test them for myself so I did not and was not planning on mentioning it
7) Groq 3 LPX, Jalapeño, Taalas etc - I disagree here
Groq 3 LPX is good on small models , but not so much on the ones that actually matter. It has massive physical interconnect bottlenecks when handling high user concurrency at long context windows. But, sure credit where credit is due they're good at very specific things like minimizing wait time which is useful for voice etc
But cerebras doesn't absolutely suck either at that...
Jalapeno is great and OpenAI get's to write all of their inference serving in pure Jalapeño ISA which makes a significant difference. But it certainly is not a chip built to be fast mode premium tier either, it's made to run a shit ton of inference on at scale at a good speed.
Taalas is the better option of the competitors but it has a load of issues of it's own mainly that it's frozen to a model architecture and that RSI will make new models come out quicker and quicker which is a major problem
TLDR: They're not there yet, and even if they were there is too much demand for one entity.
8) Margins
This is probably your strongest financial point.
Core GM went from 46.5% in Q1 to 40.6% in Q2, and Q3 is guided to roughly 38–40%.
That is not what I would point to as evidence that Cerebras is already capturing monopoly scarcity rents either
There is an important qualification though.
Management said approximately 500bps of the sequential pressure came from renting back systems to expand private-cloud capacity. Take that out and Q2 would have been around 45.6%, basically around the Q1 level.
Management also says Q3 should be the trough and expects improvement as those rented systems are replaced with lower-cost Cerebras-owned capacity.
Now the explanation is something they actually have to prove though sure. But so I wouldn't characterize the gross-margin decline as evidence that premium pricing is disappearing...
Granted again, I shouldn't use hypothetical future premium pricing as though the margins have already arrived. Demand scarcity and shareholder economics are admittedly two different things lol
9) “TPS is a dogshit metric” while writing an article about speed
No contradiction here. TPS is a dogshit cross-model universal metric.
It's not my fault that the market rewards the wrong thing lmao, I'd use the better metrics if I could they just don't exist for the most part so I have to go with what I have lol
What was I supposed to do, invent the numbers they refuse to release because they can't be benchmaxxed?
When you're comparing the same model, same weights, same reasoning configuration, same workload, then sure generation rate becomes relevant.
But the better metric I actually care about is something closer to time/cost per successful task.
10) Anthropic
I think you are severely underestimating the importance of speed for RSI development...
But I should've actually said that it was the reason why I was bearish instead of the biggest blanket statement the world has ever seen, my bad g
They clearly can compete
But if OpenAI can use Cerebras chips to train the next gen model, regardless to what Anthropic is using, they're at a severe disadvantage... And I am very confident about that
Lastly
"scheduled unlocks". come on you are in crypto.
Dude we're talking about boomers that have been working on this for 10y lmfao let them get some fucking cash they earned it
"VC cost base is $15-89, do you think they're going to hold? I don't. I think they're going to sell, maybe you're right about revenue coming in 13x the street in FY2028. above my paygrade. my base case is you'll get a very good dip to buy."
First of all mate I'll be plunge protection, so it can't happen
But if it were
I am fine holding
Let me hold my spot in piece now that I stopped being a professional gambler
PS: It is clear that I did not use enough Due Dilligence before sending this out, was under time constraint and didn't wan't to postpone but I should've. Many mistakes shouldn't have happened and honestly I am not happy that they did
But you live and you learn aye
Thank you for your feedback, it will always be welcome