Building the fastest AI inference neocloud

San Francisco
General Compute retweeted
We ran some stats on a GPU Cloud vs General Compute Both on standalone requests and an IRL agent coding session The TLDR: @general_compute is much much faster on both This was using @sambanovaAI's last gen SN40 hardware, so just wait to see how fast the SN50 is Here's the breakdown based on harness and build Standalone: TTFT 8,084ms vs 1,136ms -> 7.1x faster Decode 121 tok/s vs 1,025 -> 8.5x faster One request end to end, 34.3s vs 2.1s -> 16.3x faster The bigger bottleneck at some point for long horizon agents becomes CPU operations and storage access. Which colocating with the accelerators will help solve
1
2
13
768
General Compute retweeted
Speculative decoding is how most production inference already runs The question i get from clients isn't whether it works It's whether it works on ASICs It does, and the gains are bigger than they are on gpus The question usually carries an assumption underneath it: that software is what keeps gpus ahead, so better software eventually removes the case for dedicated inference silicon But the same technique runs on the ASIC, on hardware already built for decode The gap widens, it doesn't close nitter.net/thejessezhang/status/2…
One cool technique to improve latency in LLMs is to have a small, fast model generate tokens and a larger, slower model review the output in one pass. It'd be like having a junior employee write a report, and then having a senior employee review it in one go before it gets sent to a client. This is called "speculative decoding" and lets you generate faster while maintaining the output quality. Great write-up from the @DecagonAI team on how we've been doing this:
1
1
3
475
General Compute retweeted
A great evening at the New York Stock Exchange last Friday as part of our AI Factories series. John Furrier an insightful panel discussion with Darren Chien, managing director, APAC, at @positron_ai; and Jason Goodison, CTO of @general_compute. #theCUBE #NYSEWired
2
3
105
General Compute retweeted
1/ It takes a couple of iterations of physical deployment to figure out chip-market fit. We saw this with Trainium. Gen 1 barely got used; gen 2 actually ran some production workloads, gen 3 is prod-ready. The difference wasn't a smarter chip, it was the years of customer feedback baked into the next one. sambanova.ai/blog/sn50-premi…
3
2
26
2,075
General Compute retweeted
Full house at the @general_compute + @SambaNovaAI + @pipecat_ai voice agents hackathon at @AGIHouseSF. Thank you also to the @GradiumAI team for supporting with speech-to-text and text-to-speech APIs.
8
7
40
2,534
Let me get my popcorn
Time to place your bets — how fast do you think M3 is about to run on @SambaNovaAI 🚀
4
731
🚨 BREAKING: @Cerebas just stopped pitching wafer-scale as a monolith. we got to the supernova cerebras event this week and have some thoughts on their CS-4 launch. it's a beast. but the CS-4 story is not just about tokens per second ... (which cloaked 120b above 4.4k tps per user!) the change is the fastest decode silicon in the industry is now built to sit inside somebody else's cluster, with prefill running on whatever is best at prefill. native disaggregation out of the box - here’s the breakdown... The CS-4 incorporates: - 2.4 Tb/s off-wafer per wafer, double CS-3 - RoCE v2 over standard ethernet, no proprietary switch required - switchless wafer-to-wafer links, 5µs → 2µs - a claimed path to attention-feedforward disagg on top of prefill-decode for a machine whose defining constraint is 44GB of SRAM per wafer, I/O is the release valve. push KV off the wafer and batch goes up without adding capacity. AMD Helios and AWS Trainium are already named as prefill partners. the compute story is exciting too. Cerebras doubled per-wafer performance on the same 5nm WSE-3 die, no tape-out, by moving power conversion to 0.5mm from the processor. roughly 100x closer than a conventional GPU board. twice the power in, roughly double the clock. doubling frequency on 46,225mm² without a superlinear power blowup is hard engineering. 21.6 to 43.2 PB/s per wafer 3 wafers per rack, up from 2 750 PFLOPS, 129.6 PB/s, 7.2 Tb/s per rack over 4,400 tok/s/user on gpt-oss-120b and the rack finally looks like infrastructure. one power rack, three pluggable backpacks, 50% fewer components, no cable plant, cooling pulled out on the assumption the facility is already liquid. deployment in hours. what's still open to be solved by next gen: 44GB per wafer did not move and cannot until new silicon. @SemiAnalysis models that 1.6T params at 256k context and 256 concurrent requires roughly 40 WSE systems. basically relegates the system to large buyers. you wouldn’t want to own just one of these. 2x perf at roughly 2x power is flat perf/watt, so the 10x throughput-per-watt claim is quoting a much higher concurrency point than CS-3 was ever sold at. that's a pareto shift, which is exciting, but it's a different claim than the headline in our opinion. none of that is a knock, the roadmap is telling this story and we are excited about the visibility into next generation launches here at general compute. the takeaway is the direction they have taken. We think nobody wins inference by owning the whole rack. it's prefill and decode on the silicon each phase actually needs, assembled by operators who aren't locked to one vendor. the whole decode silicon market is converging towards disaggregated inference. CS-4 is the biggest vote yet for that architecture, from the company with the most to lose by admitting it. Final thoughts massive congratulations to the @cerebras team - more good silicon in the ecosystem is good for everyone building on it.. Can’t wait to see it in production.
1
9
712
If only someone would build a neocloud for them!
Fireworks and Baseten are the two largest inference clouds and I think of them as *an* index for open-source AI adoption. Also first time seeing SpaceXAI on the list of “fastest growing” IIRC. SambaNova!  Will wonders never cease.
5
360
🔥
This is extremely bullish for ASICs (Like SambaNova SN50) GPU engineers are desperately trying to reproduce the primitives that are built into the ASIC hardware and in effect just validating the market exists If you believe TileRT is cool, imagine it with 10x throughput, 2-4x the speed, and only 10% the energy draw, and that is the advantage you get when you use an ASIC
1
5
1,491
General Compute retweeted
Open-weight AI models are gaining traction. General Compute Co-founder & CTO Jason Goodison explains why companies are making the shift: "These open weight models are something people are using to keep their data private and also get better unit economics and just better AI in general that's optimized for their use cases.” “And that's why we're seeing a big shift to it." Supported by General Compute
1
3
13
13,748
Upper90 has committed up to $400M in debt financing to General Compute, one of the largest facilities ever raised for an ASIC cloud.   It means more capacity for us to build the fastest tokens available anywhere, sooner.   Full story in @TechCrunch today: techcrunch.com/2026/07/17/wh…
1
6
360
General Compute retweeted
There is unlimited demand for blazing fast inference. Excited to support @FPuklowski @GoodisonJason and the @fastinference team as they build the fastest inference neocloud.
3
5
38
8,371