Rodrigo Liang retweeted
One big takeaway from @hotchipsorg: agents are changing the inference game. More reasoning + more tool calls = a LOT more tokens. And across six frontier-model configs, decode made up 75–97% of inference time. So bandwidth & data movement matter. A fast chip is great. A fast system is better. Read the recap: bit.ly/4iacTGu
3
4
1,113
Rodrigo Liang retweeted
Our mixer after this week's Hot Chips was a good time, so we're doing it again! After AI Infra Summit, swing by the Hyatt Regency Santa Clara so you can meet the team, grab drinks, enjoy some food, and learn more about our open roles. September 15 | 6–9 PM Register here: bit.ly/4cjYLa2
4
12
53,481
Rodrigo Liang retweeted
As models get smarter and agentic workloads get more demanding, users expect both intelligence and speed. That’s why inference speed is becoming part of the AI experience, not just an infrastructure metric. Premium inference is here. ⚡ More in our blog: bit.ly/3S3jjMQ
1
4
9
76,246
Super excited!
Our squad just keeps getting stronger. 🦾 @mmoazami is joining SambaNova as Vice Chair of Global Strategy & Partnerships, reporting to @RodrigoLiang. Mohsen has supported some of the world's hottest technology companies, enabling global expansion, forging strategic relationships, and navigating major industry shifts. Welcome to SambaNova, Mohsen! bit.ly/4xk7lxA
1
5
618
Great to chat @jonfortt!
“How do we solve the scaling problems as AI becomes a global problem?” — @RodrigoLiang Thank you to @JonFortt for having our CEO on @CNBC Fortt Knox to discuss why enterprises are bringing AI on-prem and why we're the only proven premium AI chip inference. We appreciate the chance to share our mission! 🚀 cnb.cx/44KX3KG
3
450
Rodrigo Liang retweeted
My latest Bloomberg Interview, where I talked about AI technologies coming from China, including Kimi K3, and 01.ai's TrueNorth and Boss AI. bloomberg.com/news/videos/20…
13
38
111
42,923
Great to see so many companies get behind open weights. Open models widen access, force competition, and let organizations own what they build instead of renting it back forever. Next comes performance: full precision, fast, at scale. The whole ecosystem gets stronger for it.
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models. images.nvidia.com/pdf/Open-W…
1
10
1,034
Rodrigo Liang retweeted
AI training made GPUs the standard. AI inference is changing the equation. As Abhi shared with @FierceNetwork, minimizing data movement is the key to faster, more efficient inference—and our RDUs are purpose-built to deliver enterprise AI without requiring entirely new data centers. Thanks to @DiaMariesbeat for the coverage! bit.ly/4yzhnfr
5
4
1,134
Rodrigo Liang retweeted
Same prompt. Same model. Two stacks. At #Computex, we demonstrated disaggregated inference live: GPUs handling prefill, SambaNova RDUs handling decode, and CPUs orchestrating agent execution. The result? Up to 2X the speed of B200-only configurations 🦾
1
4
22
1,518
Rodrigo Liang retweeted
Co-founder @sarahookr on stage at @brainstormtech with @SambaNovaAI CEO @RodrigoLiang and @FortuneMagazine's @jeremyakahn. The discussion: Chips not built for inference and models that don’t learn continuously. Two CEOs, two sides of the same problem. A new era of intelligence is here.
2
12
45
3,215
Here is what happens inside Sambanova RDU.
What actually happens during AI inference? This video breaks down how RDUs, memory architecture, and multi-level parallelism work together to generate thousands of tokens in parallel across racks. Built for scalable, real-world AI inference 🦾 Learn more: sambanova.ai/products/rdu-ai…
1
9
678
Rodrigo Liang retweeted
"VC2 is the largest commercial deployment of SambaNova technology in our history, and we're proud to partner with the industry's strongest leaders." - @LipBuTan1 Excited to see @vectorcorecomp launch live w/ our RDUs powering decode alongside GPUs & CPUs. itbrief.co.uk/story/vista-la…
1
8
42
8,325
Rodrigo Liang retweeted
At @LipBuTan1's #COMPUTEX2026 keynote today, @RodrigoLiang stepped onstage with @RFS_Vista to power up the world's first disaggregated inference cloud, VectorCore Compute (VC2), launched by @Vista_Equity and Cambium Capital. Three chips ran disaggregated inference, live from the VC2 datacenter in LA: ➡️ NVIDIA B200 GPUs — prefill, high-compute burst ➡️ SambaNova RDUs — decode, high-throughput, low-latency token generation at scale ➡️ Intel® Xeon® 6 CPUs — tool execution, end-to-end orchestration “GPUs are powerful. RDUs are fast. CPUs orchestrate. But disaggregate all three — and you get speed, performance, and economics no single chip can touch. That's the unlock.” – @RodrigoLiang
10
39
1,954
Rodrigo Liang retweeted
The future of inference is disaggregated. At #COMPUTEX2026, Vector Core Compute unveiled the world's first fully disaggregated inference cloud powered by Intel Xeon and SambaNova RDUs—designed for the scale, speed, and economics modern AI demands. Excited to help bring this vision to life with @Vista_Equity and Cambium Capital.
Vector Core Compute, a new enterprise inference cloud formed by @Vista_Equity and Cambium Capital, unveil fully disaggregated inference running on Intel Xeon processors and @SambaNovaAI RDUs @RFS_Vista at Computex ms.spr.ly/6016vbWHp
2
11
39
2,193
Rodrigo Liang retweeted
Following Cerebras’ IPO, @RodrigoLiang joined @Bloomberg with @mattmiller1973 and @daniburgz to explain why inference could become the biggest business in tech. Watch here: bloomberg.com/news/videos/20…
1
2
9
27,849
Rodrigo Liang retweeted
The next AI war won’t be about training models. It’ll be about: ➡️ Inference costs ➡️ Compute shortages ➡️ Scaling AI profitably Following Cerebras’ IPO, @RodrigoLiang joined @Bloomberg with @mattmiller1973 and @daniburgz to explain why inference could become the biggest business in tech. Enterprise AI demand is exploding. The infra race is just getting started. Watch here ⤵️ bloomberg.com/news/videos/20…
2
3
28
222,122
Rodrigo Liang retweeted
UK sovereign AI is becoming reality. Together with Argyll , we’ve launched a sovereign AI inference cloud built on SambaNova’s full-stack AI platform — designed for performance, efficiency, and control without the trade-offs of traditional infrastructure. As AI adoption scales, sovereignty can’t just be a label. It has to be demonstrated through infrastructure ownership, operational control, energy efficiency, and where intelligence is deployed. This deployment delivers: ⚡ High-performance AI inference 🌍 Renewable-powered infrastructure 🔋 ~10kW rack density with air-cooled systems 🏗️ Disaggregated architecture across UK data centers 🧠 SambaNova RDUs for efficient large-scale AI This is what the next generation of AI infrastructure looks like. Read more via DatacenterDynamics ⤵️ lnkd.in/eAqaBqtr
3
5
34
294,367
Cerebras IPO signals that the market has clearly moved to inference. That’s where costs stack up, where systems break. Where scale either works… or doesn’t. AI won’t be decided by who trains the biggest model. But who can make inference economics actually work at global scale.
1
1
10
229
Rodrigo Liang retweeted
Thank you @ArtificialAnlys for verifying our speeds, as always 🦾 @MiniMax_ai M2.7 is running the FASTEST on SambaCloud. Try it now: cloud.sambanova.ai
MiniMax-M2.7 is now available across six inference providers on Artificial Analysis, with significant differentiation in speed and price @SambaNovaAI leads on speed at 435 output tokens/s, >3x faster than any other provider. @FireworksAI_HQ, @novita_labs, @togethercompute, and @GMI_cloud have all matched @MiniMax_AI's first-party API pricing, while SambaNova is 2x higher. Key takeaways: ➤ Fireworks and SambaNova are on the Pareto frontier for Speed vs. Price. At 127 output tokens/s and ~$0.22 per 1M tokens blended, Fireworks is ~2.2x faster than MiniMax's first-party API at the same blended price, whereas SambaNova delivers 435 output tokens/s but at ~2-3.5x the blended price of the other providers (depending on cache usage) ➤ SambaNova is the fastest provider at 435 output tokens/s, ~3.4x the next fastest provider (Fireworks at 127 output tokens/s). The remaining providers run substantially slower: MiniMax’s first-party API at 57 output tokens/s, Novita at 54, GMI at 41, and Together AI at 29 ➤ Cache discounts vary across providers. Fireworks, MiniMax, Novita, and Together AI offer 80% cache hit discounts, while GMI and SambaNova do not offer a discount. For cache-heavy workloads, this can materially increase the relative pricing for GMI and SambaNova ➤ Optimal provider choice depends on workload. SambaNova may be more suited to latency-sensitive deployments, albeit at a higher cost, while Fireworks may be more suitable for high-volume workloads that are not as latency-sensitive
1
5
14
2,278