Founder CEO @dmatrix_ai | AI and Cloud | Entrepreneur | Enjoy going from zero to one | Striving Yogi

Los Altos Hills, CA
Pinned Tweet
Big news: @dMatrix_AI is partnering with @nvidia. Our next-gen Raptor XPUs with our path-breaking 3D DRAM stacked memory will plug directly into NVIDIA MGX rack-scale systems via NVIDIA NVLink Fusion — delivering ultra-low latency inference for premium token services. Inference is an infinite opportunity, and it will demand a diversity of compute. Customers will be able to run a Raptor rack standalone, or alongside NVIDIA GPUs — matching the best compute to each phase of the workload. Grateful to @JensenHuang and the NVIDIA team for building an open ecosystem that makes this possible. Let's go build the future! 🔗 d-matrix.ai/announcements/d-…
1
5
54
20,520
Sid Sheth retweeted
Nvidia’s Ian Buck shouted out pending @dMatrix_AI customer announcements in his AI Infra keynote so I had to see what the hype was about. Great to see Chief Architect Sudeep Bhoja in his lab after his electric talk at Hot Chips. Emulation, testing, and bring-up done right next to Santa Clara Convention Center. d-Matrix is scaling their approach to speculative decoding pioneered with @gimletlabs where 4 H200s hold the target model and 2 corsair boards run the draft model. This has resulted in throughput and latency gains of 2-5x and customers are trying it out on a variety of model endpoints. Raptor on track for 2028 and lightning for 2030. A part of the NV Link Fusion ecosystem to watch!
3
7
32
2,239
🤝 @dMatrix_AI is adopting NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs with NVIDIA AI infrastructure, joining a growing ecosystem of partners using our rack-scale architecture for AI factory-scale deployment. By bringing together NVIDIA NVLink scale-up, Spectrum-X scale-out networking and the NVIDIA MGX rack architecture, d-Matrix can build specialized inference systems on a proven, rack-scale foundation spanning compute, racks, networking and software. Learn more: nvda.ws/3SGWCyC
4
40
289
36,310
Sid Sheth retweeted
The trillion-dollar AI chip market is taking shape, and inference is opening up a new competitive landscape. The AI infrastructure landscape is changing fast. @business for Bloomberg looks at @dMatrix_AI, @nvidia, @AMD, @Broadcom, @Google, @Amazon, @Meta, @OpenAI and others are shaping what comes next. As our CEO @sidsheth shared: “Inference is not a one-size-fits-all, so brute-forcing inference with a single chip is not going to work.” The Age of Inference is here. finance.yahoo.com/technology…
1
6
28
1,214
Sid Sheth retweeted
Kimi KDA allows D-Matrix to run a 3T parameter model in a single rack with 20-30x faster memory bandwidth than HBM and and 1k TPS. Only 54 concurrent users at 12 GB cache per user. Must be why I’ve been hearing about more new DC builds putting Corsair racks in. Way better TCO than multi-rack disaggregation.
3
7
95
9,232
Discontinuous innovation w/ our next gen ultra low-latency inference accelerator platform. Today at Hot Chips Symposium, d-Matrix co-founder and CTO Sudeep Bhoja and Aayush Ankit with Meta took the stage to present a deep dive into the architecture behind the path breaking Raptor inference accelerator, the first chip to use a 3D DRAM stacked approach for ultra low-latency agentic inference. With 3D-DRAM stacked silicon back in our labs (picture attached), the solution leverages an already proven high-volume CoW packaging technology with an alternate supply chain that is unconstrained by HBM. More here @ 🔗 ServeTheHome: servethehome.com/d-matrix-ra…
1
7
226
Sid Sheth retweeted
A closer look at Raptor™ 3D DRAM from #HotChips2026. @FiresideAlpha captured the discussion on how our 3D DRAM moves data 6–7x more efficiently than HBM4 on measured silicon, while delivering 100+ TB/s of bandwidth. Watch the clip below for the technical breakdown. #AIInference #3DDRAM #AIInfrastructure #dMatrix
Hot Chips 3D DRAM session: D-Matrix says its 3D DRAM moves data 6-to-7x more efficiently than HBM4 on measured silicon "On comparing with HBM4, like we heard in the morning, we can get about 2.4 picojoules per bit energy efficiency on HBM4 of moving data into the base die." "There's obviously additional energy burned to move the data sideways into the GPU, which is not counted here." "But using this face-to-face stacking with 3D DRAM, we are able to reduce that energy to 0.37 picojoules per bit. This is a measured number." "And so this gives us, you know, 6 to 7X energy efficiency." "So for the same amount of power you can drive the bandwidth up, and so we were able to drive the bandwidth up here to 100 terabytes per second." "And in this construction it's a single-high stack, and like I showed, even with 32 gigabytes of memory capacity we are able to solve SOTA models in a scale-up network." "So no bits are wasted."
6
9
32
3,834
Sid Sheth retweeted
At #HotChips2026, d-Matrix is presenting Raptor™, our 3D DRAM architecture built for the growing memory demands of AI inference. Raptor brings memory closer to compute, delivering massive bandwidth with significantly lower I/O energy than HBM. @ServeTheHome goes inside the architecture and the work behind Raptor: servethehome.com/d-matrix-ra… #AIInference #AIInfrastructure #3DDRAM #dMatrix
6
24
106
8,604
Sid Sheth retweeted
A milestone worth celebrating. Thank you, @Nasdaq, for recognizing d-Matrix’s acquisition of @WallarooAI on the Nasdaq Tower in Times Square. We’re excited about what’s ahead as we help customers move AI from models to production at scale. #AI #Inference #EnterpriseAI #dMatrix
8
4
34
3,332
Sid Sheth retweeted
We're excited to announce that @parasail_io is deploying d-Matrix Corsair™ inference accelerators alongside @nvidia Hopper and Blackwell GPUs to power heterogeneous AI inference. By pairing NVIDIA GPUs for compute-intensive prefill with Corsair for latency-sensitive decode, Parasail can deliver up to 10x faster interactive token generation with improved cost-performance and energy efficiency. The future of AI inference is heterogeneous. Read more: d-matrix.ai/announcements/pa… #AI #Inference #LLM #GenAI #NVIDIA #AIInfrastructure #dMatrix @Parasail_io
2
5
18
1,651
A new way of doing things. Different types of compute working together to augment user experience. New world beckons. #bettertogether
Nvidia and AI chip startup d-Matrix are combining their hardware in a new system to power AI models, the latest example of Nvidia partnering with rivals (@_pheebini / The Information) (Visit Techmeme dot com for the link and full context!)
3
143
Sid Sheth retweeted
The future of AI inference is rack scale. A full rack of SMC X-14 servers powered by d-Matrix Corsair™ accelerators and connected with JetStream™ networking. Purpose-built for high-performance, low-latency, power-efficient AI inference. Real infrastructure. Real engineering. Real systems. And yes...real racks have wires. 😉 #AI #AIInference #RackScale #AIInfrastructure #HPC #DataCenter #dMatrix
3
26
1,770
Sid Sheth retweeted
May the throughput be with you. AI doesn’t win at training. It wins at inference. Every prompt. Every response. Every real-time decision. That’s where latency matters. That’s where efficiency matters. We’re building systems designed for this moment—high throughput, low latency, and power efficiency at scale. Built for inference. Ready for production. #AI #Inference #AIInfrastructure #MayThe4th
1
1
5
383
Sid Sheth retweeted
This week we brought together leaders across AI to talk about what comes next. Real-time AI isn’t the future. It’s happening now. The Age of Inference requires a new class of infrastructure. d-matrix.ai/announcements/gi… #AI #AIInference #RealTimeAI #AIInfrastructure
5
8
484
Sid Sheth retweeted
AI inference is no longer just a compute challenge. It is a memory challenge. At d-Matrix we built our architecture around SRAM for ultra-low latency inference. Now we’re pushing further with 3DIMC, a 3D stacked DRAM approach validated by our Pavehawk chip to scale memory for the next generation of AI workloads. Read the blog: d-matrix.ai/going-vertical-w… #AIInference #Semiconductors #AIInfrastructure #MemoryArchitecture #dMatrix
3
12
412
"The future of AI hinges on real-time, energy-efficient inference," says President & CEO of @dMatrix_AI, @sidsheth. Tells @Parikshitl, India can play across the full tech stack but must expand access to compute and energy.
1
2
1
229
Sid Sheth retweeted
There’s undeniable energy in India this week as the nation goes all in on AI. Day one at #IndiaAIImpactSummit was packed with panels, press, and important conversations between @sidsheth and the leaders shaping the country’s AI future. Thank you @Elets_TechnoMedia for hosting a standing-room-only opening panel alongside @GoI_MeitY, @IndiaAIMission, NxtGen, Govt of Jammu & Kashmir, and BWSSB. India is moving fast. We’re looking forward to supporting you on the journey. #AIInference #IndiaAI #SovereignAI
1
3
301