AI @ AMD

San Francisco, CA
2 ways to win on OpenRouter, serving GLM5.3/MI355X: 1. Speed: TileRT hits 469 tps/user on AgentX, ~45% higher interactivity v. GB300/FP4: amd.com/en/developer/resourc… 2. Profits: $11.18/GPU/hr at 67 tps (SemiAnalysis AgentX, 60% utilization) => $11.18*24*265*8 = $560k profit/node/year
6
4
58
6,610
Ramine Roane retweeted
MI355X IS UP TO 1.7X BETTER 💰️PERF PER DOLLAR 💰️THAN DGX B300. The AMD Mainland China UMBP team co-designed, in collaboration with Alibaba & the @sgl_project community, a new feature in SGLang that removes the duplicated KVCache contained between local L2 DRAM & distributed L3 DRAM, allowing for up to 2x more KVCache to be stored in DRAM. This feature is called UnifiedRadixCache external cache. But importantly, this marks the trend of AMD increasingly being first-class co-designed for new features in widely used top production engines like SGLang.
16
29
263
43,386
Ramine Roane retweeted
AMD ALERT🚨🚨: On Kimi K3 2.8T, AMD has even better profit margins per gigawatt compared to B300, GB200 NVL72, and B200. This is due to AMD's lower TCO and better performance, leading to better profit margins of 53.6% on Mi355X compared to 44.3% on GB200 NVL72. This is from our open-source Agentic Inference benchmark, AgentX. Shoutout to @roaner & @EmadBarsoumPi 10x engineers for optimizing this 2.8T model!
20
71
644
61,791
Ramine Roane retweeted
On total tokens per $ TCO, a new AMD MI355x submission beats B300 at lower interactivity ranges on AgentX. Shoutout to vLLM, AMD, and LMCache engineers.
13
27
237
34,901
Ramine Roane retweeted
Kimi K3 running on AMD MI350X with SGLang. Amazing work to @AMD @sgl_project @Kimi_Moonshot this worked great out of the box!
28
108
1,443
79,584
Ramine Roane retweeted
With Kimi K3 Day-0 on vLLM: Open Frontier Intelligence for Everyone 🚀 At 2.8 trillion parameters, Moonshot AI's Kimi K3 is one of the most powerful open-weight models ever released. Starting today, you can serve it on vLLM the moment the weights are public. What K3 brings: 🧠 2.8T-parameter Mixture-of-Experts (16 of 896 experts active per token) 📚 1M-token context window 🛠️ Native multimodal understanding, including vision ⚡ Kimi Delta Attention: a hybrid of linear and full attention that makes million-token context affordable Huge thank you to @Kimi_Moonshot AI for the model release and partnership, @inferact for leading the vLLM optimizations, and to our partners at @nvidia, @AMD, and the broader vLLM community. 1/6
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: huggingface.co/moonshotai/Ki… Tech report: github.com/MoonshotAI/Kimi-K… Tech blog: kimi.com/blog/kimi-k3
17
56
410
104,368
Ramine Roane retweeted
Agent workloads bring long contexts, bursty traffic, and frequent tool calls, creating new demands for AI serving. Together with @AMD , we rebuilt the stack for Kimi K2.6 on AMD Instinct™ MI355X, with scheduler-aware multi-tier KV caching. Up to 3.2× smaller p99 TTFT and 7.7% higher total-token throughput, with no accuracy tax. Happy to see more Kimi running on AMD chips! Tech blog: amd.com/en/developer/resourc…
53
108
1,704
119,856
Ramine Roane retweeted
.@radixark brought @deepseek_ai V4 to production with DigitalOcean and @AMD. AMD Instinct™ MI350X GPU Droplets, plus joint engineering that drove a ~10x throughput gain. When the team that builds the inference engine picks your stack, that's the signal.
4
9
29
13,837
Ramine Roane retweeted
The more you embrace AI, the more you need SaaS. This is not obvious to armchair market analysts who love disruption narratives, but it is obvious to people actually running companies.
Levie now uses Salesforce 5x more than at any point before. The Box CEO @levie connected Salesforce's MCP server to Claude Code. Now he runs customer and market intelligence queries he would never have bothered pulling up by hand. The agent removes the friction. The underlying system gets queried more, not replaced. If you hold $CRM, the agentic era is an engagement tailwind, not a disruption risk - gated on whether the data platform handles the query load. Why incumbents gain from agents: podcastalpha.substack.com/p/… Source: CXOTalk - piped.video/watch?v=ylBDQHk3…
45
44
388
83,476
Ramine Roane retweeted
.@realGeorgeHotz doesn’t follow the script. From jailbreaking the iPhone at 17 and reverse-engineering the PS3 to building open-source self-driving technology at @comma_ai, he's consistently pushed the boundaries of what's possible. Now, as founder of @__tinygrad__, he’s focused on opening up the AI compute stack. He’s also one of the most candid voices on AI and how to get the most out of AMD solutions. That’s exactly why he’s joining the Advancing AI Developer Track. Register: amd.com/en/corporate/events/… #AdvancingAI #AMDevs
27
57
784
42,268
The fastest whale isn’t always the bigger one 🐳 MI355X GPU Sets a New Bar for DeepSeek Inference amd.com/en/developer/resourc…
1
1
22
4,434
Ramine Roane retweeted
🚀 Launch today: Kog generates 3,000+ output tokens/s per single request, on standard datacenter GPUs. We are bringing real-time LLM inference to hardware that companies already run in production. The speed previously associated with purpose-built silicon is now delivered on NVIDIA H200 and AMD MI300X. Today, we are opening our Tech Preview with a 2B coding model, with large frontier MoE support coming next. Try our Playground → playground.kog.ai 💥 Why that matters, and how we did it → blog.kog.ai/real-time-llm-in… 📖 Monokernel deep dive → blog.kog.ai/building-a-singl… 📖 Delayed Tensor Parallelism research → blog.kog.ai/delayed-tensor-p… read the thread 👇
16
40
267
6,165,469
Ramine Roane retweeted
AMD ALERT 🚀 MI355 is now 40% cheaper than B200 on GLM5 architecture for Single Node serving FP8 14 weeks after the initial launch of GLM5 on both non-MTP & MTP with spec decode for SGLang v0.12 for both CUDA & ROCm.  SPEED IS THE MOAT!! Great work to @AnushElangovan, @roaner, HaiShaw & his team! Next step is for MI355X to catch up to CUDA when composing production inference optimizations like FP4 & on distributed inferencing where you can gang up MI355 boxes such that per GPU performance goes up thus the cost per million tokens goes down.
7
44
456
49,183
Ramine Roane retweeted
Today, we announced more than $10B in investment across Taiwan’s ecosystem to scale advanced packaging and accelerate next-gen AI infrastructure, from 6th Gen EPYC CPUs codenamed “Venice” to our Helios rack-scale platform including Instinct MI450X GPUs, with multi-gigawatt deployments beginning in 2H 2026. Additionally, AMD and TSMC have hit another major production milestone, with Venice EPYC CPUs ramping on TSMC 2nm technology in Taiwan with future plans to ramp production at TSMC’s Arizona Fab. More on the news: bit.ly/4tJrUkR
31
95
706
93,152
Ramine Roane retweeted
Huge respect to HaiShaw, Thomas, @roaner, @AnushElangovan from @AMD for the relentless work, fusing mHC ops and RoPE hadamard transforms, and shipping new attention indexer + KV cache kernels in TileLang & Triton at remarkable speed. Honored SGLang is the stack powering it. Cheering you on for the next 5x.
SPEED IS THE MOAT: AMD ROCm software stack has improved performance by over 75x in the last 14 days since DeepSeekv4 launch. The performance comes from fusing mHC operations & also fusing RoPE hadamard transformations to reduce cpu overhead & improve HBM memory utlization. Furthermore, other kernels like the attention indexer & kvcache compressor has been written using TileLang & Triton for fast development velocity. Another 5x performance improvement is needed to catch up to single node aggregated B200 performance & then another 1.5x is needed to catch up to PD disaggregated B200 performance, which is within the realm of possibility for AMD within the next couple of weeks. Great work to HaiShaw, Thomas, @roaner, @AnushElangovan for this rapid improvement.
1
5
48
6,242
Ramine Roane retweeted
🎉 Meet MiniCPM-V 4.6 from @OpenBMB, a 1.3B edge-friendly multimodal LLM with superior efficiency. Day-0 support is now live in SGLang! ✅ Leading capability: scores 13 on Artificial Analysis Intelligence Index benchmark ✅ Strong multimodal: Matches Qwen3.5 2B-level capacity across 5 major VL benchmarks ✅ Ultra-efficient: 50%+ less visual FLOPs via LLaVA-UHD v4 and mixed 4x/16x visual token compression ✅ Mobile-ready: can be deployed across iOS, Android, and HarmonyOS Cookbook: docs.sglang.io/cookbook/auto… Run it now with SGLang!
1/5 MiniCPM-V 4.6 (1.3B) is now live 🚀🚀 High-res visual processing, optimized for consumer-grade and mobile hardware. We’ve leveraged the latest LLaVA-UHD v4 technique to cut vision encoding costs by 55%, enabling native edge deployment with extreme efficiency. 🔥 Beats Gemma4-E2B-it and Qwen3.5-0.8B across key multimodal and Artificial Analysis benchmarks — scoring higher than Qwen3.5-0.8B using just 2.5% of its token budget. ⚡ TTFT (75.7ms) 2.2x Faster than Qwen3.5-0.8B even with 3136² high-res images. 🏗️ ~1.5x Token Throughput compared with Qwen3.5-0.8B on a single RTX 4090. Try the model here: 🤗 Hugging Face: huggingface.co/openbmb/MiniC… 💻 GitHub: github.com/OpenBMB/MiniCPM-V 🔭 Modelscope: modelscope.cn/models/OpenBMB… 🌐 Web Demo: huggingface.co/spaces/openbm… 📱 App Demo: github.com/OpenBMB/MiniCPM-V…
4
24
4,201
Ramine Roane retweeted
SPEED IS THE MOAT: AMD ROCm software stack has improved performance by over 75x in the last 14 days since DeepSeekv4 launch. The performance comes from fusing mHC operations & also fusing RoPE hadamard transformations to reduce cpu overhead & improve HBM memory utlization. Furthermore, other kernels like the attention indexer & kvcache compressor has been written using TileLang & Triton for fast development velocity. Another 5x performance improvement is needed to catch up to single node aggregated B200 performance & then another 1.5x is needed to catch up to PD disaggregated B200 performance, which is within the realm of possibility for AMD within the next couple of weeks. Great work to HaiShaw, Thomas, @roaner, @AnushElangovan for this rapid improvement.
24
84
832
166,277
Ramine Roane retweeted
1/7 Running huge MoE models on affordable hardware is all the rage. Adding a new approach to the mix optimized for speed and model quality. Introducing FOMOE: Fast Opportunistic Mixture Of Experts (pronounced fomo). Runs Qwen3.5 flagship model with 397 billion parameters at 5 – 9 tok/s on a $2,100 desktop! Uses Q4_K_M quants. Two $500 GPUs, 32GB RAM, one NVMe drive. Runs on Linux.
21
21
273
23,487
Who did this?
2
2
20
3,955