Building the future of inference through @vllm_project

San Francisco, California
Inferact retweeted
Our first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode, against 450 tokens/s for our GB200 baseline, both with DSpark speculative decoding. To our knowledge, this is the first TPU inference megakernel. The whole model runs in a single Pallas kernel, and without spec decoding it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8. We are open sourcing it today. 1/2
22
92
846
233,528
Thanks for the shoutout @SemiAnalysis_ ! Full breakdown linked here: inferact.ai/blog/tpu-megaker…
ALERT ALERT ALERT 🚨 🚨 🚨 VLLM MAINTAINERS HAVE JUST SHOWN THAT TPUv7 CAN GET 700 tok/s/user,  56% BETTER PERFORMANCE THAN NVIDIA GB200 NVL72 THROUGH MEGAKERNEL OPTIMIZATION ON KIMI K3. As we said awhile ago, the TPU externalization of software is full steam ahead. This is ultra important to follow the progress of this.
21
5,000
Our first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode, against 450 tokens/s for our GB200 baseline, both with DSpark speculative decoding. To our knowledge, this is the first TPU inference megakernel. The whole model runs in a single Pallas kernel, and without spec decoding it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8. We are open sourcing it today. 1/2
22
92
846
233,528
Inferact retweeted
That's a wrap on the @vllm_project Conference at Ray Summit. Lines out the door for talks about an inference engine—that says everything about how much this community shows up. Thank you to @anyscalecompute for co-hosting with us, and we're already looking forward to the next event.
3
8
48
10,903
That's a wrap on the @vllm_project Conference at Ray Summit. Lines out the door for talks about an inference engine—that says everything about how much this community shows up. Thank you to @anyscalecompute for co-hosting with us, and we're already looking forward to the next event.
3
8
48
10,903
Kimi K3 on @vllm_project is now 2.2–2.8× faster 🚀 Inferact is proud to have co-led this optimization effort with @RedHat_AI, @NVIDIAAI, and @Huawei, spanning scheduling, KDA state handling, and custom MoE kernels. Read the technical deep dive: vllm.ai/blog/2026-09-13-kimi…
Kimi K3 serving in vLLM now delivers 2.2–2.8x throughput on our B300 benchmark vs v0.27.1. We break down the work across scheduling, KDA state handling, and MoE kernels, with benchmarks and commands to reproduce the results. Thanks to the vLLM community for pushing Kimi K3 performance forward! Read the deep dive: vllm.ai/blog/2026-09-13-kimi…
3
15
66
6,640
We’re partnering with @inferact (the team behind @vllm_project) to make TPUs a first-class citizen for open-source AI inference. Together, we’re delivering: - Upstreamed open-source TPU kernels - Native PyTorch integration via TorchTPU + so much more ↓
@googlecloud and Inferact are announcing today a partnership to make TPU a first-class citizen in @vllm_project. This partnership puts both teams on one engineering roadmap to bring TPU to the broader open model ecosystem, optimizing vLLM as the agentic production serving engine for TPU: • Production serving features and optimized kernels • A native PyTorch path via TorchTPU • Moving towards day-0 support for frontier model releases We're also launching a community program: shared TPU capacity for open-source contributors, plus dedicated review and design help from the core vLLM maintainers at Inferact. Everything this collaboration produces is open source. Read the full announcement: inferact.ai/news/google-tpu-…
8
23
154
32,072
Day-0 support for DeepSeek v4.1 Flash landed across H100, H200, B200, B300, GB200, and GB300. Thank you to @SemiAnalysis_ for verifying this independently and @NVIDIAAI for the collab. Serving recipes here: recipes.vllm.ai/deepseek-ai/…
On the Day 0 release of DeepSeekv4.1 Flash, NVIDIA vLLM works out of the box with zero issues across all 6 SKUs: H100, H200, B200, B300, GB200, GB300! Amazing work by the NVIDIA & Inferact teams! In comparison, AMD vLLM still does not work on DeepSeekv4.1 Flash, as we will describe below👇️(1/2)🧵
3
5
41
5,405
Behind this blog is months of our team's work tuning vLLM on agentic workloads and validating on @SemiAnalysis_ AgentX benchmark. We find that open source models optimized for agentic workloads reach up to 130K tokens/GPU-sec, 106× cheaper than Opus 5 API pricing. vLLM is the open source agentic production serving engine. Inferact optimizes vLLM and builds enterprise inference on top of it. 🚀
New blog is out: vLLM x AgentX: Optimizing for Real-World Agentic Serving. Agent traffic stresses every layer of the serving stack at once. This post walks the full-stack work for optimizing vLLM on Agentic workloads, including the architecture, framework, and runtime optimizations, measured on AgentX, @SemiAnalysis_'s public agentic benchmark. 🧵1/6
2
6
29
3,160
Congratulations to @HUMAIN and @MiniMax_AI on HUMAIN-M3, a frontier Arabic model now available on HUMAIN Node 🎉. @vllm_project is running the inference under the hood, and we're looking forward to more Arabic use cases with this state-of-the-art stack.
HUMAIN unveils HUMAIN-M3, a frontier Arabic language model, commissioned by HUMAIN and developed by @MiniMax_AI, now available in research preview on HUMAIN Node. Read more: humain.com/news/humain-unvei… #HUMAIN #LEAP26 #TheEndOfLimits
3
32
19,281
Inferact retweeted
Thank you to the @SemiAnalysis_ team for the shoutout and for the collaboration on AgentX 🙏 Benchmarks are only useful when they measure the workloads people actually run, and AgentX measures the real thing: multi-turn, long-context agent traffic. vLLM is the engine for production agentic workloads. For teams serving tokens at scale, revenue depends on optimized inference over long multi-turn contexts. Our AgentX deep dive blog is coming soon! Stay tuned 📖
Shoutout to the cracked team at @vllm_project that implemented recent agentic workload optimizations. (1/5)🧵
4
8
32
5,302
Inferact retweeted
On Tuesday, capping off a full day of speakers at the first vLLM Conference, the vLLM community gathered for a rooftop happy hour at sunset. The event hit capacity, and sunset over the city was a great backdrop. Thank you to @AMD and @inferact for co-hosting! More meetups to come soon.
9
7
65
6,309
Our co-founder and CEO @simon_mo_ is featured in the PyTorch Conference promo, talking about where @vllm_project is headed! 🚀 Inferact and @simon_mo_ will be presenting at the conference, come find our team at San Jose, October 20-21!
At PyTorch Conference North America 2026, hear directly from and connect with the people working on PyTorch, @vllm_project, @DeepSpeedAI, @raydistributed, Helion, and Safetensors, alongside experts working across the AI stack. Keynote speaker @simon_mo_ (@vllm_project, @inferact) says vLLM’s goal is to become “the easiest to use and most efficient inference engine.” At #PyTorchCon NA 2025, he spoke about the difficulty and cost of using large language models and continuing to push inference efficiency forward. PyTorch conferences are the open source AI community’s town square, where what’s next gets decided. Register by September 4 to save on your conference pass: hubs.la/Q04tDbKG0
24
3,308
Much of this work came out of our team at Inferact, working with @vllm_project and @SemiAnalysis_. Huge shoutout to @yifandotqiao, who led our AgentX workstream end to end, and to @esmeetu87, @lizhewen71800, Jeff Ma, Summer Yang, Nick Hill, Woosuk Kwon, and Dao Le for months of sweeps, kernels, and upstream PRs behind these numbers. @yifandotqiao is presenting this work at the vLLM Conference this week 🚀
Congratulations to @SemiAnalysis_ on the release of AgentX 1.0 🎊, an open-source multi-turn agentic coding benchmark collected from ~$3M of real traces, running on 1000+ chips and ~2MW of continuously operated compute. We are excited to see @vllm_project’s competitive performance on frontier open models: 🔷130,093 tok/s/chip for DeepSeek V4 Pro 🔷 77,079 tok/s/chip for Minimax M3 🔷 12,479 tok/s/chip for Kimi K3. The following thread covers an overview of the work from @vllm_project and @inferact: what we tested, found, and shipped upstream. This work highlights vLLM’s performance on real-world workloads and our committed focus to making vLLM an agentic-first engine. Optimizing AgentX performance meant tackling three major challenges: prefix reuse, efficient long context parallelism, and scaling performance with PD disaggregation. First, long agentic sessions stress prefix caching and KV cache offloading. Modern hybrid models have greatly reduced the required KV cache sizes, and caching every block boundary still saturates the KV cache pool which causes prefix cache thrashing across sessions. The fix was sparse retention: one state per interval-sized segment plus the latest replay boundary (vllm-project/vllm #43447, #45845), then preserving shared-prefix boundaries so the interval can go to 0 for agent sessions (#47782). That gets us >95% hit rate at 14 concurrent requests with contexts to 1M. The bigger structural change in KV cache offloading was making the shared KV pool distributed. With Mooncake Store as a first-class connector, prefill ranks can hit the prefix cache both within a worker and across workers, so we no longer have to trade cache locality against load balance to keep a cluster busy. Session-aware routing (48048) is what lets the router act on it. For a single node deployment, SimpleCPUOffloadConnector has been greatly improved to support all hybrid model architectures and across both CUDA and ROCm platforms. For DeepSeek V4 Pro on ROCm, this implementation gave +81.7% output throughput and 46.6% lower mean e2e latency versus recomputing the prefix. Kimi K3 at 2.8T barely fits on a single node, and squeezing it in leaves almost no headroom for KV cache, which is exactly what a long multi-turn session needs most. Parallelism strategy matters more here than on any other model we tested. TP8/DCP8 tops the K3 agentic frontier across configs, with B300 vLLM peaking at ~12.5k tok/s/chip at ~8 tok/s/user and GB300 NVL72 on Dynamo + vLLM holding the curve out past 200 tok/s/user. K3 also surfaced a routing bug worth pulling: vllm-project/router#194 fixes the router dropping reasoning_content, which hits any reasoning model served behind it. On MiniMax M3, B200 vLLM reaches ~44k tok/s/chip and vLLM leads TRT-LLM on throughput vs p90 TTFT. M3 and Qwen3.5 both shipped day-0. Finally, fully optimized performance requires scaling the deployment with prefill-decode disaggregation and distributed KV cache offloading. Thanks to vLLM’s MultiConnector, this is natively supported with NIXL PD connector + MooncakeStoreConnector. By rate matching to find the ideal prefill-decode ratio, we achieved 4.45x higher throughput at a 60 tok/s interactivity for DeepSeek V4 Pro on GB300 Dynamo compared to B300. Shoutout to @NVIDIAAI, who co-tuned most of these configs with us. Dynamo's router optimizations took AgentX replay time down 23.7% on the vLLM backend, the AIPerf replay harness is what made the traces runnable at all, and NIXL and their kernels sit under each respective performance point. Shoutout to @AIatAMD as well: the team’s AITER sparse-MLA decode selection led to +5.22% AgentX output throughput; the hybrid AITER/native CSA selector led to 1.21–1.76x e2e performance boost, and the team made prebuilt lmcache gfx942/gfx950 wheels and a published Mooncake ROCm wheel. Thank you to @SemiAnalysis_ for building AgentX and for the collaboration throughout. Agentic workloads are what users serve in real production, and benchmarking on real workloads is what improves vLLM and open source inference. Next up: more upstream work and a blog with a full technical deep dive later this week. Stay tuned. 🚀
1
5
31
3,505
We're excited to kick off the first vLLM Conference next week! Two days of talks and three nights of happy hours. Join an event and learn where the future of inference is heading 🚀
vLLM Conference is next week, and we have a packed schedule 🎊 📅. Here's the full list of events you should know: Mon: 🔷 4–6PM vLLM × Ray × Google Cloud Happy Hour: rsvp.withgoogle.com/events/r… 🔷 6–9PM vLLM × Dynamo Meetup: luma.com/r8o604o0 Tue: 🔷 11–11:30AM vLLM Keynote from @simon_mo_ 🔷 12–5PM vLLM Track Day 1 🔷 6–9PM vLLM × AMD Happy Hour: luma.com/6nb8coax Wed: 🔷 12–5PM vLLM Track Day 2 🔷 6–8:30PM vLLM x DigitalOcean × NVIDIA Happy Hour: luma.com/fufg0bg5 No ticket needed for the happy hours and meetups, but space is limited. To join the full event, register here: vllm.ai/events/vllm-conferen…
1
5
21
2,600
Looking forward to the event! A great way to close out the vLLM Conference.
Valuemaxx your time after Ray Summit and the vLLM Conference. Join our happy hour, lightning talks, and live demos with @nvidia, @Inferact, and @NousResearch. 🍻 Open inference, open convos, open tab. 🍻 RSVP today. 🎟️ do.co/4g0oAhE
1
1
6
1,184