Accelerating AI innovation with open platforms and community. The future of AI is open.

What does it take to serve a model that talks back, or one that generates video with audio? Most text models advance one token at a time. These don't, and they need different scheduling to match. New recap on how vLLM-Omni serves them:
Article

Voice, Video, and Diffusion Serving with vLLM-Omni

At vLLM Office Hours #57 (video, slides), Alex Brooks, Ricardo Noriega, and Nick Cao from Red Hat AI walked through the vLLM-Omni architecture, recent project updates, and the work behind serving

2
7
35
3,430
Your agents don't need a genius. They need a model that never becomes the bottleneck. NVIDIA Nemotron 3.5 Lightning: 30B MoE, 3B active, ~670 tokens/sec. Running on vLLM day 0, with quantized checkpoints from Red Hat AI ready now. Here's the writeup:
Article

Your agents don't need a genius. They need a model that runs at 670 tokens/sec.

Most of what an agent does all day is grunt work: tool calls, retrieval, validation, formatting, classification, summarization. You don't need a frontier reasoner for that. You need something fast,

6
3
40
3,951
Every vLLM replica has its own KV cache. Scale out behind a load balancer and prefix reuse collapses. Mooncake pools them cluster-wide. On agentic traces: cache hits 1.7%->92.2%, TTFT down 46x, latency down 8.6x. How it works, from vLLM Office Hours #55:
Article

Your vLLM replicas are isolated cache islands and Mooncake fixes that

Every vLLM replica maintains its own KV cache. Scaling to multiple replicas with a traditional L7 load balancer scatters requests across them, destroying prefix reuse. Duplicated prefill compute,

5
7
53
7,938
Kimi K3 open weights landed yesterday. By that afternoon, Red Hat AI Inference was serving it on one 8x B300 node. 2.8T params. Day-0 preview images exist so you can experiment the moment weights drop. 👇
Article

Run Kimi K3 on Day 0 with Red Hat AI. Here's the Exact Command.

Kimi K3 dropped open weights on July 27th. By that afternoon, we had it serving an OpenAI-compatible API on a single 8×B300 node using Red Hat AI Inference. This is the largest open-weight model ever

6
12
50
4,125