Sawyer Bowerman retweeted
LLM Compressor v0.14.0 is out, and GPTQ just got its biggest speedup since launch. A new Triton kernel makes quantization ~15x faster end to end. Batching layers that share a shape pushes that to ~30x on some MoE workloads. Even the old eager path is 1.5-2x faster. Also new: expanded MSE/iMatrix observers that beat GPTQ for NVFP4 on internal benchmarks, REAP pruning with distributed DDP, and support for GLM 5.3 and Qwen3.8. Full release notes: github.com/vllm-project/llm-…
3
18
119
6,692
Sawyer Bowerman retweeted
If your agent can act on real systems, with real credentials, over long time horizons, what actually stops it from doing the wrong thing? Not the prompt. The model is probabilistic; your operational controls can't be. @LegareKerrison kicks off AgentOps Unlocked series with the four gaps between a laptop demo and production: execution containment, identity, observability, and lifecycle governance. Built on open source you already know: Kubernetes, SPIFFE/SPIRE, OpenTelemetry. piped.video/watch?v=LJHn6a_z…
1
1
18
2,812
Using frontier LLMs for every agentic sub-task drains budget and increases latency. Check out how we look into picking more targeted models drives us to cost-effective, and "smart-enough" agentic execution. sprou.tt/1aZWuT9HcGQ
1
17
Sawyer Bowerman retweeted
Kimi K3 is a 2.8T-parameter model. We trained a DSpark speculator for it, and the speedup holds up. Single-stream math reasoning goes from ~110 to ~435 tokens/sec per user. Under concurrent load, up to ~3.5x higher output throughput at matched interactivity. The drafter is a 5B model proposing 8 tokens a step, and on math it gets ~6.4 of them accepted per round. Training a drafter for a model this big meant going multi-node: Speculators plus a new Mooncake connector streaming hidden states between vLLM inference and training over RDMA. Two GB300 nodes to serve the target, one to train. vllm.ai/blog/2026-09-15-kimi…
1
13
102
5,973
Super cool resrouce, fun to watch someone go through the whole process!
Everything you need to start self-hosting an open LLM. Run it on your own hardware. No API keys. No per-token bill. Nothing leaves your machine. The full path with @vllm_project: batch inference in Python, an OpenAI-compatible API server in one command, and quantized models that cut an 8B from ~16GB of weights to a quarter of that while keeping 98-100% accuracy. Walkthrough by @cedricclyburn.
1
28
Sawyer Bowerman retweeted
What does it take to serve a model that talks back, or one that generates video with audio? Most text models advance one token at a time. These don't, and they need different scheduling to match. New recap on how vLLM-Omni serves them:
Article

Voice, Video, and Diffusion Serving with vLLM-Omni

At vLLM Office Hours #57 (video, slides), Alex Brooks, Ricardo Noriega, and Nick Cao from Red Hat AI walked through the vLLM-Omni architecture, recent project updates, and the work behind serving

2
7
35
3,423
Sawyer Bowerman retweeted
Nothing over 33B parameters. That's the whole challenge. @NVIDIAAI and Red Hat are running an in-person Small Models Hack in Raleigh, October 17-18. Small, open models only. Build something that punches above its weight. Three tracks: 🎯 Fine-tune a small model to beat a general one at a real task 🕸️ Wire several small models into a multi-agent system 🛠️ Make one model call real tools reliably Brev credits provided to participants for GPU time all weekend. Prizes include an NVIDIA GeForce RTX 5090, NVIDIA Jetson Orin Nano Super Developer Kits, and more! 100 spots available: solo or up to teams of 3. Apply: luma.com/nvidia-redhat-small…
7
11
80
19,513
Sawyer Bowerman retweeted
Tool calling is where agentic #AI demos fail in production. Discover why schema drift, unhandled execution errors, and non-deterministic tool outputs create the "last-mile problem," and how to build resilient execution interfaces on #OpenShift AI. ➡️ red.ht/4qXaHEw
3
6
585
Surpass GPU memory ceilings in agentic workloads. Discover how vLLM's PagedAttention, prefix caching, and intelligent context pruning mitigate KV cache explosion during long-horizon autonomous reasoning loops. Thanks for taking some time to check it out! sprou.tt/1Tn9YSBX8H2
1
32
Sawyer Bowerman retweeted
vLLM Office Hours today, 2pm ET: vLLM-Omni project update and live demos. A voice assistant on Qwen Omni, diffusion intermediate steps, and video generation. Plus what's new in vLLM 0.28 with @mgoin_. Get a recurring cal invite: red.ht/office-hours
1
2
14
874
When AI agents fail, the issue is typically independent to model intelligence.... Grace Ableidinger and I did some exploration on the last-mile challenges of tool calling and how robust interfaces keep production agents reliable. Check it out! sprou.tt/1RUsdB7c1DJ
4
2
75
In regulated industries, you need documented proof of how your models hold up under adversarial pressure. Check out Red Hat's latest insights on building risk-aware AI deployments using automated red teaming to keep your organization secure! sprou.tt/1IDJKVTAqOP
32
Sawyer Bowerman retweeted
Is your CPU:GPU ratio still built for training? In agentic workloads, 50-90% of end-to-end latency is CPU-side tool processing, not GPU math (Georgia Tech + Intel). That's pushing the CPU:GPU ratio from 1:8 in training toward 1:1, sometimes 4:1. And vLLM's CPU backend already runs PagedAttention, prefix caching, and continuous batching across x86, Arm, IBM Z, and experimental Apple Silicon. @_soyr_ and @__gracecaroline on where inference compute should live: redhat.com/en/blog/cpu-back-…
3
8
47
3,151
Sawyer Bowerman retweeted
We're back with @vllm_project office hours this Thursday, August 20 🎉 We'll share: 🚀 What's new in vLLM 0.27 🤖 Running Codex and Claude Code CLIs with vLLM and open models ⚡ What's new in Speculators 📊 GuideLLM v0.7.0: benchmarking reasoning, tool calling, and agentic traces Link to join in the thread 👇
1
8
50
6,036
I feel like nowadays, performance of a model is only half the battle, especially since none of us have any capacity to wait around lol Anyways, super awesome article about how the speed and performance trade off is a fine line we need to balance on!
Your agents don't need a genius. They need a model that never becomes the bottleneck. NVIDIA Nemotron 3.5 Lightning: 30B MoE, 3B active, ~670 tokens/sec. Running on vLLM day 0, with quantized checkpoints from Red Hat AI ready now. Here's the writeup:
Article

Your agents don't need a genius. They need a model that runs at 670 tokens/sec.

Most of what an agent does all day is grunt work: tool calls, retrieval, validation, formatting, classification, summarization. You don't need a frontier reasoner for that. You need something fast,

2
39
Sawyer Bowerman retweeted
Your agents don't need a genius. They need a model that never becomes the bottleneck. NVIDIA Nemotron 3.5 Lightning: 30B MoE, 3B active, ~670 tokens/sec. Running on vLLM day 0, with quantized checkpoints from Red Hat AI ready now. Here's the writeup:
Article

Your agents don't need a genius. They need a model that runs at 670 tokens/sec.

Most of what an agent does all day is grunt work: tool calls, retrieval, validation, formatting, classification, summarization. You don't need a frontier reasoner for that. You need something fast,

6
3
40
3,951
Sawyer Bowerman retweeted
~4x faster Kimi-K3 decoding. Our new DSpark speculator takes single-stream interactivity from ~110 to ~435 tok/s/user on math reasoning, delivers ~3.5x the output throughput at matched interactivity under load, and thanks to sliding window attention (2048-token window across all 5 draft layers), our model holds steady acceptance out to 20K context across 10 LongBench domains. Check it out! huggingface.co/RedHatAI/Kimi… More details in reply:
4
25
220
19,550
Sawyer Bowerman retweeted
Meta's Muse-Glimmer, running with Red Hat AI Inference (early preview). Copy the manifest, oc apply, and you're serving the multimodal model on @vllm_project. FP8, INT4, and NVFP4 checkpoints ready on Hugging Face for production. Here's the full setup in under 2 minutes by @_soyr_:
4
7
46
2,272