@NVIDIA Sr. Research Scientist | UIUC PhD All opinions and tweets are personal. Tweets about AI Inference, CUDA and GPU systems.

Congratulations to @PrimeIntellect team! We have been helping @PrimeIntellect to run on @nvidia Dynamo for RL and several other usecases! Learn more.
Introducing Prime Inference: We've served trillions of tokens for RL and dedicated customer deployments To own your intelligence, you need to own your inference Unpacking our inference stack
2
60
4,423
Congratulations @CoreWeave on releasing Forge. @CoreWeave uses @nvidia Dynamo to perform large scale RL rollouts at scale! This is an extremely complex problem and glad we could help!
Congrats @CoreWeave on RL Rollouts! RL post-training involves a lot of back and forth: train the model, generate responses, then train again. Inference workers need to load the updated model weights each time. As models get bigger, that can leave GPUs waiting. CoreWeave’s new service uses ModelExpress and Router in NVIDIA Dynamo to speed up those reloads with minimal downtime. Working with us and @youdotcom, CoreWeave achieved 15× faster model reloads compared with its baseline while post-training Nemotron 3.5 Lightning. Check out their blog below for details
10
1,596
Excited to share our latest work with Cedana: serving a portfolio of frontier coding models on a B200 node. TLDR: @nvidia Dynamo handles inference, NeMo Switchyard intelligently routes & escalates across models, while Cedana enables stateful model switching in ~1 min! (1/2)
More enterprises are running a portfolio of models. Easy tasks go to smaller, cheaper models, while frontier models take the hard tail. Hosting that portfolio is a GPU utilization problem. The models don't fit on a single node together, and 9-34 min cold starts make switching between them impractical. Excited to show our work with NVIDIA's Dynamo team. We show how to serve a portfolio of frontier coding models on a single 8x B200 node, switching models in ~1 minute with in-flight sessions intact. cedana.com/blog/coding-model…
2
3
20
1,735
This opens up an interesting direction: dynamically bringing the right model to the workload instead of keeping every model resident on expensive GPU capacity. Learn more! (2/2)
1
133
Vikram retweeted
What does “fungible” mean for AI infrastructure? The flexibility to run different workloads: every type of AI, every phase of AI and even workloads that aren’t AI at all. One platform, built for maximum “productivity”, “durable” years after deployment, and “fungible” for every model and workload. Read the blog to learn more: nvda.ws/4hphGmQ
70
35
221
67,336
Dear PhD students! Application for NVIDIA Graduate Fellows is now open! Apply!
We’re looking for our next cohort of graduate fellows. Our 2027–2028 Graduate Fellowship Program is open to Ph.D. students worldwide with awards up to $60,000, plus mentorship and technical support. Apply by October 30: nvda.ws/4rGKDhv
1
1
105
34,496
Replying to @Pinterest
@Pinterest is using @nvidia Dynamo with vLLM and NVIDIA Blackwell GPUs to serve its vision-language models enabling them to process 25× more visual context per request while avoiding repeated image processing. (1/2)
We’re expanding our collaboration with @Pinterest to scale AI that understands images and words together. Pinterest serves 640M users monthly and every search, every recommendation, every visual moment has to be instant. With NVIDIA Blackwell and Dynamo, Pinterest's AI responds 85x faster, with 7.3x less wait time and 25x more visual context per request. Learn more: nvda.ws/3Vd8I38
1
9
1,148
Pinterest also post trained open models on NVIDIA GPUs with its own data, and is able to run at <8% of the inference transaction cost of comparable closed models. Read more!
81
The easiest way to become expert in inference is contributing to Open source frameworks like @vllm_project @sgl_project @NVIDIAAIInfra Dynamo. We’re always on the lookout for excellent developers who make impactful contributions to Dynamo through high-quality PRs.
My agent tells me that inference engineering job openings have doubled in the past 6 months. I think that's an underestimate. Demand for inference has inflected, again. So too has demand for engineers who can operate AI models in production.
14
80
1,148
58,823
Multimodal models are becoming increasingly important for computer use and everyday AI tasks. A key to serving these models efficiently is encoder and PD disaggregation. Read more about how we are enabling this in a performant way using @nvidia Dynamo!
Multimodal models put different demands on vision encoding, prefill and decoding. Separating vision encoding from the other stages can reduce resource contention and help models respond faster, but only for the right workloads. See how EPD disaggregation works, when it helps and what to consider before using it: nvda.ws/4xUFUv2
4
1,780
Vikram retweeted
GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years. AGI has arrived. Congratulations @OpenAI team. 400K GPUs coming online next.
1,860
4,746
41,827
12,206,925
Excellent to see more adoption of NIXL! We created a simplist way to transfer data between any source and destination for exactly this reason. Fast but simple to use!
Our RL stack now supports NIXL weight transfer, reducing trainer-to-inference transfer time 9x compared with NCCL: from 86 seconds down to single-digit seconds for an 800B-parameter model, and even <4 seconds in our experiments. For prime-rl users, this means over 25% more throughput end-to-end compared with our previous speed. It also clears the way for fault-tolerant, elastic inference scaling that NCCL's rigid process groups made difficult.
2
20
1,211
Vikram retweeted
Exciting day for NVIDIA and @huggingface. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI. Thank you @ClementDelangue for coming to me. NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. 🤗 blogs.nvidia.com/blog/nvidia…
1,715
3,320
27,352
6,123,152
Vikram retweeted
(1/7) Last week FastVideo released FastH3. Today we bring it local: @NVIDIA DGX Spark 💚 and @Apple Silicon 🍎through MLX. 8x speedup! FastH3 now supported on NVIDIA DGX Spark and Apple Silicon! over base @MiniMax_AI H3 for local generation! This 15 seconds 1344×768 clip was generated on two NVIDIA DGX Sparks. Fully local. @NVIDIARTXSpark @NVIDIAAI
34
46
319
101,123
Vikram retweeted
Hanging out next Thursday 9/10 with my friends from NVIDIA Dynamo and SGLang. We'll have tech talks on inference optimization at every layer of the stack. Join us! luma.com/basedynsgl
5
10
93
8,283
Vikram retweeted
Already running an inference engine? So where does NVIDIA Dynamo fit in? In five minutes, we break down how Dynamo sits around engines like @sgl_project, @vllm_project and TensorRT-LLM to scale inference across GPUs and nodes. Full video in the comments 🔽
52
41
327
30,879
Today, we are releasing Shadow engine recovery, a new feature in @NVIDIAAIInfra Dynamo that lets the engine restart up to 39x faster in k8s by using a warmed-up engine. Benchmarking on GLM-5.2 with @nvidia B200 nodes showed that Shadow engine recovery reduced (1/2)
When an LLM engine crashes, a cold restart can mean minutes of lost capacity. Shadow engine recovery, a new preview feature in NVIDIA Dynamo, keeps a standby engine warmed up and ready to take over. In our GLM-5.2 test, it restored capacity in 7.3 seconds, nearly 39x faster than a cold restart. Read all about it: nvda.ws/45OhcA0
4
4
31
4,105
failover time from 283s (cold restart) to 7.3s, dramatically improving TTFT, decode rate, and SLA adherence after process failure. All of these benefits are enabled using the GPU memory service discussed in our earlier blog: developer.nvidia.com/blog/nv… Read more! (2/2)
1
2
199
A higher KV cache hit rate doesn’t always mean lower TTFT and TPOT, or higher throughput. We studied LLM routers with different request scoring objectives under agentic and RL workloads. Here’s what we learned. 🧵
2
20
75
6,551
Over the weekend, @SemiAnalysis_ released the AgentX benchmark, providing a deeper look into the emerging characteristics of agentic AI workloads. (1/9)
📢 First ever on-silicon NVIDIA Vera Rubin performance measured on how agents actually run. ⚡ Up to 30x more throughput per megawatt and up to 35x lower token cost than GB300 NVL72. Agentic sessions are nothing like chat or summarization workloads. Context grows across hundreds of steps and can reach hundreds of thousands of tokens. NVIDIA measured Vera Rubin performance on @SemiAnalysis_ AgentX workload consisting of real-world agentic coding trajectories using DeepSeek V4 Pro model.
2
4
13
2,352
It was also exciting to see @nebiusai become the first AI cloud to adopt Groq 3 LPX. I was fortunate to interact with the Nebius team when the company was just getting started, which makes seeing this milestone particularly rewarding. (8/9)
1
110