High-performance RL post-training infra. Achieves bitwise train-inference consistency across heterogeneous engines & extreme memory efficiency for GRPO, PPO.

Singapore
Zero mismatches. Bit-for-bit aligned. 🔥 Incredibly proud of our team's work integrating with vime @vllm_project. Rock-solid RL pipelines are now a reality on @AMD ROCm. Read the deep dive below and star our repo⭐ github.com/RL-Align/RL-Kerne…
Training and rollout logprobs matched bit for bit on ROCm. The @RLKernel team integrated RL-Align/RL-Kernel with vllm-project/vime. A 200-step Qwen3-8B GRPO run on 8× @AMD MI300X recorded zero logprob mismatches between Megatron training and vLLM rollout. The strict path aligns reduction order, intermediate precision, rounding points, and math primitives across both sides. Deep dive: vllm.ai/blog/2026-09-14-rl-k…
1
2
9
159
RL-Align retweeted
RL-Kernel v0.1.0 is now available. 🎉 Bitwise train–inference consistency across parallel configurations and heterogeneous execution engines for RL post-training. With Qwen3-8B on a full vime pipeline, RL-Kernel recorded zero train–rollout LogP mismatches across all 200 training steps. 1,129 commits from 29 contributors. Highlights: 🎯 Strict numerical alignment across rollout and training engines ⚡ 68.7% higher rollout throughput and 8.2% lower end-to-end step time on NVIDIA H100 🔶 AMD Instinct MI300X: zero mismatches across 9,400,614 active-token LogP elements from 1,600 samples 🔁 End-to-end integration with vime, vLLM rollout, and Megatron-LM training 🧩 Deterministic Attention, dense FFN, LogP, reduction, and collective operators 🌐 NVIDIA CUDA and AMD ROCm support, with initial Ascend integration and MUSA adaptation underway GitHub 👇 github.com/RL-Align/RL-Kerne…
1
4
8
45,082
RL-Kernel v0.1.0 is now available. 🎉 Bitwise train–inference consistency across parallel configurations and heterogeneous execution engines for RL post-training. With Qwen3-8B on a full vime pipeline, RL-Kernel recorded zero train–rollout LogP mismatches across all 200 training steps. 1,129 commits from 29 contributors. Highlights: 🎯 Strict numerical alignment across rollout and training engines ⚡ 68.7% higher rollout throughput and 8.2% lower end-to-end step time on NVIDIA H100 🔶 AMD Instinct MI300X: zero mismatches across 9,400,614 active-token LogP elements from 1,600 samples 🔁 End-to-end integration with vime, vLLM rollout, and Megatron-LM training 🧩 Deterministic Attention, dense FFN, LogP, reduction, and collective operators 🌐 NVIDIA CUDA and AMD ROCm support, with initial Ascend integration and MUSA adaptation underway GitHub 👇 github.com/RL-Align/RL-Kerne…
1
4
8
45,082
Inference vLLM and training Megatron execution graphs never quite align. Single-operator boundaries break down, and treating kernels from one side as a drop-in contract simply doesn't work. That's why we built RL-Kernel 🚀 Today, we’re dropping our first milestone: an operator-level training-inference consistency analysis & roadmap preview for DeepSeek V4 Flash MoE. Our goal is to build reusable, hardware-agnostic infrastructure bridging both worlds. If your daily stack involves writing low-level CUDA/Triton kernels or wrestling with vLLM & Megatron MoE distributed comms, you'll feel right at home here. We're fully open-source and looking for core contributors. We'll be dropping the DS V4 roadmap in our GitHub issues over the next week or so, complete with detailed task breakdowns. External contributors are more than welcome to jump in! Grab an issue, join us, and let's ship this together! 👇 Read the Blog : rl-align.github.io/RL-Kernel… #RLKernel #DeepSeek #AIInfrastructure #RLHF #CUDA #Triton #MoE #RLAlign #ROCm
1
1
3
149