Our RL stack now supports NIXL weight transfer, reducing trainer-to-inference transfer time 9x compared with NCCL: from 86 seconds down to single-digit seconds for an 800B-parameter model, and even <4 seconds in our experiments. For prime-rl users, this means over 25% more throughput end-to-end compared with our previous speed. It also clears the way for fault-tolerant, elastic inference scaling that NCCL's rigid process groups made difficult.
7
26
189
37,231
Over the past few months, we've spent a lot effort improving our trainer and inference, reducing GLM 5.2’s training time to <4 minutes per step. However, these improvements introduced a different bottleneck: weight transfer. When training speeds up, trainer-to-inference transfer makes up a larger fraction of the overall step time. As MoE models pack in more total parameters, sync time can become a major impediment. Using NCCL, weight transfer for GLM 5.2 (a 800B model) takes up ~60–90 seconds per update: this was a large enough proportion of the step time that we've decided to optimize it.

Sep 3, 2026 · 8:07 PM UTC

1
19
3,887
Instead of NCCL, we rebuilt the weight transfer on NIXL, using remote direct memory access (RDMA) to read and write bytes directly between GPUs. This bypasses the CPU entirely for fast synchronization. But doing it on an open-source inference engine isn’t trivial: you need to know exactly where each byte land, and that runtime weight format isn't documented, but instead changes with hardware, model, and parallelism strategy.
2
15
1,456
To solve this, we adapted a concept of LazyTensor from the vLLM team: we traced the exact operations vLLM performs when loading weights and built a map from trainer bytes to their destination. This allowed us to transfer only what's needed from GPU to GPU, with minimal post-load processing. The result: less memory waste and faster transfer.
1
1
27
1,508
Because we trace vLLM's real load path instead of hand-coding it, this method works automatically for any model, kernel, or parallelism strategy. In other words, it eliminates the need to rewrite the mapping every time a new model or quantization scheme is introduced.
1
13
732
With sharding-aware RDMA transfer using NIXL, we cut GLM 5.2’s weight transfer time down to only ~9 seconds, including engine pause and resume — transferring the full 1.6TB policy between hundreds of GPUs. Experimental path allows us to cut this down further, saturating the network bandwidth at ~4 seconds per update.
1
15
1,625
Sort replies: Relevant Recent Liked