Congrats
@CoreWeave on RL Rollouts!
RL post-training involves a lot of back and forth: train the model, generate responses, then train again. Inference workers need to load the updated model weights each time. As models get bigger, that can leave GPUs waiting.
CoreWeave’s new service uses ModelExpress and Router in NVIDIA Dynamo to speed up those reloads with minimal downtime.
Working with us and
@youdotcom, CoreWeave achieved 15× faster model reloads compared with its baseline while post-training Nemotron 3.5 Lightning.
Check out their blog below for details