Preserving entropy is critical for continued training; in modern post-training recipes, entropy is often a fixed resource that gets exhausted over the course of a training run, making it difficult for the model to improve and learn on new tasks.
Adaptive entropy control methods expose entropy as a controllable hyperparameter instead of a side effect. We find that in a continued training setting, GRPO leads to entropy collapse which stalls training performance on subsequent training phases. REPO-R (Petrenko et al., 2026) can hold entropy near a specific target, preventing this collapse, and pushing performance higher in follow-on tasks when the target entropy value is properly tuned.