Why is RL so sensitive to training-inference mismatch (TIM)?
With each training step, the trainer is pulled toward the sampler, resulting in accumulating drift.
We exploit this intuition to devise a novel correction method to stabilize RL under TIM, called Score Centering.