3/
The result: with only a final-answer reward, RL solves held-out problems the base model basically never solves.
And we can see how: it first sharpens primitive skills, then builds composed procedures out of them, and reuses them as a stable toolkit.