There's no magic in "pure synthetic data" for post-training, just as there are no perpetual motion machines.
The alpha either comes from ground truth that you source or from handcrafted recipes (you are the alpha).
Ideally, you want both to increase the leverage of the source.
Making RL environments for frontier models is kinda like On Policy Self Distillation.
You're injecting some privileged information to create tasks that models wouldn't be able to conceive of otherwise.
The difference is the info is encoded in the env, rather than the prompt.