They tested this across a genuinely wide spread of data: DCLM web text, GitHub Python, C source, Metamath proofs, CIFAR-10 images, speech commands, music, even human DNA. Self-play's scaling rate came out comparable to models trained directly on that real data, in some cases even a bit faster. DNA was the outlier, scaling much slower than real DNA-specific models.