There's something missing from benchmarks. Kimi K3 often writes much better code than GPT-6 Astra. It doesn't write useless verbose tests, it doesn't give confusing explanations. To be a good software engineer, you have to be able to communicate well, and RL doesn't capture that.