The interetsing question is whether domain-specific training actually changes the progress–compute scaling law, or if it mostly just shifts the curve left by improving sample/compute efficiency at a given capability level. If it’s majorly a constant-factor gain rather than a better scaling exponent, then sufficiently strong general models may (and will) eventually eat that advantage through scale. We have seen this earlier, even in robotics (open-x embodiment / RT-X model)