Two small corrections: the post doesn't benchmark GLM-4.7-Flash on one GPU against anyone. It's only there for the SWE-rebench result + we have ran it on sequences of 39k median, not on 4k.
The fuller comparison against Axolotl, NeMo AutoModel, Megatron Bridge, MS-SWIFT and Unsloth is in the docs:
github.com/whitecircle/halo/… (multigpu too!)
That said, we should add more models and repros!! If you're seeing ~5% behind Axolotl on B200, could you share your config? We don't have B200s though, we ran everything on B300s.