idk what we're actually arguing, everyone is using stronger models to get crazy speeds I'm just stating the obvious.
Take any model on hugging face, point Fable at it, have it get it working as quick as it can under normal circumstances, let it test between vLLM and SGLang for example, see if it has mtp or dflash, (Fable / Astra / Opus 5.5) already know what these are and will proactively look.
Then once it hits a reasonable speed you like, tell it you need it 30% faster and to try all methods.
I've done this so many times now.. I got mimo flash running on my local rig this way, it discovered upstream PRs, it even fixed and patched a few of them once we started turning the knobs on kv cache.
It started to tank after it filled with context... I had it test that thoroughly and find the reason behind it. I get 4 streams at 120 tk/s at 300k context.. I didn't manually touch a single thing.. hell it even downloaded the models to my box.
Models have come a LONG way and are incredibly capable, I still feel like we're not even scratching the surface.