Megakernels are sick, but ifl comparing low batch size performance against vLLM is kind of a cop out
inference engines are optimized for high throughput & big batches! very different regime
either actually compete at large batch size or just flex absolute numbers like MBU imo
Introducing the next evolution in LLM text generation: the first fully-fledged serving system built around a decode megakernel.
Delivering up to 1.58x faster performance than vLLM. Built for North Mini Code, completely open-source.