We are excited to announce Festus Owumi (
@_enfinity) as a speaker at SysConf 2026.
His talk is titled “Distributed LLM Inference From First Principles: When One GPU Isn’t Enough.”
Distributed inference solves at least two important problems: memory capacity limits on GPUs and a need for higher serving throughput. But how do we build distributed inference?
In this talk, Festus walks us through building a distributed inference runtime from first principles, based on a learning project. It starts by reviewing the “boring” case of one model, one process and one [CPU] device. Then the talk will introduce ranks, world size and point-to-point communication.
After the above foundation, Festus will introduce more complex concepts to distribute the model such as Pipeline Parallelism, Tensor Parallelism and the techniques used to achieve them including partitioning and collectives such as AllGather or AllReduce.
#SysConf #SysConf26