10X Builder | AI Performance Engineer | Co-founded @mybridgecard(YC S22)

Making GPUs Go Brrr
Yet again, engineers coming up with complicated names for things. Tensor Parallelism,Expert Parallelism,AllGather, AllReduce. e.t.c. At the end of the day, if you follow the abstractions deep enough, you can understand most of these concepts from ideas you’re already familiar with. I’ll be speaking at #SysConf26! I won’t promise to make you an inference engineer but by the end, some of the buzzwords would sound a lot less mystical. Register!
We are excited to announce Festus Owumi (@_enfinity) as a speaker at SysConf 2026. His talk is titled “Distributed LLM Inference From First Principles: When One GPU Isn’t Enough.” Distributed inference solves at least two important problems: memory capacity limits on GPUs and a need for higher serving throughput. But how do we build distributed inference? In this talk, Festus walks us through building a distributed inference runtime from first principles, based on a learning project. It starts by reviewing the “boring” case of one model, one process and one [CPU] device. Then the talk will introduce ranks, world size and point-to-point communication. After the above foundation, Festus will introduce more complex concepts to distribute the model such as Pipeline Parallelism, Tensor Parallelism and the techniques used to achieve them including partitioning and collectives such as AllGather or AllReduce. #SysConf #SysConf26
9
12
136
6,130
Festus retweeted
LLMs keep getting bigger. Compute budgets, unfortunately, do not. This fall @Stanford, we (@AnayMehrotra @gvelegkas and Amin Saberi) are teaching MS&E 319: Efficient Generative Language Models. The course asks a simple question: given a modeling goal and a limited computational budget, how should we choose the training objective, model architecture, and inference algorithm? We’ll cover some of the main ideas, and occasionally surprising tricks, that make large language models more efficient, including efficient pre-training, mixture-of-experts architectures, attention and KV-cache compression, quantization, speculative decoding, LoRA, RLHF, DPO, and distillation. And no, “just buy more GPUs” will not be the only answer. We’ll try to post all the lecture materials and recordings as the course unfolds. web.stanford.edu/class/msand…
16
73
607
72,267
Google is rethinking the Kubernetes control plane for agentic workloads! 🤯
🥳 Excited to start revealing what we've been working on in the last few months. First, we decided to reinvent Kubernetes for agentic workloads with statefulness and fast resumption. Secondly, we are building an agentic orchestrator that will be Google's open agentic orchestrator and runtime. github.com/google/ax
2
305
This is soo sooooo coool!
We built a time machine for the web. Introducing Exa Snapshot: an index of 400 billion historical snapshots of webpages that lets you search as if it's the past. Snapshot is already being used for backtesting prediction models, RL at labs, exploring the pre-AI web, and more.
1
237
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
4,040
8,286
76,411
39,757,893
Large language models normally write one token at a time, each word waits on the one before it, which is slow and leaves the GPU mostly idle. Speculative decoding flips that: a small, fast model guesses the next several tokens, and the big model checks them all in a single pass, keeping the good ones and correcting the rest. In his talk at the 2025 edition of SysConf, Habeeb Shopeju (@HAKSOAT ), Senior ML Engineer at Autodesk Research, broke down this and the other tricks - KV caching, paged attention - that inference engines use to go faster.
1
5
34
2,004
Some early findings! DeepSeek-V4.1-Flash from @deepseek_ai has 196B Engram parameters that can be moved out of GPU HBM. I tested three placement modes: device, synchronous host, and host with asynchronous prefetch on a node with 4× @NVIDIA B200 GPUs using a pinned @sgl_project build.🧵
Replying to @jaga_prasanna
I saw the DeepSeek V4.1 Day 0 preview implementation from @sgl_project and was particularly interested in its support for keeping the Engram tables in GPU HBM or placing them in host memory with optional asynchronous prefetch: github.com/sgl-project/sglan… So I’m running a controlled benchmark across these three modes to quantify how much HBM host placement frees, the additional KV-cache capacity that it enables, the resulting host–device transfer latency and throughput costs, and whether asynchronous prefetch can recover the performance lost to offloading. Still running however!
1
519
These results come from one controlled comparison in which all three configurations were tested under matched conditions, the experiment has not yet been repeated across independent server initializations or additional hosts. Thus it is a controlled evidence of a clear memory–performance trade-off, but not yet a universal deployment recommendation. If you or your organization would like to sponsor the further study repetitions, I can test all three configurations across fresh server initializations and additional B200 hosts measuring variability, confirming reproducibility, and locating the SLO boundary needed for a stronger deployment recommendation.
1
35
You can also view the result more deeply on LLM Inference Optimization Atlas (ayobami-00.github.io/llm-inf…) This experiment is also part of a larger project I have been working on called the LLM Inference Optimization Atlas (ayobami-00.github.io/llm-inf…). I built it after working on several inference optimization studies and realizing that results are often easier to find than the conditions under which they remain true. An optimization result without its workload, hardware, constraints, and rejected alternatives is easy to misinterpret. The Atlas records inference engineering as an evidence graph, connecting each workload and hypothesis to its configurations, experimental runs, comparisons, findings, and final deployment decision.
31
👨‍🍳🚶
good luck inferencing this bad boy holy architecture
1
1
383
Deepseek has cooked again! deepseek-ai/DeepSeek-V4.1-Flash dropped some hours ago with a new Causal Encoder–Decoder architecture, native multimodality and very aggressive KV-cache compression. This might actually be a new sign of a genuinely new frontier architecture direction since the decoder-only models by GPT-2 took over. Crazy times! 🧵
1
149
Here are the benchmark results. OpenDesign has DeepSeek V4.1 Flash at 81.2 basically within touching distance of GPT-6 Astra’s 82.7 and from the technical report it's wayyyy cheaper and faster to run.
1
1
1
71
DeepSeek has made the KV cache roughly 437× smaller across generations from the deep seek v1. This means you can get close to GPT-6 Astra's performance at a lower cost, first off because you are not paying per token pricing but paying per GPU-hour and because the kv-cache size is greatly reduced you can run up to larger concurrency and longer context sizes!
1
1
38
the group project finally has one useful task: send your referral link 😭 confirmed-email referrals earn points. advertised launch prizes: $100 USDT for 1st, $50 for 2nd, $25 for 3rd. polyversemarket.com #polymarket #polyversemarket
1
2
11
Quite an interesting conversation. piped.video/watch?v=rpr_Zpbp… Listening to @dylan522p discuss disaggregated inference brought me back to Sara Hooker’s idea of the hardware lottery. 🧵
1
1
1
134
Rather than every chip supporting every model equally well, could specialized accelerators target different parts of the LLM architecture building on ideas we’re beginning to see with NVIDIA–Groq? Imagine attention on one accelerator, feed-forward on another, prefill on another, decode somewhere else, CPUs handling orchestration or lighter stages and perhaps memory-heavy or routing operations moving to entirely different hardware. Inference starts to look more like an assembly line where each stage is handled by the hardware best suited to it. Of course, this only works if we have much faster, lower-latency interconnects between these heterogeneous accelerators.
1
12
Maybe the future isn’t one chip for every model, but an elastic fabric of specialized accelerators that can be mixed and matched around the model. Could this be where AI hardware is ultimately headed?
10