I'm building an LLM inference engine from scratch, in public. Not using vLLM or TGI. The goal: re-derive every serving trick until I understand how modern inference squeezes throughput from a GPU. I call it nanoserve: github.com/pjdurden/nanoserv…
Day 76 of building an LLM inference engine from scratch. I ran my benchmark sweep as a real process expecting another crash like yesterday. It passed. The bug was in what the output left out. github.com/pjdurden/nanoserv…#buildinpublic
An American lab's open-weight pitch is now: as good as China's, for less compute.
Reflection says Beam costs less compute. Training or serving? Only one hits your bill: active params and KV bytes per token.
Show tokens/sec on one H100 node.
Your test fakes are a second copy of an API you don't own. They drift.
DeepSpeed #8730: Transformers now reads per-layer config, so the rollout cache fakes had to start providing it.
Every fake is a bet that upstream stands still.
ChatGPT's text gets a watermark in the EU. One rewrite erases it.
Text watermarks usually live in the sampler: tilt token picks toward a secret pattern, test for the skew later. Paraphrasing scrambles it.
Catches lazy copy-paste. Nobody else.
Rural data centers are getting a federal tax break. Latency won't stop the move.
A cross-country round trip is ~70ms. An agent turn decoding 2k tokens at 50 tok/s takes 40s. The network is a rounding error.
The binding constraint is power, not distance.
An eval that skips the answers it can't parse is grading on a curve.
NVIDIA/kvpress #299 keeps the Loogle key count when a prediction doesn't parse, so the miss stays in the denominator.
A failed parse is a wrong answer. Score it like one.
Every free image you generate burns real GPU time. OpenAI just found who pays for it.
One image request holds a GPU far longer than a chat reply. The free tier never covered that.
Ads next to the output aren't a feature. They're the subsidy.
Nvidia's 7-year-old Shield TV just got a $100 price hike. Blame your chatbot.
Memory makers keep moving wafers to HBM for inference GPUs. Plain DRAM gets what's left.
AI's real tax is landing on gadgets that don't run AI at all.
Sending KV cache to another GPU starts with telling the NIC which memory it may touch.
TensorRT-LLM #19680 registers Mooncake KV pools per GPU pool mapping.
Disaggregated serving is mostly bookkeeping about who owns which bytes.
Google froze its open source bug bounty. AI submissions buried it.
A plausible bug report now costs nothing to write. Checking one still costs a maintainer hours.
Same math is hitting PR queues. Slop scales. Review doesn't.
Day 75 of building an LLM inference engine from scratch. I ran my benchmark script as a real process for the first time, over a tiny checkpoint written to a temp dir. It died at boot. Three bugs, all there since Day 57.
github.com/pjdurden/nanoserv…
The worst was bug 2. Fix only the TypeError and the run finishes, then every ratio in the table compares two eager servers. A card run would have printed a number that meant nothing.
The toy checkpoint: config.json, safetensors under HF names with a tied head, and a byte-level BPE with zero merges. One byte = one token, so the test asserts the padded prompt is exactly 524 tokens. 14s on CPU. vLLM and SGLang are the systems I'm learning from. #LLM
Day 74 of building an LLM inference engine from scratch. My benchmark script now boots a third server, one running flash-decoding style split attention, and puts it next to the CUDA-graphed one. One command, three arms, same bytes. github.com/pjdurden/nanoserv…
And it counts the prompt alone, not prompt + max_tokens. A completion can hit EOS on its second token, so the budget is a row length on paper only. The client controls half the row and that half has to clear the chunk by itself.
vLLM and SGLang ship split decode in production. I'm learning what it takes to measure one honestly: the server under test sets the load, and five gates check the arm really split. 2338 tests green, first card run is one command now. #buildinpublic