Pinned Tweet
Inspired by @percyliang's CS336, I built a static performance model for LLM inference — no dynamic batching/chunking, just clean analytical bounds. Pick model × GPU × batch × seq length × parallelism(DP/TP/EP/PP)  → VRAM check + TTFT + TPOT + prefill/decode breakdown + throughput vs batch size. Covers 5 KV cache variants (GQA/MLA/SSM/sliding window/linear attention), speculative decoding, multiple quant precisions. Calibrated against ~100 public benchmarks (TRT-LLM, Splitwise, MLPerf, Koyeb…). Demo: llm-inference-calculator-del… Repo: github.com/pochenai/llm-infe…
3
11
115
12,712
Po🎏 retweeted
Just for fun, finetuned Qwen3-14B (QLoRA) for parallel constrained decoding (Jev-like) - all JSON fields decided simultaneously via KV-cache broadcast, zero autoregression: - 4 fields in ~230ms on a single RTX 3090 (4-bit) - 100% schema-valid JSON by construction + calibrated per-field confidence Adapter: huggingface.co/Foodoo1/Qwen3…
6
1
20
1,845
Introducing CUDA Rust! CUDA Rust lets you write GPU kernels natively in Rust, not just launch them from it. Two paths: cuda-oxide for SIMT kernels compiled to PTX, and cutile-rs for Tile-based programming on stable Rust. Both can catch aliasing errors at compile time. Technical blog: nvda.ws/4hm1bHS
117
648
6,238
514,046
cool~
4B parameters. ZERO distillation. 61.5% on SWE-bench Verified 🤯 Meet FrogNano 🐸: Qwen3.5-4B post-trained purely with RL on synthetic tasks from TaskPilot. Just 5 iterations × 300 tasks. Who said coding agents have to be huge? 🐸
84
Openai给内部的开发人员权限真大, Tibo经常直接给大家充值额度。 昨天参加一个OpenAI给公司内部的workshop,做了下showcase,主持人直接给我送了3个月20x pro,感动落泪
1
5
271
If this project helps you, please consider giving it a ⭐ Star! github.com/pochenai/llm-infe…
Inspired by @percyliang's CS336, I built a static performance model for LLM inference — no dynamic batching/chunking, just clean analytical bounds. Pick model × GPU × batch × seq length × parallelism(DP/TP/EP/PP)  → VRAM check + TTFT + TPOT + prefill/decode breakdown + throughput vs batch size. Covers 5 KV cache variants (GQA/MLA/SSM/sliding window/linear attention), speculative decoding, multiple quant precisions. Calibrated against ~100 public benchmarks (TRT-LLM, Splitwise, MLPerf, Koyeb…). Demo: llm-inference-calculator-del… Repo: github.com/pochenai/llm-infe…
1
156
Inspired by @percyliang's CS336, I built a static performance model for LLM inference — no dynamic batching/chunking, just clean analytical bounds. Pick model × GPU × batch × seq length × parallelism(DP/TP/EP/PP)  → VRAM check + TTFT + TPOT + prefill/decode breakdown + throughput vs batch size. Covers 5 KV cache variants (GQA/MLA/SSM/sliding window/linear attention), speculative decoding, multiple quant precisions. Calibrated against ~100 public benchmarks (TRT-LLM, Splitwise, MLPerf, Koyeb…). Demo: llm-inference-calculator-del… Repo: github.com/pochenai/llm-infe…
3
11
115
12,712
Estimator on the same 4×H200 FP8 box, 20k-in / 512-out:batch 16 seems pretty correct → 0.77s TTFT, ~1655 tok/s total (~103 tok/s/stream), close to your ~100 tok/s/stream @ 16 llm-inference-calculator-del… Theoretical maximum throughput (batch 886) → ~11.2k tok/s (~12.6 tok/s/stream), but ~40s TTFT) llm-inference-calculator-del…
here we go again: deployed a FREE public endpoint for Qwen3.8-Flash-Next 🚀 (going at +100 tok/s) No token needed, OpenAI-compatible, vision + tool calls, 262K context, thinking from xhigh → off. Light rate limiting, be nice to your neighbors 🤗 4× H200 · FP8 · SGLang cookbook · ~140 tok/s per stream · ~100 tok/s @ 16 concurrent · 0.8s TTFT Guide + chat UI 👇
151
Wow — Stanford NLP Group just reposted my LLM inference calculator 🤯 Try it: llm-inference-calculator-del… Now supports the latest architectures: Qwen3.8-Flash-Next (N-gram hybrid), Kimi K3 (2.8T), GLM-5.3-Flash, DeepSeek V4, and Tencent Hy4-preview.
Inspired by @percyliang's CS336, I built a static performance model for LLM inference — no dynamic batching/chunking, just clean analytical bounds. Pick model × GPU × batch × seq length × parallelism(DP/TP/EP/PP)  → VRAM check + TTFT + TPOT + prefill/decode breakdown + throughput vs batch size. Covers 5 KV cache variants (GQA/MLA/SSM/sliding window/linear attention), speculative decoding, multiple quant precisions. Calibrated against ~100 public benchmarks (TRT-LLM, Splitwise, MLPerf, Koyeb…). Demo: llm-inference-calculator-del… Repo: github.com/pochenai/llm-infe…
3
26
3,572
seems the main reason for charging the gas limit instead of gas spent is that Monad uses asynchronous execution: consensus first decides which transactions to include, and only then does execution happen, so this issue becomes more severe. In contrast, current clients like reth and the geth client simulate execution first before deciding whether to include a transaction in a block, and ultimately only count block space usage based on gas used, so the problem is not as serious.
CL researchers @liobaheimbach @KushalBabel and @jason_of_cs analyzed gas pricing on Ethereum and Base. The results are complex but in summary, they produce real evidence in support of: 1. Multi-dimensional gas metering (i.e. charging separately for state growth, removing it from being multiplied by current gas price which is variable). 2. Not relying on access lists in order to give reasonable UX -- optimistic parallel execution is much better. It is a terrible experience to submit a transaction that touches certain slots in simulation, and have the transaction revert due to a change in slots touched. Users shouldn't have to worry about slots, they sign high-level function calls. 3. Charging gas limit instead of gas spent, to reduce free-rider spam seen on Base where arbitrageurs repeatedly ping low bids and high offers in the hopes of executing at a good price, while halting early to avoid spending much gas. On (1), we have argued for this before but delayed implementing it to reduce complexity for Monad mainnet. Definitely worth revisiting. On (2) and (3), these behaviors are already live in Monad. This research piece might fly under the radar, but it is helpful to have rigorous scientific studies of the data to confirm intuition for the right design.
2
133
free-rider spam prevention by charging gas limit is interesting
CL researchers @liobaheimbach @KushalBabel and @jason_of_cs analyzed gas pricing on Ethereum and Base. The results are complex but in summary, they produce real evidence in support of: 1. Multi-dimensional gas metering (i.e. charging separately for state growth, removing it from being multiplied by current gas price which is variable). 2. Not relying on access lists in order to give reasonable UX -- optimistic parallel execution is much better. It is a terrible experience to submit a transaction that touches certain slots in simulation, and have the transaction revert due to a change in slots touched. Users shouldn't have to worry about slots, they sign high-level function calls. 3. Charging gas limit instead of gas spent, to reduce free-rider spam seen on Base where arbitrageurs repeatedly ping low bids and high offers in the hopes of executing at a good price, while halting early to avoid spending much gas. On (1), we have argued for this before but delayed implementing it to reduce complexity for Monad mainnet. Definitely worth revisiting. On (2) and (3), these behaviors are already live in Monad. This research piece might fly under the radar, but it is helpful to have rigorous scientific studies of the data to confirm intuition for the right design.
89
发现jeff dean也是这么理解CoT的工作机制,有点小激动 piped.video/watch?v=UTTeXZrp…
最近一直在想 CoT 为什么 work。一开始直观想到的是两个常见解释:上下文更长了、能提取出更多知识,或者多做了几次 forward pass(算力更多)。 想了想,第一个是错的,第二个方向对但没说到点上。真正说得通的是计算复杂度这条线。 - 先说上下文:固定 L 层的 Transformer,一次前向传播就只有 L 层深度。而且 prefill 是并行的——prompt 再长也只是加宽度,不加深度。塞再多东西,还是锁死在 TC⁰(电路复杂度类)里。其实早期CoT(Chain-of-Thought Prompting Elicits Reasoning in Large Language Models)指的就是这个意思,但是不work - 再说 forward pass:方向是对的,但光「多做几次」不够。Let's Think Dot by Dot:Hidden Computation in Transformer Language Models这篇paper发现,无意义的「......」填充符也能提升某些任务,但论文明确说填充符再多也留在 TC⁰ 内——同样是多次 forward pass,一个突破了、一个没有。 差别在于写回的 token 里装了什么。CoT 把真实算出的中间结果写回输入流,下一步从第 0 层重新进入 → 深度预算重置 → T×L → 这才真正跳出 TC⁰。 不过复杂度只决定了天花板。还得靠 RL 用「答案对不对」这个信号,把模型推到真正会用这块草稿纸的地方——复杂度定天花板,RL 决定你够不够得到这个天花板。
1
129
很有前瞻性的论点,推理的软件护城河正在消失,价值正沿着“推理引擎 → 服务运营 → 资本/产能 → 数据中心”这条路径不断下沉
1
134