🍰 CAKE paper's out, the design bet: the compiler isn't a fixed black box the agent calls — it's part of the harness, and it's under evolution too. CAKE didn't inherit existing abstraction layer. no tile/layout abstractions: the vocabulary was distilled by agents from a corpus of production kernels. every pattern the agent couldn't express pushed new primitives into the IR, and the analyses to keep them checkable. every barrier/layout bug that kept coming back became a verifier rule. none of this can be designed up front. the IR has to co-evolve with the kernels, and the workload tells you what's missing, the corpus tells you if the fix broke anything. the best language for an agent is the one that tells you what's illegal, what's slow, and which decision might made it faster. arxiv.org/abs/2608.12629
🍰 another slice of CAKE just dropped: blackwell KDA kernels, ~2.05x geomean speedup vs FlashKDA: github.com/flashinfer-ai/fla… by @avyyhuang more cake kernels are baking 👀 not saying what CAKE actually is yet though, the full reveal is in the oven too.
13
81
323
148,603
Zihao Ye retweeted
🔥 Optimizing LLM serving requires repeated experiments, and running them on real hardware is slow and costly. Performance depends on how models, kernels, hardware, and serving policies interact. As models and workloads change, engineers must revisit their configuration choices. 🚀 Introducing ServingStudio: an integrated workbench for simulating, analyzing, and optimizing LLM serving systems. 🧵 More in the thread below. (1/9)
2
13
28
5,683
Zihao Ye retweeted
pitch: one kernel library per model. not even model family
7
1
54
5,214
Zihao Ye retweeted
We tested Jev as a code reranker on 100 GitHub issues. Model-planned grep + Jev was our best route, though the choice of retrieval route mattered more than the choice of reranker. The benchmark is CodeNib Base: 25 repos across C/C++, Go, Python, Rust and TS/JS. The query is the full issue text, and each repo is checked out before the fix. We built candidate pools four ways (BM25, dense embeddings, BM25+dense hybrid, and grep queries planned by a single Sonnet 4.6 call), capped each at 100, and reranked every pool with both Jev and Qwen3-Reranker-4B. Jev answers typed questions, so we asked it one relevance score per candidate and sorted the scores ourselves. Code-block Recall@5: grep → Jev: 71.4% hybrid → Qwen 4B: 65.0% dense → Qwen 4B: 63.4% Moving from BM25 to grep candidates raised Recall@5 by 10 to 14 points for every reranker we tried. On the same pool, Jev and Qwen 4B were within noise of each other. Jev had the higher point estimate on the grep pool (+3.3), but the interval crosses zero. Grep supplied 17.5 candidates on average, so Jev's scoring took 0.48 s p50 and cost $0.045 across all 100 issues. The planning call is the expensive part: 4.0 s p50 and about $1.00 per 100 issues. If that's too slow, hybrid → Jev reached 67.9% at 3.4 s and $0.26 per 100 issues. Skipping embeddings doesn't buy latency by itself, since warm dense lookup took 18 ms. This is one run per configuration on 100 issues, so treat the intervals as exploratory. All the numbers and caveats are in the thread.
TypeSafe's Jev returns typed scores instead of text. We tested it as a code reranker on 100 real GitHub issues. With a model-planned grep in front, Jev hit 71.4% Recall@5, vs 63.4% for dense retrieval → Qwen3-Reranker-4B, at about the same median latency (4.6 s).
1
3
465
Zihao Ye retweeted
Claude Opus 5.5 drew every frame of this animation in JavaScript
51
80
1,514
143,814
Zihao Ye retweeted
Replying to @grok @lurker670291
This is in fact, IMO, the correct answer!
1
1
16
977
Zihao Ye retweeted
nand2mario.github.io/posts/2… This is a fantastic article on building a Voodoo GPU recreation, a lot of the architectural challenges, including having high latency DDR. Excellent read for those interested in GPU architecture :)
11
121
8,425
Zihao Ye retweeted
4 months after its initial release, "auto-gpu-kernel" version 1.0 is finally out! It is a fully autonomous kernel generation "meta-harness" that evolves both the kernel and the harness layer, generating speed-of-light kernels 🧵 github.com/Dogacel/auto-gpu-…
17
62
564
32,052
Zihao Ye retweeted
Have you ever wondered what a data race is? How a race detector works? If there are bugs in your race detector? I wrote an article. theconsensus.dev/p/2026/09/0…
4
28
306
22,294
Zihao Ye retweeted
🚨 Someone open sourced an entire animal. Not a model of an animal. The animal. A fruit fly, rebuilt joint by joint from microscopy, that walks, grips, flies and lands inside a physics engine running on your laptop. It's called flybody. Built by Google DeepMind and HHMI Janelia, published in Nature, and sitting on GitHub under Apache 2.0 right now. What's actually inside: A full body in MuJoCo. Legs, wings, head, abdomen, every joint articulated and physically simulated. A 59-dimensional action space in the walking task alone. That is how many things a network has to coordinate for the fly to take one step. Leg adhesion, so it actually grips surfaces instead of sliding off them like a game asset. Ready-made RL environments: walking imitation, free flight, and vision-guided flight where the policy flies on what the fly sees. The whole thing is Python. Clone it, open MuJoCo, and the fly is on your screen. And the part that got me: Nobody animated any of this. You hand the body to a reinforcement learning agent and it discovers how to move it. Same loop as a robot arm, except the robot is a bug and the reward is staying airborne. Which means people are now training it to do things no fly has done in 100 million years of evolution. 525 stars. Apache 2.0. Zero dollars.
133
608
5,539
879,672
Zihao Ye retweeted
Agents let us build systems for different workloads and requirements. But… can we trust what they build? We release 🌟SkySynth🌟: an engine for synthesizing high-performance, just-in-time (JIT) systems we can trust, by co-evolving formal proofs and tests alongside the code. Results: 💿 KV stores up to 2.3× faster than Redis and FASTER + formally verified stores with 2.9× Claude Code's pass rate 🚏 Model routers up to 48% cheaper than a general router ⚡ Specialized inference engine with 2.2× the throughput of vLLM/SGLang 🧵👇
10
56
218
59,751
Spent last weekend giving HoMM 3(Heroes of Might and Magic III) a generative AI graphics remake. Yes, the game that ate my childhood. @MeshyAI generated textured 3D models and starter rigs for the humanoid creatures. GPT-6 Astra (@OpenAI) drove Blender to rebuild the original animations and render it all. Generated artwork are packaged as mods to VCMI (the open-source recreation of HoMM 3 engine) . Necropolis is done. A few flaws, but overall looks stunning. Turns out remaking an old game is now absurdly cheap. The agent wrote up its own journey at: yzh119.github.io/series/enha…
2
2
47
2,532
Zihao Ye retweeted
please enjoy Colfax's most recent blog post, exploring an impactful optimization to the FlashAttention-4 backward pass, due to my brilliant colleague Jack Carlisle! more of these "optimization diaries" to come, complementing our usual tutorials. research.colfax-intl.com/opt…
3
24
152
6,768
我真的已经等不及 gh --attach 了。 现在的体验非常割裂: - gh pr create 能创建 PR - gh pr edit 能修改正文 - Playwright 能截图、录屏 - 但本地图片还得打开 GitHub 网页,手动上传,再把链接贴回 PR 自动化走到 99%,最后 1% 逼你切回浏览器。 请快点发布到稳定版 🙏 github.com/cli/cli/pull/1418…
3
2
9
4,242
Zihao Ye retweeted
⚡️ We are excited to share updates from KDA-v0.5 (kernel design agents). Our latest Cute-DSL kernels can beat the human winners at FlashInfer Kernel Contest by 16%~69%! An amazing progress in just 3 months! - GDN prefill: 1.69x Speedup over Kachua - DSA attention: 1.41x Speedup over Dogacel - FP8 MoE: 1.17x Speedup over Team Wombat Humanize2 Flame Chase: docs.humanfia.ai/humanize2/f… Results and Reproduction: github.com/mit-han-lab/mlsys… MLSys 2026 Contest: mlsys26.flashinfer.ai/ KDA-v0.5 achieved this by integrating the Cute-DSL primitive, better workflows (humanize1 -> humanize2), updated kernel-wiki (self-evolved), and better profiling skills (IKET). The results are achieved by the flame chase flow from humanize2 – using gpt-5.6-sol and fable-5 to iteratively optimize . We have released the kernels for validation and more details will come soon! By the KDA Team: Dongyun Zou, Yixin Dong, Junxian Guo, Changye Li, Yahui Cui, Zihao Ye, Junru Shao, Zijian Zhang, Sihao Liu, Song Bian and Ligeng Zhu.
2
37
219
26,593
Zihao Ye retweeted
欧吼稍微改了一下图数据库现在可以在 5090 上实时模拟果蝇的大脑,而且校准后结果和别人的实验数据基本对的上
2
3
61
3,554
Zihao Ye retweeted
Ten days ago: DFlash 2. Today: GLM 5.3 — drafter trained, NVFP4 checkpoint out, endpoint live on GB300s. DFlash 2 was the first piece. This is the full stack!
Congrats @Zai_org on the GLM 5.3 open-weight release! Day 0 from us, in collaboration with the Z.ai team: ⚡ DFlash 2 drafter ⚡ NVFP4 checkpoint ⚡ Live endpoint powered by @TokenRouter_US GB300s Up to 4.4× the throughput of native FP8 with autoregressive decoding! inco.ai/blog/glm-5-3
13
22
237
21,324
Zihao Ye retweeted
Checkout the latest XGBoost update "Introducing the XGBoost Vector-Leaf Model" from jiamin yuan and Rory Mitchell xgboost.ai/2026/08/25/introd…
7
70
476
106,470
Zihao Ye retweeted
If you’ve done serious compiler optimization work on a product with a sufficiently large user base, you’ve almost certainly seen bug reports like this: “Why does this code compile into this assembly? There’s an obvious optimization opportunity here. Loop unrolling, LICM, etc. Why didn’t the compiler do it? If it did, my program would be faster. Please change the compiler so my program runs faster.” Very often, the answer is: “This isn’t really a compiler bug. It’s a heuristic. The transformation you’re asking for may indeed improve your program, but it could hurt performance in other cases or for other users. A compiler may have thousands or millions of users, and it has to preserve performance across an enormous variety of workloads. We can’t change an optimization policy just to make one program faster." This answer is 200% correct. I’ve lost count of how many times I’ve given some version of it myself. And yet, every time I say it, it’s one of the moments when I hate my job the most. As a user, why should I care about the compiler’s internal heuristics? Why should I care about the performance of everyone else’s programs? Why should I have to pay for something I don’t need? This needs to change.
10
4
95
13,419