Learning how systems run fast

NYC
Prefix-conditioned OPD in a nutshell:
1
1
100
training MoEs requires balancing expert load at two scales: global balance prevents expert under-utilization local balance keeps expert-parallel execution efficient in our new paper, we address both: • EQB makes distributed global QB exact, improving over the approximate histogram variant used by Kimi K3 • LEI improves local balance over the standard GShard auxiliary loss at comparable quality validated on 7.5B MoEs trained up to 500B tokens. paper: arxiv.org/abs/2609.28053 🧵
5
23
132
16,205
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
80
450
3,501
713,944
Kimbo retweeted
Great work from the @TileRT_AI and @AIatAMD teams, who got GLM-5.3 to 469 tok/s single-user decode with vLLM on 8× MI355X on @SemiAnalysis_ AgentX. The run uses a disaggregated setup where vLLM handles prefill and TileRT handles latency-critical decode through vLLM's V1 connector interface. 1/2
AMD MI355X becomes the first official TileRT result no AgentX! This configuration achieves an astounding 470 TPS on GLM 5.3 (FP8). This is over 40% faster than GB300 TRTLLM using FP4! Great work to the TileRT x AMD team! (1/3)🧵
19
18
106
12,701
Kimbo retweeted
Training and rollout logprobs matched bit for bit on ROCm. The @RLKernel team integrated RL-Align/RL-Kernel with vllm-project/vime. A 200-step Qwen3-8B GRPO run on 8× @AMD MI300X recorded zero logprob mismatches between Megatron training and vLLM rollout. The strict path aligns reduction order, intermediate precision, rounding points, and math primitives across both sides. Deep dive: vllm.ai/blog/2026-09-14-rl-k…
5
9
76
6,444
Kimbo retweeted
Most open source RL environments tend to live on two extremes of a spectrum: - Simple, educational games like Wordle to learn how RL works - Complex, long-horizon tasks like Terminal Bench Science that are designed to probe the limits of frontier models We found that there's a huge gap in the middle, especially for those wanting to improve the coding capabilities of models <10B parameters. Today we're filling that gap by releasing >5k RL environments that are specifically designed for real-world data science tasks and are a great testbed for developing the next variant of GRPO or OPSD ;) Awesome release by Adithya who single-handedly built the envs, defeated dozens of challenges with async RL and figured out how to make all this run on a single GPU 🔥
Releasing SmolDataEnvs 🤗 5K+ Verifiable RL Environment tasks for hill-climbing small models in code and data science. Completely open source: environments, evals, training
7
22
254
20,171
in @stochasticchasm tradition, I am going to do a paper breakdown of the new DeepSeek Elastic Compute after reading through it. looks very nerd snipe-y and interested in the lessons they learned. arxiv.org/pdf/2609.22978
11
23
222
28,680
Today, we're releasing Prime VM Sandboxes, co-designed with our research team for large-scale agentic RL training with tens of thousands concurrent sandboxes.
Introducing Prime Sandboxes: MicroVM sandboxes purpose-built for RL training. Model training requires running tens of thousands of concurrent sandboxes, leading to complex and costly configuration. We built Prime Sandboxes for our own team. Today we're releasing them publicly.
6
11
192
10,160
Kimbo retweeted
Our first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode, against 450 tokens/s for our GB200 baseline, both with DSpark speculative decoding. To our knowledge, this is the first TPU inference megakernel. The whole model runs in a single Pallas kernel, and without spec decoding it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8. We are open sourcing it today. 1/2
23
93
853
236,639
Kimbo retweeted
What if you had 100, 1,000, or even 100,000 GPUs to scale LLM RL—what would you scale first? 🤔 One of the most fundamental knobs is batch size. But how should it grow with available compute—and where should it stop? This may sound like hyperparameter tuning. At scale, however, it becomes a question of sample efficiency, hardware utilization, and ultimately the GPU-hours required to reach a target capability. Answering it requires a principled understanding of RL scaling. So what is the right batch-size scaling law for LLM RL? I’m excited to share my recent research paper on this topic. #LLM #RL #Scaling
7
17
185
19,789
Kimbo retweeted
MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today. Less prefill, a smaller KV cache, better long-context retrieval—and we got all three at once. Compared with MiMo-V2.6's Hybrid SWA architecture: • 5.02× lower prefill FLOPs at 1M tokens • 4.5× smaller KV cache at 1M tokens • Better MRCRv2 and RULER-v2 scores, plus lower AgentPPL and LongPPL Why build a new architecture? Agentic inference is a very different workload. Each round, a short action can return a long observation that needs to be prefilled, while the context keeps growing. That puts prefill cost, KV-cache size, and retrieval accuracy on the critical path at the same time. HySparse2 tackles all three with two levels of KV sharing: • KV Bridging: Following YOCO, full-attention layers in the cross-decoder build their K/V from self-decoder hidden states. • KV Reuse: Within each hybrid block, sparse layers reuse the preceding full-attention layer's KV cache and selection indices. Two more changes: token-level selection replaces block-level selection, and a forced window of recent tokens replaces the separate SWA branch, so local and global tokens share one KV cache. Since all cross-decoder KV caches now come from the self-decoder, prefill can stop once the self-decoder finishes. Paper: arxiv.org/pdf/2609.26368
212
385
4,356
395,968
DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale Jialiang Huang, Hongxuan Tang, Jingchang Chen, Yuxuan Liu, Yixiao Chen, Yuan Cheng, Yi Tao, Jingli Zhou, Yupeng Chen, Haoyu Chen, Jiarui Wang, … arxiv.org/abs/2609.22978 [𝚌𝚜.𝙳𝙲]
6
19
1,396
Dylan Patel (@dylan522p) of @SemiAnalysis_ says open source is dying: "There's multiple Chinese model labs who are telling all the inference guys, 'Our next model's not going to be open source. We're going to license it to you.'" " Open is dying quickly, unfortunately."
Community note
This July 2026 quote (recirculated without date) claimed Chinese labs were shifting away from open source. Multiple labs have since released large open-weight models, including Qwen updates in Aug-Sept 2026. huggingface.co/blog/state-of-… rits.shanghai.nyu.edu/news/ wired.com/story/chinas-o… x.com/sourceryy/stat…
13
14
441
76,451
Kimbo retweeted
MiMo-V2.6: The Hard Road to Scaling Up RL MiMo-V2.6 is very likely one of the largest single RL runs, by compute, that any open-source model team has undertaken to date. In an era when compute is brutally scarce, we still chose to dedicate a team of several dozen people to one goal over an extended period: scaling up RL. That takes more than research conviction. It takes a vision for AGI, respect for the unknown, and the nerve to walk straight into the hardest problems. The result is a model whose potential was built through mid-training and unlocked through heavy RL. Today, it is the number one open-source model. I strongly recommend reading the technical report. I believe it will become one of those papers that Agent RL practitioners keep reopening and discovering something new in each time. In my view, the research innovations and engineering challenges behind it surpass those of DeepSeek R1, which I was partly involved in. Some will ask: why MixRL instead of MOPD? First, they are not competing choices. We ran MixRL on verifiable tasks of moderate difficulty, including code and related agentic tasks, and found that the resulting models generalize remarkably well. Second, tasks that are difficult to verify, extremely long-horizon, or simply too challenging to include in a joint RL run are trained separately. Including them would substantially reduce rollout efficiency or introduce significant rollout staleness. We then merge the resulting capabilities through MOPD. Games, 3D tasks, and tasks with subjective evaluation signals all fall into this category. There is also a third, slightly cheeky answer. Our team is flat enough and free enough of organizational silos that MixRL simply is not difficult for us. More importantly, everyone enjoys working this way. People from different domains come together every day, driven by the pursuit of AGI and intelligence that can continuously improve itself, to confront and resolve the RL bottlenecks in each field. I will always remember the RL daily update meetings from this period. They were intense and dense, with intelligence emerging in real time. To help the open-source community focus on solving real Agentic RL problems, we have released a Qwen model distilled from MiMo RL trajectories as a stronger starting point for RL, along with 7K diverse environments and a complete RL training framework. We hope these resources will help move Agentic RL research forward. MiMo-V2.6 is only the beginning. In an era when intelligence is easy to replicate, we still choose the hard road toward self-improvement and AGI. Much of what lies ahead remains unknown. But we are willing to keep investing the time, compute, and passion required to take on one hard problem after another and work each of them all the way through, until intelligence crosses into a new regime.
235
357
4,322
654,545
Kimbo retweeted
We enable full-parameter RL on TPUs: MiMo-V2.6 at 310B, plus other stable training runs of 1,000+ steps across 1,000+ TPUs. With JAX, scaling up is a config change, not a rewrite. We built on that with optimized vLLM inference for faster rollouts and full bitwise trainer–sampler agreement in validation. Trainer and sampler share one TPU ICI fabric. All 310B MiMo-V2.6 parameters transfer in <2 seconds.
6
24
106
48,579
Kimbo retweeted
Very intriguing discussion here to rethink policy gradient: we always want to optimize E_p[R], so we would need the IS techniques that introduced all variance. Score centering is taking a different view: let’s just optimize E_q[R], as q (sampling distribution) is what we finally care about. As samples are indeed from q, no IS is needed. And since we cannot optimize q (inference engine/old policy) directly, we just use p’s grad as a surrogate to q’s grad (yes its biased), as explained in the STE estimate here. A change of perspective (from optimizing p obj to q obj) gives a lot of simplicity!
Actually this argument is not limited to linear policy. Consider the gradient w.r.t. logits. s = ∇_z log p(a) = e_a - p \bar{s} = E_{a~q}[e_a - p] = (q - p) centered score s - \bar{s} = e_a - q Thus, score centering is a STE: z_p + sg(z_q - z_p).
2
8
95
13,488
Kimbo retweeted
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself 🤗: hf.co/papers/2609.22068
7
24
261
87,728