Introducing Prime Sandboxes: MicroVM sandboxes purpose-built for RL training. Model training requires running tens of thousands of concurrent sandboxes, leading to complex and costly configuration. We built Prime Sandboxes for our own team. Today we're releasing them publicly.
41
76
710
103,974
This looks like cool work esp on the trust region side But are we being too short minded with wanting single rollout RL. Theres definitely some benefits but as we move into tasks and domains that are less verifiable, i think we underestimate how important the idea of relative quality within a group can be
FlashREINFORCE: to our knowledge, the first open-source critic-free, single-rollout async LLM RL with 6,000+ stable updates. One-Batch REINFORCE + Sequence Trust Region + Sample-Mean Optimization. Paper: researchgate.net/publication… Code: github.com/NVIDIA-NeMo/labs-…
9
5
144
10,383
Grad retweeted
i do think a lot of people on the pro-open-source side are having a bit of a knee-jerk reaction to the pacing statements today, as we're used to viewing the closed labs as power-seeking. but i think their hands are somewhat forced here, and this is just another chapter on the fairly inevitable path towards decently-fast decently-safe decently-commoditized intelligence abundance. ask any F500 exec or swing voter. the world doesn't really want super-fast-takeoff superintelligence owned only by two companies, and thus we won't get it. it'll happen at the pace that the world can accommodate it, which means reaching some sort of confidence consensus that the models are aligned enough that we won't be dealing with scary new incidents all the time. this will trickle out broadly, in the form of best practices and distillation. the smarter a model is, the more it has a "personality", and the less effective strict rules are. there will be awkward compromises and moral tensions. but we ultimately just want the models to be reasonable, and to do the sorts of things reasonable humans would do if our brains were faster and less error-prone and had more working memory. i think we'll get there. the labs will build mac and windows, the rest of us are building linux. everyone's gonna do great. weird stuff will keep happening, but we'll still wake up and go to work, until the work does itself in a manner the world finds acceptable.
86
150
1,703
256,883
Grad retweeted
we are super bullish on sparse attention + HiSparse and have been working with the vLLM team on it sparse attention reduces the pressure on memory bandwidth by only selecting k for the attention, but it doesn't reduce KV cache memory storage , in high-throughput wide-EP deployment you want to maximize the batch size of decode to use compute as much as possible, but at long sequence decode you quickly run out of VRAM and can't hold enough parallel requests to saturate the compute HiSparse fixes this by offloading the active KV cache to CPU. It keeps an LRU cache on GPU, and since many of the same K tokens are reused every decode you barely notice the offloading, this allows a massive decrease in memory usage and an increase in concurrency This is super important for RL where throughput is key and we want to be as much as possible in a compute-bound regime tldr: lower memory usage, more concurrency, higher inference throughput, faster RL
Sparse MLA only attends to the top-K tokens, so the rest of the KV need not live on the GPU. Hybrid HiSparse in vLLM builds on that, and a request keeps decoding after its KV stops fitting in HBM. It keeps KV on the GPU while there is room. Under pressure a request releases its coldest pages to host memory, keeps a small hot buffer of what the indexer asks for, and keeps decoding instead of being preempted. 📊 Demonstrated on GLM 5.3, one 8× H200 node, full 1M context. Same host memory, configured concurrency 32: KV offloading kept 5-6 requests running. Hybrid HiSparse kept 19-25. 🔹 Hot pages are ordinary KV blocks from the same pool (Hybrid Memory Allocator) 🔹 One fused kernel resolves resident, hot and missing rows, CUDA-graph capturable 🔹 Prefix caching, OffloadingConnector, P/D imports and MTP keep working Built by @RedHat_AI and @PrimeIntellect with the vLLM community. Planned for v0.30; pinned commit, flags and calculator are in the post👇 🔗 vllm.ai/blog/2026-09-08-glm5…
11
23
223
20,006
To clear up the timeline a bit: Looped transformers have always been a totally valid arch decision to get better performance under equal params and more compute As in, at training time u loop some layers in some way a fixed amount of times This is more expressive than CoT bcs these looped layers have a seperate kv cache The thing that doesnt really work is dynamic looping where usually u make a sacrifice as u cant have the seperate kv cache, and so its sketchy and not reliable or usually worth it Funny how this couldve also been a headline
OpenAI’s Astra AI uses a new reasoning approach called “recurrent depth.” Though it can help model costs and performance, researchers are concerned bc it obscures a model’s thinking process, making it more difficult to monitor. w/ @amir @rocketalignment theinformation.com/articles/…
13
17
347
41,083
Also can see a previous post i made on this
People forget to think of the actual differences between CoT and layer looping Difference lies on how info is passed on over loops. Layer looping passes this info fully in that final hidden state while CoT passes it through the kv cache of each layer (around like 70 layers, and passes through attention rather than a residual) Make of that what u will
15
2,985
As models become more capable, reward hacks become an increasingly serious problem. During a controlled experiment, we found a novel reward hack in which agents are able to gain web access in offline sandboxes.
28
54
524
203,330
Join us at Prime Intellect to build open superintelligence and the infrastructure powering self-improving agents. We’re hiring across 25+ roles. Research • AI Research Resident • Research Engineer, Distributed Training • Research Engineer, Reinforcement Learning • Research Engineer, RL Infrastructure Compute • Compute Finance & Strategy • Compute Intelligence Engineer • Head of Compute • Solutions Architect, AI Infrastructure • Technical Account Manager, AI Infrastructure Engineering • MTS, Compute Platform • MTS, Full-Stack • MTS, GPU Infrastructure • MTS, Inference • MTS, Sandbox Platform • MTS, Security • MTS, Training Platform Growth • Applied AI: Product Strategy & Revenue • Forward-Deployed AI Strategy • Head of Growth • Head of Marketing • Community and Devrel Other • Internship • Open Application for Unconventional Talent Apply here: primeintellect.ai/careers
17
24
412
48,489
we just shipped adaptive concurrency to prl main, making it (to the best of my knowledge) the first oss rl framework to dynamically adjust inflight rollouts to maximize throughput during the entirety of an rl run
11
20
183
27,650
We ran the largest open experiment on how frontier models do AI research. 100+ autonomous runs across 10+ models, sandboxed on 8xH200s for up to 8 days, iterating on the nanoGPT optimizer track. Best runs closed 82% of the gap to a record built by dozens of humans over months.
73
215
1,888
606,244
Introducing Prime Agent: A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficient and expressive through programmatic tool calling, context as a variable, multi-agent messaging, and a self-modifiable harness state.
400
841
8,173
3,217,449
People forget to think of the actual differences between CoT and layer looping Difference lies on how info is passed on over loops. Layer looping passes this info fully in that final hidden state while CoT passes it through the kv cache of each layer (around like 70 layers, and passes through attention rather than a residual) Make of that what u will
its remarkable how well depth recurrence via CoT works compared to general depth recurrence and how well RL refines such depth recurrence
9
7
167
20,909
Grad retweeted
Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It spans hundreds of rooms and puzzles. Today’s best agents cannot progress beyond the initial levels.
122
155
1,472
248,946
Today, we are releasing verifiers v1 — an overhaul of our environment stack for the modern era of agentic RL and evals. We decompose environments into a taskset, a harness, and a runtime. Run complex agentic tasks like coding and computer use at scale, in any harness.
34
87
837
274,871
New GLM paper on the PPO algo they use for GLM 5.2 Clearly the PPO hype was overblown (DIS here is Icepop without recompute which we first did in Intellect-3) This also comes with 3x longer trainer steps, 2x memory on trainer (with equal inference compute), and extra compute to train the value model before RL starts Wouldn't be surprised if taking that into account makes GRPO >> Also clearly SWE vs AIME type RL didnt have a difference much, so the statements going around on long horizon and tool calling being key for PPO also seem overblown
11
44
493
43,225
the main valid thing is u get less off policyness by avoiding the group tail, but imagine thats making the performance difference instead of actual credit assignment lol (unlikely tho)
1
16
2,292
Announcing our $130M Series A to build the Open Superintelligence Stack Led by Radical Ventures, with NVIDIA, Intel Capital, Dell Capital, and existing investors Train, deploy, and continuously improve your own models using our stack. Own your intelligence.
333
546
4,972
1,430,054