Introducing Real-time RL. In the real world, time isn't free. The environment keeps "moving" even when you're computing your next action. We show how RL agents can learn to adaptively think in real-time games. 1/🧵
17
84
720
77,836
Aneesh Muppidi retweeted
Introducing Real-Time EXPO-FT – Fast and Reliable RL for Real-Time VLA Policies! Real-Time EXPO-FT unlocks π0.5 on challenging dynamic tasks, such as balancing a ball on a plate and striking a ball into the goal (1/6)
6
22
201
48,129
Aneesh Muppidi retweeted
Excited to share Context-Sharded Block Parallelism (CSBP) and Turbo-dLLM, our new optimized distributed training library for Diffusion LLMs! Our distributed parallelism strategy unlocks significant training efficiency for diffusion LLMs, with speedup gains growing with context length 🚀 On 8x H100 GPUs, Turbo-dLLM accelerates DFlash2 speculative-decoder training by 2.48x at 512K and 7.59x at 1M context length. Simply pip install turbo-dllm⚡️ or check out scalingintelligence.stanford…. Honored to work with @TarunSures41845, @PranshuChatur11, @hangoo_kang, @pshroff_ , @KumbongHermann, and advisor @Azaliamirh 🌟!
Diffusion LLMs and speculative decoding promise much faster agents. Yet agents need long contexts, and training on them is painfully slow. Introducing Context-Sharded Block Parallelism (CSBP), a new distributed parallelism strategy unlocking significant training efficiency for diffusion LLMs, with speedup gains growing with context length 🚀 ⚡ 7.59× faster DFlash2 speculative decoding drafter training ⚡ 1.61× faster block diffusion fine-tuning ⚡ 1.33× faster autoregressive → block diffusion adaptation With the same GPU hours, models trained with CSBP score higher on SWE-bench Verified and Terminal-Bench Lite 👑 Open-sourced in Turbo-dLLM, our new optimized distributed training library. Advised by @Azaliamirh and with an amazing team: @PranshuChatur11 @hangoo_kang @pshroff_ @ishanskhare @KumbongHermann
25
20
66
2,725
Aneesh Muppidi retweeted
Excited to share SceneAgent 🤖, our agentic pipeline for scene simulation from 3D captures! Made by Luke Hollis and Tianxing Fan! We use a combination of semantic features and predictive per-Gaussian physics to transform individual complex objects or entire scenes for robotics policy training and evaluation. Our pipeline automates comprehensive details converting input images, video, lidar, or generated 3DGS scenes like from World Labs Marble. It automates: 0. 3DGS processing from input data at 30k and 60k training steps, with automated per-image correction/calibration 1. Infer semantic features for each gaussian 2. Segment objects and infill background 3. Bake predictive physics materials for each object (rigidity, friction, density, etc). 4. Decompose objects into individual parts 5. Articulate movable pieces and joints 6. Generate similar 3d meshes with different geometries, textures, physics properties, and articulations Connected to autoresearch tools, we believe that this will solve an important step in the time-intensive data problem for robotics training and policy evaluation. Paper and code coming soon, more at computationalrobotics.seas.h…
5
20
145
8,921
Aneesh Muppidi retweeted
Simulating the consequences of an experiment before running it is a holy grail of AI for biology. How can we get there? Our new paper, “World models for biomedicine,” which I led together with @njwfish, is out now in @CellCellPress’s special issue on AI in biology. We explore how biomedical world models could help scientists navigate the enormous space of possible biological interventions. In principle, they could: 🧪 Simulate how biological systems respond to interventions never previously tested 🎯 Search for interventions that drive cells, tissues, or patients toward desired states 🧭 Plan experimental campaigns over long time horizons 🤖 Power autonomous, closed-loop AI scientists However, realizing this vision requires much more than calling a model a “world model.” What exactly defines a biomedical world model, and what capabilities distinguish it from related models? Where do today’s models already show promise, and where do they fall short? What new data and architectures are needed? What evaluations would demonstrate that a world model works? We tackle these and other questions in this latest paper from @marinkazitnik's lab: cell.com/cell/fulltext/S0092… More below 🧵👇🏽 1/8
19
55
276
17,626
Aneesh Muppidi retweeted
is Astra really this good at VLA? I spent hours trying a zero-shot demo with the SO-101 arm and this is all I got
i gave astra a robot, a paint brush, and a camera then asked it to paint the golden gate bridge in real life! it figured out how to control the robot, and progressively got better throughout its attempts. the timelapse is sick
19
16
173
45,559
Aneesh Muppidi retweeted
Introducing OM-1, our first robot foundation model, zero-shot generalizing to any robot: table-top arms, industrial arms and humanoids. - learned directly from human manipulation data - no teleop/robot data - close to human-level dexterity and efficiency - multi-robot collab
153
479
3,499
1,749,048
Aneesh Muppidi retweeted
We built a small "primitive" of an artificial brain. Think of primitives as Lego bricks for a different kind of intelligence, or more scientifically, something closer to the neural priors we're born with. The really cool thing: the primitive wasn't trained. No learned policy. No reinforcement learning. No hand-coded obstacle avoidance (obviously). Just ~300 neurons reacting to a changing environment in real time through their own neural dynamics (that's all). As we better understand some of the math and dynamics behind biological brains, we are now on a path to a different kind of robotic intelligence: running locally, without the need of millions of data points or traditional split between training and inference, and ultimately, able to react and learn continuously in world that changes. We've only scratched the surface, but while our first papers go through peer review and our experiments continue, we'll be sharing more about primitives and two other core findings that gets us closer to continual learning. Follow if intrigued. Note: We haven't solved continual learning yet, but now we're confident we can solve the basic dynamics behind it within the next year or so. In the video (1x): ~300 no-pretrained neurons responding and adapting in real time using only their neural dynamics to control directly an e-puck body to avoid obstacles.
16
11
78
6,075
Aneesh Muppidi retweeted
We’ve all been amazed by how capable Astra is—whether at solving Millennium Prize–level mathematics, reconstructing scenes, or controlling robots. A common pattern across these examples seems to be that, before Astra, many humans had already spent years attacking these problems—or closely related ones—creating a large body of knowledge, techniques, and examples. That accumulated human effort may be what eventually enables the tipping point where Astra becomes better, perhaps even substantially better, than individual humans at solving the problem. By contrast, for problems that are not yet well defined, whose goals are ambiguous, or are so ill posed that very few people have seriously studied them, Astra still seems much less capable. There are research problems in my own group that I simply cannot imagine asking Astra to solve directly—not necessarily because the underlying mathematics is harder, but because we cannot yet formulate the problem clearly. This creates an interesting dilemma. If you work on a hot and practically important problem, there is a good chance that Astra will soon be better at solving it than you are. If you work on something extremely niche, you may remain better than Astra—but the problem itself may not matter very much. Perhaps the best research strategy, then, is what great research has always been: find an important problem that has not even been properly defined yet; discover the right way to formulate and frame it; and then use systems like Astra to help solve it. The most valuable human contribution may increasingly shift from solving well-defined problems to discovering which problems should exist in the first place—and framing them in a way that, once solved, produces unexpectedly large impact.
3
8
103
5,903
Aneesh Muppidi retweeted
AI has solved Navier-Stokes. What would it take for AI to make similar advances in life sciences? In our latest preprint, led by @AdaFang_, we argue that, while fields like mathematics and programming can cheaply evaluate many AI-generated candidates at scale, "closing the loop" of hypothesis ➡️ experiment ➡️ revised hypothesis remains a central bottleneck in AI-driven biomedical discovery. We outline a vision for autonomous AI scientists that: 🧠 Reason over scientific knowledge, multimodal data, competing hypotheses, and uncertainty across multi-loop campaigns lasting days to months. 🔬 Learn under sparse and delayed feedback, combining soft verifiers (e.g., simulations, predictive models, and biological world models) with hard verifiers such as wet lab assays, robotic experiments, organoids, and clinical studies. 🎯 Allocate limited experimental budgets to tests with the greatest expected information gain. We suggest that achieving this vision will require advances in long-horizon reasoning, process-based evaluation, uncertainty-aware exploration, self-driving laboratories, and human-on-the-loop oversight. Read our paper here: preprints.org/manuscript/202… More details below from @marinkazitnik. 👇🏽
Closed-loop AI scientists AI can generate hypotheses much faster than we can experimentally verify them. @AdaFang_ In our latest piece on Closing the Loop in AI-Driven Biomedical Discovery we outline a path for closed-loop AI scientists, preprints.org/manuscript/202… AI scientists can generate hypotheses, propose experiments, and analyze data, and several have produced findings that laboratories then confirmed. What would it take to close that loop more generally?
3
11
87
12,579
Aneesh Muppidi retweeted
World models have emerged as one of the biggest directions in physical AI. At the same time, RL fine-tuning is unlocking capabilities in frontier models beyond what pretraining can achieve on its own Can we get the best of both worlds? We propose Q-Learning with World Models (QWM) (1/7)
7
56
515
32,120
Aneesh Muppidi retweeted
𝗚𝗼𝗼𝗱 𝗺𝗮𝗻𝗶𝗽𝘂𝗹𝗮𝘁𝗶𝗼𝗻 𝗽𝗼𝗹𝗶𝗰𝗶𝗲𝘀 𝗺𝗮𝘆 𝘀𝘁𝗮𝗿𝘁 𝘄𝗶𝘁𝗵 𝗯𝗲𝘁𝘁𝗲𝗿 𝘀𝘁𝗮𝘁𝗲𝘀, 𝗻𝗼𝘁 𝗯𝗲𝘁𝘁𝗲𝗿 𝗿𝗲𝘄𝗮𝗿𝗱𝘀. We sample diverse, physically feasible contact states as starts + goals for RL. The behaviors that emerge are surprisingly dynamic and reactive: including recovery, regrasping, and adapting through contact. 👇 (1/n)
4
32
264
30,781
Aneesh Muppidi retweeted
We’re excited to announce DiG-bench, a new benchmark for discovery! Over the last few weeks we’ve been testing frontier AI models on our novel discovery games and seeing how they score. Each game is a text-based environment, so they probe discovery capabilities in the natural domain of language models, rather than requiring additional, potentially confounding, visual understanding. TL;DR frontier models have improved a lot over the last few months. But they are still stumped by some surprisingly simple problems, even in their native text domain. With @cocosci_lab (@Princeton) @MITCoCoSci (@MIT) @SchmidhuberAI (@KAUST_News) @misovalko (@Inria) @tri_dao (@PrincetonCS) @RMBattleday @zebkDotCom @FraserGreenlee @akaijsa @ClareMaguire @TimMuller1 @kubicek_ales @physicscat0x7d @SukritSumant @thoughtchannel_ (1/5)
44
360
1,191
1,192,631
Aneesh Muppidi retweeted
Pretraining has worked remarkably well across domains We show this doesn’t hold for Q-functions in online RL from a pretrained policy — and propose IPE, a more effective way to learn Q-functions for online RL fine-tuning (1/6)
5
29
306
85,107
Aneesh Muppidi retweeted
I think I have found a proof of the Forsythe Conjecture for s=2 using Codex, open since 1968. A detailed check by a leading expert, who has called the proof "highly likely to be correct," is in progress.
11
33
541
57,746
Aneesh Muppidi retweeted
1/ We are happy to announce the largest known dataset of expert trajectories for physics-based tasks, containing over 11M unique levels in Kinetix.
3
15
149
17,497
Congrats @yiding_song and @DozenDucc!! This is such a brilliant team and awesome direction!
Introducing Waddle Labs: Claude Code for robots. Connect our API to your robot and enter a prompt, then our agents write code to achieve the task in 20 minutes. @yiding_song @theWaddleLabs
9
2,462
Aneesh Muppidi retweeted
We've passed 5 million users!! It's crazy to see the compounding value of focusing on one goal... To give every human superintelligence and make superintelligence more human We could not do any of it without your support. Thank you so much for being there since day 1. We're just getting started :)
Design Arena has surpassed 5 million users! Since launch, we've grown from 6 arenas to 30+ and seen our community create millions of designs. What started as a way to compare AI-generated design has become one of the clearest views into how people actually use AI models in the real world. Thank you to everyone building and sharing. We can't wait to see what you design next!
22
11
268
50,926