Robotics | CV | ML

Seattle
Frontier models, especially Astra, is performing very well on robotics taks. This opens up many questions. Are these models trained on in-domain data? Do they behave differently from BC models? What does this mean for the research field in the coming years?
GPT-Astra outperformed all open-source VLA/WAM baselines on a subset of our zero-shot MolmoSpaces-v1 benchmark! traces: orayyan.com/ms-results
2
212
Max argus retweeted
GPT-Astra outperformed all open-source VLA/WAM baselines on a subset of our zero-shot MolmoSpaces-v1 benchmark! traces: orayyan.com/ms-results
4
16
183
27,395
Max argus retweeted
Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects. Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below🧵
16
41
276
29,098
Max argus retweeted
We're releasing MolmoMotion, a 3D motion forecasting model. Given one or a few video frames, 3D points on an object, & an instruction like "Put the white bowl on the table," MolmoMotion predicts where those points will go over the next few seconds in a shared 3D world frame. 🧵
12
62
371
188,990
Max argus retweeted
We believe in the representation of 3d point trajectory as an alternative form of world modeling. it is universal to all object, more compact than pixel, and can directly be consumed by many downstream task. p.s. the key for us to generate high quality real world data is we respect objectness and do hard work in object grounding -- which allows us to properly filter and smooth noisy trajectories coming out of off the shell 3d tracking methods.
We're releasing MolmoMotion, a 3D motion forecasting model. Given one or a few video frames, 3D points on an object, & an instruction like "Put the white bowl on the table," MolmoMotion predicts where those points will go over the next few seconds in a shared 3D world frame. 🧵
3
12
63
137,286
Max argus retweeted
And that's a wrap on a fantastic ICRA 2026! 🎉 Incredible run for MolmoBot — clean sweep on workshops, winning Best Paper at all three we entered: Synthetic Data for Robot Learning, Beyond Teleoperation, and VLA Pipelines. 🤖
1
5
28
5,174
Max argus retweeted
Our paper MolmoB0T won top honor at the SDRL workshop today! Congrats to my co-authors @ab_deshpande, Maya, Snehal, @RanjayKrishna, @shahdhruv_ (plus others, you know how it goes), and we'll see you as well at the VLA Pipelines and Beyond Teleoperation workshops on Friday 🚀
7
25
1,895
Max argus retweeted
MolmoBot, our open robotic manipulation suite trained entirely in simulation, now has code, training data, a data generation pipeline, & evals all available. This puts our robotics models within reach of any research lab—no extensive real-world data collection required. 🧵
8
36
232
54,297
Jensen approves! Hercules efforts from @YejinKim4 @omarrayyann, Max Argus & team! This has a decent chance of becoming a super important benchmark fo robotics going forward. Check out this @RoboPapers episode with the MolmoSpaces folks.
Benchmarking, evaluating, and developing robotics code is difficult, and part of this is because no simulator really reflects the diversity and scale of real embodiments. Enter MolmoSpaces from AI2: a massive open ecosystem with a range of 230,000 handcrafted and procedurally-generated home environments, including 48,000 manipulable objects. Crucially, MolmoSpaces provides simulation environments which work for both navigation and manipulation. We talked to the team: @YejinKim4, @omarrayyann, and Max Argus, to tell us more. Watch Episode 69 of RoboPapers, with @micoolcho and @DJiafei, now!
1
3
7
1,264
Max argus retweeted
Today, a step forward in open robotics - our results show that sim-to-real zero shot transfer for manipulation is possible. MolmoBot is our open model suite for robotics, trained entirely in simulation on MolmoSpaces.🧵
10
40
283
65,567
Max argus retweeted
Are you sure your training data is actually synced? Egocentric camera sees a hand grasping an orange, but the wrist cam shows nothing and tactile reads zero contact. Your policy is learning from broken data and doesn't even know it. In Physical AI, multi-modal sync is everything. → Egocentric: 30fps → Wrist: 30fps → Tactile: 100Hz Different devices, different clocks, slightly different rates. The drift starts small. Barely noticeable frame by frame. But over a 4-minute episode, that tiny difference compounds into seconds of misalignment. And you had no way to even check. Until now. We built the Sync Quality Dashboard. One score tells you if your data is clean. Then go deeper. Clock offset, drift rate, jitter, frame drops, per-stream correction. All visible, all measurable. In a 4-min episode, accumulated clock drift reached 7.5 seconds by the end of the recording. After correction: 9.0ms. That's the difference between "roughly aligned" and "actually aligned." Visually confirm vision-to-vision, vision-to-tactile alignment frame by frame. No more "trust me, the data is fine." We don't just collect multi-modal demos. We ship a quality assurance layer so you can verify every episode before it touches your model. All data in @LeRobotHF format. Ready to train. Verified in sync. Stop guessing. Start verifying.
5
11
104
10,250
Max argus retweeted
MolmoSpaces-Bench leaderboard is now live! Test your generalist policies to see how they compare across tasks and environments. Feel free to reach out if you need help setting it up. molmospaces.allen.ai/leaderb…
2
5
37
2,051
Max argus retweeted
𝐃𝐫𝐞𝐚𝐦𝐙𝐞𝐫𝐨 𝐢𝐬 #𝟏 𝐨𝐧 𝐛𝐨𝐭𝐡 𝐌𝐨𝐥𝐦𝐨𝐒𝐩𝐚𝐜𝐞𝐬 𝐚𝐧𝐝 𝐑𝐨𝐛𝐨𝐀𝐫𝐞𝐧𝐚 🏆 𝗪𝗵𝗮𝘁 𝗺𝗮𝗸𝗲𝘀 𝘁𝗵𝗶𝘀 𝗻𝗼𝘁𝗮𝗯𝗹𝗲: DreamZero-DROID is trained 𝑓𝑟𝑜𝑚 𝑠𝑐𝑟𝑎𝑡𝑐ℎ using only the DROID dataset. No pretraining on large-scale robot data, unlike competing VLAs. This demonstrates the strength of video-model backbones for generalist robot policies (VAMs/WAMs). More broadly, training 𝑜𝑛𝑙𝑦 on real data and evaluating on (1) transparent, distributed benchmarks like 𝐑𝐨𝐛𝐨𝐀𝐫𝐞𝐧𝐚 or (2) scalable sim-benchmarks like 𝐌𝐨𝐥𝐦𝐨𝐒𝐩𝐚𝐜𝐞𝐬 is an exciting step toward fairer and more reproducible evaluation of generalist policies, one that the community can hillclimb together to measure progress. Special thanks to the Ai2 MolmoSpaces team (@notmahi @omarrayyann @YejinKim4 Max Argus) and the RoboArena team (@pranav_atreya) for helping with the set-up and getting these evaluations! Special shout out to @youliangtan @NadunRanawakaA @chuning_zhu, who led these efforts from the GEAR side :) + We also release our DreamZero-AgiBot checkpoint & post-training code to enable very efficient few-shot adaptation. Post-train on just ~30 minutes of play data for your specific robot, and see the robot do basic language following and pick-and-place 🤗(See YAM experiments in our paper for more detail). ++ We also provide the entire codebase & preprocessed dataset to replicate the DreamZero-DROID checkpoint. 🌐 dreamzero0.github.io 💻 github.com/dreamzero0/dreamz… RoboArena: robo-arena.github.io/leaderb… MolmoSpaces: molmospaces.allen.ai/leaderb…
5
30
184
41,005
Max argus retweeted
MolmoSpaces also comes with 42M+ grasps that cover 48K+ objects across 250K+ scenes, allowing large-scale functional trajectory generation in MuJoCo and IsaacSim.
4
33
316
19,096
Max argus retweeted
MolmoSpaces provides singular scale and diversity. We built a benchmark that puts that scale to use. MolmoSpaces-Bench evaluates zero-shot policies across thousands of environments previously unseen to them under systematic variation, providing insights that go beyond a success rate % More Below:
Introducing MolmoSpaces, a large-scale, fully open platform + benchmark for embodied AI research. 🤖 230k+ indoor scenes, 130k+ object models, & 42M annotated robotic grasps—all in one ecosystem.
6
16
158
16,314
Max argus retweeted
Introducing MolmoSpaces, a large-scale, fully open platform + benchmark for embodied AI research. 🤖 230k+ indoor scenes, 130k+ object models, & 42M annotated robotic grasps—all in one ecosystem.
10
101
717
98,503
Excited to present “Climate-sensitive Urban Planning through Optimization of Tree Placements” w/ Ferdinand Briegel, Max Argus, @envmet & @ThomasBrox at the #NeurIPS2023 @ClimateChangeAI workshop. Tl;dr: we optimize urban tree locations for improved outdoor human thermal comfort.
2
4
8
749