Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects. Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below🧵
17
42
276
29,076
Are we experiencing a GPT moment in robotics? Our MolmoSpaces benchmark seems to suggest so. Check out @omarrayyann's thread to see the details 👇
GPT-Astra outperformed all open-source VLA/WAM baselines on a subset of our zero-shot MolmoSpaces-v1 benchmark! traces: orayyan.com/ms-results
3
4
42
18,171
GPT-Astra outperformed all open-source VLA/WAM baselines on a subset of our zero-shot MolmoSpaces-v1 benchmark! traces: orayyan.com/ms-results
4
16
183
27,386
It used fewer reasoning tokens than other VLM APIs (though the counts might not be directly comparable). I couldn’t run the full benchmark for these API models to include them on our leaderboard (molmospaces.allen.ai/leaderb…) because of the API cost :)
1
7
968
I think distilling the model’s spatial knowledge into a raw action space policy would be an interesting direction. its good vision understanding makes this verifiable enough to happen in an automated closed-loop way
8
798
Training for the robot olympics with MuJoCo’s new surfacevel for geoms
11
30
321
25,444
Omar Rayyan retweeted
A solid step toward humanoid loco-manip, super cool!
Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects. Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below🧵
1
3
32
5,768
Learning mobile manipulation is hard because collecting data for mobile manipulation is hard. But seems like @omarrayyann has found a promising way to crack this problem!
Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects. Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below🧵
4
23
3,352
Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects. Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below🧵
17
42
276
29,076
We follow this sim-to-real recipe to train two policies: a single-object Fetch policy that works across scenes, and an object-conditioned policy that takes the target object's name as an instruction. Both are two-system architectures with the decoupled SONIC controller acting as the low-level controller.
1
8
1,041
Check out our website for more videos and interactive simulation episodes: orayyan.com/fetchman. This project was done with advice from @YuchenCui1 and great collaborators @max_argus, Zhi Li, Chang Yu, Yuxin @chenfanfujiang. Paper: arxiv.org/abs/2608.17027 Code (soon): github.com/omarrayyann/Fetch…
1
12
1,061
We will be in the #RSS2026 poster session at 6:30 PM – stop by with all of your questions!
MolmoSpaces provides singular scale and diversity. We built a benchmark that puts that scale to use. MolmoSpaces-Bench evaluates zero-shot policies across thousands of environments previously unseen to them under systematic variation, providing insights that go beyond a success rate % More Below:
3
16
1,786
I’m at RSS in Sydney 🇦🇺 to present: - MolmoSpaces, Tuesday 6:30pm - Contact Anchored Policies, Wednesday 4:00pm - a (tbr) work on training general visual loco-manipulation policies in sim. If you work on anything from learning with off-domain data to sim-evals, let’s chat!
1
3
23
2,514
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data? Introducing Do as I Do: establishing a much needed correspondence between human videos and dexterous robot data. Some fun insights below: 🧵
11
65
364
96,607
Wall-OSS is now the #1 policy on the zero-shot MolmoSpaces evals. A lot of details in their paper, I recommend checking it out.
We are open-sourcing Wall-OSS-0.5. Pretrain Once, Act Anywhere. Wall-OSS-0.5 is a VLA model for real-world robotic manipulation, exploring whether pretraining alone can produce robot capabilities directly testable on physical hardware before task-specific fine-tuning. Key technical highlights: • Gradient-bridged co-training • Vision-Aligned RVQ Action Tokenizer • Action-Space Supervision • DMuon distributed optimizer In zero-shot real-robot evaluation, the pretrained checkpoint achieved task-progress scores above 80 on multiple tasks, including Block Sorting, Fruit Sorting, Ring Stacking, and Rope Tightening. Paper, code, blog, and uncut videos: x2robot.com/oss#resources
2
10
43
7,434
Omar Rayyan retweeted
Robotics is still data starved. Collecting high-quality robot demonstrations remains brutally slow and expensive. Introducing COBALT: A cloud-native teleoperation platform designed for large-scale robot learning. We are democratizing data collection by leveraging the hardware everyone already owns: the smartphone All you need is to download an app (today)! Read on for more!
29
52
399
104,563