Filter
Exclude
Time range
-
Minimum likes
Replying to @SOTA_kke
they can also just add them to the pool of rl envs they use for post-training. maybe they already do that
5
159
I think distilling the model’s spatial knowledge into a raw action space policy would be an interesting direction. its good vision understanding makes this verifiable enough to happen in an automated closed-loop way
8
807
It used fewer reasoning tokens than other VLM APIs (though the counts might not be directly comparable). I couldn’t run the full benchmark for these API models to include them on our leaderboard (molmospaces.allen.ai/leaderb…) because of the API cost :)
1
7
970
All models used the same closed-loop harness in @robocurve’s results. they get exocentric/egocentric rgb observations from the DROID setup and robot state, and output cartesian poses. history is cleared after each of the 200 episodes.
2
2
15
2,168
GPT-Astra outperformed all open-source VLA/WAM baselines on a subset of our zero-shot MolmoSpaces-v1 benchmark! traces: orayyan.com/ms-results
4
16
183
27,397
Replying to @kaiwynd
how much did this cost
3
716
Training for the robot olympics with MuJoCo’s new surfacevel for geoms
11
30
321
25,448
Replying to @JulienBlanchon
Thanks! code will be out soon. but yes, per iteration I sample 64 tasks and roll each 8 times from the same starting state. noise is injected at a single random flow step and logp is taken there. and thanks for the suggestion, I will be trying it out.
1
2
159
Check out our website for more videos and interactive simulation episodes: orayyan.com/fetchman. This project was done with advice from @YuchenCui1 and great collaborators @max_argus, Zhi Li, Chang Yu, Yuxin @chenfanfujiang. Paper: arxiv.org/abs/2608.17027 Code (soon): github.com/omarrayyann/Fetch…
1
12
1,061
We follow this sim-to-real recipe to train two policies: a single-object Fetch policy that works across scenes, and an object-conditioned policy that takes the target object's name as an instruction. Both are two-system architectures with the decoupled SONIC controller acting as the low-level controller.
1
8
1,041
Alongside FetchMan, we're releasing FetchMan-Bench, a simulation benchmark for humanoid policies that expands MolmoSpaces-Benchmarks toward humanoids, starting with the Fetch task. Thousands of evaluation rollouts can be run across scenes to give a quick, reliable read on how a humanoid policy is doing.
1
5
898
Post-training FetchMan with large-scale Flow-GRPO in sim lets the policy break beyond the confines of its synthetic demonstrations, using only a simple sparse reward on our Fetch task. We turned to RL because performance saturates at a ceiling that more of the same synthetic simulation data can't push past.
1
1
13
1,151
Limited diversity in simulated scenes and assets has long been a bottleneck for vision-based sim learning, preventing policies from generalizing to the real world. FetchMan is first pre-trained on large-scale synthetic loco-manipulation data generated across MolmoSpaces' 150,000+ scenes, with humanoid-specific domain randomization axes. The pipeline generated 350,000 episodes in two days.
1
11
1,631
Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects. Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below🧵
16
41
276
29,101
I’m at RSS in Sydney 🇦🇺 to present: - MolmoSpaces, Tuesday 6:30pm - Contact Anchored Policies, Wednesday 4:00pm - a (tbr) work on training general visual loco-manipulation policies in sim. If you work on anything from learning with off-domain data to sim-evals, let’s chat!
1
3
23
2,514
are we thinking of the same thing
1
4
127
Replying to @chris_j_paxton
I think people should start adding "state-based" in the corner of video demos if they're not visual. similar to how it's now standard to add the playback speed (eg. x4). especially if there's manipulation involved, otherwise it might be deceiving
5
70
Wall-OSS is now the #1 policy on the zero-shot MolmoSpaces evals. A lot of details in their paper, I recommend checking it out.
We are open-sourcing Wall-OSS-0.5. Pretrain Once, Act Anywhere. Wall-OSS-0.5 is a VLA model for real-world robotic manipulation, exploring whether pretraining alone can produce robot capabilities directly testable on physical hardware before task-specific fine-tuning. Key technical highlights: • Gradient-bridged co-training • Vision-Aligned RVQ Action Tokenizer • Action-Space Supervision • DMuon distributed optimizer In zero-shot real-robot evaluation, the pretrained checkpoint achieved task-progress scores above 80 on multiple tasks, including Block Sorting, Fruit Sorting, Ring Stacking, and Rope Tightening. Paper, code, blog, and uncut videos: x2robot.com/oss#resources
2
10
43
7,434
Replying to @NicholasEPfaff
cool work!
2
28
Replying to @TX_Leo_Wang
Congrats on the release. quick question, did you ablate feeding the policy the spatial observations (ict tokens) only without any rgb? judging by the "human rgb" vs "human rgb + ict" gap it seems like that's all what's being used?
1
4
274