๐Ÿค–๐Ÿ”„๐ŸŒŽ | Senior undergrad, Yao class @Tsinghua_Uni | RA @uwcse | RS @wuji_global | Prev. RA @mldcmu.

Seattle, WA
๐Ÿค– World Action Models predict the future to act better. However, WAM research remains fragmented in code, ad hoc in design, and unprincipled at scale. ๐Ÿ’ก Today, we release ๐—ข๐—ฝ๐—ฒ๐—ป๐—ช๐—”๐— : ๐—”๐—ป ๐—ข๐—ฝ๐—ฒ๐—ป, ๐— ๐—ผ๐—ฑ๐˜‚๐—น๐—ฎ๐—ฟ ๐—˜๐˜…๐—ฝ๐—น๐—ผ๐—ฟ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ง๐—ผ๐˜„๐—ฎ๐—ฟ๐—ฑ๐˜€ ๐—ฆ๐˜†๐˜€๐˜๐—ฒ๐—บ๐—ฎ๐˜๐—ถ๐—ฐ ๐—ช๐—ผ๐—ฟ๐—น๐—ฑโ€“๐—”๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐— ๐—ผ๐—ฑ๐—ฒ๐—น ๐—ฃ๐—ฟ๐—ฒ๐˜๐—ฟ๐—ฎ๐—ถ๐—ป๐—ถ๐—ป๐—ด to change this. ๐ŸŒ openwam-official.github.io/ ๐Ÿงต1/14
6
40
263
2,264,756
๐Ÿค– World Action Models predict the future to act better. However, WAM research remains fragmented in code, ad hoc in design, and unprincipled at scale. ๐Ÿ’ก Today, we release ๐—ข๐—ฝ๐—ฒ๐—ป๐—ช๐—”๐— : ๐—”๐—ป ๐—ข๐—ฝ๐—ฒ๐—ป, ๐— ๐—ผ๐—ฑ๐˜‚๐—น๐—ฎ๐—ฟ ๐—˜๐˜…๐—ฝ๐—น๐—ผ๐—ฟ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ง๐—ผ๐˜„๐—ฎ๐—ฟ๐—ฑ๐˜€ ๐—ฆ๐˜†๐˜€๐˜๐—ฒ๐—บ๐—ฎ๐˜๐—ถ๐—ฐ ๐—ช๐—ผ๐—ฟ๐—น๐—ฑโ€“๐—”๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐— ๐—ผ๐—ฑ๐—ฒ๐—น ๐—ฃ๐—ฟ๐—ฒ๐˜๐—ฟ๐—ฎ๐—ถ๐—ป๐—ถ๐—ป๐—ด to change this. ๐ŸŒ openwam-official.github.io/ ๐Ÿงต1/14
6
40
263
2,264,756
๐Ÿ“ฆ We release OpenWAM as a full stack empowering WAM research: - Modular infrastructure for training, inference, deploy, and eval. - OpenWAM-ฮฑ pretrained and post-trained weights. - Study checkpoints for exploring the design choices. - Evaluation protocols and data recipes. ๐Ÿฅž We hope researchers can use the stack to reproduce the results, adapt OpenWAM-ฮฑ to new tasks and embodiments, and test new ideas for representations, worldโ€“action interaction, and embodied pretraining. ๐Ÿ”ญ There is much more to explore. We hope OpenWAM-ฮฑ serves as a strong, reproducible baseline and OpenWAM-Infra makes that next round of experiments easier to build. ๐ŸŒ Project: openwam-official.github.io/ ๐Ÿ“„ Paper: arxiv.org/abs/2609.07398 ๐Ÿ’ป Code: github.com/OpenWAM-Official/โ€ฆ ๐Ÿค— Models & data: huggingface.co/OpenWAM ๐Ÿงต13/14
1
1
276
๐Ÿ™ OpenWAM is a large undertaking, made possible by the dedication and collective effort of the entire team. Thank you to everyone who contributed to this project, couldn't have done it without y'all! ๐Ÿคž Iโ€™m grateful for the opportunity to have co-lead this project with @YuranWang_rise. Special thanks to the amazing adviors: @zhaohang0124, @linshaonju, @HaoDong123, and @LotusSapphire for their advice and guidance throughout the project. ๐Ÿงต14/14
2
256
I still donโ€™t get why people think code as policy is so compelling. Language models and their harness can be very useful for robots, but writing programs that directly output for example, eef actions, seems to me not the best interface.
5
2
46
4,077
Siqiao Huang retweeted
Hey, maybe you are looking for our STARFlow-V (starflow-v.github.io/starfloโ€ฆ) ๐Ÿค”๐Ÿ˜ฒ and it is not diffusion.
Most โ€œautoregressive videoโ€ is not autoregressive. Nobody pretrains BERT, distills a GPT-like student, and calls that native. Video does the equivalent every day: train a bidirectional short-clip DiT, then add a causal mask and a few-step sampler to produce a faster version. I wrote a short blog on what native would actually require: waynejin0918.github.io/how-fโ€ฆ ( Thanks to the TML blog for providing stylistic references, and to grok4.6 for helping me polish my content. )
2
7
67
10,235
"Weโ€™re using language as a crutch to help the deficiencies of our vision systems to learn good representations from images and video." Yann Lecun, Lex Fridman Podcast #416. Sorry, this thread is very cool, but I still believe in the long run, building models with intrinsic understandings of the world requires a more vision-centric approach than modern LLMs.
You can just RL a coding model to paint with javascript btw
2
2
72
6,751
This one is impressive. "Behavior Prompting", or in-context learning of robotics from seeing demonstrations has a long history, and it's a good time to go down this memory lane. (1/n)
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
2
1
43
3,624
This year is a especially noise year for this field. There are multiple research more or less related to this coming out everywhere: - Behavior Prompting Policy: behavior-prompting.github.ioโ€ฆ - HumanEgo: humanego-ai.github.io/ - RoboTTT: Context Scaling for Robot Policies: arxiv.org/abs/2607.15275 - WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time: arxiv.org/abs/2607.06988 5/n
1
3
248
I would end by a couple closing remark: 1. In-context learning for robotics is a long-standing problem; 2. Scaling context length (& format diversity) is a commonly-suggested advise for robotics. However, does/should this icl/reasoning capability happen at the VLA layer rather than the VLM layer? Since VLA layer inference needs to be served at a certain frequency for closed-loop control, and naive attention's compute scale quadratically with context length; 3. If you combine this half-a-year+ training with DYNA-2's curve, the takeaway should not be let's scale to 1B hours of robot data & 1 million gpus for compute, but there's some fundamentally wrong with the current paradigm, and there's something crucial missing in the robot learning repertoire. What an exciting time to be alive. 6/6
6
678