๐Ÿค– World Action Models predict the future to act better. However, WAM research remains fragmented in code, ad hoc in design, and unprincipled at scale. ๐Ÿ’ก Today, we release ๐—ข๐—ฝ๐—ฒ๐—ป๐—ช๐—”๐— : ๐—”๐—ป ๐—ข๐—ฝ๐—ฒ๐—ป, ๐— ๐—ผ๐—ฑ๐˜‚๐—น๐—ฎ๐—ฟ ๐—˜๐˜…๐—ฝ๐—น๐—ผ๐—ฟ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ง๐—ผ๐˜„๐—ฎ๐—ฟ๐—ฑ๐˜€ ๐—ฆ๐˜†๐˜€๐˜๐—ฒ๐—บ๐—ฎ๐˜๐—ถ๐—ฐ ๐—ช๐—ผ๐—ฟ๐—น๐—ฑโ€“๐—”๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐— ๐—ผ๐—ฑ๐—ฒ๐—น ๐—ฃ๐—ฟ๐—ฒ๐˜๐—ฟ๐—ฎ๐—ถ๐—ป๐—ถ๐—ป๐—ด to change this. ๐ŸŒ openwam-official.github.io/ ๐Ÿงต1/14
6
40
267
2,266,269
OpenWAM is three releases in one: - ๐Ÿงฑ ๐—ข๐—ฝ๐—ฒ๐—ป๐—ช๐—”๐— -๐—œ๐—ป๐—ณ๐—ฟ๐—ฎ provides a modular substrate for the systematic study of WAMs. - ๐Ÿ”ฌ ๐—ข๐—ฝ๐—ฒ๐—ป๐—ช๐—”๐— -๐—ฆ๐˜๐˜‚๐—ฑ๐˜† poses three central questions around WAM pretraining, and investigated them through controlled experiments. - ๐Ÿš€ ๐—ข๐—ฝ๐—ฒ๐—ป๐—ช๐—”๐— -๐—”๐—น๐—ฝ๐—ต๐—ฎ is a fully-open WAM pretrained on ~6,400 hrs (518.5M frames) of mixed egocentric human + robot data built on the insights from OpenWAM-Study, and delivers excellent performance across both simulation and real-world tasks. ๐Ÿงต2/14
1
7
1,521
๐Ÿงฑ ๐—ข๐—ฝ๐—ฒ๐—ป๐—ช๐—”๐— -๐—œ๐—ป๐—ณ๐—ฟ๐—ฎ provides the common infrastructure for this exploration. โญ๏ธ The key underlying design principle is ๐™ข๐™ค๐™™๐™ช๐™ก๐™–๐™ง๐™ž๐™ฉ๐™ฎ. OpenWAM-Infra separates data loading, visual encoding, video backbones, worldโ€“action architectures, and attention masks into composable modules, supporting 4 visual encoders, 5 video backbones, 6 architecture variants, and 4 attention masks. ๐Ÿฆˆ These components share training, inference, deployment, and evaluation interfaces, so a new design can be simply plugged in while reusing the same surrounding pipeline. In addition, we integrate 8 simulation benchmarks and support deployment on 5 real robots. ๐Ÿงต3/14
1
2
864
๐Ÿ”ฌ Building on the substrate of OpenWAM-Infra, we systematically study three questions. โ‰๏ธ Q1. What upstream knowledge should a WAM inherit? ๐Ÿ“š We study two sources of this upstream knowledge: the generative backbone and the representation space. ๐Ÿ”ฎ On the generative backbone side, we find using video generation models with increasing parameter counts often lead to better success rate, despite coming from different model families and lineages. ๐Ÿฅ‘ On the representation space choice, we find directly switching vision encoders to representation encoders (DINOv3, V-JEPA 2.1) does not fruit well. However, when combined with a S-VAE layer that compresses high-dimensional representations to low-dimensional vectors, the resulting system yields better success rate than reconstructive encoders (Flux.2 VAE) with no native temporal compression. But still, these models slightly underperform Wan-VAE, which is by default co-designed with the DiT backbone, as well as including native temporal compression. ๐Ÿคฒ The finding opens up more choices for WAMs: representation encoders can work well when their features are made compact enough for worldโ€“action learning. Latent dimensionality and temporal compression deserve attention alongside the choice of backbone. ๐Ÿงต4/14
1
1
1,043
โ‰๏ธ Q2: How do we create synergy between world and action learning during training? ๐Ÿ  The architecture study favors dedicated capacity for action prediction. By comparing single- dual- and tri-system architectures, we find that giving the world and action streams their own parameters is essential, and we carry forward this dual-system joint self-attention architecture into OpenWAM-Alpha. ๐ŸŒŠ We then vary the information flow by altering cross-modality attention masks. In our experiments, allowing the action stream to access world features reaches 92.39% average success, versus 87.41% when the streams are isolated. Allowing only the world stream to see actions gives 87.63%. ๐Ÿ– Together, these results motivate a dedicated action branch with explicit access to the world stream. We revisit whether that access should become fully mutual after embodied pretraining later in the study. ๐Ÿงต5/14
1
1
485
๐Ÿงฌ Worldโ€“action interaction also depends on how the two streams are denoised at inference. ๐Ÿ–๏ธ We compare inference schedules in which video leads, action leads, or both advance together. For this experiment, training covers independently sampled world and action noise levels, so we can test different paths through the joint denoising process. ๐Ÿ”€ Synchronized joint denoising performs best in our setting, leading at 93.0% success, compared with 88.5โ€“92.3% for the alternatives. ๐Ÿงต6/14
1
1
401
โ‰๏ธ Q3: What does embodied pretraining add, and how should we combine human and robot data? ๐Ÿ“ˆ We study this under a fixed 600-hour data budget, followed by the same downstream fine-tuning. Evaluation uses RoboTwin2.0 Cleanโ†’Randomized, separating in-domain performance from out-of-domain generalization. ๐Ÿšฉ For mixed egocentric + robot pretraining, the main gain is OOD: approximately +12 percentage points, while the ID improvement is less than 1 point. ๐Ÿคฒ Robot trajectories provide action grounding; egocentric human data broadens world coverage. One-stage co-training combines these sources effectively, with results close to the two-stage ego-then-robot recipe and a simpler training flow. ๐Ÿฅท We also revisit the attention mask: after embodied pretraining, mutual worldโ€“action visibility is consistently preferred in our comparisons. โžก๏ธ These findings guide the final recipe: co-train ego and robot data in one stage, and let the two streams mutually accessible. ๐Ÿงต7/14
1
2
370
๐Ÿš€ ๐—ข๐—ฝ๐—ฒ๐—ป๐—ช๐—”๐— -๐—”๐—น๐—ฝ๐—ต๐—ฎ brings those findings together in one pretrained model. - World stream: Wan2.2-TI2V-5B with Wan2.2-VAE. - Action stream: a dedicated 1B-parameter ActionDiT. - Coupling: joint self-attention with mutual worldโ€“action visibility. - Inference: synchronized joint denoising. - Data: one-stage co-training on egocentric human and robot data. ๐ŸŽฅ We pretrain on 518.5M frames, roughly 6,369 hours. An 80-D unified action space gives the different embodiments a shared format with fixed slot semantics. ๐ŸŽง During training, we sample world and action noise levels independently. At inference, they follow the synchronized path selected in the study. ๐Ÿงช Starting from the pretrained checkpoint, we adapt and evaluate OpenWAM-ฮฑ across a wide range of simulation benchmarks and real-robot tasks. ๐Ÿงต8/14
1
1
461
๐Ÿ“Š In simulation, OpenWAM-ฮฑ achieves top-tier performance across 8 benchmarks spanning 5 embodiment categories: single-arm, bimanual, mobile single-arm, mobile bimanual, and dexterous-hand manipulation. ๐Ÿ’ช OpenWAM-ฮฑ achieves strong performance across all benchmarks, including state-of-the-art results on EBench. Compared to representative VLA and WAM baselines, OpenWAM-ฮฑ consistently delivers outstanding performance. ๐Ÿงต9/14

Sep 9, 2026 ยท 3:08 PM UTC

1
1
327
๐ŸŒ On real-world evaluation, we evaluate on three setups: bimanual, single-arm and dexterous bimanual manipulation. For bimanual manipulation, we use RoboDojo-Real, a comprehensive real-world benchmark spanning 18 tasks and 3 embodiments. ๐Ÿ† OpenWAM-ฮฑ ranks #๐Ÿญ on the RoboDojo-Real leaderboard, nearly twice the success rate of the next-best model (ฯ€0.5). ๐Ÿงต10/14
1
1
317
๐Ÿคณ On a single-arm real robot, OpenWAM-ฮฑ outperforms LingBot-VA (representative WAM baseline) and ฯ€0.5 (representative VLA baseline) on success rate averaged across the 6 evaluated tasks. ๐Ÿงต11/14
1
1
301
โœ‹๐Ÿฟ Can the pretrained model adapt to a dexterous embodiment that was absent from pretraining? ๐Ÿ”Œ We fine-tune OpenWAM-ฮฑ on four tasks using a Wuji hand mounted on a Tianji arm. This platform and its action space do not appear in the pretraining mixture. ๐Ÿ•บ We evaluate both in-domain conditions and OOD variations in background, layout, lighting, and objects, with the applicable variations defined for each task. OpenWAM-ฮฑ outperforms ฯ€0.5 across all four tasks in both the ID and aggregated OOD results. โš™๏ธ More results on the website: openwam-official.github.io/ ๐Ÿงต12/14
1
1
294
๐Ÿ“ฆ We release OpenWAM as a full stack empowering WAM research: - Modular infrastructure for training, inference, deploy, and eval. - OpenWAM-ฮฑ pretrained and post-trained weights. - Study checkpoints for exploring the design choices. - Evaluation protocols and data recipes. ๐Ÿฅž We hope researchers can use the stack to reproduce the results, adapt OpenWAM-ฮฑ to new tasks and embodiments, and test new ideas for representations, worldโ€“action interaction, and embodied pretraining. ๐Ÿ”ญ There is much more to explore. We hope OpenWAM-ฮฑ serves as a strong, reproducible baseline and OpenWAM-Infra makes that next round of experiments easier to build. ๐ŸŒ Project: openwam-official.github.io/ ๐Ÿ“„ Paper: arxiv.org/abs/2609.07398 ๐Ÿ’ป Code: github.com/OpenWAM-Official/โ€ฆ ๐Ÿค— Models & data: huggingface.co/OpenWAM ๐Ÿงต13/14
1
1
283
๐Ÿ™ OpenWAM is a large undertaking, made possible by the dedication and collective effort of the entire team. Thank you to everyone who contributed to this project, couldn't have done it without y'all! ๐Ÿคž Iโ€™m grateful for the opportunity to have co-lead this project with @YuranWang_rise. Special thanks to the amazing adviors: @zhaohang0124, @linshaonju, @HaoDong123, and @LotusSapphire for their advice and guidance throughout the project. ๐Ÿงต14/14
2
263
Sort replies: Relevant Recent Liked