๐ค World Action Models predict the future to act better. However, WAM research remains fragmented in code, ad hoc in design, and unprincipled at scale.
๐ก Today, we release ๐ข๐ฝ๐ฒ๐ป๐ช๐๐ : ๐๐ป ๐ข๐ฝ๐ฒ๐ป, ๐ ๐ผ๐ฑ๐๐น๐ฎ๐ฟ ๐๐
๐ฝ๐น๐ผ๐ฟ๐ฎ๐๐ถ๐ผ๐ป ๐ง๐ผ๐๐ฎ๐ฟ๐ฑ๐ ๐ฆ๐๐๐๐ฒ๐บ๐ฎ๐๐ถ๐ฐ ๐ช๐ผ๐ฟ๐น๐ฑโ๐๐ฐ๐๐ถ๐ผ๐ป ๐ ๐ผ๐ฑ๐ฒ๐น ๐ฃ๐ฟ๐ฒ๐๐ฟ๐ฎ๐ถ๐ป๐ถ๐ป๐ด to change this.
๐ openwam-official.github.io/
๐งต1/14
6
40
267
2,266,269
OpenWAM is three releases in one:
- ๐งฑ ๐ข๐ฝ๐ฒ๐ป๐ช๐๐ -๐๐ป๐ณ๐ฟ๐ฎ provides a modular substrate for the systematic study of WAMs.
- ๐ฌ ๐ข๐ฝ๐ฒ๐ป๐ช๐๐ -๐ฆ๐๐๐ฑ๐ poses three central questions around WAM pretraining, and investigated them through controlled experiments.
- ๐ ๐ข๐ฝ๐ฒ๐ป๐ช๐๐ -๐๐น๐ฝ๐ต๐ฎ is a fully-open WAM pretrained on ~6,400 hrs (518.5M frames) of mixed egocentric human + robot data built on the insights from OpenWAM-Study, and delivers excellent performance across both simulation and real-world tasks.
๐งต2/14
1
7
1,521
๐งฑ ๐ข๐ฝ๐ฒ๐ป๐ช๐๐ -๐๐ป๐ณ๐ฟ๐ฎ provides the common infrastructure for this exploration.
โญ๏ธ The key underlying design principle is ๐ข๐ค๐๐ช๐ก๐๐ง๐๐ฉ๐ฎ. OpenWAM-Infra separates data loading, visual encoding, video backbones, worldโaction architectures, and attention masks into composable modules, supporting 4 visual encoders, 5 video backbones, 6 architecture variants, and 4 attention masks.
๐ฆ These components share training, inference, deployment, and evaluation interfaces, so a new design can be simply plugged in while reusing the same surrounding pipeline. In addition, we integrate 8 simulation benchmarks and support deployment on 5 real robots.
๐งต3/14
1
2
864
๐ฌ Building on the substrate of OpenWAM-Infra, we systematically study three questions.
โ๏ธ Q1. What upstream knowledge should a WAM inherit?
๐ We study two sources of this upstream knowledge: the generative backbone and the representation space.
๐ฎ On the generative backbone side, we find using video generation models with increasing parameter counts often lead to better success rate, despite coming from different model families and lineages.
๐ฅ On the representation space choice, we find directly switching vision encoders to representation encoders (DINOv3, V-JEPA 2.1) does not fruit well. However, when combined with a S-VAE layer that compresses high-dimensional representations to low-dimensional vectors, the resulting system yields better success rate than reconstructive encoders (Flux.2 VAE) with no native temporal compression. But still, these models slightly underperform Wan-VAE, which is by default co-designed with the DiT backbone, as well as including native temporal compression.
๐คฒ The finding opens up more choices for WAMs: representation encoders can work well when their features are made compact enough for worldโaction learning. Latent dimensionality and temporal compression deserve attention alongside the choice of backbone.
๐งต4/14
1
1
1,043
โ๏ธ Q2: How do we create synergy between world and action learning during training?
๐ The architecture study favors dedicated capacity for action prediction. By comparing single- dual- and tri-system architectures, we find that giving the world and action streams their own parameters is essential, and we carry forward this dual-system joint self-attention architecture into OpenWAM-Alpha.
๐ We then vary the information flow by altering cross-modality attention masks. In our experiments, allowing the action stream to access world features reaches 92.39% average success, versus 87.41% when the streams are isolated. Allowing only the world stream to see actions gives 87.63%.
๐ Together, these results motivate a dedicated action branch with explicit access to the world stream. We revisit whether that access should become fully mutual after embodied pretraining later in the study.
๐งต5/14
1
1
485
๐งฌ Worldโaction interaction also depends on how the two streams are denoised at inference.
๐๏ธ We compare inference schedules in which video leads, action leads, or both advance together. For this experiment, training covers independently sampled world and action noise levels, so we can test different paths through the joint denoising process.
๐ Synchronized joint denoising performs best in our setting, leading at 93.0% success, compared with 88.5โ92.3% for the alternatives.
๐งต6/14
1
1
401
โ๏ธ Q3: What does embodied pretraining add, and how should we combine human and robot data?
๐ We study this under a fixed 600-hour data budget, followed by the same downstream fine-tuning. Evaluation uses RoboTwin2.0 CleanโRandomized, separating in-domain performance from out-of-domain generalization.
๐ฉ For mixed egocentric + robot pretraining, the main gain is OOD: approximately +12 percentage points, while the ID improvement is less than 1 point.
๐คฒ Robot trajectories provide action grounding; egocentric human data broadens world coverage. One-stage co-training combines these sources effectively, with results close to the two-stage ego-then-robot recipe and a simpler training flow.
๐ฅท We also revisit the attention mask: after embodied pretraining, mutual worldโaction visibility is consistently preferred in our comparisons.
โก๏ธ These findings guide the final recipe: co-train ego and robot data in one stage, and let the two streams mutually accessible.
๐งต7/14
1
2
370
๐ ๐ข๐ฝ๐ฒ๐ป๐ช๐๐ -๐๐น๐ฝ๐ต๐ฎ brings those findings together in one pretrained model.
- World stream: Wan2.2-TI2V-5B with Wan2.2-VAE.
- Action stream: a dedicated 1B-parameter ActionDiT.
- Coupling: joint self-attention with mutual worldโaction visibility.
- Inference: synchronized joint denoising.
- Data: one-stage co-training on egocentric human and robot data.
๐ฅ We pretrain on 518.5M frames, roughly 6,369 hours. An 80-D unified action space gives the different embodiments a shared format with fixed slot semantics.
๐ง During training, we sample world and action noise levels independently. At inference, they follow the synchronized path selected in the study.
๐งช Starting from the pretrained checkpoint, we adapt and evaluate OpenWAM-ฮฑ across a wide range of simulation benchmarks and real-robot tasks.
๐งต8/14
1
1
461
๐ In simulation, OpenWAM-ฮฑ achieves top-tier performance across 8 benchmarks spanning 5 embodiment categories: single-arm, bimanual, mobile single-arm, mobile bimanual, and dexterous-hand manipulation.
๐ช OpenWAM-ฮฑ achieves strong performance across all benchmarks, including state-of-the-art results on EBench. Compared to representative VLA and WAM baselines, OpenWAM-ฮฑ consistently delivers outstanding performance.
๐งต9/14
Sep 9, 2026 ยท 3:08 PM UTC
1
1
327
๐ On real-world evaluation, we evaluate on three setups: bimanual, single-arm and dexterous bimanual manipulation. For bimanual manipulation, we use RoboDojo-Real, a comprehensive real-world benchmark spanning 18 tasks and 3 embodiments.
๐ OpenWAM-ฮฑ ranks #๐ญ on the RoboDojo-Real leaderboard, nearly twice the success rate of the next-best model (ฯ0.5).
๐งต10/14
1
1
317
๐คณ On a single-arm real robot, OpenWAM-ฮฑ outperforms LingBot-VA (representative WAM baseline) and ฯ0.5 (representative VLA baseline) on success rate averaged across the 6 evaluated tasks.
๐งต11/14
1
1
301
โ๐ฟ Can the pretrained model adapt to a dexterous embodiment that was absent from pretraining?
๐ We fine-tune OpenWAM-ฮฑ on four tasks using a Wuji hand mounted on a Tianji arm. This platform and its action space do not appear in the pretraining mixture.
๐บ We evaluate both in-domain conditions and OOD variations in background, layout, lighting, and objects, with the applicable variations defined for each task. OpenWAM-ฮฑ outperforms ฯ0.5 across all four tasks in both the ID and aggregated OOD results.
โ๏ธ More results on the website: openwam-official.github.io/
๐งต12/14
1
1
294
๐ฆ We release OpenWAM as a full stack empowering WAM research:
- Modular infrastructure for training, inference, deploy, and eval.
- OpenWAM-ฮฑ pretrained and post-trained weights.
- Study checkpoints for exploring the design choices.
- Evaluation protocols and data recipes.
๐ฅ We hope researchers can use the stack to reproduce the results, adapt OpenWAM-ฮฑ to new tasks and embodiments, and test new ideas for representations, worldโaction interaction, and embodied pretraining.
๐ญ There is much more to explore. We hope OpenWAM-ฮฑ serves as a strong, reproducible baseline and OpenWAM-Infra makes that next round of experiments easier to build.
๐ Project: openwam-official.github.io/
๐ Paper: arxiv.org/abs/2609.07398
๐ป Code: github.com/OpenWAM-Official/โฆ
๐ค Models & data: huggingface.co/OpenWAM
๐งต13/14
1
1
283
๐ OpenWAM is a large undertaking, made possible by the dedication and collective effort of the entire team. Thank you to everyone who contributed to this project, couldn't have done it without y'all!
๐ค Iโm grateful for the opportunity to have co-lead this project with @YuranWang_rise. Special thanks to the amazing adviors: @zhaohang0124, @linshaonju, @HaoDong123, and @LotusSapphire for their advice and guidance throughout the project.
๐งต14/14
2
263













