[1/9] What happens when you treat vision as a first-class citizen during multimodal pretraining? To find out, we studied the design space of training Transfusion-style models that input and output all modalities, from scratch. Here is what we learned about visual representations, data, world modeling, architecture, and scaling behavior! Paper: arxiv.org/abs/2603.03276 Website: beyond-llms.github.io/ @TongPetersb, @DavidJFan, @__JohnNguyen__, @ellisbrown, @GaoyueZhou, @JasonQSY, @boyangzheng, @webalorn, @han_junlin, @rob_fergus, @NailaMurray, @gh_marjan, @ml_perception, Nicolas Ballas, @_amirbar, Michael Rabbat, Jakob Verbeek, @LukeZettlemoyer, @koustuvsinha, @ylecun, @sainingxie
13
60
310
76,942
[2/9] Not-So “Unified” Landscape. Most current methodologies just finetune pretrained LLMs. The emphasis is more on preserving language, rather than treating vision as a first-class citizen and truly learning from all modalities simultaneously. This also confounds the science: are we studying the effect of multimodal training or language initialization? We trained from scratch to truly understand the dynamics of multimodal pretraining.

Mar 4, 2026 · 4:57 PM UTC

1
1
16
1,663
[3/9] Towards Simplicity. We also wanted to simplify the architecture. Using separate encoders for visual understanding vs. generation adds unnecessary complexity. We prove that you only need ONE encoder to excel at both. Just use RAE (e.g. SigLIP 2 or WebSSL). Raw pixel with x-pred is a promising future direction.
1
1
13
1,258
[4/9] The “Modality Tax” myth. Does adding vision destroy your language modeling? Mostly not. We find that raw video is very compatible with language modeling. The performance drop often associated with multimodal training actually comes from the distributional gap of image captions, not from visual inputs themselves! Adding multimodal data just slightly hurts your OOD text generalization, but with large benefits for tasks such as visual understanding and generation, and even world modeling.
1
1
12
996
[5/9] World Models <-> Multimodal Models. World models don’t need specialized architectures nor large domain-specific data! By formatting actions directly as text tokens and training with general multimodal data, the model can naturally learn tasks such as navigation, and be used for MPC planning. No latent actions, no specialized action heads. We even show that the model can qualitatively generalize OOD using free-form language like “get out of the shadow!”, unlocking a nearly infinite set of actions, unlike existing works such as Genie. This is an insight that I'm personally very excited about :)
1
2
18
1,566
[6/9] Multimodal MoE Works! It’s well-known that MoE provides effective yet efficient scaling for LLMs. But what about unified multimodal models? We show that multimodal MoE benefits greatly from higher granularity and sparsity, especially when paired with semantic visual representations such as RAE. We also see that MoE naturally learns modality specialization from the data, flexibly allocating capacity per input, unlike hard-coded designs such as modality-specific FFN.
1
1
11
795
[7/9] MoE Bridges the Vision / Language Scaling Asymmetry. We provide the first IsoFLOP analysis for both vision and language within the same model. Our analysis shows that vision and language scale differently; vision is more data hungry while language is more parameter hungry - it’s hard to be compute optimal for both. But MoE saves the day! MoE bridges the gap so that the data requirements of vision and language become nearly identical. This means that MoE is not just a way to efficiently scale, but also an architectural necessity for building unified multimodal models.
1
1
12
743
[8/9] What's Next? Ultimately, we hope this work pushes the community towards unifying multimodal and world models, to build systems that truly understand and reason about the physical world. We aren’t quite there, but this is our first step towards this vision. @ylecun
1
1
10
667
[9/9] On a personal note: I grew a lot from this project. More challenging than the technical hurdles was navigating organizational dynamics, advocating the research vision, and securing the compute resources. A huge thanks to the incredible team that made this a reality, to FAIR for the support, and especially to @TongPetersb + @__JohnNguyen__. We started this journey together and stuck by each other’s side through all the highs and lows. This work builds upon all of our collective experiences in and long-term beliefs about visual representations + multimodal modeling. I see this as just the first step in a long-term research agenda. I couldn't ask for a better partnership!
2
1
16
1,332
Also many thanks to the numerous people who gave feebdack on the work and supported us through this journey! @JimmyTYYang1, @sharut_gupta, @JiachenAI, @jiawzhao, @Hu_Hsu, @em_dinan, @XiaochuangHan, @TusharNagarajan, @TianhongLi6, @jilin_14, @Aniket_d98, @shushengyang, @jihanyang13, @WeijiaShi2, @gene_ch0u, Daniel Bolya, just to name a few
16
1,228
Sort replies: Relevant Recent Liked