We live in a multimodal world. We see, talk, act, and dream.
Yet most LLMs still start with language pretraining. Why not train them natively with multimodal I/O from scratch?
Because it’s SUPER HARD, adding modalities triggers training instability, design complexity, and often modal competition
So what’s the path forward?
Introducing: Towards Physics of Multimodal Pretraining (junlinhan.github.io/projects…)
We unpack the underlying mechanics of multimodal pretraining across 4 aspects: Knowledge Flow, Modality Synergy, Early Unification, and Recipe.
Aug 6, 2026 · 9:41 PM UTC
28
171
1,031
109,343
[1/5] Knowledge flow:
Knowledge flow is highly asymmetric and concept-dependent. We demonstrate this in both real-world data and controlled synthetic environments:
Language is a universal booster for all visual tasks.
Understanding is a strong prior for generation, significantly boosting visual creation.
Generation acts as a latent prior. it doesn’t yield instant zero-shot gains for understanding, but sharply accelerates low-level concept acquisition.
1
1
23
3,773
[2/5] Modality Synergy:
Joint training isn't automatically synergistic; interaction depends heavily on task complexity and architecture. Key findings across our experimental axes:
Task complexity: Simple tasks foster cross-modal synergy, while complex tasks trigger competition.
Architecture sharing: Share Attention & Norms to drive cross-modal interaction and synergy; Decouple FFNs to eliminate capacity bottlenecks and parameter competition.
Encoder generalization: These architectural guidelines (and knowledge flow results) consistently hold across diverse vision representations (continuous vs. discrete, pixel-level vs. semantic).
1
1
20
2,356
[3/5] Early Unification:
Timing matters. Early unification with joint training allows modalities to better co-evolve.
Simultaneous early fusion significantly outperforms late alignment or sequential training.
Delayed integration causes "Vision Laziness": Pretraining on text first creates strong language priors, causing the model to perceive and learn fewer visual signals.
1
1
28
1,782
[4/5] Recipe:
Translating physical mechanics into a practical scaling recipe:
Data recipe: Allocating as little as 5% of the token budget to visual generation achieves competitive generative modeling while preserving strong language & understanding capabilities.
Scale Verification: Validated on multiple 13.5B MoE models trained across 2T tokens , proving that early fusion + asymmetric data blend + MoE scales efficiently.
2
1
23
1,662
[5/5] This is likely the final paper of my PhD.
When I began my PhD around the release of GPT-4V (Oct 2023), the AI landscape was very different. My entire doctoral journey has been a progression from computer vision research to cross-modality interaction, and ultimately to native multimodality.
I hope our exploration serves as a small step towards advancing multimodal intelligence. I still believe vision is a fundamental component of intelligence, and future models should pursue natively multimodal co-evolution, even if it's harder.
Please visit our webpage for more details: junlinhan.github.io/projects…
Joint work with @TongPetersb , @DavidJFan , @MinghaoChen23, @philiptorr , @filippos_kok, and @ml_perception .
1
2
36
1,627
























