We live in a multimodal world. We see, talk, act, and dream. Yet most LLMs still start with language pretraining. Why not train them natively with multimodal I/O from scratch? Because it’s SUPER HARD, adding modalities triggers training instability, design complexity, and often modal competition So what’s the path forward? Introducing: Towards Physics of Multimodal Pretraining (junlinhan.github.io/projects…) We unpack the underlying mechanics of multimodal pretraining across 4 aspects: Knowledge Flow, Modality Synergy, Early Unification, and Recipe.

Aug 6, 2026 · 9:41 PM UTC

28
171
1,031
109,343
[1/5] Knowledge flow: Knowledge flow is highly asymmetric and concept-dependent. We demonstrate this in both real-world data and controlled synthetic environments: Language is a universal booster for all visual tasks. Understanding is a strong prior for generation, significantly boosting visual creation. Generation acts as a latent prior. it doesn’t yield instant zero-shot gains for understanding, but sharply accelerates low-level concept acquisition.
1
1
23
3,773
[2/5] Modality Synergy: Joint training isn't automatically synergistic; interaction depends heavily on task complexity and architecture. Key findings across our experimental axes: Task complexity: Simple tasks foster cross-modal synergy, while complex tasks trigger competition. Architecture sharing: Share Attention & Norms to drive cross-modal interaction and synergy; Decouple FFNs to eliminate capacity bottlenecks and parameter competition. Encoder generalization: These architectural guidelines (and knowledge flow results) consistently hold across diverse vision representations (continuous vs. discrete, pixel-level vs. semantic).
1
1
20
2,356
[3/5] Early Unification: Timing matters. Early unification with joint training allows modalities to better co-evolve. Simultaneous early fusion significantly outperforms late alignment or sequential training. Delayed integration causes "Vision Laziness": Pretraining on text first creates strong language priors, causing the model to perceive and learn fewer visual signals.
1
1
28
1,782
[4/5] Recipe: Translating physical mechanics into a practical scaling recipe: Data recipe: Allocating as little as 5% of the token budget to visual generation achieves competitive generative modeling while preserving strong language & understanding capabilities. Scale Verification: Validated on multiple 13.5B MoE models trained across 2T tokens , proving that early fusion + asymmetric data blend + MoE scales efficiently.
2
1
23
1,662
[5/5] This is likely the final paper of my PhD. When I began my PhD around the release of GPT-4V (Oct 2023), the AI landscape was very different. My entire doctoral journey has been a progression from computer vision research to cross-modality interaction, and ultimately to native multimodality. I hope our exploration serves as a small step towards advancing multimodal intelligence. I still believe vision is a fundamental component of intelligence, and future models should pursue natively multimodal co-evolution, even if it's harder. Please visit our webpage for more details: junlinhan.github.io/projects… Joint work with @TongPetersb , @DavidJFan , @MinghaoChen23, @philiptorr , @filippos_kok, and @ml_perception .
1
2
36
1,627
Sort replies: Relevant Recent Liked
Replying to @han_junlin
Nice work! Congratulations Junlin!
1
2
164
Thank you XuDong!! TV2TV and ReconAlign are super inspiring to me!
1
122
Replying to @han_junlin
Love it!! I've always felt a pull towards 'raising' my agents via digital stimuli and open exploration - this seems like the next level up
1
2
122
Thanks so much!! Moving from passive perception to active exploration ( in an open world) is such a cool direction!
1
79
Replying to @han_junlin
This is super super cool
1
1
161
Thank u so much Georgia!!! See u around in Oxford!
1
96
Congrats Junlin!
1
1
128
Thank u Tyler!! Learned a lot from u!
1
105
Replying to @han_junlin
Cool work! Just added this to my weekend reading list :)
2
3
378
🥹thanks Nan! Well then enjoy your weekend? 😆
1
330
Replying to @han_junlin
one more thing, The CLEVR experiment was the most interesting part for me. However, I have question about that experiment, zero-shot transfer fails both ways - that makes sense; But after fine-tuning, generation priors accelerate understanding while understanding barely helps generation. Why does the asymmetry only work in one Direction?? Is it because generation learns denser features?
1
2
519
Hey Manpreet, thanks for the great question! My understanding is that in the fine-tuning experiments, we are evaluating lower-level concepts (like color and shape) where generation priors can be more helpful. On the other hand, understanding helps generation in higher-level concepts, such as spatial relationships and counting (and help in understanding the text prompts).
1
2
435
Replying to @han_junlin
Impressive!
1
2
177
Thank u very much Jikun!
1
176
Replying to @han_junlin
Great work! I'm ready to follow it.
1
1
305
Thank you Yu!! Sorry, we don't have plans to open-source it, but I think we are still far away from understanding multimodal training. There are lots of things worth exploring!!
1
299
Replying to @han_junlin
really solid and impressive
1
1
360
Replying to @han_junlin
solid work. can't wait to dive into the details🫡
1
2
122
Thank u Qihang!! Hope details are clear enough!
99
Replying to @han_junlin
Impressive!
1
1
162
Thanks DY!
154
Replying to @han_junlin
Hey! Super cool work! we have also been working on this and would love to chat with you if possible!
1
1
192
Yes would love to! Lets' DM
1
175