We live in a multimodal world. We see, talk, act, and dream. Yet most LLMs still start with language pretraining. Why not train them natively with multimodal I/O from scratch? Because it’s SUPER HARD, adding modalities triggers training instability, design complexity, and often modal competition So what’s the path forward? Introducing: Towards Physics of Multimodal Pretraining (junlinhan.github.io/projects…) We unpack the underlying mechanics of multimodal pretraining across 4 aspects: Knowledge Flow, Modality Synergy, Early Unification, and Recipe.
28
171
1,031
109,357
Nice work! Congratulations Junlin!
1
2
164
Replying to @XDWang101
Thank you XuDong!! TV2TV and ReconAlign are super inspiring to me!

Aug 7, 2026 · 6:17 PM UTC

1
122
Sort replies: Relevant Recent Liked