We show several interesting findings in the paper. TL;DR: When multimodal models are trained from scratch, knowledge flows asymmetrically, task complexity determines whether modalities synergize or compete, and late vision alignment leads to “vision laziness.”
We live in a multimodal world. We see, talk, act, and dream. Yet most LLMs still start with language pretraining. Why not train them natively with multimodal I/O from scratch? Because it’s SUPER HARD, adding modalities triggers training instability, design complexity, and often modal competition So what’s the path forward? Introducing: Towards Physics of Multimodal Pretraining (junlinhan.github.io/projects…) We unpack the underlying mechanics of multimodal pretraining across 4 aspects: Knowledge Flow, Modality Synergy, Early Unification, and Recipe.
5
341
Minghao Chen retweeted
Can video generation models do for vision what LLMs did for language? Introducing GenCeption from @GoogleDeepMind: one feed-forward video model for various vision tasks — SOTA, data-efficient, and emerging behaviors (ECCV 2026) 🌐 genception.github.io (1/8)
17
99
468
118,032
AutoPartGen code is now available! 🚀 We are excited to finally release a reimplementation of AutoPartGen, as promised! Check out the code, models, project page, and paper here: 💻 Code: github.com/facebookresearch/… 🤗 Models: huggingface.co/facebook/auto… 🌐 Project page: silent-chen.github.io/AutoPa… 📎 Paper: arxiv.org/abs/2507.13346 We sincerely apologize for the long delay, and we greatly appreciate your patience and continued interest in our work!
4
26
209
95,671
It has been a long journey to this code release. Huge thanks to all the co-authors for their help and contributions in making this possible: @jianyuan_wang , @ovrdr , @t_monnier , Hyunyoung Jung , @dilin_wang , Rakesh Ranjan , @irolaina , Andrea vedaldi
3
426
Minghao Chen retweeted
Congrats to the team for wining CVPR Best Paper Award!! 🏆 Come to our oral session (Mile High Ballroom 13:00-14:15) and poster (16:00-18:00) today for more details 🚀
A SINGLE encoder + decoder for all the 4D tasks! We release 🎯 D4RT (Dynamic 4D Reconstruction and Tracking). 📍 A simple, unified interface for 3D tracking, depth, and pose 🌟 SOTA results on 4D reconstruction & tracking 🚀 Up to 100x faster pose estimation than prior works
22
37
297
43,529
Introducing VGG-Omega, our new 4D foundation model and the next step beyond VGGT! 🥳 Over the past year, we explored a broad range of ideas, directions, and approaches to push the boundaries of spatial intelligence. We hope VGG-Omega sparks fresh inspiration, new insights, and exciting future directions for the community! 🧐 Huge thanks to all the co-authors for the amazing collaboration: @jianyuan_wang , Shangzhan Zhang, @n_karaev , Johannes Schönberger, @monsieurlabatut , @p_bojanowski , @davnov134 , Andrea Vedaldi, Christian Rupprecht
Introducing VGGT-Ω: scaling feed-forward reconstruction across static and dynamic scenes, and studying whether the learned geometric representations transfer beyond reconstruction.
10
19
180
658,726
Minghao Chen retweeted
🚀 Introducing Articraft, a coding agent for articulated 3D asset creation. Articraft writes code, executes it, receives validation feedback, and refines the result into simulation-ready 3D assets with parts, joints, and motion. We’re also releasing Articraft-10K: 10,000+ articulated objects across 250 categories, unlocking large-scale interactive scenes for robotics simulation and physical AI. 🔗 Project page: articraft3d.github.io/ 💻 Code: github.com/mattzh72/articraf…
22
108
742
191,947
Minghao Chen retweeted
We made an interactive client-server viewer for LagerNVS with @JonathonLuiten! You can now interactively explore scenes from just a photo capture - no optimization, no 3D Gaussians, just load your images, run the model on a cloud GPU and stream the renders to your local browser. Check out the video below for some spaces I recently captured in Oxford, London and beyond!
5
26
175
17,358
Check out our recent work on generalizable real-time novel view synthesis. 🥳🥳 It suggests that explicit 3D representations may not be necessary, as long as the right structural biases are learned.
🍺 LagerNVS (CVPR 2026) 🍺 LagerNVS is a generalizable, feed-forward, real-time Novel View Synthesis network which - performs rendering in real time, - generalizes to in-the-wild data, - works with and without known source cameras, - sets a new state-of-the-art among deterministic methods, - can be paired with a diffusion decoder for generative extrapolation. LagerNVS shows that 3D biases are useful for Novel View Synthesis but explicit 3D representations are not required to achieve them. We use 3D biases in (1) architecture design and (2) pre-training: (1) In NVS with explicit 3D representations (3DGS, NeRF) reconstruction is typically difficult and slow, but rendering is much faster and simpler. We mimic this process in the network design: we use a large (1B params) encoder and a small, lightweight decoder (ViT-B). This allows increasing the network capacity while still achieving real-time rendering. (2) The encoder, initialized from VGGT, was pre-trained with 3D reconstruction objectives, making the initial features 3D aware. Both substantially improve performance. Project page: szymanowiczs.github.io/lager… Code: github.com/facebookresearch/… Paper: arxiv.org/abs/2603.20176 Models: huggingface.co/collections/f… Work done with @jianyuan_wang @MinghaoChen23 Christian Rupprecht and Andrea Vedaldi
3
18
3,119
Minghao Chen retweeted
A SINGLE encoder + decoder for all the 4D tasks! We release 🎯 D4RT (Dynamic 4D Reconstruction and Tracking). 📍 A simple, unified interface for 3D tracking, depth, and pose 🌟 SOTA results on 4D reconstruction & tracking 🚀 Up to 100x faster pose estimation than prior works
18
70
444
119,765
🥳Excited to present AutoPartGen at #NeurIPS2025 in Mexico City, our new work on part-level 3D generation! Come chat about the next frontier of 3D generation at my poster! 📅 Wed, December 3, 6:30PM-9:30PM 📍 Hilton Mexico City Reforma (Foyer) Paper: arxiv.org/abs/2507.13346 Project page: silent-chen.github.io/AutoPa…
2
8
364
🚀Check out our new paper on improving VLMs’ understanding of physics implausibility! We contribute: 1⃣ A trajectory-aware attention fine-tuning recipe that helps vision encoders of VLM grasp motion and temporal cues. 2⃣ A challenging benchmark to test VLMs’ ability to detect and reasoning about physics-implausible events. 🔗 Project: sam-motamed.github.io/projec… 🤗 Dataset: github.com/insait-institute/… 📜 Paper: arxiv.org/abs/2510.07550Huge Thanks to @sammtmd for leading the paper and dataset to top quality!
🚀📣 Excited to share our new paper: TRAVL — A Recipe for Making Video-Language Models Better Judges of Physics Implausibility! ⭐ GitHub: github.com/insait-institute/… 👨🏻‍💻 Project page: sam-motamed.github.io/projec… 🎥 TL;DR: Video generative models look increasingly realistic, but still break physics (teleportation, gravity violations, impossible deformations). If we want multimodal systems that reason about the physical world, we first need reliable judges of plausibility. Can VLMs become that plausibility judge? 🤔 We introduce: • 🍳 TRAVL — a drop-in trajectory-aware attention recipe that keeps VLM backbones frozen while making vision tokens richer in motion and scene detail. • 🗂️ TRAVL Training Dataset — 3,482 short videos with 19,708 physics-focused Q/A pairs (balanced real + implausible). • 🧩 ImplausiBench — a 300-video benchmark (150 real / 150 implausible) with matched first frames & adversarial MCQs to blunt language shortcuts; we report both Human and LLM-as-judge scores. 🏆 Result: LLaVA-NeXT + TRAVL achieves the best performance on the implausible split in our study — outperforming Gemini 2.5 Pro and GPT-4o! 💡 Why it matters: Current VLMs often miss physics violations because full-video encoding explodes token counts → training resorts to sparse frames + heavy pooling (losing motion cues), and most training data is only real footage. TRAVL fixes this by injecting spatial self-attention (within frames) + trajectory-guided temporal attention (across frames), so models track continuity, contact-before-motion, and no-teleport constraints — the stuff physics is made of. 🧰 Use it today: • Method: Add TRAVL on top of a frozen VLM to enrich motion & scene context without blowing up tokens. • Data: Fine-tune with our curated 3,482 videos / 19,708 QAs focused on physics plausibility. • Eval: Rigorously benchmark VLMs' physics understanding with ImplausiBench! Please leave us a ⭐if you find this work interesting! 👨🏻‍💻Project webpage: sam-motamed.github.io/projec… 🤗 Code and Datasets: github.com/insait-institute/… 📜 arXiv: arxiv.org/abs/2510.07550 👥 Authors: @sammtmd , @MinghaoChen23 , Luc Van Gool, Iro Laina @INSAITinstitute | @Oxford_VGG
3
377
Minghao Chen retweeted
Thrilled and honored to receive the Best Paper Award at #CVPR2025! Huge thanks to my fantastic collaborators @MinghaoChen23, @n_karaev, Andrea Vedaldi, Christian Rupprecht, and @davnov134. Could not be there without you!
39
17
470
37,908
This is amazing! We did it!
Many Congratulations to @jianyuan_wang, @MinghaoChen23, @n_karaev, Andrea Vedaldi, Christian Rupprecht and @davnov134 for winning the Best Paper Award @CVPR for "VGGT: Visual Geometry Grounded Transformer" 🥇🎉 🙌🙌 #CVPR2025!!!!!!
4
2
50
3,399
Check out this amazing project from Stan! The results looks super impressive!
⚡️ Introducing Bolt3D ⚡️ Bolt3D generates interactive 3D scenes in less than 7 seconds on a single GPU from one or more images. It features a latent diffusion model that *directly* generates 3D Gaussians of seen and unseen regions, without any test time optimization. 🧵👇 (1/9)
1
5
735
Check out our new 3D reconstruction work, VGGT! 🥳🥳🥳 VGGT predicts all key 3D attributes (cameras, point maps, depths, point tracks) from arbitrary input views in seconds!
Introducing VGGT (CVPR'25), a feedforward Transformer that directly infers all key 3D attributes from one, a few, or hundreds of images, in seconds! No expensive optimization needed, yet delivers SOTA results for: ✅ Camera Pose Estimation ✅ Multi-view Depth Estimation ✅ Dense Point Cloud Reconstruction ✅ Point Tracking Project Page: vgg-t.github.io/ Code & Weights: github.com/facebookresearch/…
7
76
5,741
🥳Excited to share my recent work at Meta, "PartGen: Part-level 3D Generation and Reconstruction with Multi-View Diffusion Models", which aims at compositional/part-level 3D generation and reconstruction from various modalities. Project page: silent-chen.github.io/PartGe…
3
48
228
35,186
Excited to share my paper "DGE: Direct Gaussian 3D Editing by Consistent Multi-view Editing" on fast 3D editing at Milan for #ECCV2024 Come to talk and chat with me at Poster # 287 on Friday morning! Project page: silent-chen.github.io/DGE/ Codes: github.com/silent-chen/DGE
14
72
5,554
Minghao Chen retweeted
Releasing Flex3D, a two-stage pipeline for generating high-quality 3D assets in a feed-forward manner, as a further step toward high-quality 3D generation and reconstruction. Project page: junlinhan.github.io/projects…
3
14
113
9,500