🚀📣 Excited to share our new paper: TRAVL — A Recipe for Making Video-Language Models Better Judges of Physics Implausibility!
⭐ GitHub:
github.com/insait-institute/…
👨🏻💻 Project page:
sam-motamed.github.io/projec…
🎥 TL;DR:
Video generative models look increasingly realistic, but still break physics (teleportation, gravity violations, impossible deformations). If we want multimodal systems that reason about the physical world, we first need reliable judges of plausibility. Can VLMs become that plausibility judge? 🤔
We introduce:
• 🍳 TRAVL — a drop-in trajectory-aware attention recipe that keeps VLM backbones frozen while making vision tokens richer in motion and scene detail.
• 🗂️ TRAVL Training Dataset — 3,482 short videos with 19,708 physics-focused Q/A pairs (balanced real + implausible).
• 🧩 ImplausiBench — a 300-video benchmark (150 real / 150 implausible) with matched first frames & adversarial MCQs to blunt language shortcuts; we report both Human and LLM-as-judge scores.
🏆 Result:
LLaVA-NeXT + TRAVL achieves the best performance on the implausible split in our study — outperforming Gemini 2.5 Pro and GPT-4o!
💡 Why it matters:
Current VLMs often miss physics violations because full-video encoding explodes token counts → training resorts to sparse frames + heavy pooling (losing motion cues), and most training data is only real footage.
TRAVL fixes this by injecting spatial self-attention (within frames) + trajectory-guided temporal attention (across frames), so models track continuity, contact-before-motion, and no-teleport constraints — the stuff physics is made of.
🧰 Use it today:
• Method: Add TRAVL on top of a frozen VLM to enrich motion & scene context without blowing up tokens.
• Data: Fine-tune with our curated 3,482 videos / 19,708 QAs focused on physics plausibility.
• Eval: Rigorously benchmark VLMs' physics understanding with ImplausiBench!
Please leave us a ⭐if you find this work interesting!
👨🏻💻Project webpage:
sam-motamed.github.io/projec…
🤗 Code and Datasets:
github.com/insait-institute/…
📜 arXiv:
arxiv.org/abs/2510.07550
👥 Authors:
@sammtmd ,
@MinghaoChen23 , Luc Van Gool, Iro Laina
@INSAITinstitute |
@Oxford_VGG