1/
Text-to-video models generate beautiful videos, but they still struggle to follow complex prompts.
Relations like left/right, towards/away, on top of, behind, or even multi-stage instructions are often generated incorrectly.
In our new paper, accepted to SIGGRAPH Asia 2026, we introduce CVG, an inference-time guidance method that improves compositional video generation without retraining or architectural changes.