My first PhD paper is out! ๐
"What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?"
tl:dr: JEPA-WMs for robotics: learn dynamics on top of visual encoders, optimize actions towards goal ๐
w/ @JimmyTYYang1, Jean Ponce, @AdrienBardes, @ylecun
14
110
946
148,024
1. ๐ช๐ต๐ฎ๐ ๐ฎ๐ฟ๐ฒ ๐๐๐ฃ๐-๐ช๐ ๐?
JEPA-WMs learn a dynamics predictor in the embedding space of a frozen visual encoder.
At planning time, we sample action sequences, rollout in latent space, and optimize trajectories that reach the goal ๐ฏ
1
16
3,213
2. ๐ข๐๐ฟ ๐๐๐๐ฑ๐
We ran a comprehensive study across 8 environments: navigation ๐งญ (Maze, Wall, Push-T) and manipulation ๐ฆพ (Metaworld, Robocasa, real Franka arm)
Goal: find what really matters for planning success โ
1
14
2,066
3. ๐ฃ๐น๐ฎ๐ป๐ป๐ถ๐ป๐ด ๐ผ๐ฝ๐๐ถ๐บ๐ถ๐๐ฒ๐ฟ
CEM Lโ wins overall ๐
Gradient-based (Adam, GD) excel on smooth landscapes (Metaworld) but fail on navigation/contact-rich tasks due to local minima
We also introduce a Nevergrad interface - competitive with less tuning!
1
11
1,820
4. ๐ ๐๐น๐๐ถ๐๐๐ฒ๐ฝ ๐ฟ๐ผ๐น๐น๐ผ๐๐
Training with 2-step rollout (predicting from own predictions) beats pure teacher forcing
But too many rollout steps hurts on simpler tasks! ๐
Real-world data (DROID) can benefit from more
1
1
10
1,593
5. ๐๐๐ก๐ข > ๐ฉ-๐๐๐ฃ๐ ๐ณ๐ผ๐ฟ ๐บ๐ฎ๐ป๐ถ๐ฝ๐๐น๐ฎ๐๐ถ๐ผ๐ป
Frozen image encoders (DINOv2, DINOv3) consistently outperform video encoders (V-JEPA, V-JEPA-2)
Why? Fine-grained object segmentation matters for precise control ๐
2
4
25
2,201
6. ๐ฃ๐ฟ๐ผ๐ฝ๐ฟ๐ถ๐ผ๐ฐ๐ฒ๐ฝ๐๐ถ๐ผ๐ป ๐ต๐ฒ๐น๐ฝ๐
Adding proprioceptive input and in the goal conditioning consistently boosts performance on simulated tasks
Precise goal distance โ less oscillation around target ๐
1
11
1,420
7. ๐ฃ๐ฟ๐ฒ๐ฑ๐ถ๐ฐ๐๐ผ๐ฟ ๐ฎ๐ฟ๐ฐ๐ต๐ถ๐๐ฒ๐ฐ๐๐๐ฟ๐ฒ
AdaLN conditioning outperforms feature/sequence conditioning on average ๐๏ธ
1
9
1,348
8. ๐๐ผ๐ป๐๐ฒ๐
๐ ๐๐ถ๐๐ฒ
Training maximum context length matters: Wโฅ2 is crucial (model needs to infer velocity!) but too long hurts โ fewer unique trajectories seen during training ๐
1
12
1,304
9. ๐ ๐ผ๐ฑ๐ฒ๐น ๐๐ฐ๐ฎ๐น๐ถ๐ป๐ด
Larger encoders DON'T help on simple simulated tasks ๐ค But DO on real-world data: bigger encoder = better!
Deeper predictor also helps on real-world data (DROID/Robocasa)
Scaling benefits depend on task complexity ๐
1
11
1,274
Huge thanks to my co-authors @JimmyTYYang1, Jean Ponce, @AdrienBardes, and @ylecun for their guidance ๐
Special thanks to @garridoq_ and @DanielDugas14 !
Work done at @AIatMeta & @Inria
๐ Paper: arxiv.org/abs/2512.24497
๐ป Code + data + checkpoints: github.com/facebookresearch/โฆ
Jan 12, 2026 ยท 3:46 PM UTC
3
9
106
25,447












