My first PhD paper is out! ๐ŸŽ“ "What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?" tl:dr: JEPA-WMs for robotics: learn dynamics on top of visual encoders, optimize actions towards goal ๐Ÿ‘‡ w/ @JimmyTYYang1, Jean Ponce, @AdrienBardes, @ylecun
14
110
946
148,024
1. ๐—ช๐—ต๐—ฎ๐˜ ๐—ฎ๐—ฟ๐—ฒ ๐—๐—˜๐—ฃ๐—”-๐—ช๐— ๐˜€? JEPA-WMs learn a dynamics predictor in the embedding space of a frozen visual encoder. At planning time, we sample action sequences, rollout in latent space, and optimize trajectories that reach the goal ๐ŸŽฏ
1
16
3,213
2. ๐—ข๐˜‚๐—ฟ ๐˜€๐˜๐˜‚๐—ฑ๐˜† We ran a comprehensive study across 8 environments: navigation ๐Ÿงญ (Maze, Wall, Push-T) and manipulation ๐Ÿฆพ (Metaworld, Robocasa, real Franka arm) Goal: find what really matters for planning success โœ…
1
14
2,066
3. ๐—ฃ๐—น๐—ฎ๐—ป๐—ป๐—ถ๐—ป๐—ด ๐—ผ๐—ฝ๐˜๐—ถ๐—บ๐—ถ๐˜‡๐—ฒ๐—ฟ CEM Lโ‚‚ wins overall ๐Ÿ† Gradient-based (Adam, GD) excel on smooth landscapes (Metaworld) but fail on navigation/contact-rich tasks due to local minima We also introduce a Nevergrad interface - competitive with less tuning!
1
11
1,820
4. ๐— ๐˜‚๐—น๐˜๐—ถ๐˜€๐˜๐—ฒ๐—ฝ ๐—ฟ๐—ผ๐—น๐—น๐—ผ๐˜‚๐˜ Training with 2-step rollout (predicting from own predictions) beats pure teacher forcing But too many rollout steps hurts on simpler tasks! ๐Ÿ“‰ Real-world data (DROID) can benefit from more
1
1
10
1,593
5. ๐——๐—œ๐—ก๐—ข > ๐—ฉ-๐—๐—˜๐—ฃ๐—” ๐—ณ๐—ผ๐—ฟ ๐—บ๐—ฎ๐—ป๐—ถ๐—ฝ๐˜‚๐—น๐—ฎ๐˜๐—ถ๐—ผ๐—ป Frozen image encoders (DINOv2, DINOv3) consistently outperform video encoders (V-JEPA, V-JEPA-2) Why? Fine-grained object segmentation matters for precise control ๐Ÿ”
2
4
25
2,201
6. ๐—ฃ๐—ฟ๐—ผ๐—ฝ๐—ฟ๐—ถ๐—ผ๐—ฐ๐—ฒ๐—ฝ๐˜๐—ถ๐—ผ๐—ป ๐—ต๐—ฒ๐—น๐—ฝ๐˜€ Adding proprioceptive input and in the goal conditioning consistently boosts performance on simulated tasks Precise goal distance โ†’ less oscillation around target ๐Ÿ“
1
11
1,420
7. ๐—ฃ๐—ฟ๐—ฒ๐—ฑ๐—ถ๐—ฐ๐˜๐—ผ๐—ฟ ๐—ฎ๐—ฟ๐—ฐ๐—ต๐—ถ๐˜๐—ฒ๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ AdaLN conditioning outperforms feature/sequence conditioning on average ๐Ÿ—๏ธ
1
9
1,348
8. ๐—–๐—ผ๐—ป๐˜๐—ฒ๐˜…๐˜ ๐˜€๐—ถ๐˜‡๐—ฒ Training maximum context length matters: Wโ‰ฅ2 is crucial (model needs to infer velocity!) but too long hurts โ€” fewer unique trajectories seen during training ๐Ÿ“
1
12
1,304
9. ๐— ๐—ผ๐—ฑ๐—ฒ๐—น ๐˜€๐—ฐ๐—ฎ๐—น๐—ถ๐—ป๐—ด Larger encoders DON'T help on simple simulated tasks ๐Ÿค” But DO on real-world data: bigger encoder = better! Deeper predictor also helps on real-world data (DROID/Robocasa) Scaling benefits depend on task complexity ๐Ÿ“ˆ
1
11
1,274
10. ๐—ฅ๐—ฒ๐˜€๐˜‚๐—น๐˜๐˜€ ๐ŸŽ‰ Our best model predicts object interaction and beats DINO-WM & V-JEPA-2-AC! Key recipe: โ€ข DINOv2/v3 encoder, Large for real-world data โ€ข AdaLN predictor + 2-step rollout โ€ข Proprioception if available โ€ข CEM Lโ‚‚ planner
2
17
2,011
Sort replies: Relevant Recent Liked
Interesting point about context length in training! Finding the sweet spot is key for model performance.
1
178
Agreed, finding the right balance is key for optimal training results.
1
189