Building AI that learns by interacting with the world. Associate Professor @ MIT, leading the Scene Representation Group (scenerepresentations.org).

Cambridge, Massachusetts
I am open-sourcing DeckWerk, a multi-platform slide editor with stellar video support, real-time collaboration, and Agent support. I used to be on Linux but had to switch to Mac b/c there was no good enough slide editor - @OmarchyLinux & @dhh inspired me to vibe-code my own, so I'm finally back on Linux! Install from source via GitHub (installers coming soon!): github.com/vsitzmann/deckwer…
18
38
374
38,671
Accepted at NeuriPS - meet us in Sydney!!
Introducing MilliVid, our new method for long-context video generation! MilliVid creates videos that are consistent over long time spans, without using retrieval heuristics or 3D maps! (1/n) davidcharatan.com/millivid/#
3
51
4,401
This paper has been accepted at NeurIPS! Meet the team in Sydney!
We discovered that our latest Dataset Distillation project can be used to create some beautiful synthetic images based on an artist's body of work! Come see us at the @eccvconf Art Gallery starting today! Explanation and some of my favorites in thread below: 1/ (Claude Monet)
3
19
2,200
I am open-sourcing DeckWerk, a multi-platform slide editor with stellar video support, real-time collaboration, and Agent support. I used to be on Linux but had to switch to Mac b/c there was no good enough slide editor - @OmarchyLinux & @dhh inspired me to vibe-code my own, so I'm finally back on Linux! Install from source via GitHub (installers coming soon!): github.com/vsitzmann/deckwer…
18
38
374
38,671
It's entirely vibe-coded with something like ~100 hours of my time put in (my weekend and night project) and many many more agent hours. I use it exclusively now instead of Keynote (which in the year of our lord 2026 still does not allow you to crop videos!!). There are definitely still bugs and rough edges, but I have used it in ~7 talks so far and it has not let me down, both remote and in-person - I think it's ready for prime-time :)
1
1
22
1,952
Oops the release build was 2 weeks old - if you installed it earlier, pls reinstall it to get the latest version 😅
2
978
Vincent Sitzmann retweeted
We recently came out with our blog post, Does Scaling Web-Video Pre-training Help Real Robots Do Real Work? This is a pretty exciting moment for robotics, because as far as I am aware, it is the first evidence that robot foundation models can scale on internet video data; not merely on teleoperation, UMI data, or egocentric data, but on the data you can truly find anywhere. It’s the necessary requirement for reaching massively powerful models.
At Rhoda, we care deeply about the science of pre-training for robotics. In one of the most rigorous studies of its kind, over thousands of trials and hundreds of hours of robot evaluations, we show how scaling web-video pre-training leads to better real-world robot performance.
1
4
52
6,599
Vincent Sitzmann retweeted
Stop by the #ECCV2026 Art Panel in Malmömässan C1 at 13:30 today (Saturday) to get a sneak peak at the tech behind these art distillations!
We discovered that our latest Dataset Distillation project can be used to create some beautiful synthetic images based on an artist's body of work! Come see us at the @eccvconf Art Gallery starting today! Explanation and some of my favorites in thread below: 1/ (Claude Monet)
1
4
18
2,096
Tongzhou has carefully evaluated the actual real-world scaling laws of Rhoda's web-scale video generative model - very cool study!
Does scaling pre-training on general web video improve a complex manipulation task in real deployment? We scale model size and pre-training compute, and test on one industrial task. Yes. The better a pre-trained model predicts web video, the better its post-trained policy. 🧵
1
30
4,510
Based on his pioneering work on dataset distillation, my student George has noticed that his most recent method (to be released soon, stay tuned!), besides being SOTA in dataset distillation, also creates stunning synthetic "composites" of an artist's body of work. Check it out! He was invited to display these at ECCV's art gallery and they came out amazingly :)
We discovered that our latest Dataset Distillation project can be used to create some beautiful synthetic images based on an artist's body of work! Come see us at the @eccvconf Art Gallery starting today! Explanation and some of my favorites in thread below: 1/ (Claude Monet)
1
8
130
11,188
I think this is my personal favorite :)
Replying to @GCazenavette
Thanks @elluba and @YSiglidis for organizing! You can find the gallery at the expo in Malmömässan near the t-shirt pickup. It's open to browse, and there will be guided tours at 11:00 and 17:00 on Thursday (today) and Friday and at 11:00 on Saturday. 4/ (Georgia O'Keefe)
4
1,121
I agree with Phil's take here: The progress of LLMs on controlling robots is quite interesting. Intuitively, this makes sense: controlling a robot is not so different from computer use, and an agent that is good at computer use is probably also good at controlling a robot and vice versa - credit to my student @RyuHyunwoooo to pointing that out to me originally!
Recently, there have been a lot of impressive demos of AI agents, like Claude, controlling robots. I wrote a short blog post with my thoughts on the advent of these "robot-use agents." web.mit.edu/phillipi/www/wri… I think it's an important change in the trajectory of robotics!
8
22
191
42,709
One nuance is that it wouldn't surprise me if OpenAI had added a lot of robot and computer data to their training set, so it's not clear how much of this is "emergent" behavior!
2
1
22
3,237
On latency / speed concerns: I think LLMs and video models will become really fast over the last decade. It's good to remind ourselves that this technology has not been around for long! Even if the large models may not crack the hundreds of Herz that may be required for dexterous, dynamic control, I believe that smaller models that are actively learned via interaction with the real world just might.
8
1,423
At 10:20 am Malmö time, I will be speaking at the "X-Reason" workshop in the Palisades South room (turn right at registration, then down the stairs) about whether intermediate representations are important for embodied intelligence! #ECCV2026
2
37
4,123
This looks like a really cool product, congrats to the @theworldlabs team!
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
5
1
54
6,016
I was a guest on the @MIT_CSAIL podcast to discuss robots, AI, and what the near future might look like. @klgiven did a great job steering the conversation, and I think much of our chat is accessible to non-experts - hope you find it interesting! bit.ly/4zruZdt Also on Spotify: open.spotify.com/episode/5QX…
1
7
96
5,624
Congrats - this looks super exciting, Russ & team!!
In January, I started "building something new" with an incredible team. Today I finally get to share some first details about what we've been building. We've called it Walden Robotics (waldenrobotics.com). I thought long and hard about my own reasons for starting this company. It's not only about the robots. It's also about people. I've tried to capture those thoughts in my first Walden blog post: waldenrobotics.com/news/why-… It's been an incredible ride so far. Within just a few months of forming the company, we were already operating a general-purpose robot with an end-to-end policy in production in one of the most important factories in North America. It's amazing at how much I've already learned from that experience. There is a lot of work to do, but the mission has never been so clear. Please help me welcome Walden Robotics into the world. And stay tuned for more updates! piped.video/watch?v=fewvZrck…
2
1
32
11,259
We are open-sourcing a fine-tuned video policy + pre-trained IDM! In our paper, we demonstrate that this paradigm has the potential for plug-and-play manipulation across embodiments - very exciting :)
Robot learning is moving beyond policies built for one robot, one scene, one task. At MIT, we’re exploring a different path: turning video world models into embodiment-agnostic robot policies. Introducing VERA: a 14B video-to-action system that controls robots across embodiments, skills, and environments. From zero-shot pick-and-place on a real Panda arm to contact-rich cube reorientation with a 16-DoF robotic hand. Different robots. Different environments. Different tasks. Same video planner. Same weights. We’re open-sourcing everything so you can fine-tune VERA for your own robot setup too. Deep dive in the thread: 🔗 vera.csail.mit.edu 🧵 (1/7)
1
8
127
17,763