Building at the frontier of embodied intelligence.

Palo Alto, California
Pinned Tweet
To bring generalist intelligent robots to the real world, we have to overcome the data scarcity problem. At Rhoda, we are solving it by reformulating robot policies as video generation. Today, we introduce the Direct Video-Action Model (DVA)
19
38
223
75,811
At Rhoda, we care deeply about the science of pre-training for robotics. In one of the most rigorous studies of its kind, over thousands of trials and hundreds of hours of robot evaluations, we show how scaling web-video pre-training leads to better real-world robot performance.
13
30
219
45,573
The result is clear: better pre-training produces better robot policies. And the gains aren’t limited to headline performance. Better pre-training also makes policies more robust when robot post-training data is sparse.
1
1
14
3,217
At Rhoda, we tackle real-world problems through fundamental research. Interested in building the future of robotics? Feel free to apply via our careers page below. rhoda.ai/careers
12
1,219
Today marks week one of our new blog series. First one's from @tongzhou_mu . Will scaling web-video pre-training lead to better robot policies?
Does scaling pre-training on general web video improve a complex manipulation task in real deployment? We scale model size and pre-training compute, and test on one industrial task. Yes. The better a pre-trained model predicts web video, the better its post-trained policy. 🧵
2
7
42
4,692
Rhoda AI retweeted
i made this right after our launch @RhodaAI with some leftover b-roll - watch with sound on 🔉! we've been silent for a while, but we have many exciting things to share in the future 👀
8
5
73
15,799
Rhoda AI retweeted
Nobody in hard tech has a stronger track record of turning frontier technology into something legacy industries will buy than Jagdeep Singh (@startupjag) That's why the CEO and co-founder of @RhodaAI is our next Dirty Jobs speaker. Infinera (acq. Nokia for $2.3B) QuantumScape (public) Now Rhoda: robots that learn a task from just 10 to 20 hours of data while competitors cite 70K hours and up. Telecom carriers. Automakers. Factory operators. These are three industries with long sales cycles, expensive failure modes, and little appetite for unproven technology. Jagdeep has spent his career getting them all to say yes. In just 30 minutes on stage, he'll share how you take hard tech from concept to deployable reality, and what it takes to get bought by industries that hate risk. One day. September 23rd in San Francisco. 300 founders, operators, and partner-level investors in physical AI. Want to join us? Link to apply in the comments
2
7
22
2,366
Rhoda AI retweeted
Rhoda AI is betting the future of robotics starts with internet video, not lab data. We visited the company’s headquarters to see how its Direct Video Action model learns physics from hundreds of millions of clips.
6
22
17,179
Rhoda AI retweeted
🙏 We’re incredibly grateful to everyone who joined our @RhodaAI party last night at @CVPR. The turnout exceeded anything we expected, and it was a pleasure meeting so many researchers and builders from the vision / robotics community. Thank you for all the great conversations!
1
2
45
4,406
Rhoda AI retweeted
I'll be at CVPR next week (6/3–6/7). If you’re working on or exploring opportunities in video models for robotics (research or engineering), happy to chat 🤖 We’re also hosting a Rhoda party Thursday night with many of our technical team in town. DM me for an invite 🍻
14
5
58
8,120
Can a large foundation video model run as a real-time robot policy at the edge, on a single RTX 5090? • ✅ No quantization • ✅ No distillation • ✅ Full denoising (all the way from noise to clean video) We just proved it's possible. 👇🎬
12
32
220
39,849
How? Existing video models aren't optimized for real-time inference. Instead of fine-tuning off-the-shelf video models, we co-design inference-aware model architectures and model-aware inference optimizations from the ground up.
1
19
2,821
Teaching a robot a new task typically means stopping operations, collecting teleoperated demonstrations, and retraining. That process takes hours at a minimum. We wanted to know if we could collapse it to seconds — from a single human demo, on the fly, no retraining required. Early research preview: we can.
9
14
91
8,826
How it works: we train on paired human demo and robot execution data. Because our DVA, FutureVision, has long-context visual memory built in (nitter.net/RhodaAI/status/2042302…), we prepend the full human video into the model's context and predict robot actions closed-loop. The model watches a human do something once and understands what to do next.
Here’s something we’ve never seen done before. Real-world tasks are long and ambiguous. Solving them requires visual memory and state tracking. Most robot policies only see the last few frames. Ours doesn't. We put our DVA, FutureVision, to the perfect testbed: the shell game 🐚. The DVA nails it.
1
6
2,211
Here’s something we’ve never seen done before. Real-world tasks are long and ambiguous. Solving them requires visual memory and state tracking. Most robot policies only see the last few frames. Ours doesn't. We put our DVA, FutureVision, to the perfect testbed: the shell game 🐚. The DVA nails it.
8
37
239
91,606
How? Our DVA implements robot policy as future video generation. Given the context, the model generates future videos (bottom left) predicting not just the correct cup to pick up, but even the appearance of the hidden object. Native training on long, continuous videos gives the model built-in long-context memory.
1
10
2,496