Excited to share Do as I Do! We turn everyday human videos into physically consistent robot data that can be directly executed in the real world. This was a fun collaboration with @bhawna_paliwal_ and @willjhliang, with lots of moving parts. More details in Mahi's thread below👇
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data? Introducing Do as I Do: establishing a much needed correspondence between human videos and dexterous robot data. Some fun insights below: 🧵
4
12
32
4,789
Haritheja retweeted
One human demonstration. Any multi-fingered hand. Zero-shot sim-to-real visuomotor policy. morphometricimitation.github… Collaborators: @he_siming @ckwolfeofficial @HaozhiQ @LeaMue27 Shankar Sastry, Claire Tomlin, @JitendraMalikCV
12
20
167
30,028
Haritheja retweeted
Achieving human-like dexterity is the next frontier for robotics, and yet dexterity data is often subtly hard to scale. Real-world dexterity data, including things like finger-pose estimates, is often slightly off, making it physically invalid and hard to execute on real hardware and hard to learn from. DO AS I DO is an algorithm for reconstructing and retargeting monocular RGB videos to robot hands, outperforming the state of the art and even working from generated videos. @bhawna_paliwal_ @HarithejaE @willjhliang and @notmahi join us to tell us more. Watch Episode 107 of RoboPapers, with @chris_j_paxton and @DJiafei, now!
13
70
22,441
Haritheja retweeted
Full episode dropping soon! Geeking out with @bhawna_paliwal_ @HarithejaE @willjhliang @notmahi on Do as I Do: Dexterous Manipulation Data from Everyday Human Videos do-as-i-do.com/ Co-hosted by @chris_j_paxton @DJiafei
2
6
22
1,854
Grounded API is live. - SOTA on hand-tracking benchmarks (< 1 cm) - SOTA on SLAM benchmarks - In-the-wild ego data -> enriched data in minutes - Integration with @huggingface @LeRobotHF & @rerundotio - Built for @BitRobotNetwork RoboCap suite Technical report & more↓
48
78
447
170,574
Haritheja retweeted
I didn't get what GPT-6 meant for robotics until I actually tried it. GPT-6 Astra just does physical ICL out of the box. we drop a recording of a human doing a novel task into 𝗰𝗼𝗱𝗲𝘅 app. Prompt it to drive a robot arm the same way. It just works on the first pass!
54
143
1,481
238,679
Haritheja retweeted
Meet FetchMan: a vision-based humanoid policy trained entirely in simulation that transfers zero-shot to diverse real-world scenes and objects. Simulation has produced impressive locomotion policies that transfer to the real world. We wanted to see how far the same recipe goes for vision-based loco-manipulation. More below🧵
17
42
276
29,066
Haritheja retweeted
Excited to share SPD: simulation pre-training for dexterity. We pre-trained a policy in simulation and fine-tuned with less than 2 hours of real data (with @sarthakkamat)
17
71
537
138,183
Shout-out to Kevin for carrying on the lab tradition of live demos!! 🦾🦾
We flew our robot to Seattle for the @meta_aria Summit to prove HUG really works. Thanks to everyone who brought objects! Ft @BillyYYan building YOR; PC @irmakkguzey @mangahomanga 🌐 HUG: grasping.io 💻 Code: github.com/KevinyWu/hug 🤖 Robot: yourownrobot.ai/
10
2,371
Haritheja retweeted
Your policy doesn't need 7B params. It simply needs dense features. Introducing Patch Policy: pretrained ViT + small transformer beats OpenVLA-OFT with 0.7% of its params, and trains on a 5090. Here it inserts a cable (~2mm tol), and does it again as we unplug mid-rollout. 🧵
11
76
673
133,684
Haritheja retweeted
Pretrained ViTs see the world in rich, dense detail. Most policies pool it to a single vector before acting, discarding most of it. We introduce Patch Policy: a minimal architectural extension that enables transformer-based policies to consume dense tokens directly, no billion-param VLM required. It outperforms a fine-tuned 7B VLA by 18% with ~0.7% of its parameters, enabling robust, precise manipulation.
17
40
280
55,043
I'm at #RSS2026 – presenting MolmoSpaces on Tuesday & CAP (Contact Anchored Policies) on Wednesday. I've been thinking a lot about what robotics 2-5 years from now looks like: beyond teleop and position control. If you're interested about anything from robot free/human data or sim evals to force/torque controlled dexterous hands, let's chat! 🧵
2
9
95
11,895
Haritheja retweeted
How can generalist policies adapt to new challenges at deployment using skills they already have? We optimize VLA *prompt inputs* with reinforcement learning, enabling efficient real-robot adaptation on complex tasks where existing methods struggle. 🧵 semantic-action-rl.github.io
3
20
87
20,247
Haritheja retweeted
🤖 How can we teach dexterous robots to perform precise, contact-rich assembly? Introducing Play2Perfect: first learn to play with objects, then perfect the policy for tight insertion, multi-part assembly, and screwing. Sound on! 🔊 🧵👇
10
53
287
67,885
Excited to release Do As I Do: a pipeline that turns everyday RGB human videos into dexterous robot manipulation trajectories! Most prior work has been narrow, consisting of just lab recorded demos, egocentric-only, or assuming a closed set of objects. We develop a modular pipeline that can handle Internet, egocentric, exocentric, AND generated videos with virtually any rigid object. Also check out Mahi's post below!
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data? Introducing Do as I Do: establishing a much needed correspondence between human videos and dexterous robot data. Some fun insights below: 🧵
3
20
142
14,173
Haritheja retweeted
We can convert human videos to robot hand-object interaction trajectories in 4D. Enjoy! Paper: arxiv.org/abs/2606.19333 Website: do-as-i-do.com Code: github.com/malik-group/do-as… Authors:@bhawna_paliwal_,@HarithejaE,@willjhliang, @pabbeel , @notmahi , @JitendraMalikCV
13
85
815
73,791
Haritheja retweeted
Replying to @notmahi
@notmahi @HarithejaE solving cross-embodiment without a target embodiment - very neat!
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data? Introducing Do as I Do: establishing a much needed correspondence between human videos and dexterous robot data. Some fun insights below: 🧵
1
5
407
Haritheja retweeted
This project looks super useful for anyone doing human to robot learning research! Gave the repo a star ⭐️
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data? Introducing Do as I Do: establishing a much needed correspondence between human videos and dexterous robot data. Some fun insights below: 🧵
1
4
20
1,996
Haritheja retweeted
Introducing Do as I Do 👀, a framework to transform everyday human videos into 100s of dexterous robot demos. Co-led with @bhawna_paliwal_ and @HarithejaE, and check out @notmahi's thread! Here’s a little preview of our dexterous manipulation results. More about how we produce them from human reconstructions in this mini-thread! 🧵 nitter.net/notmahi/status/2067640…
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data? Introducing Do as I Do: establishing a much needed correspondence between human videos and dexterous robot data. Some fun insights below: 🧵
8
36
225
28,709
Haritheja retweeted
Enabling learning motion directly from videos rather than using them for action supervision is a superior method and likely more scalable. While it is early this line of work suggests replicating the playbook that made robots walk. --> Real videos provide state supervision (not action) --> retargeting provides reference trajectories. --> RL tracks these trajecotries. This is a very good example of the separation of the "What" and the "How"
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data? Introducing Do as I Do: establishing a much needed correspondence between human videos and dexterous robot data. Some fun insights below: 🧵
7
10
51
10,749