Pierre Sermanet retweeted
At @UMA_Robots , we are convinced that latent-space world models will unlock scaling for robotics 1 position left for the roles 🧪 Research Scientist: app.dover.com/apply/UMA/e223… 🚀 Research Engineer: app.dover.com/apply/UMA/1bb6…
5
10
120
16,175
The SPRIND grant will accelerate UMA’s effort to develop latent world models for humanoid robots. Throughout my career I’ve kept the objective of scaling up learning using minimal supervision and efficient data collection. This led me to propose self-supervision methods in latent space such as Time-Contrastive Networks, which demonstrated 3rd person imitation of a human entirely without labels: sermanet.github.io/imitate/ One key component was co-training different modalities, viewpoints and embodiments into a single latent space, entirely from raw unlabeled videos. The dimensions of the world discovered by the models ended up naturally aligning in latent space. This latent alignment across viewpoints and embodiments is what unlocked imitating humans directly from 3rd-person raw pixels, similarly to the concept of “mirror neurons” firing when observing someone else performing a task. We later developed self-supervised methods with @coreylynch for generating actions by pulling together in latent space the representations of (start, goal) image pairs and sequence of images [start, …, goal]. We combined this label-free learning method with an efficient data collection method: playing. Play is an efficient way to discover the world by leveraging existing knowledge and skills. Humans and animals use it to learn about the world and practice in advance. We called this Learning Latent Plans from Play: learning-from-play.github.io… We subsequently augmented this approach with language representations so that you could control the latent space via language conditioning. But most of the learning was coming from self-supervising on video, language labels accounted for less than 1% of data: sermanet.github.io/language-… I still believe that learning the world in latent space by mostly self-supervising on raw unlabeled data is the way to go. We now have stronger unsupervised learning methods, more capable hardware and better data acquisition means, it is an exciting time to keep pushing that vision. If that vision speaks to you, we are hiring for our Embodied World Models team: Research Scientist: app.dover.com/apply/3a4a12c9… Research Engineer: app.dover.com/apply/3a4a12c9…
Join our team to work on Embodied World-Models with @MustafaShukor1 (co-author VL-JEPA with @ylecun ) Link to apply in comments 😇
3
7
36
9,977
At UMA, we build and own the full stack, from hardware to software. This allows us to bake safety in at all levels, rather than bolt it on as an afterthought. Because trust is everything when robots share our world. Reach out if you want to be part of this journey.
3
21
68
18,786
Pierre Sermanet retweeted
Excited by what we’ve achieved in the last 9 months. We went from a model that could merely pick objects to one that can run autonomously for hours on challenging, high-precision tasks.
AI developed at @UMA_Robots Single neural network, End to End It operates robustly for hours, with precision, adapting to what it sees.
8
27
2,780
In house embodied AI designed for robust and continuous operation.
AI developed at @UMA_Robots Picking and Scanning items Deformable, Slippery, Thin, Fragile We mess with it at the end :)
8
32
5,713
Pierre Sermanet retweeted
Starting with the fundamentals Prototype Version 0 AI, Software, Hardware A small team, 9 months Designed and assembled in Paris at @UMA_Robots
87
121
884
244,377
At UMA we're pushing the boundaries of embodied AI with our own humanoid robot.
Unveiling our Northstar design @UMA_Robots A body that can navigate our space Hands that can use our tools Approachable, competent, calm A robot you feel at ease having at your workplace and home
1
5
30
2,344
Pierre Sermanet retweeted
Tl;dr After 10+ years pushing the boundaries of RL at @GoogleDeepMind , @Google Brain and then Gemini, I’m looping back to startup mode and joining @UMA_Robots to push the frontier of human-centric robotics. (1/4)🧵
17
32
275
24,684
Pierre Sermanet retweeted
One of the first AgiBot X2 humanoid robots just landed in our lab 🤖 I like its articulated neck, back handle, battery location, anthropomorphic wrist and white soft cover in a sort of Polyurethane foam. Impressive delivery speed and support from the @AGIBOTofficial team!
23
55
321
34,260
Pierre Sermanet retweeted
I recently moved from @Tesla_AI to @UMA_Robots. 🧵(1/5)
Made with AI
23
14
100
33,946
I’m really happy to share that we’re launching UMA. Together with @RemiCadene, @alibert_s, @therobotstudio, and an exceptional founding team, we’re building general-purpose mobile and humanoid robots. If you want to be part of this adventure, reach out at uma.bot Throughout my career, I have been obsessed with scalable learning and data acquisition methods that require little to no labels. Back in 2005 with @ylecun, we were self-supervising our “deep” 2-layer network to do long range vision using short range stereo information, this was running live onboard our robot. However, because our deep model was so slow, the robot would crash constantly, so I designed a decoupled fast & far architecture for robust navigation, allowing fast control to coexist with slow long horizon thinking, much like systems 1 & 2 in modern humanoids. My PhD was focused on making deep learning work for computer vision, including unsupervised feature learning with @koraykv, writing and open-sourcing a C++ deep learning library with @soumithchintala, and open-sourcing one of the first deep learning vision systems. I came back towards robotics at @Google Brain and @GoogleDeepMind, where I pushed for entirely label-free methods on real robots. In 2017, @coreylynch and I managed to make our robot imitate human motion by co-training self-supervision across sim and real domains jointly, without any labels. With @imkelvinxu and @svlevine , we showed that unsupervised visual reward learning could be used for RL in the real world. In 2020, Corey and I developed the first manipulation VLA, which was trained with very few language labels thanks to self-supervision on play data (playing is an efficient way to demonstrate and practice a broad set of skills and is essential for human development). I was never satisfied with the status quo of top-down data collection, where researchers decide a few tasks to collect data on. Instead, I believed that we should let the data speak: tasks should be automatically discovered bottom-up (scalable and general) from cheap and continuous data collection, with a sprinkle of more expensive data and labels. In 2022, I explored long-horizon reasoning for robotics using scalable automatic labeling augmentations for VQA tasks and studied the economics of different data collection schemes. Most recently, I developed approaches to scalably discover laws of robotics from real data (images, hospital reports, sci-fi literature) in a broad and bottom-up fashion, which improved robot behavior over top-down approaches like Asimov’s laws. All these experiences nourished my vision for UMA as Chief Scientist, I’m incredibly excited to put everything together and so grateful I get to contribute to this incredible moment in human history. Picture: Yann supporting UMA as an advisor and investor, with the team in Paris a couple weeks ago.
44
53
798
188,355
After 11 incredible years at @Google Brain and @GoogleDeepMind, I’m turning the page to start something new in robotics (currently in stealth 👀). I joined Brain in 2014, drawn by the momentum around robotics and machine learning, and I’ll always be grateful to @V_Vanhoucke for creating such a unique space for robotics research and for bringing together some of the brightest minds I’ve ever met. During my time there, I had the joy of exploring ideas that once felt like science fiction: self-supervised imitation from video, unsupervised visual reward models, the first vision-language-action model for manipulation, long-horizon reasoning, and even using sci-fi and Asimov’s laws to improve robot behavior. Huge thanks to my managers Vincent, Anelia, Kanishka, and to all my amazing collaborators: you made this journey unforgettable. After 20 years in robotics, it’s clearer than ever that the field is maturing fast and I feel lucky to keep building through this moment. More soon 🤖
28
15
460
88,563
Great to see more robotics projects in Europe🇪🇺
I am starting a venture on top of LeRobot! We’re at a pivotal time. AI is moving beyond digital to the physical world. Embodied AI will change our surroundings in ways we can barely imagine. This technology holds the potential to empower everyone. It must not be controlled by just a few. This conviction led me to propose an ambitious open-source AI robotics project to Thom, Clem, and Julien back in 2024. Hugging Face, home to a community of millions of AI builders and a team of experts who brought us transformers, datasets, and the Hugging Face Hub, was the perfect place to launch LeRobot. I’m incredibly grateful for all the support that allowed me to build LeRobot alongside an amazing team and community. In such a short time, we built one of the most adopted open-source robotics platforms, used by startups, universities, and research labs. It is helping countless people take their first steps in robotics. Together, we’ve even assembled the world’s largest open robotics dataset. And this is only the beginning for LeRobot! Building on this momentum, I now feel the urgency to start something new on top of LeRobot. It will push the limit of what robots are capable of and commoditize them within society. Like LeRobot, it will start in Paris, leveraging its vibrant international AI scene. Stay tuned! As LeRobot continues to expand, it’s now in the best possible hands with @AractingiMichel, @pepijn2233 and Steven Palma taking the lead. Watching the team deliver exceptional results over the last weeks has been one of the most rewarding experiences. Their creativity, dedication, and capability to ship fast is proving just how strong the team is today! I am extremely grateful to the many people who contributed to making LeRobot at Hugging Face and within its powerful community. Many thanks to Thom, Clem, Julien, Simon, Rob, Michel, Pepijn, Steven, Gloria, Adil, Martino, Caroline, Marine, Mishig, Guillaume, Pablo, Lysandre, Arthur, Quentin, Florent, Brigitte, Victor, Marina, Mustafa, Francesco, Jess, Jade, Ville, Leo, Max, Julien, Alexander, Flavien, Raphael, Adina, Tao, Dana, Batu, Olivier, Matthieu, Eugene, Theo, Guilherme, Hynek, Loubna, Clémentine, Merve, Vaibhav, Anna, Jeff, Adrien, Emily, Johanne, Adrien and others. There are too many of you to be all named! Thanks again and see you soon!!! :) ~ Remi
4
29
8,848
Pierre Sermanet retweeted
I am starting a venture on top of LeRobot! We’re at a pivotal time. AI is moving beyond digital to the physical world. Embodied AI will change our surroundings in ways we can barely imagine. This technology holds the potential to empower everyone. It must not be controlled by just a few. This conviction led me to propose an ambitious open-source AI robotics project to Thom, Clem, and Julien back in 2024. Hugging Face, home to a community of millions of AI builders and a team of experts who brought us transformers, datasets, and the Hugging Face Hub, was the perfect place to launch LeRobot. I’m incredibly grateful for all the support that allowed me to build LeRobot alongside an amazing team and community. In such a short time, we built one of the most adopted open-source robotics platforms, used by startups, universities, and research labs. It is helping countless people take their first steps in robotics. Together, we’ve even assembled the world’s largest open robotics dataset. And this is only the beginning for LeRobot! Building on this momentum, I now feel the urgency to start something new on top of LeRobot. It will push the limit of what robots are capable of and commoditize them within society. Like LeRobot, it will start in Paris, leveraging its vibrant international AI scene. Stay tuned! As LeRobot continues to expand, it’s now in the best possible hands with @AractingiMichel, @pepijn2233 and Steven Palma taking the lead. Watching the team deliver exceptional results over the last weeks has been one of the most rewarding experiences. Their creativity, dedication, and capability to ship fast is proving just how strong the team is today! I am extremely grateful to the many people who contributed to making LeRobot at Hugging Face and within its powerful community. Many thanks to Thom, Clem, Julien, Simon, Rob, Michel, Pepijn, Steven, Gloria, Adil, Martino, Caroline, Marine, Mishig, Guillaume, Pablo, Lysandre, Arthur, Quentin, Florent, Brigitte, Victor, Marina, Mustafa, Francesco, Jess, Jade, Ville, Leo, Max, Julien, Alexander, Flavien, Raphael, Adina, Tao, Dana, Batu, Olivier, Matthieu, Eugene, Theo, Guilherme, Hynek, Loubna, Clémentine, Merve, Vaibhav, Anna, Jeff, Adrien, Emily, Johanne, Adrien and others. There are too many of you to be all named! Thanks again and see you soon!!! :) ~ Remi
99
84
759
124,309
So how would AI-powered robots behave if we dropped them into Science Fiction literature? And can we generate useful robot constitutions 📜 from Sci-Fi? We generate the first Sci-Fi-inspired constitutions and large-scale ethics benchmark scifi-benchmark.github.io @GoogleDeepMind
1
2
11
3,700
We are grateful to Sci-Fi authors for the potential positive real-world impact as we find that Sci-Fi-inspired robot constitutions yield some of the most aligned behavior when measured against realistic scenarios in the ASIMOV benchmark asimov-benchmark.github.io
1
1
918
While the findings are positive, it does not mean that AI cannot be manipulated for negative outcomes. By releasing this benchmark and dataset, we hope to further safety research and ethical deployment of AIs and robots. Authors: @psermanet, @Majumdar_Ani, @vikassindhwani
2
827