Combining real-time interactivity, task understanding, and full-body action prediction on a humanoid is so, so hard. Here's an example where we bring all of these together in Gemini Robotics 2 🤖🧠
14
31
243
41,568
Combining real-time interactivity, task understanding, and full-body action prediction on a humanoid is so, so hard. Here's an example where we bring all of these together in Gemini Robotics 2 🤖🧠
14
31
243
41,568
Under the hood, this requires: - Action model that coordinates whole-body movement and fine motor skills simultaneously, while following instructions and recovering on-the-fly - Embodied reasoning (ER) agent with SOTA video understanding, long-horizon reasoning about user intent, and interruption handling
1
2
8
1,285
This was a huge team effort @GoogleDeepMind! hard to fathom the amount of work that goes into making full systems like this, but it's incredibly rewarding when things come together, and it's so fun to see how responsive the models are :)
1
5
675
Excited to share Gemini Robotics 2! 6 months ago, I wrapped up my PhD at Stanford to join GDMR full-time. Since then, I’ve learned so much from the team, while working on video understanding for our ER agent and improving the robustness & instruction-following of our action models. Still lots of work ahead, but excited about our progress towards a general robot brain! ✨
One brain. For any robot. 🤖 We’re launching Gemini Robotics 2: our next-generation physical AI bringing full body intelligence to humanoids, advanced dexterity, multi-robot teamwork and more.
1
3
52
3,571
Annie Chen retweeted
Introducing Gemini Robotics ER 2, our latest robotics embodied reasoning model based on Gemini. So much progress from our last ER model (and baseline Gemini 3.6 Flash). Very excited to see more robots do useful things with this model!
122
172
2,125
177,722
Annie Chen retweeted
For decades, we’ve dreamed of robots that can seamlessly step into our world and lend a hand. Today, we take a major stride toward making that dream a reality: Introducing Gemini Robotics 2 from @GoogleDeepMind, the intelligence layer powering the next generation of truly adaptable robots. This major advance unlocks intelligent whole-body control, advanced dexterity, and even multi-robot collaboration 🤯. Ok but... how does a robot actually "think"? Real-world tasks take time and planning. To manage that complexity, our new embodied reasoning model, Gemini Robotics ER 2, acts as the robot’s high-level brain, enhancing the robot’s capabilities to: — Observe the environment — Reason about the actions needed to complete the task — Coordinate with the vision-language-action model to carry out actions — Track progress until the job is done This setup allows robots to execute complex multi-step workflows, self-correct if a step fails, and adapt to completely novel situations. Learn more about Gemini Robotics ER 2 (and our two other brand new models) here: goo.gle/4x4E8q6
313
781
5,248
785,562
Annie Chen retweeted
75
203
2,250
391,885
New work led by @riadoshi21 on training a single VLA policy for multi-robot collaboration Excited about all the new kinds of tasks this can unlock, when robots can coordinate and work together!
🤔 Can we train one VLA policy to control multi-robot teams without any explicit communication? ✨ Introducing CHORUS: a single policy for decentralized, multi-embodiment collaboration 🧵⬇️
1
2
9
1,186
Annie Chen retweeted
We’re rolling out an upgrade designed to help robots reason about the physical world. 🤖 Gemini Robotics-ER 1.6 has significantly better visual and spatial understanding in order to plan and complete more useful tasks. Here’s why this is important 🧵
139
412
2,568
558,176
And it's even better in-person! Got to see Memo live a few weeks ago and it's such a great design :) Love the gloves, seems to enable a scalable path to high quality data. Huge congrats to the team, especially @tonyzzhao so impressed by your resilience over the past 2 years!
Today, we present a step-change in robotic AI @sundayrobotics. Introducing ACT-1: A frontier robot foundation model trained on zero robot data. - Ultra long-horizon tasks - Zero-shot generalization - Advanced dexterity 🧵->
6
10
168
22,505
How should an RL agent leverage expert data to improve sample efficiency? Imitation losses can overly constrain an RL policy. In RL via Implicit Imitation Guidance, we show how to use expert data to guide more efficient *exploration*, avoiding pitfalls of imitation-augmented RL
6
37
229
27,188
DGN achieves up to 2-3x performance improvement over existing methods that do RL with expert data.
1
8
1,004
Paper link: arxiv.org/pdf/2506.07505 Super fun project co-led with @perryadong @AlecLessing and with @chelseabfinn. Excited to further push sample efficiency for RL with better exploration priors!
10
871
🚨 We’re thrilled to announce our ICCV 2025 Workshop: MMRAgI – Multi-Modal Reasoning for Agentic Intelligence! 🚨 🌐 Homepage: agent-intelligence.github.io… 📥 Submit: openreview.net/group?id=thec… 🗓️ Submission Deadline (Proceeding Track): June 24th 2025 23:59 AoE 🗓️ Submission Deadline (Non-Proceeding Track): July 24th 2025 23:59 AoE AI Agents are evolving fast — but true intelligence needs reasoning across modalities. Vision, language, audio… it’s time to unify them. From digital and virtual agents to wearable and physical embodiments, agentic intelligence is reshaping how AI interacts with the world. As agents increasingly engage in 3D perception and geo-centric reasoning, bridging modalities with spatial understanding is more critical than ever. 💡 Join us to explore the frontiers of multi-modal agents: • Reasoning with MFM-powered agents • Applications in OS copilots, Scientific Agents, Digital Agents, Virtual Agents, Wearable Agetns and Embodied Agents! • Challenges in alignment, evaluation, efficiency, and robustness 📝 Call for Papers is now OPEN! 📅 Workshop: Oct 19–20 2025 Whether you work on models, methods, or applications — we want to hear from you! #ICCV2025 #MMRAgI #MultimodalAI #AIagents #LLM #MFM #EmbodiedAI #3DVision
9
23
14,658
Annie Chen retweeted
How can robots problem solve in novel environments? We combine high-level reasoning with VLMs with low-level controllers to allow test-time problem solving. Paper & code: anniesch.github.io/vlm-pc/
How can robots autonomously handle ambiguous situations that require commonsense reasoning? *VLM-PC* provides adaptive high-level planning, so robots can get unstuck by exploring multiple strategies. Paper: anniesch.github.io/vlm-pc/
7
27
151
20,341
How can robots autonomously handle ambiguous situations that require commonsense reasoning? *VLM-PC* provides adaptive high-level planning, so robots can get unstuck by exploring multiple strategies. Paper: anniesch.github.io/vlm-pc/
1
18
92
24,333
Leveraging VLMs in this way allows a robot to handle (zero-shot!) a wide range of complex real-world situations that wide range of complex scenarios that would otherwise require environment-specific engineering or human guidance
1
4
423
Thanks to wonderful collaborators @AlecLessing, @tangerinecoder, Govind Chada, @smithlaura1028, @svlevine, @chelseabfinn! I’ll be presenting this on Thursday at #ICRA2025 in Atlanta! Let me know if you’re around :)
4
397