Official account for OpenDriveLab @hkuniversity and Beyond. We do cutting-edge research in Robotics, Autonomous Driving. Email: contact@opendrivelab.com

Hong Kong
Sydney, come sundown. 🌃 OpenDriveLab x @archon_robotics are throwing an afterparty! TopTalents for Embodied AI, a social mixer at W Sydney, harbour views, golden hour, and the people building the next gen of Embodied Intelligence. 📅 Jul 14 · 7–10PM 📍 W Sydney 🎟️ Apply: luma.com/0ml7kuwb #RSS2026 #EmbodiedAI #Sydney #WorldModel #FoundationModel
1
4
7
2,450
🚀 Join us at the #CVPR 2026 Workshop: From Labs to Life: Embodied Intelligence in the Wild (opendrivelab.com/cvpr2026/wo…) As embodied AI moves into the real world, we ask: how can agents perceive, reason, and act reliably in the wild? Featuring invited talks from: Hao Su, Zhiyu Huang, Jiahui Lei, Yilun Du, Rika Antonova, Jiatao Gu. 📅 9:00 AM - 5:30 PM, June 3, 2026 📍 Four Seasons 1, Colorado Convention Center. #CVPR #EmbodiedAI #PhysicalAI #Robotics
1
5
673
📢📢📢 Call for Contributions @ RSS 2026 Towards Robust Execution of Long-Horizon Whole-Body Control Tasks 🧑‍💻👩‍💻🧑‍💻 Speakers: - Javier Alonso-Mora (TU Delft) - Leslie Pack Kaelbling (MIT) - Shan Luo (King's College London) - Hamidreza Kasaei (University of Groningen) - Roberto Martín-Martín (UT Austin) - Fan Shi (NUS) 📝📝📝 Call for Contributions: We invite researchers to share their work with the community through submissions to the workshop in a variety of formats beyond traditional papers, including reports, demos, video, and etc. Submissions may include research papers or reports, but we equally welcome alternative formats such as videos demonstrating systems in action, demos, interactive artifacts, or other creative presentations of research ideas. We particularly welcome ongoing, preliminary, or exploratory work. 🌍🌍🌍 Website: opendrivelab.com/rss2026/wor… #RSS2026 #AI #Embodied
2
4
772
🚀 MM-Hand 1.0 Tech Report released: a 21-DoF multi-modal modular dexterous robotic hand with remote tendon-driven actuation. Motors are relocated outside the hand, freeing space for tactile sensors, joint encoders, in-palm stereo vision, and maintainable modular fingers. MM-Hand achieves 25N fingertip force through 1m tendon-sheath transmission and supports closed-loop joint control for dexterous manipulation research. Paper: arxiv.org/abs/2604.17245 Page: opendrivelab.com/MM-Hand #Robotics #DexterousHand #EmbodiedAI
2
34
273
16,187
【4/5】🚘Production-validation
1
1
176
【3/5】📺Visualization ① Our Behaviour World Model seeds from real-world cases to synthesize what your dataset is missing — cut-ins, sudden braking, edge-case interactions at scale. ② Post-training with World Engine: before vs. after. ③/④ The post-trained E2E model deployed in WorldEngine.
1
1
130
【2/5】🎉Highlights - Data-driven long-tail discovery: Failure-prone scenarios are automatically identified from real-world driving logs by the pre-trained agent itself — no manual design, no synthetic perturbations. Photorealistic interactive simulation via 3DGS: Each discovered scenario is reconstructed into a fully controllable, real-time-renderable simulation environment with independent dynamic agent manipulation. - Behavior-driven scenario generation: Leverages Behavior World Model (BWM) to generalize and synthesize diverse traffic variations from existing long-tail scenarios, expanding sparse safety-critical events into a dense, learnable distribution. - RL-based post-training on synthesized safety-critical rollouts substantially outperforms scaling pre-training data alone — competitive with a ~10× increase in pre-training data. - Production-scale validation: Deployed on a mass-produced ADAS platform trained on 80,000+ hours of real-world driving logs, reducing simulated collision rate by up to 45.5% and achieving zero disengagements in a 200 km on-road test.
1
1
162
UMI made robot data collection intuitive. 🤖 TAMEn takes it further — bringing vision and touch into a unified, closed-loop learning system. 🌏 opendrivelab.com/TAMEn 📑 arxiv.org/abs/2604.07335 🔗 github.com/OpenDriveLab/TAME… ✨ What’s new: - Dual-mode data collection (MoCap ↔ VR) - Online replayability check - AR-in-the-loop + real-time tactile feedback (tAmeR) - Self-evolving pyramid data pipeline 🚀 Results: 34% → 75% success rate on bimanual tasks. This marks a shift from usable data to self-improving data engines. TAMEn turns robots from "blind operators" into tactile-aware, evolving collaborators. #EmbodiedAI #Tactile #Robotics #Bimanual #TAMEn
20
95
9,314
🚀 Introducing SMASH — the world’s first high-dynamic humanoid robot for outdoor table tennis, fully autonomous with onboard perception. No motion capture needed—SMASH uses only onboard sensors for stable, full-body human–robot interaction in real-world settings. Embodied intelligence is stepping out of the lab. 🔗mmlab.hk/Smash/ #HumanoidRobots #TableTennis #RobotLearning #SportsTech #AI #OpenDriveLab #Algorithms
5
21
79
20,959
Proud to announce the strategic partnerships with 3 Leading Embodied AI companies. Together with Unitree, Noitom Robotics, and BrainCo, HKU’s Embodied Intelligence Joint Lab is live at our Zhangjiang base. We’re in this for the long game: turning embodied intelligence into a durable, shared foundation for what comes next. Builders: let’s collaborate. Press Release: hku.hk/press/news_detail_289… @HKUniversity @hkudatascience @HKU_CDS @YiMaTweets @francislee2020 @UnitreeRobotics @noitomrobotics @BrainCo_Tech #EmbodiedAI #Robotics
3
17
3,508
📣 #Recognition #ResearchAward As we usher in the Year of Horse, our team recognizes outstanding members from the past year in the exceptional contribution of areas. Congratulations!!! Let’s rock ‘n’ roll in 2026 🕶️🍾🎆🥂#opendrivelab
2
2
10
1,567
【4/5】The ultimate test? Tai Ping Shan. 🏔️ Steep slopes, long horizons, pitch-black night - all in one. SparseVideoNav handles it like a champ. 🔥
1
2
157
【3/5】💡To combat sim2real gap, the largest real world VLN dataset has been collected. We will release the whole dataset to benefit the community !!! (Estimate 2026 Q3)
1
2
81
【2/5】Through four stages of training combining with our sparse design, SparseVideoNav can now perform inference within 1 second - and predict 20 seconds of sparse future! 🚀 - Sparsification: 1.4x training speedup & 1.7x inference speedup - Stage1 T2V→I2V: 2x training speedup - Stage2 History compression: 1.4x inference speedup - Stage3 Distillation: 9.6x inference speedup - Stage4 Action Learning
1
2
108
【1/5】❓Why do Vision-Language Navigation (VLN) agents need "step-by-step" language instructions? No.❌ We only want to give a intent. 📍 Meet SparseVideoNav: a new paradigm that uses video generation to navigate like a human. Tell the destination. Imagine the path. Reach the goal. 🌄🤖 Paper: arxiv.org/abs/2602.05827 Page:opendrivelab.com/SparseVideo… Github:github.com/OpenDriveLab/Spar…
1
9
30
2,365
🦾 【MM-Hand 1.0】Open-Source, Lightweight, High-DoF, Multimodal, Modular Design for Easy Disassembly and Modification 🔗 Details: mmlab.hk/research/MM-Hand 🎁 Beta Program: forms.gle/QhaHGCigY6buuSNfA 📮 Inquiries: research@mmlab.hk #MMHand #DexterousHand #Robotics #OpenSource #EmbodiedAI #TactileSensing
2
11
50
4,816
[3/5] Problem: Expensive Iteration Collect new data → Retrain everything → Repeat Slow yet expensive. How? Model Arithmetic: • Train only on new data • Merge via weight interpolation • Merged model > full-dataset model Models trained separately preserve distinct modes.
1
1
391
[2/5] Problem: Distribution Mismatch Training data ≠ Model behavior ≠ Real-world execution This gap causes failures. Solution → Mode Consistency: • DAgger for failure recovery • Augmentation for coverage • Inference smoothing for clean execution
1
499
🧥 Live-stream robotic teamwork that folds clothes. 6 clothes in 3 minutes straight. χ₀ = 20hrs data + 8 A100s + 3 key insights: - Mode Consistency: align your distributions - Model Arithmetic: merge, don't retrain - Stage Advantage: pivot wisely 🔗 mmlab.hk/research/kai0
2
7
29
15,271
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control @ HKU | AgiBot | FDU | SII page:opendrivelab.com/WholeBodyVL… arXiv:arxiv.org/pdf/2512.11047 🔥WholeBodyVLA enables large-space, long-horizon humanoid loco-manipulation!
4
509
5/n 🏆 #1 on the NAVSIM official leaderboard huggingface.co/spaces/AGC202…
1
4
195
4/n 🌍Visualization of how we simulate diverse scenes from a single real-world scenario.
1
4
274
3/n 🦊🐰My favorite comic presents SimScale in a Zootopia-like style, created with NotebookLM.@NotebookLM @NanoBanana
1
6
340
2/n 🔥 SimScale Highlights 🏗️ Scalable Real-World Simulation Pipeline Powered by high-fidelity neural rendering and reactive environments, SimScale adopts disturb-and-plan strategy to generate pseudo-expert demonstrations—producing diverse driving scenarios directly usable for training at scale. 🚀 Efficient Sim-Real Co-training Strategy Compatible with any end-to-end policy. Supports expert-based or reward-driven training. Mixed sampling fully leverages simulation while preserving faithful alignment to human driving behavior. 🏅 Boosting Any End-to-End Policy Regression, diffusion, scoring-based—SimScale improves robustness and generalization of policies across all architectures. With multi-expert ensemble, it reaches #1 on the NAVSIM official leaderboard 🏆. 🔬 First Time!Revealing Scaling Effect of Simulation 💡Crucial findings: exploratory pseudo-experts are more efficient, multimodal modeling sparks data expansion potential, reward-driven learning trains effectively, and sustained simulation gains across real data scales.Simulation
1
4
508
1/n 🎉 SimScale: Learning to Drive via Real-World Simulation at Scale 🤖 An innovative real-world simulation pipeline and a real-sim co-training strategy that significantly boost the robustness and generalization of any end-to-end planner. 📈 For the first time, we reveal the scaling effect of simulation data in autonomous driving: with zero extra real-world data, simply scaling up simulation alone can keep improving model performance!
1
6
29
11,530
🚀 New Dataset Alert! 🚀 "FreeTacMan: Robot-free Visuo-Tactile Data Collection System for Contact-rich Manipulation" We leverage FreeTacMan🤖 to collect a large-scale multimodal dataset✨, comprising over 3000k paired visual–tactile images with end-effector poses, 10k demonstration trajectories across 50 diverse contact-rich manipulation tasks. 💡 Breakthrough Moments: ✅ Data-collection System: An in-situ, robot-free, real-time tactile data-collection system to excel at diverse tasks efficiently. ✅ Visuo-Tactile Dataset: A large-scale, high-precision visuo-tactile manipulation dataset with over 3000k visuo-tactile image pairs, more than 10k trajectories across 50 tasks. ✅ Policy Learning Enhanced by Tactile Pretraining: Imitation policies trained with our visuo-tactile data outperform vision-only baselines by 50% on average. 📊 Dataset: huggingface.co/datasets/Open… 🚀 Website: opendrivelab.com/freetacman 📜 Paper: arxiv.org/abs/2506.01941 📺 Video: opendrivelab.github.io/FreeT… 🛠️ Hardware Guide: Hardware Guide 💻 Official Repo: github.com/OpenDriveLab/Free… #Robotics #Dataset #DataCollection #Tactile #VisuoTactile #FreeTacMan
1
11
1,802
CCAI 9025 Field Trip Day 2 just redefined our understanding of #EmbodiedAI. Our Shenzhen tech expedition took us deep into Astribot @Astribot_Inc and SLAI (Shenzhen Loop Area Institute). 🙌 Highlight: symposium where experts shared next-gen insights that'll resonate for months. Truly transformative discussions.🦾🦿 Huge kudos to the TA Group for orchestrating such an inspiring event! You absolutely crushed it! 🚀 #HKU #hkudatascience #CCAI9025 #HKUCommonCore @hkudatascience
【🚀 Field Trip Alert! 🤖】 Get ready for an exciting 2-day journey with HKU IDS & Team CCAI9025, led by Prof Hongyang Li! Explore cutting-edge robotics in Hong Kong and Shenzhen with expert-led tours and discussions. 🤖】 👉Details: mmlab.hk/CCAI9025 Don’t miss out!
6
5
1,846
🚀 Join us at  #ICCV2025 for a full-day workshop: “Learning to See: Advancing Spatial Understanding for Embodied Intelligence” 🗓️ October 19 • 📷 Room 312 Meet our incredible lineup of speakers: @MattNiessner @jiadeng @pulkitology @KaterinaFragiad @YunzhuLiYZ @imankitgoyal Kun Zhan from @Li_Auto_ More details and the full schedule are available at: opendrivelab.com/iccv2025/wo…
1
8
37
5,555
Connecting theory with practice! 🧠 That's a wrap on our insightful field trip! 🤖 Our students had an amazing day with talks at Huawei HK, a tour of the Hong Kong Science Park, and a lab visit and panel discussion at CUHK. ✨ Thank you to our incredible hosts for welcoming us and offering a glimpse into the future of robotics. And ofc, a huge kudos to our dedicated TAs for all their hard work!🤗 @hkudatascience #HKU #HKUIDS #CCAI9025 #OpenDriveLab
【🚀 Field Trip Alert! 🤖】 Get ready for an exciting 2-day journey with HKU IDS & Team CCAI9025, led by Prof Hongyang Li! Explore cutting-edge robotics in Hong Kong and Shenzhen with expert-led tours and discussions. 🤖】 👉Details: mmlab.hk/CCAI9025 Don’t miss out!
4
9
1,650
# GO-1: The Embodied Foundation Model is Now Fully Open-Sourced! Following the release of the AgiBot World embodied intelligence dataset with over one million real-world robot samples in March, the embodied foundation model **GO-1** is now officially open-sourced on GitHub! In robotics, model accuracy is important, but true success depends on every step of the end-to-end system—from hardware to data, training to deployment. Any detail gone wrong can lead to complete failure. Over the past year, we've systematically summarized our key learnings and insights into a technical blog. We hope the release of GO-1 and our blog can provide valuable inspiration and reference for researchers and developers worldwide! GitHub: github.com/OpenDriveLab/AgiB… Hugging Face: huggingface.co/agibot-world/… Blog: opendrivelab.com/OpenGO1/ Arxiv: arxiv.org/abs/2507.06219 | arxiv.org/abs/2503.06669
5
932
🚀 New Video Alert! 🚀 📹 Video: piped.video/Ah-xYnST0yw "FreeTacMan: Robot-free Visuo-Tactile Data Collection System for Contact-rich Manipulation" We are pleased to introduce FreeTacMan, a human-centric and robot-free visuo-tactile data collection system for high-quality and efficient robot manipulation! 🤖✨FreeTacMan achieves multiple improvements in data collection performance compared to prior works, and enables effective imitation policy learning—policies trained with our visuo-tactile data achieve an average 50% higher success rate than vision-only approaches.
6
1,065
#ICCV2025 DetAny3D: Detect Anything 3D in the Wild Can your 3D detector handle novel objects & unseen cameras from just a single image? DetAny3D can. 👁️ Monocular Input 🗂️ Box / Point / Text Prompts 🎯 Zero-shot Generalization 🌐 Works even on mobile snapshots & YouTube driving videos! 🔍 Key Innovations: ✅ 2D Aggregator: aligns multi-FM features ✅ 3D Interpreter with Zero-Embedding Mapping: stable 2D→3D transfer ✅ Jointly predicts camera intrinsics and 3D bounding boxes 📈 Up to +21.02 AP3D in novel categories! 📦 Detect Anything, in Any Image, Across Any Domain. 📜 Paper: arxiv.org/abs/2504.07958 💻 Code: github.com/OpenDriveLab/DetA…
16
120
8,639
🚀The AgiBot World Challenge @ IROS2025 starts now! More details on opendrivelab.com/challenge20… Two Tracks 🤖Manipulation (online & onsite): Train models to tackle complex real-world tasks in diverse environments, such as microwave operation and supermarket packaging. The test server is now open! 🌐World Model (online): Predict the evolution of visual perspectives based on action sequences, requiring participants to work with real-robot data and simulate various robotic interactions. ⏰ Key Dates Online Challenge: June 25th to September 1st Grand Finals: October 19th in Hangzhou
1
398
🤔 How to reliably simulate future driving scenarios under a wide range of ego behaviors, especially for rare and non-expert ones? 😭 Challenge of data shortage: Non-expert data with hazardous actions are scarce and unsafe to gather in the physical world. Without such data in training, the learned world model would struggle to follow hazardous actions. 💡Key insight: We can gather such data flexibly and effectively within a driving simulator! By incorporating simulated, non-expert data in training together with the real-world corpus, the resulting world model can successfully transfer the dynamics of simulated situations into real-world scenarios. Introducing ReSim, a driving world model that enables Reliable Simulation of diverse open-world driving scenarios under various actions, including hazardous non-expert ones. A Video2Reward model estimates the reward from ReSim’s simulated future. The high-fidelity prediction, accurate action-following, and reward estimation abilities of ReSim can be applied to various driving applications. Title: ReSim: Reliable World Simulation for Autonomous Driving arXiv: arxiv.org/abs/2506.09981 Project page: opendrivelab.com/ReSim Team work by @jiazhi_yang2024, Kashyap Chitta, Shenyuan Gao, Long Chen, Yuqian Shao, Xiaosong Jia, @francislee2020, Andreas Geiger, @YueXiangyu, and @ilnehc
1
7
14
2,244
The IEEE / CVF Computer Vision and Pattern Recognition Conference @CVPR is being held soon at the Music City Center, Nashville TN, USA. Many members of the MMLab team at HKU @HKUniversity @hkudatascience will attend CVPR in person. Meet us on-site - we'd love to connect, chat, and exchange ideas! Check mmlab.hk/event/cvpr2025 for more.
1
8
1,865
🚀 New Paper Alert! 🚀 FreeTacMan: Robot-free Visuo-Tactile Data Collection System for Contact-rich Manipulation We are pleased to introduce FreeTacMan, a human-centric and robot-free visuo-tactile data collection system for high-quality and efficient robot manipulation! 🤖✨ FreeTacMan achieves multiple improvements in data collection performance compared to prior works, and enables effective imitation policy learning—policies trained with our visuo-tactile data achieve an average 50% higher success rate than vision-only approaches. 💡 Breakthrough Moments: ✅ Visuo-Tactile Sensor: A high-resolution, low-cost visuo-tactile sensor designed for rapid adaptation across multiple robotic end-effectors. ✅ Data-collection System: An in-situ, robot-free, real-time tactile data-collection system to excel at diverse tasks efficiently. ✅ Policy Learning Enhanced by Tactile Pretraining: Imitation policies trained with our visuo-tactile data outperform vision-only baselines by 50% on average. 🚀 Website: opendrivelab.com/FreeTacMan 🛠️ Hardware Guide: docs.google.com/document/d/1… 📜 Paper: arxiv.org/abs/2506.01941 💻 Repo: github.com/OpenDriveLab/Free… #Robotics #DataCollection #Tactile #VisuoTactile #FreeTacMan
2
11
1,633
🚀 #MTGS is now open-sourced! It manages to leverage multi-traversal data for scene reconstruction with better geometry. We utilize the nuPlan dataset with extensive multi-traversal data. 📷 github.com/OpenDriveLab/MTGS
8
35
2,617
🚀 Ready for the #IROS2025 challenge? We've got you covered! This briefing session includes everything you need: task, data, baseline, metrics & more. 🌍 Two identical sessions will run for different time zones. Don't miss it! [Asia/Europe] 28 May 2025, 17:00 (UTC+8) us06web.zoom.us/j/8840270584… [America] 28 May 2025, 18:00 (UTC-7) us06web.zoom.us/j/8634217231…
1
3
701
💥 Forget slow autoregression and skip rigid full-sequence denoising! Nexus is a next-gen predictive pipeline for realistic, safety-critical driving scene generation. What’s new? ✅ Decoupled diffusion → fast updates, goal-driven control ✅ Noise-masking training → inject goals, respond to changes ✅ 540h Nexus-Data → the largest corpus of safety-critical scenes ✅ –42% displacement error, +20% planning boost Why it matters? 📌Building robust autonomous agents needs rare, risky, real scenes. Nexus makes them - efficiently, on demand. 📑 Page: opendrivelab.com/Nexus/ 📜 Paper: arxiv.org/abs/2504.10485 💻 Repo: github.com/OpenDriveLab/Nexu…
11
47
4,441
🏙Proud to support the advancement of autonomous driving in #Shanghai. As part of a collaborative initiative, we are honored to contribute to the city's innovation ecosystem through collaborative efforts with key stakeholders. #ShanghaiInnovation #SmartCity #AutonomousDriving
6
580
Why This Matters Robots today are stuck in narrow tasks. UniVLA unlocks: 🔸 Scalability: Train on YouTube-style videos. 🔸 Efficiency: 21x less pretraining compute than OpenVLA. 🔸 Generalization: Works on unseen tasks/robots with minimal data. 🧵4/5
1
1
416
The Core Idea Traditional robot policies rely on labeled action data, which is expensive and limits scalability. UniVLA solves this by: 🔹 Learning task-centric latent actions from raw videos (no labels!). 🔹 Disentangling task-relevant motions from distractions (e.g., camera shakes). 🔹 Enabling cross-embodiment transfer - train once, deploy anywhere! 🧵 2/5
1
2
269
#RSS2025 UniVLA: Learning to Act Anywhere with Task-centric Latent Actions UniVL is a game-changer for generalist robots! 🤖✨ By learning task-centric latent actions from videos, UniVLA achieves SOTA performance across manipulation & navigation tasks - even with 1/20 the pre-training cost and 1/10 downstream data of prior work! 📜 Paper: roboticsproceedings.org/rss2… 🔗 Official Repo: github.com/OpenDriveLab/UniV… 🧵 1/5
1
11
58
4,816
🚀 New Paper Alert! 🚀 UniVLA: Learning to Act Anywhere with Task-centric Latent Actions We’re excited to introduce UniVLA, a game-changer for generalist robots! 🤖✨ By learning task-centric latent actions from videos, UniVLA achieves SOTA performance across manipulation & navigation tasks - even with 1/20 the pre-training cost and 1/10 downstream data of prior work! 🔑 Key Wins ✅ Scalable Leverages web-scale videos (even human demos!) without action labels. ✅ Efficient Outperforms OpenVLA by 18.5% on LIBERO (also beats pi0 on LIBERO-10 by 6.8%) with minimal fine-tuning. ✅ Generalist Deploys to any robot via lightweight (~12M param.) action decoding. 📌 Why it matters? Robots can now learn from any video source, bridging embodiment gaps and unlocking real-world versatility. 📜 Paper: roboticsproceedings.org/rss2… 💻 Official Repo: github.com/OpenDriveLab/UniV… #Robotics #VLA #UniVLA
19
79
7,272
🤖 Spoiler alert! 👀 At IROS 2025 in Hangzhou, robots will compete, live on stage. 🔗 Details will be released soon (very soon) on opendrivelab.com/challenge20… #IROS2025 #humanoidrobot
1
8
1,188
⏰ The #CVPR2025 submission deadline is approaching fast — don't forget to upload your results by May 10th! Right after CVPR, the #ICCV2025 season begins. 🤖 Coming in early May: The AgiBot World Challenge! In this challenge, we are going beyond the lab and into the real world — deploying them on real robots, live at #IROS2025. 🔗 Learn more at opendrivelab.com/challenge20…
2
6
904