Postdoc @berkeley_ai. Researcher @IBMResearch. PhD @TelAvivUni. Working on Multimodal Robot Foundation Models and Structured Physical AI.

Berkeley, CA
If you want a vision encoder for dexterous manipulation, what should be the most important part to model? ๐Ÿค” Current standard models like CLIP, SigLIP, and DINOv2 have an incredible grasp of semantics and spatial details. But they lack the action-centric structure needed for downstream visuomotor control. But collecting annotated robotic trajectories at scale is SUPER expensive and largely unrealistic. So, how do we bridge this gap? We introduce CAIP (Contrastive Action-Image Pre-training) โฌ‡๏ธ ๐Ÿ”ธ Action-centric upstream: we align visual observations with action chunks through a contrastive objective. ๐Ÿ”ธ Human video as a proxy: we represent 3D human hand poses analogously to robotic end-effector actions, tapping into a massive source of human demonstrations. ๐Ÿ”ธ Massive scale: pre-trained on over 32,000 hours of manipulation video, driving both sample efficiency and robust generalization. ๐Ÿ”ธ Hardware proven: achieves a 76% average success rate on a real-world Dexmate Vega bimanual @DexmateAI manipulator with dual 22-DoF Sharpa Wave hands @SharpaRobotics . ๐Ÿ”ธ State-of-the-art: significantly outperforms strong baselines like DINOv2, SigLIP, MVP, and Qwen3.5 ViT across complex tasks, even under unexpected lighting changes and visual distractors. ๐ŸŒ Project: caip-encoder.github.io/ ๐Ÿ“ Blog: nitter.net/roeiherzig/status/2063โ€ฆ ๐Ÿ“„ Paper: arxiv.org/abs/2606.17256 ๐Ÿ’ป Model: huggingface.co/yuvansharma/cโ€ฆ
4
22
146
26,502
If Codex is down, does #ICLR allow a 1-day grace period? These days, it's worse than if OpenReview is down.
We are aware that codex is down and are working hard to bring back normal service.
4
758
Yep.
there are many many people who studied physics and came to the conclusion that working on machine learning would probably create more scientific progress than studying physics including many of the original chatgpt crew like john schulman, liam fedus, amodei & jared kaplan at ant
366
ื›ืœ ื”ืคื™ื“ ืฉืœื™ ืžืœื ื‘ื™ืฉืจืืœื™ื ืฉืžืชืœื”ื‘ื™ื ืžืžื™ื•ื– ืžืฉื•ื ืžืงื•ื ื•ื‘ืœื™ ืฉื•ื ืงืฉืจ. ืื™ื–ื” ืžื•ื–ืจ ื–ื”, ื‘ืžื™ื•ื—ื“ ืฉืืฃ ืื—ื“ ืขื•ื“ ืœื ื”ืชื—ื™ืœ ื‘ืืžืช ืœื”ืฉืชืžืฉ ื‘ื–ื” ืคื”.
ืชื•ืš ื›ืžื” ื–ืžืŸ Muse ืžื’ื™ืข ืœ 100 ืžืœื™ื•ืŸ ืžืฉืชืžืฉื™ื?
7
1
2,675
I've been playing with @Muse recently, but still haven't figured out the best way to use it. What are the good use cases people are actually using it for? For example, it refuses to control my WhatsApp and send automatic messages, but it sends on Instagram, which is really not useful.
3
665
The kind of dataset whose impact only becomes obvious 1โ€“3 years later. HUGE RESPECT TO THE TEAM BEHIND IT.
Full ABC release ๐Ÿ”ค ! Train, test, and evaluate your models. The full suite of simulation tasks and data has also been released! Excited to see what people build ๐Ÿ”จ๐Ÿค–
1
10
2,333
Every few years the field agrees the review system is broken. Then submissions double. ICLR just passed 50,000 ๐Ÿคฏ At some point we have to cap how many papers an author can submit โœ‹๐Ÿ“„...Nobody has more than 5 conference-worthy ideas in a single cycle.
15
6
171
30,383
๐Ÿšจ Call for participation: the Childโ€“Robot Dexterity Challenge at the CoRL 2026 ๐—ฆ๐—ฃ๐—œ๐—ก ๐—ช๐—ผ๐—ฟ๐—ธ๐˜€๐—ต๐—ผ๐—ฝ! Bring your robot, or run on ours. Apply by Oct 5. ๐Ÿง’๐Ÿค– Robots and a child will take on the same everyday tasks, side by side, live on stage. The demonstration offers a simple way to see what todayโ€™s robots can already do and where thereโ€™s still room to improve. We make it easy to join: - Bimanual YAM setup (MolmoAct 2 style) + a dexterous Sharpa hand setup, provided on stage. - Bring your own hardware (subject to approval). ๐ŸŒ Website: spin-workshop.github.io/#conโ€ฆ ๐Ÿ—ฃ๏ธ Speakers & panel from Stanford, Columbia, TU Darmstadt, and AMI Labs. ๐Ÿ“ Nov 12 ยท Austin, TX #CoRL2026
2
10
33
5,711
Possible tasks: grasping, opening bottles, insertion, folding, and pouring. Every task will be child-appropriate. Weโ€™ll share the final list and completion guidelines in advance.
1
3
378
Academic and industry teams are welcome. Bring approved hardware, use a provided setup, or contribute hardware, data, or sponsorship. Apply by Oct 5: roeiherz@gmail.com, amirb4r@gmail.com @_amirbar , or csferrazza@berkeley.edu. @carlo_sferrazza
5
336
Fun fact: fewer than 0.8% of researchers are one coauthorship hop from Jitendra Malik ๐Ÿ˜€ alphaxiv.org/number/@jitendrโ€ฆ
2
1
22
8,602
Have we reached the point where hitting your token limit is a valid reason to ask for a conference deadline extension?
As ICRA drawing close, it feels like GPT-6 Astra will be the new baseline all reviewers will be asking for, but the tokens ainโ€™t cheap ๐Ÿ˜….
3
801
Achieving a perfect score on this benchmark suggests that the model could potentially have been trained using that specific data. What insights did we gain from this?
GPT-6 Astra is the most significant leap in robotics Iโ€™ve seen in the past few years. It cracked RoboLab with a near-perfect score. Solid infrastructure + scaling ultimately outperformed the heuristics explored in small-scale studies. Weโ€™re definitely on the brink of physical RSI. Source: anonymous-report-421.github.โ€ฆ
3
1
17
3,222
Interesting, but for what itโ€™s worth, I donโ€™t think Anthropic or OpenAI will solve robotics. Iโ€™d put my money on robotics companies, but time will tell.
Iโ€™m sensing deep despair in academics over the past week. Astra, Fable, Muse are zero-shotting benchmarks in robotics & world models. Uneasy pill to swallow, but this is what step jumps in progress looks like.
11
2,655
ื”ื“ืจืš ื”ื™ื—ื™ื“ื” ื”ืคืจืงื˜ื™ืช ืฉื™ืฉ ื”ื™ื ืœืคืชื•ื— ืืช ื”ืžื•ื“ืœื™ื ืœืฆื™ื‘ื•ืจ. ืจืง ื‘ืฆื•ืจื” ื”ื–ืืช, ืขื•ืœื ื”ืžื—ืงืจ ื™ื•ื›ืœ ืœื—ืงื•ืจ,ืœื”ื‘ื™ืŸ ืžื” ื”ื™ื›ื•ืœื•ืช, ื•ืื™ืš ืœื”ื’ืŸ.
ืฉื•ื‘ ืœื ื ืชืขื•ืจืจ ื‘ื–ืžืŸ? ืœืคื™ ืžื ื›ืดืœ ืื ืชืจื•ืคื™ืง, ื”ืžื•ื“ืœื™ื ืฉืœ AI ื™ื”ื™ื• ืžืกื•ื›ื ื™ื ื‘ืชื•ืš ืฉื ื”. ืœื ืžื‘ื™ืŸ ืืช ื”ื”ืชืขืœืžื•ืช ื”ืžื•ื—ืœื˜ืช ืžื”ืกื›ื ื”. ื•ื™ืฉ ื”ืžื•ืŸ ืžื” ืœืขืฉื•ืช! ืœื ืฉืžืขืชื™ ืืฃ ืžื•ืขืžื“ ืฉืžืฆื™ืข ืœื”ืงื™ื ืจืฉื•ืช ื”ื’ื ื”, ืœื‘ื•ื“ื“ ืชืฉืชื™ื•ืช, ืœืคืชื— ืืœื˜ืจื ื˜ื™ื‘ื•ืช ื•ืขื•ื“ ืฆืขื“ื™ื ืจืœื‘ื ื˜ื™ื™ื. ื”ืื ืื™ื–ื• ืžืคืœื’ื” ืชืฆื™ื’ ืชื›ื ื™ืช? ืœื ืกื‘ื™ืจ:(
2
5
1,344
Roei Herzig retweeted
Want to see how far a robot policy can go on just a handful of demos? At #ECCV2026 in Malmรถ today I'm presenting Robotic Steering: adapt a VLA by tuning only the attention heads that matter. Same performance, fewer parameters, and it holds up on a real robot! ๐ŸŽค Spotlight Oral, 9:10, Palissad (Malmรถ Arena) ๐Ÿ–ผ๏ธ Poster #344, 10:30โ€“12:30, Exhall Project Page: chancharikmitra.github.io/roโ€ฆ
1
7
280
Sorry to burst the bubble, but GPT-3.5 also did ICL out of the box. davidyyd.github.io/roboprompโ€ฆ
I didn't get what GPT-6 meant for robotics until I actually tried it. GPT-6 Astra just does physical ICL out of the box. we drop a recording of a human doing a novel task into ๐—ฐ๐—ผ๐—ฑ๐—ฒ๐˜… app. Prompt it to drive a robot arm the same way. It just works on the first pass!
5
2
87
12,000
That's a really interesting perspective, as Michael always provides. However, this feels a bit like a "computer vision" viewpoint. In the field of Physical/Embodied AI, I believe the area is much more exploratory, with a lot yet to be discovered and done...
What is the role of academic computer vision research in the age of increasingly powerful large models? Is GPT-6 Astra a step change? How can a researcher have an impact today in academia? These are the questions I ask myself as I head off to ECCV 2026, a conference Iโ€™ve attended since 1992. One of my papers this year is VIGA, a method that takes an image as input and outputs a 3D Blender scene that represents that image. This is a classical inverse-graphics task and VIGA was the first method to solve it using an agentic approach. The idea is now several years old and the first version of the paper was rejected. This delayed publication significantly. After it was accepted at ECCV, it was quickly surpassed by people using Claude Code for the same purpose. Today GPT-6 Astra blows away all previous results. But we still head off to ECCV to tell the community about our invention that is now fully out of date. The way academic work often progresses is that one reads recent papers, notices that they have limitations, comes up with a new idea, explores this, publishes it, etc. Any published paper I read today is based on ideas that are at least a year old. And those ideas were based on the literature of the time, which was also a year old. That means that any paper I see at ECCV is likely two years out of date. In AI today, two years means your work is likely irrelevant. At CVPR this summer I noticed that many authors have not gotten the message. They continue to work on โ€œoldโ€ problems that have a long history. This history is based on assumptions about how the โ€œvision problemโ€ will be โ€œsolvedโ€. The truth is that it is being solved in a very different way and many of these problems are no longer relevant. Another group of papers focuses on very niche problems where large models likely fail because of insufficient data or lack of business interest. The impactful papers were largely from industry and had long author lists and massive data+compute behind them. These papers were also out of data, describing systems that had been released months before, but at least they served to provide the community with more complete documentation and analysis of commercial systems. So what should academics do? First, we need to put aside the tools weโ€™ve used for years and start from scratch. Every project should start by trying really hard to solve the problem with existing tools. I would like to see every paper begin with a detailed experimental analysis of how existing models perform and why they fail (if they do). This gives the kind of insight we need today. Then, assuming current models fail, the solution should provide some fundamental insight that will outlive the next release of such models. Reviewers today still focus on technical novelty. This pushes people to focus on tweaking architectures rather than clearly moving the field forward. Papers need to be judged based on their novel insight and not their novel technical contribution. This is a real shift in thinking but it focuses us on what matters - progress of the field. If we want there to be a โ€œfieldโ€ of computer vision, then it canโ€™t become a marginal backwater, focusing on esoteric problems. If you havenโ€™t tried using Astra (or whatever comes next) to solve your problem, then you have not done your homework. This omission should be seen as negatively as not having a previous work section. Concretely, I think papers should include a new section analogous to โ€œRelated Workโ€ where that related work is current models and how they perform on the task. Reviewers should start asking for this and expecting authors to be able to articulate their insights about the limitations of existing large models. I'm interested in your thoughts.
3
534
One last thought before I go to sleep: How long until OpenAI builds Raccoon City? I always knew Capcom knew what they were talking about.
1
373
What I read is "We don't care about image generation at this point" ๐Ÿ˜…
Images 2.5 is here. I don't think it can solve super difficult math problems, but it is really good and we hope you enjoy it. openai.com/index/introducingโ€ฆ
1
1
1,303
None of this is clear to me... "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our modelsโ ." Does this basically mean that all the research I conduct using Codex belongs to OpenAI?
Very sad to see Levent double down on the plagiarism accusation. I hope my friends at @AnthropicAI stand up to this internally. It should be clear by now what the truth is.
22
9
312
11,761