We're graduate students, postdocs, faculty and scientists at the cutting edge of artificial intelligence research.

Berkeley, CA
Berkeley AI Research retweeted
One human demonstration. Any multi-fingered hand. Zero-shot sim-to-real visuomotor policy. morphometricimitation.github… Collaborators: @he_siming @ckwolfeofficial @HaozhiQ @LeaMue27 Shankar Sastry, Claire Tomlin, @JitendraMalikCV
8
18
142
22,310
Berkeley AI Research retweeted
Formal verification × agents 🚀 Our IDS paper was accepted as a NeurIPS oral! We study how agents can co-evolve code + proofs, achieving 3x the success rate of Claude Code on distributed systems specs. IDS Paper: arxiv.org/abs/2605.23109 Code: github.com/skydiscover-ai/sk… And also check out our recent work on the next challenge: ensuring the spec itself captures human intent. 👇 nitter.net/istoica05/status/21009…
I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions. Video attached. TLA+ also works well. I sometimes combine Lean and TLA+ to look for issues around data flow, concurrency, and state mgmt. I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted. Is formal verification the future of coding (or at least, bug finding)?
21
40
268
41,865
Tonight we kick off the 3rd season of the @UCBerkeley AI & Society Initiative series with an interdisciplinary panel on cybersecurity & AI focused on the recent Hugging Face OpenAI and related incidents, 5:30-7:00 PM in @BerkeleyCDSS Gateway Building. Co-organized this year by @irenetrampoline, @alsuhr, and me, the event is a co-production of Law + EECS/CDSS. Featuring - @cchio (Co-Founder & CEO of Coverbase.AI) - @hoofnagle (Faculty Director, Berkeley Center for Law & Technology / Professor at @BerkeleyLaw) - Starchy Grant (Principal Systems Administrator @EFF) Even if you can't make it in person, sign up to join the community and receive our substacks and the recording. - RSVP: lnkd.in/gcE6v7bC - Substack : lnkd.in/g6NY4Yqh - Website: lnkd.in/gc3xXA98 See below for a list of our fall lineup so far, with more to come! See you in Gateway or online!
4
10
2,633
Berkeley AI Research retweeted
Introducing TANGO 💃🏻, our #CoRL2026 work on whole-body VLA for humanoid navigation. TANGO maps RGB directly to 29-DoF actions, enabling coordinated whole-body navigation through cluttered 3D spaces. Trained entirely in simulation, zero-shot in the real world. 🤖🚀
💡 Humanoid navigation is more than a 2D problem. How can a robot use its whole body to traverse cluttered 3D spaces? 💃🏻 Introducing our #CoRL2026 work TANGO, a whole-body VLA for humanoid navigation that directly predicts 29-DoF joint actions from RGB observations, enabling zero-shot real-world navigation with coordinated arms, torso, and gait. [1/6] 🌐 tango-vla.github.io 📄 arxiv.org/abs/2609.09158
5
5
50
9,426
Berkeley AI Research retweeted
Please help spread the word! We are recruiting multiple Postdoctoral Fellows as part of the recently launched Bakar Computational Biomedicine Initiative (BCBI) at UC Berkeley and UCSF. BCBI website: bcbi.berkeley.edu/ Apply by Nov 1, 2026: berkeley.infoready4.com/#fre… (1/n)
4
65
152
20,383
Berkeley AI Research retweeted
We just trained a 2.3B MoE Hybrid Mamba model that competes with a Llama-3.2-3B dense using less than 1% of the pre-training FLOPS. We may not have a lot of resources in academia but we make up for it with creative students who find ways to train innovative new architectures across whatever hardware they can find. If anyone wants to help, we would love to scale this research to the next rung of the training ladder.
We pretrained a 2.3B MoE (360M active) Hybrid Mamba-2 that lands within a few points of Llama-3.2-3B using <1% of its pretraining FLOPs. No dedicated cluster. The run hopped between H100s, A100s, V100s (yes, V100s) and TPU v5p/v6e on a single codebase. Meet Rigel 🧵
1
6
74
7,902
Berkeley AI Research retweeted
We pretrained a 2.3B MoE (360M active) Hybrid Mamba-2 that lands within a few points of Llama-3.2-3B using <1% of its pretraining FLOPs. No dedicated cluster. The run hopped between H100s, A100s, V100s (yes, V100s) and TPU v5p/v6e on a single codebase. Meet Rigel 🧵
22
81
633
81,334
Berkeley AI Research retweeted
Interestingly, one attention head can be important across five ICL task families! Our original thread gave an in-depth mechanistic analysis of addition ICL (quick video recap 👇). We've now generalized this to four more ICL task families, spanning both arithmetic tasks and semantic tasks! Our paper “Understanding In-context Learning of Addition via Activation Subspaces” has been accepted to COLM 2026! 🎉 See more details in thread. (1/N)
3->5, 4->6, 9→11, 7-> ? LLMs solve this via In-Context Learning (ICL); but how is ICL represented and transmitted in LLMs? We build new tools identifying “extractor” and “aggregator” subspaces for ICL, and use them to understand ICL addition tasks like above. Come to @interplaywrkshp at COLM to learn more!
6
15
54
13,117
Berkeley AI Research retweeted
How do we run RL with real-time chunking (RTC)? In this work we figured out how to use a small RL policy with a large robot foundation model, where the RL policy observed more recent images (due to faster inference) and steers the policy toward better behaviors! A fun collaboration with Siemens, led by Brian Zhu, Momen Khalil, Emanuele Poggi from Siemens and @ehharrison4 from Berkeley, with lots of amazing contributors!
Asynchronous VLA inference reduces inference delay, but breaks the Markovian assumption necessary for RL fine-tuning. How can we enable RL fine-tuning of VLAs with async inference? We introduce ARLI: Asynchronous RL with Intermediate Information! async-rl-intermediate-inform… (1/n)
7
54
467
42,741
Berkeley AI Research retweeted
Asynchronous VLA inference reduces inference delay, but breaks the Markovian assumption necessary for RL fine-tuning. How can we enable RL fine-tuning of VLAs with async inference? We introduce ARLI: Asynchronous RL with Intermediate Information! async-rl-intermediate-inform… (1/n)
3
10
134
62,442
Berkeley AI Research retweeted
Check out our paper at COLM 2026. We make (world-model based) controller language-instructable. It needs no manual labels with rollouts + our pos-hoc annotation method. The result is a flexible controller interface that any VLM/human can plug into. zinengtang.github.io/instruc…
5
47
7,786
Berkeley AI Research retweeted
ABC is fully out and headed to CoRL 2026! ABC-130K is the largest open teleop dataset to date: 3,500 hours, 130K+ episodes, 195 tasks, collected on an $8K bimanual setup. The full stack is open: hardware, training code, sim, and eval, so academic labs can do this research on equal footing. Beyond the data, we ran the baseline science for the community to stack on: • Sim-to-real correlation (r = 0.91 on task progress) • Which offline metrics matter? • Scaling law in compute (training steps and batchsize) • Conditioning the policy Project page: abc.bot/ Github: github.com/amazon-far/abc
Happy to announce the full ABC release! We’re also excited that ABC was accepted to CoRL 2026! Check out our website for code, 400+ hours of sim data on 24 tasks, and 5,850 labeled policy-evaluation episodes. @arthurallshire @Cinnabar233 @ritvik_singh9 @redstone_hong @davidrmcall
1
9
66
16,116
Berkeley AI Research retweeted
Agent harnesses seem to make a lot of difference, at least for cost, even on open source coding benchmarks. Melissa’s research project digs into why.
Does your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge: 1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost 2️⃣A simple harness can be competitive 3️⃣The native harness isn’t always the best. Millions of people are using coding agents, but the impact of harness choice remains unclear. (1/n) More details in the thread. 🧵
11
19
184
40,595
Berkeley AI Research retweeted
Does a model need its native harness? We evaluated seven models across the Claude Code, Codex, and Pi harnesses and found a surprising result: harness choice had little effect on task success but a substantial effect on cost. Sometimes, a simple harness is all you need!
Does your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge: 1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost 2️⃣A simple harness can be competitive 3️⃣The native harness isn’t always the best. Millions of people are using coding agents, but the impact of harness choice remains unclear. (1/n) More details in the thread. 🧵
26
16
157
26,223
Berkeley AI Research retweeted
🚨 Call for participation: the Child–Robot Dexterity Challenge at the CoRL 2026 𝗦𝗣𝗜𝗡 𝗪𝗼𝗿𝗸𝘀𝗵𝗼𝗽! Bring your robot, or run on ours. Apply by Oct 5. 🧒🤖 Robots and a child will take on the same everyday tasks, side by side, live on stage. The demonstration offers a simple way to see what today’s robots can already do and where there’s still room to improve. We make it easy to join: - Bimanual YAM setup (MolmoAct 2 style) + a dexterous Sharpa hand setup, provided on stage. - Bring your own hardware (subject to approval). 🌐 Website: spin-workshop.github.io/#con… 🗣️ Speakers & panel from Stanford, Columbia, TU Darmstadt, and AMI Labs. 📍 Nov 12 · Austin, TX #CoRL2026
2
10
33
5,711
Berkeley AI Research retweeted
Does your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge: 1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost 2️⃣A simple harness can be competitive 3️⃣The native harness isn’t always the best. Millions of people are using coding agents, but the impact of harness choice remains unclear. (1/n) More details in the thread. 🧵
179
187
1,349
353,016
Berkeley AI Research retweeted
New paper evaluating LLMs' professional knowledge across occupations! We develop a scalable approach to benchmarking occupation-specific knowledge by turning online authoritative sources into test questions. Really enjoyed working on this with @abhishekn and @ShreyasK1609!
🚨🚨🚨 New paper alert 🚨🚨🚨 To understand LLMs impact on work, we first need to examine their capabilities across the entire occupation distribution. Doing this in a scalable way is hard and costly. Presenting - "ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge" a new working paper with @serinachang5 and @ShreyasK1609 that takes this challenge head on! paper: arxiv.org/abs/2609.12366 dashboard: orqabench.org
6
36
7,424
Berkeley AI Research retweeted
New work in @ScienceTM by Meghana Kamineni and a great team led by @ZhiYu_ACGT and @pnatarajanmd! We used deep learning-extracted features from imaging and genomics to understand the role of spleen in coronary artery disease.
Our study led by M. Kamineni & with @ZhiYu_ACGT uses imaging and genomics to prioritize a role for the spleen in coronary artery disease science.org/doi/10.1126/scit… @ScienceTM
1
4
10
6,413