Pinned Tweet
😈 WebGym is now 300% faster with better performance. We introduce AsyncWebRL, with 300% end-to-end throughput and better performance over WebGym (which uses REINFORCE) with async GRPO. Two algorithmic takeaways we learnt: (1) do asynchronous multi-step agentic RL with the **decoupled loss**, and (2) **don’t do normalization over step-number**, instead use the same weighting over each step across trajectories. Details below 🧵: Paper: arxiv.org/pdf/2606.05597 Website: asyncwebrl-website.github.io… Code: github.com/microsoft/webgym 1/11
1
7
41
6,417
Jack Bai retweeted
Introducing CUA-Lite 🧵 — an open platform for computer-use agents. Training and benchmarking CUAs (Computer-Use Agents) requires four core pieces: 1️⃣ Agents — the models and the scaffolding that drives them 2️⃣ Environments — runtime/sandboxes for agents to interact with, tasks & verifiers/graders 3️⃣ Traces — records of agent trajectories 4️⃣ Frameworks — to evaluate, SFT & RL-train agents Today, all four are fragmented. Every agent ships with its own implementation, often in a separate repo — there is no unified way to run them all. Every environment exposes its own interface and action space, often requiring an expensive VM sandbox for each verifiable task. Traces come in incompatible formats. And without common standards across the stack, every project ends up rebuilding its own tooling/framework for eval, SFT, and RL. CUA-Lite unifies the stack: → One standardized interface & action space for agents and environments → One standardized format for agent traces → One framework for evaluation, SFT & RL → Across desktop, browser & mobile And open resources plug straight in, creating the largest open collection of CUA agents, environments and traces, all in a unified format: 🤖 10+ CUAs, including GPT, Claude, Gemini, Qwen, Muse-Glimmer, UI-TARS 🌐 15+ benchmarks, including OSWorld, WebArena & AndroidWorld ⚡ Optional VM-free sandboxes with 30K+ verifiable tasks for training 📚 10+ trace datasets, freely available on Hugging Face, including public datasets converted into the standardized format and fresh rollouts from frontier open-weight CUAs Led by @BerkeleyRDI , our goal is for CUA-Lite to become a community-driven, open-source ecosystem for computer-use agents. Join the community and contribute today: bring an environment (runtime/sandbox + tasks + verifier), traces, or an agent, and plug it into CUA-Lite!
18
34
212
82,990
Jack Bai retweeted
Behavioral cloning mystery seohong.me/blog/behavioral-c… I wrote a new blog post about "mysteries" in behavioral cloning that appear with real-world robot data (e.g., overfitting is "good"). I also tried to demystify them and shared my thoughts!
38
136
1,023
159,309
Jack Bai retweeted
on side note, the fun of working in RL theory is that results like this are still discovered from time to time, where a few lines explain (to experts) significance + proof machinery, and we are like "how come we didn't realize this before"
Cool paper from @LarsvanderLaan3 and @nathankallus: you can use MLE to estimate density-ratio in MDP w/o completeness. Usually you need latter b/c bellman op and projection contract in diff geometries, but in this case it's both KL. Very cool insight. arxiv.org/abs/2607.05375
1
1
50
9,547
So glad to see that WebGym is being used as an environment in training Qwen-AgentWorld!!
1
2
21
2,944
😈 WebGym is now 300% faster with better performance. We introduce AsyncWebRL, with 300% end-to-end throughput and better performance over WebGym (which uses REINFORCE) with async GRPO. Two algorithmic takeaways we learnt: (1) do asynchronous multi-step agentic RL with the **decoupled loss**, and (2) **don’t do normalization over step-number**, instead use the same weighting over each step across trajectories. Details below 🧵: Paper: arxiv.org/pdf/2606.05597 Website: asyncwebrl-website.github.io… Code: github.com/microsoft/webgym 1/11
1
7
41
6,417
Takeaways: 1. Use async RL for end-to-end RL, even if it is multi-step. It’s significantly faster than the sync version due to near-zero bubble time (both CPU and GPU run at 90-100% utility all the time), if well optimized. 2. Always use decoupled loss under asynchronous setting. Center your trust region around the **proximal policy**, not **behavior policy**. 3. Avoid **step-number normalization** in GRPO under multi-step agentic settings. This is both theoretically lame (as explained in Dr. GRPO) and empirically leading to unexpected behaviors (as it causes unnecessarily long trajectories and responses). 10/11
1
1
3
109
This is a great collaboration with @RuiYang70669025, @ye_chenlu, @spncrwhthd, Professor @aviral_kumar2 , and Professor Tong Zhang, and we had great discussions with @Anikait_Singh_ and Professor @nanjiang_cs. There are many more loss variants you can tune to improve the end-to-end asynchronous performance under multi-step agentic setting. Check out the website and paper for the complete results. Code is fully open-sourced. 11/11 The End.
3
142
Check out Rui's paper on a fully open framework for training visual web agents with online multi-turn RL on live websites, where a tiny 4B model hits open-source SOTA and even rivals proprietary systems!
Excited to introduce OpenWebRL 🚀 A fully open framework for training visual web agents with online multi-turn RL on live websites. Today’s strongest web agents often rely on proprietary infrastructure, large-scale closed data, and costly API-based evaluation. OpenWebRL aims to make this training stack open, reproducible, and cost-friendly. With only 0.4K SFT trajectories and 2.2K online RL tasks, OpenWebRL-4B sets a new open-source SOTA on live-web benchmarks including Online-Mind2Web, DeepShop, and WebVoyager, outperforming much larger open-source agents and even matching proprietary systems. We release the full stack: data, models, training framework, and judge models. Paper: arxiv.org/abs/2606.02031 Project: openwebrl.github.io/ Code: github.com/OpenWebRL/OpenWeb… Models & Data: huggingface.co/OpenWebRL
1
9
1,244
Tomorrow at ExHall A & F Poster Location: 471, 5pm-7pm, we'll present WebGym, together introducing two recently pre-released works (AsyncWebRL and OpenWebRL) as a surprise :) WebGym: webgym-website.github.io/web… OpenWebRL: openwebrl.github.io AsyncWebRL: asyncwebrl-website.github.io
I will be attending CVPR in Denver on June 5 and June 6, presenting our work WebGym, the largest yet visual web agentic RL framework, along with many interesting observations we found during RL. Come by if you're interested! Track the event at cvpr.thecvf.com/virtual/2026… :)
2
9
709
I will be attending CVPR in Denver on June 5 and June 6, presenting our work WebGym, the largest yet visual web agentic RL framework, along with many interesting observations we found during RL. Come by if you're interested! Track the event at cvpr.thecvf.com/virtual/2026… :)
4
1,154
Jack Bai retweeted
Today we launch Recursive. We are building AI that discovers knowledge automatically and improves itself recursively, an open-ended process that will fundamentally change how science and technology advance. Our 25 top researchers and engineers in San Francisco and London bring diverse expertise spanning agentic AI scientists, architecture and algorithm design, world models, optimization, and interpretability, united by a shared conviction that this is the most important problem we could be working on today. If you are interested in joining, please send your resume to talent@recursive.com. Follow us at @Recursive_SI!
89
150
1,365
175,508
Congrats, Dr. Tong!
I defended my thesis today! Sincere thanks to my advisors @sainingxie @ylecun and committee members: @mengyer @YiMaTweets @LukeZettlemoyer @liuzhuang1234. I could not have wished for a better PhD life, and I want to thank everyone who was part of this journey. Slides Link: tsb0601.github.io/data/defen…
1
6
1,361