Autoscaling RL environments and benchmarks

On the internet
Vibrant Labs retweeted
At @evo__hq we hillclimbed a smarted router setup for agentic workloads ( ITSM bench by @vibrantlabsai ) that achieves SOTA results while being 20x cost effective vs frontier models - beating as well the newly released models of Grok4.6 and Deepseek v4 pro in both quality and cost For enterprise agentic workloads, what we have repeatedly found is that frontier models still do not get that last mile reliably and are prohibitively costly. It becomes necessary to hillclimb on your specific data distribution, schema, policy and traces. This is what we do at @evo__hq
2
8
33
13,581
Vibrant Labs retweeted
This has been a concept that I have been tinkering for a while and is missing in current open environments to help agents generalize beyond benchmarks. DM if you like to work on this with us.
Article

Agent Evals Need Partial Observability

Most RL environments expose too much information to the agent than what’s observable in real world deployments. A lot of tool-use benchmarks give agents the full tool list, a single clean policy file,

2
3
12
402
Vibrant Labs retweeted
8/ Enterprise Worlds is open source, starting with ITSMBench. ITSMBench builds on the seed data and action space from ServiceNow’s open-sourced EnterpriseOps-Gym. Big thanks to the @ServiceNowRSRCH team for making that foundation available. Repo: github.com/vibrantlabsai/Ent… Leaderboard: enterpriseworlds.vibrantlabs…
2
2
176
Vibrant Labs retweeted
1/ Today, we’re open-sourcing Enterprise Worlds: executable environments for training and evaluating AI agents on realistic enterprise workflows. First release: ITSMBench, an IT service management benchmark built around multi-turn tasks. For multi-turn agent evaluation, τ-bench from Sierra has become an important standard. But most multi-turn benchmarks still focus on consumer-facing workflows: airline, retail, bookings, support-style interactions. Enterprise work has a different shape. Blog: vibrantlabs.com/research/ent… Leaderboard: enterpriseworlds.vibrantlabs…
3
7
31
9,311
Vibrant Labs retweeted
1/n In last week’s Paper Club, we focused on three papers that give insights on how to generate post-training data that meaningfully improves agent capabilities. This is the central question that @ VibrantLabs exists to solve. - PlanBench-XL (from @UofIllinois), which focuses on realistic tool envs - TMax (from @allen_ai, and @uwcse), which focuses on synthesizing and training against harder tasks, and - Autodata (from @AIatMeta), an agent data scientist loop that synthesizes data and can be meta-optimized.
2
3
12
874
Vibrant Labs retweeted
We had a blast last week when we hosted the first Hot Takes game night (+ dinner) focused on autoscaling RL envs. Carefully chosen guests were instructed to bring the most controversial opinions they could to a discussion on post-training. As per usual, only technical practitioners, no VCs. We had conversations on how to improve diversity of synthetic envs, distribution collapse, whether it’s even possible to do entirely autoscaling and RSI. We host these specific research-focused events regularly, with different focuses and activities. If you’re working on any of the following problems, we’d love to include you in the next one: - Unsupervised environment design - Efficient RL training for multi-turn tool use - Self-evolving benchmarks - Autonomous AI research - Open-Endedness
2
1
13
4,837
Vibrant Labs retweeted
At @VibrantLabsAI, we’ve always been a research-minded team internally, so it felt completely natural when we started doing a regular, organized Paper Club as a team. What we didn’t expect was how much interest we’d get from that over the past few months from folks outside the team. Every Paper Club, we post our notes from the discussion on Twitter/LinkedIn and mention the authors and the orgs involved with relevant papers. We often end up speaking with those authors before and after our discussion, and nowadays, we even work with some of them on a regular basis to help us autoscale RL envs. This week, we decided to formalize our Paper Club a little further by adding a dedicated section to our site where you can see all of the notes and papers discussed: vibrantlabs.com/paper-club Hope anyone who is following along enjoys it.
2
1
11
765
Vibrant Labs retweeted
Replying to @VibrantLabsAI
evals at scale with @browserbase sounds good to me
1
2
281
Vibrant Labs retweeted
we are releasing ecom-bench which tracks how agents perform in basic e-commerce tasks with DOM vs CUA modalities with the stagehand harness from @browserbase . there were a couple of interesting takeaways
1/n For browser agents, a major bottleneck in evaluation is truthful scoring on the live web. A task is only as good as your ability to confirm the agent actually did it, on a real site whose state keeps moving and that the agent can potentially misreport. So we took matters into our own hands. Today, we're releasing Ecom Bench on @PrimeIntellect: 40 shopping tasks on real Shopify storefronts, each run in a live @browserbase browser and graded by a deterministic verifier. vibrantlabs.com/research/eco…
1
2
4
289
Vibrant Labs retweeted
Everyone building browser agents eventually comes to the same divergence: should the agent read the DOM or look at the page visually like a human? At @VibrantLabsAI, we ran this experiment on Ecom Bench (with 40 verified shopping tasks on live storefronts) and the answer was somewhat complicated. Labs like @yutori_ai and @AnthropicAI are thinking deeply about the trade-offs between DOM and CUA (see their "bitter lesson for web agents" and “demystifying evals” posts, respectively). Full results, cost and latency breakdowns, and the failure-mode analysis are in the thread below and the post. Blog: vibrantlabs.com/research/eco… Env on @PrimeIntellect (built w @browserbase): app.primeintellect.ai/dashbo…
1/n For browser agents, a major bottleneck in evaluation is truthful scoring on the live web. A task is only as good as your ability to confirm the agent actually did it, on a real site whose state keeps moving and that the agent can potentially misreport. So we took matters into our own hands. Today, we're releasing Ecom Bench on @PrimeIntellect: 40 shopping tasks on real Shopify storefronts, each run in a live @browserbase browser and graded by a deterministic verifier. vibrantlabs.com/research/eco…
1
2
7
474
Vibrant Labs retweeted
static benchmarks can't tell you if your agent can actually do human tasks online. @VibrantLabsAI built one that can: their web agent eval runs on live shopify stores + deterministic verifiers via @browserbase. run it yourself↓
1/n For browser agents, a major bottleneck in evaluation is truthful scoring on the live web. A task is only as good as your ability to confirm the agent actually did it, on a real site whose state keeps moving and that the agent can potentially misreport. So we took matters into our own hands. Today, we're releasing Ecom Bench on @PrimeIntellect: 40 shopping tasks on real Shopify storefronts, each run in a live @browserbase browser and graded by a deterministic verifier. vibrantlabs.com/research/eco…
1
2
11
1,877
Vibrant Labs retweeted
Evaluating web agents on the actual web is hard. @VibrantLabsAI did it right: live Shopify stores, deterministic verifiers, all running on @browserbase. Open for everyone to run now ↓
1/n For browser agents, a major bottleneck in evaluation is truthful scoring on the live web. A task is only as good as your ability to confirm the agent actually did it, on a real site whose state keeps moving and that the agent can potentially misreport. So we took matters into our own hands. Today, we're releasing Ecom Bench on @PrimeIntellect: 40 shopping tasks on real Shopify storefronts, each run in a live @browserbase browser and graded by a deterministic verifier. vibrantlabs.com/research/eco…
1
4
147
Vibrant Labs retweeted
Replying to @VibrantLabsAI
Less saturated benchmarks that showcase the gap in current model capabilities are necessary to continue pushing the frontier Love this from @VibrantLabsAI with browserenv, excited to see more benchmarks in different verticals!
2
3
108
1/n For browser agents, a major bottleneck in evaluation is truthful scoring on the live web. A task is only as good as your ability to confirm the agent actually did it, on a real site whose state keeps moving and that the agent can potentially misreport. So we took matters into our own hands. Today, we're releasing Ecom Bench on @PrimeIntellect: 40 shopping tasks on real Shopify storefronts, each run in a live @browserbase browser and graded by a deterministic verifier. vibrantlabs.com/research/eco…
7
10
33
9,570
13/ Ecom Bench is available now for training and evals: prime env install ecom-bench DOM mode runs the full 40-task set, and grounding="cua" flips the same tasks to pixel grounding for the comparison above. @PrimeIntellect (built on @browserbase): app.primeintellect.ai/dashbo… Full writeup: vibrantlabs.com/research/eco…
1
1
6
182
14/ We build these environments so that evaluating and training web agents on real, verifiable tasks stays open to any lab, regardless of who owns the harness. If your team is interested in the post-training data pipelines we work on at @VibrantLabsAI, you can reach us at team@vibrantlabs.com!
1
6
168