Supercharge Your LLM Application Evaluations 🚀 Github: github.com/vibrantlabsai/rag… Discord: discord.gg/5djav8GGNZ

We're hosting our first public paper club session this week. This is part of a series we have been doing for while internally and now want invite more people who are researching on autoscaling RL envs. this week we'll be covering 2 papers luma.com/qb05spn8
1
1
5
319
1/n For browser agents, a major bottleneck in evaluation is truthful scoring on the live web. A task is only as good as your ability to confirm the agent actually did it, on a real site whose state keeps moving and that the agent can potentially misreport. So we took matters into our own hands. Today, we're releasing Ecom Bench on @PrimeIntellect: 40 shopping tasks on real Shopify storefronts, each run in a live @browserbase browser and graded by a deterministic verifier. vibrantlabs.com/research/eco…
7
10
33
9,570
For years the goal has been to replicate the human brain. We've gotten remarkably far. But the brain was never designed. It was grown — by a simple process, running in a rich enough world, for a very long time.
1
3
4
217
ragas retweeted
1/n We at @VibrantLabsAI study the literature, old and new, on autoscaling RL environments so we can build better ones. This week in Paper Club, we went a level deeper: not just what makes a good environment, but what makes a good learning environment, and how to verify that training is working at all. We covered: - AI-GAs - Agent Psychometrics - The Universal Verifier
4
4
17
1,857
ragas retweeted
(1/n) One of the big challenges on our roadmap at @VibrantLabsAI is scaling agent benchmarks while maintaining a strong reward signal. Good verifiers are the bedrock of usable agent benchmarks. Bad verifiers can inflate model failure rates, and they often hide actual capability gaps. This week, in Paper Club, we covered SWE-bench Verified and OSWorld-Verified, which both focus on this topic. We also covered ComputerRL and BenchGuard for some of the techniques they proposed.
1
3
17
1,227
ragas retweeted
Today, let’s talk about Helix, our internal system at @VibrantLabsAI for scaling task and verifier mining. We’re also providing a deeper look into how we’re using Helix to generate post-training data for @yutori_ai 's latest computer-use model, n1.5. Our two most notable unlocks were: 1. An autonomous task miner that suggests tasks within a specified pass rate band for the target model. We already use this miner for many other use cases beyond n1.5’s e-commerce needs. 2. A rigorous QA pipeline that allows a human to quickly review tasks with a visual interface (similar to reviewing PRs in GitHub) and course-correct accordingly. If you’re interested in scalable task & environment generation let us know what you think. We’ve already brought this to a variety of domains outside of computer-use agents. vibrantlabs.com/blog/mining-…
5
15
1,760
ragas retweeted
(1/n) A month ago, researchers at @LTIatCMU @CarnegieMellon published Gym-Anything, a paper that directly addresses our core thesis (@VibrantLabsAI). We believe the biggest limiting factor in agent advancement is generating post-training data (envs, tasks, verifiers) at scale.
1
3
21
2,118
ragas retweeted
(1/n) I believe that scaling post-training data autonomously will be the next big unlock in frontier model performance, and I’m willing to bet my company on that prediction. That’s why it’s so important that we at @VibrantLabsAI spend time reviewing SOTA research from other teams that are autoscaling RL environments.
2
2
22
1,519
ragas retweeted
(1/n) Today, we’re releasing Cloning Bench. Labs are paying 6-7 figures for clones of web apps to do web/computer use-based RL training. At @VibrantLabsAI , our fundamental goal is to automate the creation of RL environments. For web/CUAs, one way that we do that is by using coding agents and custom harness to automatically generated the simulation environment. We tested Codex, Gemini, Claude Code, and GLM using our harness on their ability to recreate a Slack workspace and benchmarked their performances. We have published our methods, results and analysis here today: vibrantlabs.com/blog/cloning…
7
12
141
13,223
ragas retweeted
There are 3 elements to improving models: 1) Architecture 2) Compute 3) Data No one is changing (1), (2) is actively being solved by the compute giants. Now what’s left is (3), which has effectively become 2026’s “pickaxes in a gold rush.” Today, the choke point is fully human-created data. We at @VibrantLabsAI believe AGI will not be achieved by human data alone, so we’re laser-focused on synthesizing as much as possible to advance models to the next frontier.
1
4
8
701
ragas retweeted
Last week, we did an internal deep dive into enterprise environments/benchmarks like τ²-𝐁𝐞𝐧𝐜𝐡 and 𝐂𝐨𝐫𝐞𝐂𝐫𝐚𝐟𝐭. This type of high-fidelity RL env is becoming increasingly popular as frontier labs push their models into more and more agentic capabilities.
1
2
9
736
ragas retweeted
Writing prompts by hand is just guessing. Let DSPy optimizer find the best one for you. Here is a full report that: > build your traditional rag with @OpenAI and @qdrant_engine > optimizes prompts with @DSPyOSS MIPROv2 optimizer > trace everything with @weave_wb > evaluate the baseline and optimized RAG system using @ragas_io 🔗 report: wandb.ai/ai-team-articles/ds…
2
6
216
ragas retweeted
PA Bench - our first public benchmark on multi-tab web agents in on first page of HN now 🔥
4
7
487
ragas retweeted
How good are coding agents at cloning web apps? Check out how Claude Code (+ our harness) clones a Slack workspace completely from scratch using only recordings of the real version. 🔥
1
2
7
495
The top frontier labs are paying tiny startups millions of dollars for RL environments: newsletter.semianalysis.com/… Since most experts agree that RL post-training is causing the next wave of major model advancements, the data budget for these labs has grown more than anyone could have predicted. Browser use is a major vertical, and clones of popular consumer/enterprise websites (think: Amazon, Salesforce, Epic, etc.) are in high demand. Many companies in this space are using overseas human labor to build these environments. At Vibrant Labs, we’re instead taking the approach of automating the creation of post-training data and environments. We built out a harness that uses coding agents to clone any given web application given screen recordings of workflows we want to train on. So with all of the hype around building clones of websites, we decided to do a benchmark. Later this week, we will release Cloning Bench, a benchmark that utilizes our harness and state-of-the-art coding agents (Codex, Claude Code, etc.) to benchmark how well they perform at web cloning tasks. Stay tuned for more.
2
6
704
ragas retweeted
Browser agents are becoming a hit at the consumer level, as most ad-hoc tasks people do daily through browsers can now be automated. But are the models actually good at doing any of it reliably? To evaluate this, we built PABench - a personal assistant benchmark requiring 2+ tabs to complete real-world tasks. (1/n)
1
2
6
517
We are releasing our first public benchmark: PA Bench . PA Bench is a first-of-its-kind web/computer-use agent benchmark focused on the types of tasks normally done by a Personal Assistant (especially multi-tab, long-horizon workflows).
1
3
6
2,045
OS-Genesis: 1/ Most agent scaling is throttled by the cost of human time. OS-Genesis took a much more scalable path by using Reverse Task Synthesis. Instead of recording a user completing a task, they started from a terminal state and worked backwards to hypothesize the intent.
1
2
3
191
You are about to deploy your first RAG application. It works fine on your local environment - but you aren't sure if it will perform the same once the real users start using it. That's a sign of a weak link. Now, some obvious questions follow - How do we ensure consistent performance in production? What metrics should we track? How do we verify whether the RAG system’s responses are actually good? The answer is simple: you need a proper evaluation framework to systematically measure, validate, and improve your RAG application. The image below, you can see that we are using the RAGA's framework to make sure our RAG system/application produces contextually relevant, high-quality, and ethical responses. As you can see the RAG pipeline, only the responses are not evaluated, the RAGAS framework evaluates the system at multiple stages. Retriever quality is assessed using contextual precision, recall, and relevance by comparing the query with the retrieved contexts. Generator quality is evaluated using answer relevancy and faithfulness by grounding the generated response against the retrieved documents, enabling holistic RAG evaluation. So next time while deploying your RAG application, make sure you use any of the evaluation frameworks to safeguard your application with proper responses. I have created a simple hands-on guide to evaluating RAG applications in minutes using RAGAs framework. The link to the guide is in the comments. This is my hands-on guide to evaluating RAG applications in minutes using @ragas_io - youtu.be/-69Fx8F9ma4
1
2
126
LangChain Community Spotlight: 🧠 HMLR: Long-Term Memory for AI Agents Made by the LangChain Community HMLR adds long-term memory to AI agents via LangGraph drop-in. Perfect RAGAS scores on hardest benchmarks using GPT-4.1-mini, maintains context across days/weeks without token bloat. 📦 pip install hmlr 🔗 github.com/Sean-V-Dev/HMLR-A…
2
16
130
11,019