The top frontier labs are paying tiny startups millions of dollars for RL environments:
newsletter.semianalysis.com/…
Since most experts agree that RL post-training is causing the next wave of major model advancements, the data budget for these labs has grown more than anyone could have predicted. Browser use is a major vertical, and clones of popular consumer/enterprise websites (think: Amazon, Salesforce, Epic, etc.) are in high demand.
Many companies in this space are using overseas human labor to build these environments. At Vibrant Labs, we’re instead taking the approach of automating the creation of post-training data and environments. We built out a harness that uses coding agents to clone any given web application given screen recordings of workflows we want to train on.
So with all of the hype around building clones of websites, we decided to do a benchmark. Later this week, we will release Cloning Bench, a benchmark that utilizes our harness and state-of-the-art coding agents (Codex, Claude Code, etc.) to benchmark how well they perform at web cloning tasks. Stay tuned for more.
ALT Cloning Bench coming soon