welcome to the lab. from the researchers at @scale_AI

Scale Labs retweeted
Glad to share that 3 papers I contributed to at @ScaleAILabs have been accepted to NeurIPS 2026, spanning post-training and agent evaluation 🎉 🔍 Reward Hacking in Rubric-Based Reinforcement Learning We conducted a systematic study of reward hacking in rubric-based RL, distinguishing failures caused by weak verifiers from failures in rubric design, and identified an interesting stopping criterion. arxiv.org/abs/2605.12474 🧠 Not Every Rubric Teaches Equally We introduce a policy-aware curriculum for rubric-based RL that emphasizes the criteria most useful for learning at each stage of training. arxiv.org/abs/2605.20164 🛠️ SWE Atlas We introduce a benchmark for evaluating coding agents beyond issue resolution, covering codebase Q&A, test writing, and refactoring. arxiv.org/abs/2605.08366 Really proud of the team and the work behind these. Looking forward to sharing more and catching up with everyone at NeurIPS in Sydney!
5
7
36
1,330
Scale Labs retweeted
Excited to share three @ScaleAILabs papers accepted to NeurIPS 2026 in the Evals and Datasets track! SWE Atlas: Evaluating coding agents beyond issue resolution, across codebase Q&A, test writing, and refactoring—with attention to both correctness and engineering quality. arxiv.org/abs/2605.08366v1 MCP-Atlas: Evaluating agents’ ability to discover and use tools across real MCP servers to complete realistic, multi-step tasks. arxiv.org/abs/2602.00933 Beyond Truthfulness (MASK): Separating honesty from accuracy by testing whether models contradict their own stated beliefs when pressured to lie. arxiv.org/abs/2503.03750v2 Agent evals and benchmarks are important as capabilities increase, safety becomes critical, and reliability remains an open question. Excited to continue this work with our customers, and grateful to all my co-authors and collaborators for the hard work behind these benchmarks!
7
3
25
783
Scale Labs retweeted
"MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers" has been accepted to NeurIPS 2026!! I had the privilege to be co-first author on the paper and to work with virtually all of the frontier labs to enable them to report numbers on our benchmark!
6
7
29
2,520
Scale Labs retweeted
🎉 Excited to share that “Reward Hacking in Rubric-Based Reinforcement Learning” has been accepted to #NeurIPS2026! Reward hacking in post-training is a particularly interesting and important problem to work on. We study how models can learn to game rubric-based rewards during post-training, satisfying the rubric without necessarily improving the broader response quality. We characterize these failures at the criterion level and show that stronger verifiers help, but don’t fully eliminate reward hacking. Link: arxiv.org/abs/2605.12474 Check out some of our recent post-training research projects at @ScaleAILabs: RGSD: arxiv.org/html/2606.12507 POW3R: arxiv.org/abs/2605.20164
4
7
26
826
Scale Labs retweeted
🎉 Excited that “Not Every Rubric Teaches Equally” is accepted to NeurIPS 2026! 🎉 POW3R turns static rubric rewards into policy-aware training signals, emphasizing the criteria that actually separate good and bad rollouts for the current state of the model. Link - arxiv.org/pdf/2605.20164 Our team at @ScaleAILabs has been cooking on some really exciting RL/post-training research lately. Check it out 👀 RGSD - arxiv.org/abs/2606.12507 Reward Hacking with Rubrics - arxiv.org/abs/2605.12474
4
7
50
2,241
We're releasing HLE-Diamond: a refined version of Humanity's Last Exam, built with @CAIS. A year of review and community feedback went into refining this subset to make it more reliable for measuring frontier models. Top model tested 60.6% overall. We expect HLE-Diamond to carry signal for the next 6-12 months.
4
15
96
6,751
Over a year ago, we released SWE-Bench Pro. We refreshed the benchmark today to improve its quality based on community feedback. Our update is meant to ensure the benchmark is clean and accurately measures model capabilities on SWE-Bench Pro tasks. In our update, we observe a ~20% performance drop on our held out private example set of 272 tasks compared to the public leaderboard. We further identify a HARD subset of 51 tasks that is discriminative of the frontier model performances from the rest of the pack. Refreshing benchmarks, and sharing what we learn along the way, does more to advance model evaluation than deprecating them outright. That's why we're sharing this updated version: it addresses issues we identified over time, and we want the broader community to benefit from those findings.
SWE-Bench Pro V2 is live. What’s new: 🧵
3
2
30
4,070
SWE-Bench Pro V2 is live. What’s new: 🧵
19
8
153
203,381
211 tasks were improved to have better dependency support for running OSS harnesses. We removed 89 invalid tasks, moving down to 642 tasks across 11 repos.
2
23
9,131
600+ proposals in 🤯 We're reaching out to top contributors to start building in our public repo. We’ve also welcomed @SchmidhuberAI, a pioneer of RSI, as a senior advisor to RSI Bench. New blog on our setup and verification pipeline: rsi-benchmark.com/blog/verif…
Launching rsi-benchmark.com: The work of AI R&D has always belonged to humans. For the first time, though, it no longer seems certain that it always will. Recursive self-improvement is within a line of sight. It may still be far, but it is close enough that we should start measuring it.
4
9
122
91,272
Scale Labs retweeted
AI safety doesn’t translate one-to-one. We partnered with the Korea AI Safety Institute to develop ROK-FORTRESS, our first benchmark together, testing how language and geopolitical context affect AI safety. Across 14 frontier models, we found that translation alone can miss meaningful differences in model behavior. scale.com/blog/korea-ai-safe…
5
6
44
14,626
New research from Scale Labs: SteerDuplex and SteerBench measure how well voice models follow instructions for how a response should sound and be delivered. +44.5 points on audio steering over the best open baseline. Pause barge-ins cut from 26.5% to 9%. Plus a sharp result on why timing rewards alone collapse into silence. ⬇️
Full-duplex voice models can listen and speak at the same time, but can they change how they speak when you ask them to? Introducing SteerDuplex + SteerBench: the first benchmark that measures steerable full-duplex dialogue. 65.1% audio-steering APR vs 20.6% for the best open baseline. 🧵
1
2
16
1,907
GPT-6 Astra is a step function on DrugDiscoveryBench at 68.7%, our benchmark for early-stage drug discovery computational workflows. Muse Spark 1.3 and Opus 5 follow. @TuXinming, @KavanaghOscar and I will talk about what it measures tomorrow at 10 am PT! Register below ⬇️
1
9
37
6,135
Scale Labs retweeted
New models this week are much more codebase-aware: they check and break their own tests, and like senior/experienced engineers, they even care about code hygiene. As part of that, I want to shout out some major gains we saw on our SWE Atlas leaderboard @ScaleAILabs. @AnthropicAI is tied for 1st on all three SWE Atlas leaderboards, with the latest Opus and Fable models clustered at the top. But Fable 5.1 feels like a qualitatively different model: more consistent, more efficient, and noticeably more codebase-aware. In Test Writing, it's the first model we've seen that actively tries to mutate code and break its own tests to check for robustness. In Refactoring, it consistently cleans up old leftover code without being asked. On SWE Atlas, no model has cracked 70 percent yet. There's a lot of work these agents still can't do.
1
8
36
1,641
An AI agent can top a benchmark and still be undeployable. Introducing READY (Reliable Enterprise Agent Deployment), the first evaluation framework for qualifying AI agents for deployment on real enterprise workflows. We’re inviting the research community to build it with us. The testbed is open for contributions today (link below).
2
7
47
2,373
Two things this makes visible that benchmarks can't: → Confidence granularity is a deployment property. Routing means picking a cutoff. You can only be as fine-grained as the scores the system reports. Some offered a handful of distinct values; the best offered 36. → You can tune a routing policy until it looks great on the cases you've already seen. Deployment is a bet that the same reliability holds on cases you haven't. READY's qualification step is built around that gap.
1
3
241