Scale AI retweeted
An honor to meet with His Highness the Crown Prince of Kuwait to discuss how @Scale_AI can support Kuwait’s digital transformation and develop Kuwaiti AI talent.
ممثل حضرة صاحب السمو الشيخ مشعل الأحمد الجابر الصباح أمير البلاد سمو ولي العهد الشيخ صباح خالد الحمد الصباح يبحث مع الرئيس التنفيذي لشركة (Scale AI) العالمية فرانسيس ديسوزا آفاق التعاون المشترك في مجالات تطبيقات الذكاء الاصطناعي والبيانات الضخمة والتحول الرقمي - سموه استقبله والوفد المرافق له في مقر وفد دولة الكويت الدائم لدى الأمم المتحدة على هامش أعمال الدورة الـ81 للجمعية العامة في نيويورك - جرى استعراض سبل الاستفادة من حلول التكنولوجيا المتقدمة في دعم وتطوير القطاعات الحكومية والاقتصادية الحيوية في دولة الكويت وتأهيل الكوادر الوطنية وتنمية قدراتها - #الكويت #السعودية #فلسطين #القدس #غزه #اليمن #لبنان #السودان #سوريا #مصر #العراق #الأردن #الضفة_الغربية #حماس #ترمب #ترامب #أمريكا #تركيا #الرياض #ايران #ليبيا #صباح_الخير #يوم_الجمعة #ليلة_الجمعة #ساعة_استجابة #جمعة_مباركة #مساء_الخير #الخميس_الونيس #مضيق_هرمز #خليجي٢٧ #خليجي27 دول الخليج الدول العربيه الولايات المتحده الشرق الاوسط مجلس الامن الامم المتحدة الرئيس الفنزويلي مدريد صوت المطر صوت الرعد مطر الليل أبستين
5
26
2,877
Scale AI retweeted
The nation that can measure the AI frontier will be the one that steers it. Our future should not be steered by speculative warnings. We have to be driven by evidence. That was the clearest takeaway from this week's AI conversations at the UN General Assembly. Governments need to get serious about testing frontier AI, and nowhere is that more urgent than in Washington. Getting serious means three things: → Clear responsibility for testing across agencies → Access to models before they're deployed → Funding for the experts and tools to do the work Scale AI's cyber research shows what testing can reveal. We've seen AI agents carry out an attacker's instructions while still finishing the user's task, so the user never knew anything went wrong. We've also seen models refuse to help people defend their own systems. Both are failures. Real testing has to measure the harm AI can cause and the help it fails to give. Government shouldn't have to rely solely on AI developers to explain these failures. It needs its own testing capacity, with the expertise to ask hard questions and the tools to answer them. Scale has worked with the U.S. government for years to test the risks and capabilities of AI models, and we're accelerating that work in the months ahead. Here's what we think America should do next: scale.com/blog/steer-the-ai-…
3
9
34
1,428
We’re partnering with @googlecloud to make it easier for enterprises to put AI to work. Our joint reference architecture is a blueprint bringing the Scale GenAI Portfolio to Google Cloud, with Gemini Enterprise integration. Excited to give builders a proven foundation to put AI to work faster. scale.com/blog/deploying-ent…
5
1
52
8,149
Today, we released SWE-Bench Pro V2, a refreshed public split of SWE-Bench Pro. The frontier keeps moving, and the standard for measuring it should move with it. We’ll keep strengthening our leaderboards, incorporating feedback, and building more rigorous evals for what comes next.
SWE-Bench Pro V2 is live. What’s new: 🧵
14
8
137
46,285
AI safety doesn’t translate one-to-one. We partnered with the Korea AI Safety Institute to develop ROK-FORTRESS, our first benchmark together, testing how language and geopolitical context affect AI safety. Across 14 frontier models, we found that translation alone can miss meaningful differences in model behavior. scale.com/blog/korea-ai-safe…
5
6
44
14,628
Scale AI retweeted
This week I joined @RoyalFamily, govt ministers and leaders to discuss how AI can best serve the public good. AI can expand opportunity and solve some of our toughest challenges. It’s essential that we pursue that potential responsibly, in a way that benefits people broadly. Rigorous, independent testing - like we do at @Scale_AI - is an essential step to earning public trust. Governments and independent evaluators need to work together to deeply and continually assess AI's capabilities, opportunities, and risks. That allows us to make sound decisions about its development and use. Conversations like this one are a step in the right direction.
“I really am enormously grateful for the presence in this room of many of our greatest technological minds, and indeed, for the thoughtful stewardship demonstrated by those gathered here today as you consider the fundamental principles that should – and must – guide the development of Artificial Intelligence in the years ahead.” Today at Dumfries House, The King convened AI pioneers, Government representatives and thought leaders to discuss how AI can be developed and deployed in ways that benefit society.
2
4
30
3,785
Cue the makeover montage.
Still shipping nonstop. Here’s our new look
4
1
34
23,124
Scale AI retweeted
We’re moving to a multi-model world and that’s a good thing. Enterprises will use multiple models, optimized for different scenarios. Thanks for having me on @andrewrsorkin and @BeckyQuick
Scale AI CEO @fdesouza sees the future of AI including more models, and particularly more narrowly-tailored technology. "We want models that are specific for different scenarios," he says. cnb.cx/4xBuTOX
1
6
43
6,857
Scale AI retweeted
How do you know if an agent is enterprise-ready? NEWS: Today Scale AI is introducing READY, our Reliable Enterprise Agent Deployment benchmarks, to answer that question. READY measures agents on reliability, the human oversight required to hit a defined reliability target, and the resulting cost, together. labs.scale.com/blog/ready
2
17
52
5,094
We looked into how organizations are actually making AI work, and found only a small fraction are seeing measurable impact. 8 moves came out of it, and these are the patterns we consistently see behind real AI impact. Check out the full report: scale.com/six-percent
7
1
34
11,944
Scale AI retweeted
📣Call for contributions + co-authorship! RSI Bench is our ongoing effort to evaluate whether AI agents can develop the capabilities required to advance AI R&D. Got a frontier AI research problem on your mind? Submit it to RSI Bench, and set the standard the whole industry uses to measure progress! Contributors receive: → Co-authorship on the RSI Bench research paper → $2,000 per accepted task for the initial 50 tasks → Modal compute credits to build + iterate → Access to the RSI Bench research community Register below for full requirements.
Launching rsi-benchmark.com: The work of AI R&D has always belonged to humans. For the first time, though, it no longer seems certain that it always will. Recursive self-improvement is within a line of sight. It may still be far, but it is close enough that we should start measuring it.
13
30
297
72,472
Lots of talk over the past week about American data companies selling training data to Chinese AI labs, and why that leaves the US behind. Up front: Scale doesn't do this work, and we've turned down revenue over it. A lot of it comes down to "isn't this just labeling?" Fair question, and the answer is no, but I want to explain why… Frontier post-training data has very little in common with annotation at this point. It looks more like curriculum design for training models. You're deciding which problems a model should struggle with, how a reasoning trace should be structured, the difference between a right solution and a lucky one, and what to reward. Tasks, RL environments and Verifiers are that same judgment but in executable form. Which means what’s being sold is a set of research decisions, already made and already validated. That's also why it's not a small business for anyone involved, and it’s why the few companies like us in the space are doing well. We build frontier data pipelines from our research and build OTS datasets to match because buyers want ready-made inventory and capability gains. Not coincidentally, Chinese labs' biggest jumps are in exactly the domains where sellable environments exist: agentic coding, math, tool use, etc. A few things about this that I don't think got much attention: ➡️ Compute controls assume compute is the binding constraint. But RL post-training is bound by reward signal, not just FLOPs - a lab that buys verified environments doesn't waste compute on failed exploration, so it reaches the same capability on far fewer chips. That's a substitution the export regime doesn't account for. ➡️ Functionally, this hands labs distillation on easy mode, minus the legal exposure. The verifier confirms which teacher outputs are actually right, so you distill from verified outputs. And dense reward signals cut the compute RL burns on failed exploration. ➡️ The judgment being sold is American expert judgment. The people writing these rubrics, the domain PhDs and engineers and professionals, generally have no visibility into who the end buyer is. They think they're contributing to work they'd endorse. ➡️ As recently reported some of the same vendors hold US government contracts. That's not a hypothetical conflict, and it should be getting more scrutiny. We shouldn’t assume bad intent from anyone. This is a supply chain that grew faster than our ability to think about what the rules should be. If you're buying data, "what am I getting" is only half the diligence. Where else that pipeline goes is the other half, particularly from companies that talk publicly about American AI leadership.
2
8
72
4,303
American companies should not be selling data to Chinese AI labs. @scale_AI doesn’t do this work, and we have turned down revenue because of it. The companies that do are undermining American AI leadership and risking national security.
The AI arms race isn’t just being fought over Nvidia chips. Chinese labs are matching OpenAI and Anthropic by purchasing the exact same data from US vendors like @mercor and Surge AI (who also work with the US federal govt) forbes.com/sites/annatong/20…
13
27
220
230,906
Scale AI retweeted
Big News: I’m joining @scale_AI as CEO, starting August 10. Scale sits at a rare intersection, working with the top AI labs to push the frontier while helping enterprises and governments actually deploy AI they can verify and trust. I’m excited to lead a company with a mission to develop reliable AI systems for the most important decisions. More to come.
52
36
656
345,353
Proud to support this on behalf of @scale_AI Open models are critical to American AI leadership and to building reliable AI. Great to join so many across the industry in signing on.
3
13
88
7,937
Allegra Larche, Senior Manager of Enterprise Forward Deployed Engineering at Scale, on how our FDE teams are building reliable AI for industries that matter most.
9
12
98
20,064