Assistant Professor @WisconsinCS. Chief scientist @SnorkelAI. Working on machine learning & information theory.

Madison, WI
Fred Sala retweeted
Will appear at Neurips. Its impact so far tho is the real validation.
Once again, SlopCodeBench is at the top of hacker news:
2
20
1,855
can an AI generate novel philosophical ideas? there is no "unit test" for philosophy - and we need new methods to provide rigorous evaluation for AI performance in philosophy we are thrilled to support PhilosophyBench and partner with @michaelcheng76 @StanfordAILab via @SnorkelAI's Open Benchmarks Grants. if you're a philosopher and would like to participate, please reach out!
1/ Introducing PhilosophyBench from @StanfordAILab @StanfordHCI, the first independent, large-scale benchmark for evaluating AI’s philosophical capabilities. philosophybench.org
2
12
35
2,463
Will be at NeurIPS '26! Congratulations to @jiayuwang111 and @ming5_alvin
Agent skills are powerful. But can they help us with agent orchestration and routing? We show they can! 🎼 Introducing SkillOrchestra - a skill-aware alternative to RL-based agent orchestration: 💰 Higher efficiency 📉 Orders-of-magnitude lower training cost 🧠 More interpretable decisions TL;DR: Stop training bigger orchestrators and start modeling skills. Paper: arxiv.org/abs/2602.19672 [1/n]
1
6
48
2,190
Fred Sala retweeted
fighting code slop at home and abroad
13
3
79
6,230
How do we safely build and supervise AI at the capability frontier? Build environments that push models beyond what they can do now, give experts tools to inject insights into pipelines, and build the RSI engine. Couldn’t be more excited for the next chapter at @SnorkelAI!
I'm excited to announce @SnorkelAI's $350M Series E at $3.5B, led by @insightpartners and @S32_VC. We've grown 18x+ in the last 12 months since launching our Data-as-a-Service offering, passing $375M ARRR this week. As AI advances to superhuman capabilities, AI data & environment development must advance with it - and basic staffing and crowdsourcing approaches are not enough. AI progress now requires deep research and technology work that combines human expertise with specialized AI in compounding ways. @SnorkelAI is building the RSI data engine and frontier data lab for this next phase. We're honored to have the support of existing investors Addition, @lightspeedvp, @GreylockVC, @GVteam, P7, Factory, @WellsFargo, Walden Catalyst Ventures, and new investors @ThirdPointLLC, @MarchCPs, @BlumbergCapital, @AllegisCapital, @Frontlinevc, and @standard_vc. – @SnorkelAI started as a research project a decade ago at @StanfordAILab. Our thesis was simple: AI progress would become increasingly data-centric – and therefore data development should be studied as a true research and technology problem, not just a staffing and crowdsourcing one. Today, as AI capabilities verge on superhuman, building the data and environments to safely measure and train AI is becoming too hard for even the smartest human experts to do alone. Only humans and AI agents, collaborating together in compounding ways, can meet the accelerating needs of the frontier, and keep humans in the driver’s seat of AI progress for decades to come. At @SnorkelAI, we are building the data lab to define the shape of this new “Data 2.0” frontier, and the new paradigms of human-computer interaction needed to advance it. Our key focus is building the RSI engine for data, where specialized AI models accelerate and improve human expert output, and in turn, scaled human supervision is used to continuously evaluate and improve these models – creating a powerful compounding loop to keep pace with an accelerating RSI frontier. With this round of funding, we are also doubling down on our commitments to support data development for open benchmarking and evaluation (more news here soon!); an increasingly diverse ecosystem of general and specialized intelligence; and a path to safe, well-aligned AI built on robust training and evaluation data. Data development will guide and drive the next stages of AI – and must do so in a human-centric, AI accelerated, open, diverse, and safe way. We are excited to support this mission in the next decade of research ahead at @SnorkelAI. More thoughts here: snorkel.ai/blog/data-2-0-and…
3
10
62
2,493
Fred Sala retweeted
Congrats to all, including our own @CaiLinrong !
Congrats to graduate students Linrong Cai, Peter Halmos, Nicolaas Kaashoek, Stephen Newman and Anchengchen Zhou on being named 2027 @SiebelScholars! 🎉 The fellowships are awarded to students for outstanding academic achievement and leadership. bit.ly/4xW81Jv
1
3
21
7,291
Evaluation isn't just accuracy—it's latency, throughput, cost, and whether we can trace a latent decision path the way we would an if-else. As it moves online and gets more iterative, the drawbacks of LLM judges are hard to ignore. Our earlier work, PAJAMA, shares the same spirit
3
3
14
1,410
Seeing a lot of the benchmark folks respond to this… some say the findings resonate with real problems worth addressing… some thankful for the audit findings… and some posing real questions about who audits the auditors… Im going to pull some responses and start a thread:
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
15
2
12
1,544
You can slop-o-write a paper. And you can submit it. And at a flip of an AI coin can even get it in. But can you write one that’s actually fun, cool, a little nuts? We need more of them.
45k submissions for @iclr_conf so far 💀💀
9
11
267
27,092
prediction: more labs will self-identify as data research labs the overton window has shifted - data was perceived as "plumbing," but now it's sayable (including at frontier labs) that the data is the research!
5
10
78
4,335
Fred Sala retweeted
I like your phrasing @ConorBronsdon better than mine :). Main point though is that to deploy and improve in real production settings, you must be able to measure - and measurement is all about good (benchmark) data!
Every team says their agent works. Almost none can say what "works" means in a number. @ajratner of @SnorkelAI on defining a capability before you try to measure it:
1
7
17
1,493
Fred Sala retweeted
Excited for our @SnorkelAI Frontier Data Summit! Most importantly to this moment (when robust, independent evaluation is appropriately top of mind): the core of this summit is a poster session with leading open benchmark developers, sponsored by @Snorkel Open Benchmarks Grants!
Frontier Data Summit: Oct 8, SF. @fchollet, @sanmikoyejo, @nikogrupen, @ajratner, @alexgshaw, @StevenDillmann, and others onstage plus 25+ posters with the teams behind Agent's Last Exam, OSWorld 2.0, Terminal Bench Science, Terminal Bench, CollusionBench, PostTrainBench, T² Scaling Laws, SlopCodeBench, LONGRUN, SREGym, BizBench, and more.
2
17
46
3,796
Fred Sala retweeted
astra + fable 5.1 on slopcodebench nitter.net/i/broadcasts/1pKdRDoWD…
2
7
67
11,628
Claude Fable 5.1 now tops the Senior SWE-Bench leaderboard, edging out Fable 5 on the pass^3 tie-breaker (more consistently shipping quality code). Interestingly, its best effort tier is Medium. Combined with cheaper cache reads, it matches Fable 5's perf at ~50% of the cost.
3
7
33
6,250
Fred Sala retweeted
Terminal-Bench-Science 0.1 is the #1 featured benchmark on @AnthropicAI’s new Claude Fable release 🚀 terminal-bench-science.ai
Replying to @claudeai
Across our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.
13
20
136
8,965
Fred Sala retweeted
It would be great if there were a benchmark that measured exactly this.
We’re sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe: 1. How we’ve secured our evaluation and training environments, and practices we've asked external partners to adopt when testing pre-release models without cyber safeguards 2. An update on our alignment assessment 3. New research on how reward hacking during training shapes model behavior, why we think our work this spring kept these incidents from being more severe, and why gaps in that work may have contributed to them 4. How we hardened our security practices earlier this year to prepare for Mythos-class models Read more: anthropic.com/news/improving…
5
13
141
59,645
Fred Sala retweeted
Terminal-Bench-Science first release v0.1 is out! It's been a fun challenge in my @SnorkelAI internship to shape TB Science with @vincentsunnchen and @StevenDillmann! Very excited to drive this forward with Cornell AI for Science folks at our NSF institute @MaterialsAI! I'm bullish that TB Science will define the next generation of coding agents---a "Claude Code for Science"---we'll likely see frontier labs using TB Science as a compass to drive scientific discovery. Benchmark is hard for frontier LLMs at release. But looking beyond the task difficulty leaderboard, TB Science is good at the basics of benchmark building: - Task breadth and realism. Real computational workflows across physical, life, and mathetical sciences. The hardest part in working on TB Science this summer has been to get STEM scientists to work on this. There are very few scientists with the high domain expertise---with LEAN, FORTRAN, pymatgen, opencv etc. code---to build good tasks. Highly skilled grad students, postdocs, and faculty at the best universities worldwide. I spent a lot of 'office hours' helping scientists onboard and scour through agent evaluations. It's been a steep learning curve on speaking a vocab that both programmers and STEM experts understand. - Task sandboxing. Last year at Cornell we adopted @harborframework to reproducibly evaluate agents---it's amazing how well it scales. Would love to see it adopted across agent eval + training research infra. - Task instructions, oracle solutions, deterministic verifiers. Bad benchmark tasks underspecify instructions "write me a PDE solver" and overspecify verifiers to only allow the human-written solution. @StevenDillmann and I have done a lot of back and forth on defining the fine line. Something I've learned is that, for science tasks, it's really up to the scientist to specify what they think is reasonable. We should defer to their judgement as they'll end up using a good "Claude Code for Science"🙏 Coding agents have become insanely good at coding as a "skill" but can they understand scientific knowledge as "memory"? This year will be pretty exciting for scientific progress w/ AI!
TB-Science advances both (A) how we benchmark science and (B) the science of benchmarking 😀 - tasks built from real computational workflows across the natural sciences (e.g. life, physical, earth) - an extremely high quality bar with multi-stage expert review/adjudication in partnership with domain experts Congratulations to @StevenDillmann for leading this awesome work - we're honored to partner @SnorkelAI !
6
8
34
1,816
TB-Science advances both (A) how we benchmark science and (B) the science of benchmarking 😀 - tasks built from real computational workflows across the natural sciences (e.g. life, physical, earth) - an extremely high quality bar with multi-stage expert review/adjudication in partnership with domain experts Congratulations to @StevenDillmann for leading this awesome work - we're honored to partner @SnorkelAI !
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains. An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%. 1/n 👇
2
8
41
4,834