1
13
134
tom cunningham retweeted
[1/5] The 8 most valuable data points labs should share to help measure RSI: First, RSI would likely accelerate growth in AI capabilities. Thus, companies should report performance on diverse benchmarks for the latest internally deployed models.
9
40
227
44,864
What impact is AI having? I'm of the view that we should be focussing more on aggregate outputs, relatively less on inputs or intermediate proxies. - Inputs: time spent across different activities; money spent on tokens. - Intermediate proxies: working papers; lines of code, commits & pull requests; self-reported speedup. - Final outputs: total cyber exploits discovered; total math problems solved; algorithmic efficiency. The disadvantage of inputs & intermediate proxies is that they're very hard to interpret -- totally possible that they move, but outputs don't, and vice versa. Lines of code could explode without value changing; and vice versa. A disadvantage of final outputs is (1) it's harder to run experiments; (2) it's hard to attribute to whether the change is AI or not. But in some domains the trend-break is so large that it's *obviously* AI.
New post with Nate Rush: Have we seen an acceleration in discoveries? Many plots & some tentative conclusions: 1. Cyber: ⤴️ sharp acceleration 2. Math: ↗️ some acceleration 3. Algorithms: ➡️ no clear acceleration
5
9
81
18,467
BTW since we posted this a month ago there are big updates on curve-bending: - 1 Millenium problem solved, and another rumored solved. - A single NanoGPT contribution as big as 1 year of progress.
1
9
591
tom cunningham retweeted
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
459
431
3,634
651,912
tom cunningham retweeted
This a.m., I walked my two youngest kids (ages 7 and 4) to school. I held their little hands as we crossed streets, and gave my youngest a giant hug at his PreK4 classroom door. Then, after walking a block towards home, I opened this app to a terrifying stream of tweets. Of course, I have heard about "alignment" and "AI safety" and "x risk" concerns for years. I took these concerns at face value, especially when I heard them directly from folks at the labs working closest with the technology. And I was grateful so many smart people are working on them. And that's about it. (My work focuses on AI's impact on jobs.) But something has shifted, dramatically, for me in the last few weeks. Between the Hugging Face incident, the damning independent @METR_Evals reports that followed, and the alarming chorus of calls (pleas? shouts? SOS signals?) from inside the labs that we are careening toward potential catastrophe, all of this has made me feel, viscerally, how truly dangerous this moment is. Clearly I am not alone in this light bulb moment. It is one thing for someone in a tropical climate to try and conceptualize snow. It is another thing entirely to be knee deep in it, frozen down to your toes. Somehow hearing the details of the HF incident felt to me more like trudging through snow than theorizing about a possible blizzard; it really hit home. Which is why, standing a block from my children's elementary school in Capitol Hill this morning, I felt sick to my stomach reading my phone. I kept going back to @EvanHub's > 0.1 risk of AI killing us all. TEN PERCENT. *Greater* than ten percent. (To be fair to him, he was clear his worry isn't today's models, but what comes next.) As a parent, it is extraordinary the lengths I will go to keep my children safe against lethal risks with infinitesimal chances of happening. Bike to school? My children *must* wear a helmet. A national study found 2 bicycle-related deaths per million children <16 in states with helmet laws, vs 2.5 per million in states without. So: 0.000002. More than half the states passed a law over that. When I was pregnant, I religiously avoided soft cheese, hyper aware of the risks of listeria, even as @ProfEmilyOster's fabulous books walked me through the actual math. The joint FDA/FSIS risk assessment puts the odds of listeriosis at roughly one in 5 million per serving of soft cheese. 0.0000002. This morning, as I microwaved oatmeal for my kids' breakfast, I made sure I avoided a plastic bowl, to avoid some future cancer risks or something else we'll discover plastic causes. As a parent, I will go to extraordinary lengths to keep my kids' safe. Because that is literally my number one job as a parent. And yet, here I am, outside my kids' school this morning, scrolling through tweet after tweet about the AI arms race, the blistering pace of technological progress, the bee-line for trillion dollar IPOs, and our professed (confessed?) inability to make sure this exceedingly capable technology we are racing to create and unleash into the world won't escape our control and destroy us. Or, more bluntly, kill our children. To be sure, we are tied in knots trying to resolve the game theory dynamics with China, and there are all sorts of competitive pressures and unresolved tensions. But OBVIOUSLY, we need a proactive plan, with real coordination, effective regulation, and an ability to do what my 11-year-old bluntly stated as the clear solution if we are at risk of serious harm: just turn it off. We also need to figure out better ways of talking about this with the American public, outside of the AI-pilled Twitterverse. I just polled my college roommates, sisters, book club, mom friends, none of whom are anywhere close to the AI conversation. Suffice to say, this AI safety conversation has *not* entered the mainstream. Most people I pinged are not following it closely, beyond several listening to @kevinroose on the Daily a few days ago. The newspaper headlines don't speak to them, and they aren't clicking. It seems like a niche AI issue. Others said it all sounded hyperbolic, Y2K all over again. One friend said bluntly, "people get data centers and jobs. They like to get paid and have water." Even those who tuned it to the Daily episode found the issue hard to understand, and overwhelming to wrap their heads around. Terms like "superintelligence" and "alignment" are getting lost on people. Several people told me the paper clip example used everyday language people can actually understand. Another friend said she has never seen a clear explanation of *how* AI could kill us all. Here's my two cents. I think these issues are *exceedingly* important and hugely consequential to literally every person in this country. We need an actual plan, and we need policymakers to step up, asap. Right now, I hear only crickets in DC, where I live. For that to happen, this needs to be a kitchen table conversation, not just an x debate among AI insiders who speak in technical terms that leave normal people behind. This moment demands brilliant communicators -- journalists, experts, activists, concerned citizens, writers, researchers, filmmakers, creatives-- who can relate this to people in ways they can wrap their heads around, who can connect this to what matters in their lives (ahem, keeping their kids alive!), and who can lay out practical steps that policymakers can do to safeguard humanity. In other words, this can't just be talked about on @dwarkesh_sp, it needs to be everywhere, from the View to People Magazine. And it needs to be tackled here, in my town, in Washington, DC. One of the streets I crossed this morning with my kids on the way to school was East Capital. I told my kids to look left, and see if they could spot the capital building. There were too many trees this morning blocking the view.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
255
398
2,130
589,030
tom cunningham retweeted
I appreciate this as an initial step toward more transparent reporting on RSI. But there is still much further data we need to fully understand RSI. In particular, OAI disclosed some evidence about the inference compute usage, number of experiments/researcher, and to a lesser extent, how researchers interact with models of different capabilities (see the green boxes in the diagram). However, the evidence they provided is just on the input side (tokens, lines of code). We urgently need to know how these inputs get turned into algorithmic improvements and [probably internally deployed] AI capabilities. Even just a simple time series on those outcome variables would be extremely valuable. You can then know, as AI helps researchers do more AI R&D, how the corresponding algorithm efficiency grows (the blue arrow C -> A dot). Right now we are missing this key puzzle piece.
Today we're releasing data on models accelerating research at OpenAI. Recursive self-improvement could be the most important contributor to AI capabilities over the next few years, but by default it will only be seen inside a few frontier AI labs. Being transparent is more urgent than ever, so we can inform the public discussion on whether and how to pace model development. I ask other AI companies to do the same. openai.com/index/research-ac…
9
16
124
19,763
tom cunningham retweeted
We develop a mechanism design framework for AI alignment and control: arxiv.org/abs/2609.01595 It’s largely conceptual but we offer stylized applications to failure modes (sandbagging, alignment faking), safety practice (scalable oversight, peer prediction), and a way to think about the value of alignment, interpretability, capability, and control.
20
110
572
125,273
I think most domains look like this at the moment: the returns to expenditure on agents diminish much more quickly than the returns to expenditure on human labor: (1/n)
35
99
706
185,932
AI can one-shot a lot of tasks, but if it *can't* one-shot a task (even given ample compute), then trying over and over again doesn't seem to help. RIP to AI's learning trajectories, but humans are built different.
1
246
Extremely happy that Elasticity has been getting advice & help from @yelizarovanna and Olivia Benoit.
We are excited to welcome Anna Yelizarova (@yelizarovanna) and Olivia Benoit as advisors to the Elasticity Institute! We are thrilled to have their expertise as we grow Elasticity to include more research and seminars on the economics of AI.
1
23
2,560
tom cunningham retweeted
We’re hiring in SF! Encode cut its teeth in DC. People said AI politics was a bad bet; we made unlikely friends, won the fight to keep states in play, and helped pass the Big 3 of AI laws (CA SB 53, NY RAISE, IL SB 315). It’s time to double up in SF. Pacing the Frontier (pacingthefrontier.com), which we helped organize, was recently signed by 1,378 frontier lab employees and endorsed by both OpenAI and Anthropic. The last time we felt this rush of opportunity was Sacramento in 2024. AI risk has gone from an online culture war to a fact of life. SF is full of people who want to act, lab employees especially, but don’t see a place to plug in. We need someone to help us build the grid. Encode’s lab engagement has been running on adrenaline and personal relationships. That got @_NathanCalvin and me here, but it’s time to level up. Misalignment is getting scary. “Pacing” has real legs now. Researchers are using their voices to influence lab leaders and alert the world. Many of those researchers say that soon, their jobs - and the voice their jobs give them - will be unrecognizable or won’t exist. We want to hire someone who is: (1) Mission-aligned. You want to ensure AI doesn’t cause catastrophe (or kill everyone) and the future is the best it can be (2) Extremely organized, high-functioning, excited to add structure so ad hoc actions compound (3) Pragmatic, able to pivot when crazy things happen (if the last few months are any indication...) (4) Hard to typecast - not because you lack beliefs, but because you see the best in many kinds of people. The role will start with supporting lab engagement and expand. It’ll be insanely fun (we take the work seriously, not ourselves). You’ll visit DC regularly as AI’s political salience takes off and get to know a team that feels like family. If you’re interested, email me by September 4 at sneha@encodeai.org with the subject line “SF Special Projects” and a few sentences each on (1) a relevant project you’ve owned and (2) what you think Encode SF should do next. If someone you know could be a good fit... I want to meet them!
9
51
249
45,531
tom cunningham retweeted
“Dario believes frontier AI is too powerful to distribute; we believe it is too powerful to centralize.” The problem is that both are true.
Some thoughts on Dario’s post: 1. Dario does not actually address Gavin Baker’s account of what he said – something he could easily deny if it were inaccurate. 2. Dario claims his critics live in a “bubble” where all regulation equals regulatory capture. He calls this an overly simplified view and notes that “Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people.” This argument is a straw man. Of course treating all regulation as capture would be overly simplified – but almost no one holds that view. I have repeatedly argued for strong antitrust enforcement to keep industries competitive, especially Big Tech. If Anthropic continues toward monopoly or duopoly status, I would be among the first to demand those rules apply. 3. Regulatory capture is not vague or in the eye of the beholder. Nobel laureate George Stigler defined it as regulation acquired by an industry and designed and operated primarily for its benefit. Stigler challenged the traditional view that government regulation arises from a benevolent state protecting the public from market failures. Rather, industry groups have concentrated stakes and pour resources into influencing regulators, whereas the public’s stake is diffuse and unorganized. The revolving door between companies and the agencies that regulate them compounds the problem. Anthropic understands these dynamics: it has hired multiple senior Biden AI-policy officials and built a substantial government-affairs operation plus a network of aligned organizations to push its preferred frameworks at state and federal levels. 4. Dario has consistently pushed for a new federal agency to review and approve frontier models prior to release – a proposal framed variously as an “FDA for AI,” an “FAA for AI,” and most recently a “FINRA for AI.” I call it a “DMV for AI” because a review process modeled on the FAA or FDA (which takes years) or FINRA (which issues rules for a staid industry widely seen as protecting incumbents) will create long queues as AI models wait for testing and approval. This process will only become more labyrinthine as rules accumulate to prevent theoretical harms. This would handicap the U.S. relative to China, which will not adopt the same constraints. It would also undermine Anthropic’s own business model, whose pricing power depends on remaining ahead of open models. Whatever Dario states today, it is difficult to believe the company would simply accept outcomes that erase that advantage. 5. Anthropic is on track to become one of the most valuable companies in history, with the resources to navigate any approval process and shape the rules while competitors wait. Dario wants open models under heavier scrutiny – he has called them dangerous in Senate testimony, criticized them for not being centrally monitored or withdrawn, and linked them to IP theft. He says he has never sought a ban, but he could achieve a similar result by insisting that identical rules apply to both open and closed models. The U.S. risks becoming an island of costly closed models while the rest of the world races ahead with broader choice. 6. Dario acknowledges that AI is structurally centralizing but attributes this mainly to chips and scaling laws. Access to compute matters, but the deeper risk is who decides which capabilities are available to whom. His preferred pre-deployment testing and FAA/FINRA-style oversight would place that gatekeeping power in a federal bureaucracy working hand-in-glove with a small number of frontier labs – reinforcing centralization rather than countering it. 7. The second part of Dario’s post assumes we have amnesia about Anthropic’s well-orchestrated campaigns hyping AI fears. His May 2025 claim that AI would wipe out 50 percent of entry-level knowledge jobs within five years still lacks supporting evidence fifteen months later. Similarly Anthropic breathlessly promoted its heavily contrived “blackmail” study on 60 Minutes. Yet Dario blames public negativity on a long-standing loss of trust in institutions rather than his own messaging. 8. These narratives have done more than anything to shape public fear. People are left asking the same question Mark Zuckerberg posed: why race to build a future you describe in such negative terms? Thomas Sowell’s "The Vision of the Anointed" captures the mindset – elite intellectuals convinced that only they are enlightened enough to control the outcome. As Zuckerberg notes, concentrating power in the hands of an enlightened few has rarely produced the promised results; the practitioners turn out to be less enlightened in practice than in self-conception. 9. Gavin Baker summarized the disagreement cleanly on our pod: Dario believes frontier AI is too powerful to distribute; we believe it is too powerful to centralize. Dario appears to believe, sincerely, that safety and progress are best served by centralizing authority in a marriage of corporate and state power. The weight of human history gives us reason to fear that outcome.
5
25
220
7,239
New post with Nate Rush: Have we seen an acceleration in discoveries? Many plots & some tentative conclusions: 1. Cyber: ⤴️ sharp acceleration 2. Math: ↗️ some acceleration 3. Algorithms: ➡️ no clear acceleration
11
38
244
68,153
I forgot to add a really important qualification -- this is only analyzing what's happening in public. It's quite possible that algorithmic progress is accelerating in private.
1
25
1,648
Come work with us, the going's tough & the tough are going.
In the last 6 months, METR raised commitments of around $71 million. This will fund ambitious projects: studying autonomous capabilities, tracking recursive self-improvement, evaluating monitoring systems, conducting risk assessments, investigating AI incidents, and more.
4
15
180
14,286
Especially for people who have lab experience, lots of useful work you can do at METR.
14
572