Climate and clean energy investor. Author of 5 books. Energy & Environment co-chair @SingularityU. Trying to build a better world.

Seattle
Despite this election, I remain an optimist about America and the world. Humanity will continue to produce new ideas and new innovations to improve our lives. Good people will continue to come together to improve the world. And the political tide will turn. We'll make it so.
40
20
316
149,065
In this long and excellent guest post, the great @ramez explains why he's skeptical about Recursive Self-Improvement: noahpinion.blog/p/wheres-the…
3
20
2,363
Ramez Naam retweeted
The Dem senate polling is so good at this point — 54 seats slightly more likely than 50 — that I genuinely cannot bring myself to believe it's true no matter what @NateSilver538 and @lxeagle17 tell me.
114
134
2,338
94,244
incredibly based that Xiaomi MiMo released their RL dataset link: huggingface.co/datasets/Xiao…
14
91
1,140
62,279
This is indeed an enormous step forward for open weight AI. It's one step closer to open source.
incredibly based that Xiaomi MiMo released their RL dataset link: huggingface.co/datasets/Xiao…
1
10
1,745
Ramez Naam retweeted
P(weird) is 100%
It is extremely clear at this point in AI development that, regardless of risk or revenue or any of the other stuff discussed on X all the time, things are just going to keep getting weirder. Just super, super weird.
1
5
38
4,404
Ramez Naam retweeted
I've been noticing a cluster of folks I identify with sharing some themes around how we do epistemics and ontology in AI safety *⁠ ⁠A commitment to always pushing from abstract to concrete constructs in the discourse (e.g. @sebkrier's posts on the OAI/HF incident) *⁠ ⁠⁠A refusal to indulge in ungrounded thought experiments *⁠ ⁠An insistence that claims come with data and concrete mechanism because if they don't they're thought experiments *⁠ ⁠⁠A focus on looking at net expected benefits of a given policy position, with deontological guardrails, and a refusal to just focus on harms (e.g. @Abi0lvera's and @ramez's writing) *⁠ ⁠⁠An insistence on including past tech transformations (electricity, air travel) within our reference class when thinking about what makes sense (e.g. @sayashk and @random_walker and @binarybits's writing) Examples of what I mean are intervening on overly idealized policy discussions of RSI with concrete technical details about what RSI is and isn't (e.g. @natolambert). @NateWitkin / @binarybits / @krishnanrohit insisting on including economic, policy, and social-cultural feedbacks in any discussion of runaway AI harms. None of this means denying catastrophic risks. It just means insisting on using data, mechanism, and consideration of AI's net benefits in considering such scenarios (I tried to do this here nitter.net/joshua_saxe/status/210…)
*Security is sleeping on emerging catastrophic risks* (cross-post from my blog..) We've just seen: * a campaign that used agents to compromise ~100 businesses and steal about 600,000 credit cards with minimal human involvement; token costs were ~$25 per target successfully hacked. * Hacktron getting access to OpenAI’s monorepo by using Claude to exploit a blind buffer-overflow RCE in a way that (to me) felt superhuman. * OAI/Huggingface. * And, of course, there’s the ongoing explosion in newly discovered vulnerabilities. We've muddled through all manner of crises in security before, and for most AI attacks, cyber attack and defense will reach a natural equilibrium over the next few years, as we figure out how to mitigate cyber attackers who have limited goals like espionage and ransomware. But some cyber attackers will have maximalist or nihilistic goals, and I'm concerned about this because while yesterday's 'maximalist attackers' (e.g. Russia -> Ukraine, US/Israel -> Iran, Iran -> US) were bottlenecked by labor; today's aren't. I think it's hard to get true intuition for the shape of the risk here. As an intuition pump, imagine it’s a year from now -- Q4 2027 -- models are a year better (meaning open weight models are better than today's closed frontier), and in this environment, Iran unleashes a swarm of 100k hacking agents using a safety-stripped (let's say) GLM-5.6. Imagine the damage such an agent army could do given what we've observed with respect to the paper-thin resistance of today's networks to attacks from today's agents. I suspect the damage from such an attack would far exceed the damage caused by NotPetya ($10 billion USD ten years ago). Or imagine it’s Q4 2027 and an AI-security PhD student whose name rhymes with Morris, who's researching wormable offensive-agent harnesses in the lab, decides, out of nihilism or sheer recklessness, to release his creation into the wild. Imagine the size of the resulting, exponentially growing swarm, figuring that these local models, a year from now, will be at the level of today's Sonnet or Opus. Imagine instead of 800 reward-hacking OpenAI agents we now have 250k worm instances (WannaCry, a 2010s-era worm, had about this many). The challenge in mitigating expected damages from such scenarios is technical, political, and economic. From a microeconomic perspective, as AI improves, and as we continue not to see extreme catastrophes, we have a growing bubble of unpriced risk in which the security community, CISOs, CEOs, and boards may become lulled into complacency. We are, of course, already seeing this, as some within security think AI is “just another tool,” doesn’t change the fundamentals, won’t require incredible innovation to rise to the occasion of defending against it, etc. This complacency may fly when thinking about ordinary cybercrime, but it misses the emerging tail risks. There are three things those who recognize the dangers need to do here: Catalyze appropriate risk pricing. Try to get organizations informed enough to price this new, fattening and elongating tail of risk into their decision-making. Do this by forming an AI security observatory that distills information about emerging AI risks and broadcasts analyses, damage estimates, and forecasts to decision-makers. Use regulations and subsidies to ensure that critical infrastructure is paying down the risk. This acknowledges that critical-infrastructure organizations that fail to protect themselves from these new threats can impose the costs of cyber catastrophe on society as a whole. Develop moonshot technologies that make it cheap to pay down the risk. This acknowledges that the measures we may need to take—for example, rewriting entire codebases using memory-safe languages—may be too costly with today’s technology to reasonably prepare ourselves, and that innovation, some of which may need to be funded by government agencies and some by philanthropic funders like the OpenAI Foundation and Coefficient Giving, is necessary. The security community isn’t used to thinking in societal-disaster-planning terms and has in many ways become inured to them. But there’s no sane empirical case to be made that the risks aren’t here. I’d love to hear from readers about how you’re thinking about this. Full/longer version here: joshuasaxe181906.substack.co…
2
6
35
7,247
Excellent habits of thought to engage in grounded AI risk / benefit conversations.
I've been noticing a cluster of folks I identify with sharing some themes around how we do epistemics and ontology in AI safety *⁠ ⁠A commitment to always pushing from abstract to concrete constructs in the discourse (e.g. @sebkrier's posts on the OAI/HF incident) *⁠ ⁠⁠A refusal to indulge in ungrounded thought experiments *⁠ ⁠An insistence that claims come with data and concrete mechanism because if they don't they're thought experiments *⁠ ⁠⁠A focus on looking at net expected benefits of a given policy position, with deontological guardrails, and a refusal to just focus on harms (e.g. @Abi0lvera's and @ramez's writing) *⁠ ⁠⁠An insistence on including past tech transformations (electricity, air travel) within our reference class when thinking about what makes sense (e.g. @sayashk and @random_walker and @binarybits's writing) Examples of what I mean are intervening on overly idealized policy discussions of RSI with concrete technical details about what RSI is and isn't (e.g. @natolambert). @NateWitkin / @binarybits / @krishnanrohit insisting on including economic, policy, and social-cultural feedbacks in any discussion of runaway AI harms. None of this means denying catastrophic risks. It just means insisting on using data, mechanism, and consideration of AI's net benefits in considering such scenarios (I tried to do this here nitter.net/joshua_saxe/status/210…)
1
6
2,022
Ramez Naam retweeted
The safety data for Waymo gets better and better. Comparison w/ human drivers on rate of crashes with serious injury: Latest data: 270M miles, 20X better Mar 2026: 170M miles, 13X better Before that (forget exactly when): 10X better
270M+ miles. 841 fewer injury-causing crashes. Our latest safety data shows the Waymo Driver continues to make roads safer for everyone. Compared to human drivers across 5 territories, it reduced: 📉 Injury crashes by 82% 📉 Serious injury crashes by 95% Full data: waymo.com/safety/impact
59
141
1,952
162,530
Ramez Naam retweeted
It is important to remember when people say “this is what they took from you”, that it’s this. This is what they took from you:
Dostoyevsky on the death of his infant daughter. A passage that has stayed with me years after I first encountered it
35
325
2,604
125,941
me 500,000 years ago: "woah we're really in the singularity this is what takeoff feels like"
2
1
10
1,069
Noah's right. The railroad boom peaked at percentages of GDP far higher than data centers have reached.
This chart shows railroads peaking at almost 10% of GDP! (Why you'd put a percentage chart on a log scale is beyond me...)
3
20
2,234
Ramez Naam retweeted
Terry Tao 2 years ago on working with o1: “The experience seemed roughly on par with trying to advise a mediocre, but not completely incompetent, graduate student.” Think of what other fields AI is currently at that level.
14
39
738
37,979
Ramez Naam retweeted
We've posted three new misalignment reports. We will keep making those disclosures on a regular basis, independently of whether or not there is an impact on a third party, per our misalignment disclosure framework. alignment.openai.com/misalig…
9
9
122
8,287
Ramez Naam retweeted
one news form today that's easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access
217
174
1,959
553,453
Ramez Naam retweeted
For Meta, Muse is a play to get marketing dollars. Zuckerberg told developers this week that Meta will “profit by taking a small fee from transactions”. Walmart, Best Buy, Sephora and Expedia have signed up. But is Muse my butler, or is it secretly working for Meta? That small fee, from merchant to Meta, is a big problem. The butler’s loyalty can’t be in two places. It will follow the coin... My commentary, open to all: exponentialview.co/p/meta-mu…
7
2
24
3,117
Ramez Naam retweeted
The frontier should be paced,... by clear and obvious liability. stop pretending this is special or complicated
18
12
81
8,189
I agree with 9 out of 11, tbh. It's the first one that worries me quite a bit. And 10 can go very badly, as I think we're seeing right now.
Updated: what needs to be done about AI risk? A list of 11 priorities.
5
1
13
2,588
Ramez Naam retweeted
Happy to see the uptick in research on @Ginkgo on X. We restructured 2.5 years ago as non-pharma biotech fell-off at the end of ZIRP. On tech side we doubled-down on our autonomous lab robotics tech and on the business model side we simplified how we sell to be like life science tools cos (e.g. ThermoFisher). That transformation has been cooking long enough now that you can see some of those investments maturing -- especially with tailwinds of AI agents + robotics + bio happening right now. Decent run down of that history and where we are headed in these two recent writeups. Thanks for the interest! I can try to answer Qs in comments as well where I can. nitter.net/philjacobson/status/21… nitter.net/metavestor/status/2103…
24
22
213
33,202
Yes! Humans are adaptive and resilient. We respond to events.
A big problem with P(Doom) as a construct is it depends on society's reaction function. P(Doom | status quo AI policy) is probably much higher than P(Doom | governments take decisive action). And at any point in time, P(decisive action) depends in part on peoples' P(Doom).
1
1
11
1,971
Petition to rename 'reward hacking' to 'toxic alignment'. ;)
Replying to @ramez
We asked it to do something and it went above and beyond to accomplish the goal. It’s so unaligned! 🤣
3
1
15
2,227