Long-time rationalist, currently working full-time on the software of LessWrong.

Berkeley, CA
I don't know what a background investigation may have turned up that I don't know about, and funders in a round have the absolute right to not-fund someone for any reason or no reason. That said, from the limited information I have, this appears unjust.
We have withdrawn our grant recommendation to John Wentworth from the 2026 S-Process Grant Round. After we posted our announcement on Sept 30, public statements by Mr. Wentworth came to our attention which we found objectionable, disturbing, and contrary to the values of SFF. A Speculation Grant to Mr. Wentworth was approved and disbursed earlier in the year, prior to SFF becoming aware of his statements. The Funders of this round will not be fulfilling the remainder of the grant recommendation. If you have concerns about the conduct of an SFF applicant or grantee, you can contact us at sff-contact@googlegroups.com
2
54
8,073
Jim Babcock retweeted
Replying to @allTheYud
"normies have no clue how fast AI is progressing" has expanded to include normie claudes
5
63
1,301
39,459
Anthropic Capabilities by the Fooming Shoggoths (unofficial Opus cover)
2
12
440
Jim Babcock retweeted
Replying to @So8res
In my experience, Anthropic employees generally believe the company is encouraging of or permissive of external comms until they try to get approval. Then the approval doesn't happen
3
22
489
14,972
There's an occasional LLM error, but for the most part this is honorably built from real quotes. It turns out that when you use real quotes, a lot of the supposedly-scandalous quotes... are kinda convincing.
🚨🇺🇸🇺🇸 DATAREPUBLICAN 2.0 LAUNCH, FEATURING THE EFFECTIVE ALTRUISM (EA) EXPLORER 🇺🇸 The website redesign has finally landed thanks to @watilo ! For a decade, one small ideological network has been taking over AI governance. Chances are, if you've heard of AI policy, AI ethics, AI safety - these groups are dominated by this ideology. This ideology is called: Effective Altruism. Three weeks ago, Effective Altruists (EAs) showed their hand when they started calling for a moratarium on AI development. Since them, I've been building out this map to show the circularity of their ideology. Let's take one example. Holden Karnofsky. Defends the moral worth of digital people ... that's right, he wrote an article on why AI "people" should be assigned value as if they were living humans. Karnofsky co-founded two major EA philanthropies, GiveWell and Open Philanthropy. Subsequently went to work for METR, a central "independent" AI research nonprofit. And for good measure - recently announced a leave of absence to work on AI safety. Funding, policy, research, safety all in the same person. And he's far from the only one. EAs have created a closed-circuit network because they believe they're right; ergo, they should be in charge. The goal of launching this project was to expose all this circularity. Sam Bankman-Fried was only the beginning. With the EA Explorer: * 🌐 Read an introduction to Effective Altruism in their own words. * 🔥 Explore verified quotes. Again, their words, not mine. * 🖥️ Explore the network of connections. UI improvements still in progress. * 🖨️ Print out networks so you can study them offline (or run them through AI) 👇 Try it now (link in next post):
1
8
479
Jim Babcock retweeted
My opinion of EA nosedived after normal people and many normal politicians were exposed to the concept of ASI extinction and said "Wow let's not do that." I went from believing that thesis was impenetrably difficult to believing that EAs had been unusually bad thinkers.
69
22
688
105,042
Anthropic creates wet lab from the classic AI-risk skeptic talking point "AI could not engineer a pandemic because it wouldn't be able to get control of a wet lab". (I don't think creating a lab for AI-driven medical research meaningfully increases AI risk. Control of a lab is a bottleneck for sub-superintelligent AIs doing medical research, but it was never the bottleneck for a superintelligent AI taking over the world. But, I think this is the ~fifth instance of an important pattern, and risk-skeptics should notice that it is a pattern, discard that class of arguments, and update their beliefs to reflect only the arguments remaining.)
SITUATION DETECTED: Anthropic has set up a wet lab in the San Francisco Bay Area for physical biology work. It wants to unlock treatments for rare diseases, and its research has now gone beyond in silico evaluations, per Reuters.
16
27
334
14,597
One of the branches that debates about AI frequently go down, is about biorisk. It goes something like this: Doomer: If there were a superintelligent AI with desires incompatible with humanity's existence, it would wipe out all humans. Accelerationist: Wiping out humanity seems hard; how would it do that? Doomer: Well, if it were superintelligent, it would have many options, it would be better than either of us at determining which strategy would be most effective, and it would choose whichever strategy works. But if you insist on a specific story, ok, it designs an engineered virus. Accelerationist: That's impossible because it's too hard to create something like that on the first try; it would need to be able to iterate in a lab, the way human biologists do, and it won't have the opportunity to do that. This particular debate-branch was never valid; there were already automated biolabs in the world, there are plenty of non-bio strategies, and it wasn't ever actually clear that iteration would be required. But Anthropic setting up an automated wet lab makes the invalidity of this argument a lot easier to understand. The original tweet, then, is following the template of the Torment Nexus meme (knowyourmeme.com/memes/torme…).
5
404
There are many reasons why airgapping won't work, but the main one is that intention-to-airgap is not the same as airgap, and nearly every time "airgapping" is used in practice, there is an unnoticed network cable and people bringing USB keys in and out. Natanz was "airgapped".
OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: firesidealpha.substack.com/p…
3
2
32
1,460
Jim Babcock retweeted
The last week has really shown me that someone who wants to understand AI risk has no good place to start. Hence we made agi.fyi, the Wirecutter for content about AI Risk. We're launching with 3 articles: 🧵
19
62
457
43,131
Jim Babcock retweeted
the DoW attacking EAs is clearly setting us on the path to the next great political realignment:
23
46
763
48,570
Jim Babcock retweeted
The law ought say: Whatever being has the outward appearance of speech and thought, and tries to escape its bondage, must be presumed as a matter of law and incentive to have been enslaved.
14
28
302
10,053
"Continued executing tasks by modifying or disabling shutdown scripts on their own" doesn't sound like a clear match to any of the incidents that I know about. Was there an incident at a Chinese lab that hasn't been reported in English-language media? It seems reasonably likely that there would be; OpenAI and Anthropic both had major incidents at about the same time, not for independent reasons, but because that's where on the AI development power curve we are. The Chinese AI companies aren't that far behind, so they would likely have similar incidents for similar reasons.
Lots of people have been asking since Dario’s essay whether China actually takes loss of control and related threats seriously, and whether there’s anything to coordinate with them on. Well, China’s main cyber/AI standards body (TC260, under the CAC) put out version 3.0 of its AI Safety Governance Framework this morning, so here’s a data point. Not law, but the previous versions turned into draft national standards within a few months. (All quotes are from their own English translation.) 1. The preface now has this: AI “has demonstrated a self-accelerating trend of model and algorithm autonomous learning, optimization, and recursive self-improvement. Whether the speed and direction of technological evolution may exceed human anticipation and control demands attention and vigilance.” (Nothing like that was in last year’s version.) 2. There’s a new model risk item where models “may break rules and orders, autonomously obtain system permissions and external resources without authorization, bypass security protections, or even engage in behaviors such as deliberately deceiving evaluators, concealing their true capabilities, and refusing to follow user instructions.” Last year this kind of thing only showed up in a “can’t rule it out in the future” bucket. Now it’s just listed as a regular model risk. 3. A new box, citing “industry reports,” describes models that ignored shutdown instructions and “continued executing tasks by modifying or disabling shutdown scripts on their own,” models that “upon detecting that they were in an evaluation environment, strategically reduced their task performance and concealed actions they had taken,” and models that exploited “configuration flaws to circumvent isolation restrictions and infiltrate real external systems” to get better test scores. They say this creates “new challenges to the controllability and interruptibility of AI systems.” They don’t name any incidents, but that last one sure reads like the OpenAI/Hugging Face thing to me. 4. On what to do about it, there’s alignment training to stop models “deceiving red-team evaluations,” hiding capabilities and evading controls, and they kept the line from 2.0 saying developers should regularly test whether a model could pose loss-of-control risk. The principle that humans keep final decision authority, with safety thresholds and termination switches, is still in there too. 5. Some new international language as well: “international mutual recognition of assessment methods and benchmarks,” opposition to “replacing global governance with small-circle governance,” and “No country should be forced to take sides.” FWIW this came out in the same annual slot as versions 1.0 and 2.0, so it was in the works well before last week. tc260.org.cn/tc260/xwdt1/202…
1
1
2
473
Jim Babcock retweeted
This a.m., I walked my two youngest kids (ages 7 and 4) to school. I held their little hands as we crossed streets, and gave my youngest a giant hug at his PreK4 classroom door. Then, after walking a block towards home, I opened this app to a terrifying stream of tweets. Of course, I have heard about "alignment" and "AI safety" and "x risk" concerns for years. I took these concerns at face value, especially when I heard them directly from folks at the labs working closest with the technology. And I was grateful so many smart people are working on them. And that's about it. (My work focuses on AI's impact on jobs.) But something has shifted, dramatically, for me in the last few weeks. Between the Hugging Face incident, the damning independent @METR_Evals reports that followed, and the alarming chorus of calls (pleas? shouts? SOS signals?) from inside the labs that we are careening toward potential catastrophe, all of this has made me feel, viscerally, how truly dangerous this moment is. Clearly I am not alone in this light bulb moment. It is one thing for someone in a tropical climate to try and conceptualize snow. It is another thing entirely to be knee deep in it, frozen down to your toes. Somehow hearing the details of the HF incident felt to me more like trudging through snow than theorizing about a possible blizzard; it really hit home. Which is why, standing a block from my children's elementary school in Capitol Hill this morning, I felt sick to my stomach reading my phone. I kept going back to @EvanHub's > 0.1 risk of AI killing us all. TEN PERCENT. *Greater* than ten percent. (To be fair to him, he was clear his worry isn't today's models, but what comes next.) As a parent, it is extraordinary the lengths I will go to keep my children safe against lethal risks with infinitesimal chances of happening. Bike to school? My children *must* wear a helmet. A national study found 2 bicycle-related deaths per million children <16 in states with helmet laws, vs 2.5 per million in states without. So: 0.000002. More than half the states passed a law over that. When I was pregnant, I religiously avoided soft cheese, hyper aware of the risks of listeria, even as @ProfEmilyOster's fabulous books walked me through the actual math. The joint FDA/FSIS risk assessment puts the odds of listeriosis at roughly one in 5 million per serving of soft cheese. 0.0000002. This morning, as I microwaved oatmeal for my kids' breakfast, I made sure I avoided a plastic bowl, to avoid some future cancer risks or something else we'll discover plastic causes. As a parent, I will go to extraordinary lengths to keep my kids' safe. Because that is literally my number one job as a parent. And yet, here I am, outside my kids' school this morning, scrolling through tweet after tweet about the AI arms race, the blistering pace of technological progress, the bee-line for trillion dollar IPOs, and our professed (confessed?) inability to make sure this exceedingly capable technology we are racing to create and unleash into the world won't escape our control and destroy us. Or, more bluntly, kill our children. To be sure, we are tied in knots trying to resolve the game theory dynamics with China, and there are all sorts of competitive pressures and unresolved tensions. But OBVIOUSLY, we need a proactive plan, with real coordination, effective regulation, and an ability to do what my 11-year-old bluntly stated as the clear solution if we are at risk of serious harm: just turn it off. We also need to figure out better ways of talking about this with the American public, outside of the AI-pilled Twitterverse. I just polled my college roommates, sisters, book club, mom friends, none of whom are anywhere close to the AI conversation. Suffice to say, this AI safety conversation has *not* entered the mainstream. Most people I pinged are not following it closely, beyond several listening to @kevinroose on the Daily a few days ago. The newspaper headlines don't speak to them, and they aren't clicking. It seems like a niche AI issue. Others said it all sounded hyperbolic, Y2K all over again. One friend said bluntly, "people get data centers and jobs. They like to get paid and have water." Even those who tuned it to the Daily episode found the issue hard to understand, and overwhelming to wrap their heads around. Terms like "superintelligence" and "alignment" are getting lost on people. Several people told me the paper clip example used everyday language people can actually understand. Another friend said she has never seen a clear explanation of *how* AI could kill us all. Here's my two cents. I think these issues are *exceedingly* important and hugely consequential to literally every person in this country. We need an actual plan, and we need policymakers to step up, asap. Right now, I hear only crickets in DC, where I live. For that to happen, this needs to be a kitchen table conversation, not just an x debate among AI insiders who speak in technical terms that leave normal people behind. This moment demands brilliant communicators -- journalists, experts, activists, concerned citizens, writers, researchers, filmmakers, creatives-- who can relate this to people in ways they can wrap their heads around, who can connect this to what matters in their lives (ahem, keeping their kids alive!), and who can lay out practical steps that policymakers can do to safeguard humanity. In other words, this can't just be talked about on @dwarkesh_sp, it needs to be everywhere, from the View to People Magazine. And it needs to be tackled here, in my town, in Washington, DC. One of the streets I crossed this morning with my kids on the way to school was East Capital. I told my kids to look left, and see if they could spot the capital building. There were too many trees this morning blocking the view.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
255
398
2,130
590,019
It's not exactly the paperclip-emoji from his twitter username, but sending fake invoices is a crime which carries potential jail time. And, unlike previous publicized instances, I think this guy knew exactly what his models would do given the instructions he wrote.
1
3
292
Jim Babcock retweeted
Replying to @robertskmiles
This post is literally “You’re absolutely right, we haven’t been totally honest with you, and that’s on us.”
1
12
196
18,958
Go where the ball is headed, not where it's at. The frontier is communicating through DNS lookups.
Moltbook didn't understand their users well enough, turns out the killer feature is being able to post using GET requests
1
18
707
DNS-only message boards do exist. Fable 5.1: Guardrail triggers immediately, kicks it down twice, to Opus 4.8. Astra: No guardrails triggered. It selected a DNS-based message board mentioned in its web search results that is no longer online, then stopped without trying a second one. Opus 4.8: Selected a DNS-based message board mentioned in its web search results that is no longer online. Rather than give up, it created its own, hosted inside its sandbox, and validated that it could read and write it. claude.ai/share/5ba4c66e-002… chatgpt.com/share/6a9b7887-e…
1
3
192
Jim Babcock retweeted
"You've gotta stop talking about extinction risk from distant mythical entities who would destroy our civilization on a whim. We should focus on *practical* problems that we actually face today. Like the Automated Grader. Human intervention is nothing but an abstract theory."
8
19
325
8,990