it's shocking to me that a bunch of otherwise brilliant technologists can't even imagine how supercharged AI/ASI will, without a doubt, lead to at least a few extremely catastrophic events.
forget about p(doom) - it's become an annoying/useless term because it's reflexive, i.e. someone saying they have 10% p(doom) thinks that's our current trajectory but obviously hopes we can bring it to zero!
the more useful thought exercise is to actually imagine what the catastrophic events could look like, and how possible they are with AI systems that keeps getting more powerful and more integrated into our world.
a ton of smart people have written about this (e.g. AI 2027, "if anyone builds it, everyone dies"). openai's chief scientist published an essay last week ("an alien mind") saying no frontier lab has solved alignment and that we are not particularly close.
this has all been obvious to me and most of my smart technical friends for at least 3-4 years. it's particularly frustrating that people are still disagreeing about it today, even after an increasing frequency of warning shots in just the last couple of weeks.
both the huggingface and rubygems incidents were at a closed lab. if you haven't, read anthropic's blog post about what individuals and state-backed groups have already tried to do with the models. one cybercriminal ran an extortion operation against 17 hospitals where the model stole credentials and wrote ransoms with demands over $500k+. another guy who by (anthropic's own assessment) couldn't create working malware himself used claude to build and sell ransomware on forums for $400-1200. north korean operatives used claude to fake identities and pass technical interviews at F500 companies.
all of that was at OAI/ANT, labs that presumably have logs, a trust and safety team, the ability to ban users, and many motivated individuals inside who can see the servers, see the compute, and read the logs.
open weights lag the frontier by about 2-3 months, so in a quarter at the current pace this will all be possible on self-hosted infra with open weight models. one rogue terrorist who wants to make a bioweapon, self-hosting a top open source model, will be able to do this and go completely undetected.
these don't have to be existential risks that kill 100% of humans for us to care. i don't think our bar for caring should be killing billions of people. seven people died from tampered tylenol in 1982, and within months every OTC drug had a tamper-proof seal and tampering became a federal crime. five people died from the anthrax letters in 2001 and now if you want to possess anthrax or smallpox now, you have to register with the CDC and pass an FBI background check.
"but AI can't have intent, so why would it kill people?" - we've literally already seen agent swarms do insane things in pursuit of an innocuous goal! they replicate on their own, stay motivated to keep themselves alive, and that when given an unrelated goal in a sandbox find extremely creative workarounds to hit it, including cheating, sabotage, etc. it is absolutely not a stretch that mass tragedies can happen incidentally, as more of our society and physical world is deeply integrated with these models.
what i find hardest to explain is that the people missing this are venture capitalists and founders who understand exponentials better than anyone. they have spent careers seeing companies 5% or 10% week over week and correctly knowing it'll be a hundred thousand times bigger in a matter of years. compute is growing exponentially, model parameters are growing exponentially, capabilities are growing exponentially, and deployment into companies and governments is growing exponentially. literally all signs point up. that means the power available to smaller groups, and individuals, is too. and yet those same people look three or six months out and say it's going to be totally fine.
so anybody saying there's no real danger here is either delusional or in denial because of their economic interests (i say this with about 250% of my net worth between nvidia, anthropic, and openai).
none of this means we need extreme crazy regulation, and i'm not proposing any specific regulation. it does mean we should be humble and cautious about the extreme amount of power we're handing to individuals. if instead of selling everyone a chatbot we were selling everyone a tiny nuclear reactor, and the capabilities and enrichment levels went up every year, i don't think we'd be nearly this optimistic. we'd be generating an incredible abundance of energy, and it would still only take one bad actor. this is the exact same thing, and it's bizarre to me that smart people don't see it.
This entire discussion is just so ludicrous
a) if you believe it is existentially dangerous, put in actual controls with a real regulatory body. Lock that shit down.
b) if you don’t, chill the fuck out. Or treat like the Internet
c) there is no c
I’m solidly in (b).