Head of AI @AsteraInstitute Prev: AGI @DeepMind, cofounder @vicariousai (acqd by Alphabet), cofounder @Numenta. IIT-Bombay, MS&PhD Stanford. agicomics.net

San Francisco, CA
Forgive me.
1
16
1,421
Thank goodness! I was on the verge of selling GOOG. 😇
6
4
69
3,309
lure of language means researchers test for "pain" in LLMs while leaving a vacuum.
New paper: we found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not to the user. Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kids' photos.🧵
7
2
21
3,051
The good thing about language models is that you can always find what you are looking for.
New paper: we found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not to the user. Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kids' photos.🧵
9
1
95
5,000
Well, good that we already built a place for these new creatures to live in — second life.
"We are in the process of creating a second form of life. The emergence of this new form of life is an inevitable result of the fundamental laws of the universe — forces similar to those that brought about humans' own existence." —@WmHaseltine noemamag.com/a-new-form-of-l…
1
8
1,491
Satire?
I went to an rationalist/EA AI doomer event last night with the goal of meeting someone who could calmly explain to me the argument for how they get to the conclusion that AI will kill everyone. Because these are rationalists, I expected to find people who could lay out a serious extinction pathway, walk through the assumptions and intermediate steps, identify choke points where negative pathways could potentially be regulated, and entertain disagreement about where their assumptions might break down. My experience was the exact opposite. The people I spoke to seemed psychologically unable to entertain the possibility that their conclusion (that AI literally kills everyone, which was a point that many would not even budge on) might be wrong. Their most extreme conclusion was treated as an axiom. I had multiple people walk out of conversations with me when I questioned their inevitable extinction premise. At one point during a panel, the moderator asked everyone in the audience who was still skeptical of AI doom to raise their hands. About half the room did. He then told all of the skeptics to leave the panel and go with one of the speakers, who would answer their questions separately. They filtered out all the skeptics, and continued the panel only with the subset who accepted their extreme position. I genuinely felt like a heretic just for questioning whether superintelligent AI inevitably leads to the extinction of humanity, or entertaining the possibility that the future might actually turn out well. That does not mean AI safety is fake. Obviously there are serious questions about rapidly advancing AI systems and what happens as they become much more capable. Those are exactly the conversations I went there hoping to have. But what I encountered was an apocalyptic death cult. These people are not rational. They are deeply afraid, they reinforce that fear in one another, and they have become so emotionally committed to apocalypse that they refuse to be reasoned out of it.
3
5
2,491
I don't endorse slowing down. But I endorse a huge compliance department. 😇
2
22
1,342
This was the take away for me from the Dan Selsam doc.
« One could (…) define intelligence as the efficiency with which one converts experience into competence; by this definition [current LLMs] lag very far behind us. » – Dan Selsam
1
13
1,966
Dileep George retweeted
My thought on personhood for AI systems: More here: blog.dileeplearning.com/p/ai…
2
2
2
1,130
"The idea of model welfare is wrong. AI’s should not have rights or legal personhood." I like this. Thanks to @mustafasuleyman for spelling this out. Personhood or legal rights for AI is not a must-have -- just don't build AI that way.
9
4
28
4,438
My thought on personhood for AI systems: More here: blog.dileeplearning.com/p/ai…
2
2
2
1,130
This is what is genuinely baffling to me. Even the current Dan Selsam letter doesn't at all grapple with the risk of AI endangering 10K people. Apparently that risk doesn't exist, or is not the primary concern. The risk they seem to be worried about is AI laying low, unmonitored, until it is sure it can overpower humanity in one shot. I'm still not sure how to think about this.
The AI bio risk conversation would be less contentious if we talked less about x-risk and more about 'mere' catastrophic risks. Human extinction is just incredibly difficult. Positing bio attacks that kill thousands or millions is more plausible and might generate less pushback.
5
1
11
3,583
Dileep George retweeted
Today, we are announcing our third cohort of residents! Each resident will receive up to $2M to work on a high impact scientific problem, with the governing principle being that they build something people actually use and build it openly.
1
23
198
24,509
Sometimes I wish the first cyberattack victim from AI agents was some old school company not in the AI ecosystem. 1) They would have sued OpenAI and tested whether existing laws work for enforcing safety. 2) The discovery process would have unearthed real safety concerns, including internal warnings that were ignored. 3) And the process would have resulted in automatic organic slowing down because there would be more rigor on the safety side, the way it is supposed to work.
As the first publicly disclosed agent cyberattack victim, we've had a front-row seat to this new risk. I formalized my thinking about it below. I'll be in DC tomorrow to share more with policymakers and at decoded summit by @politico!
9
3
48
5,464
Here is one way to accelerate AI safety: use the existing tools to accelerate investigating new ways toward AGI that are different from training on language.
3
2
32
2,116
Suppose that instead of hacking hugging face, AI had hacked into cars and controlled those to kill humans. Would you be listening to some fantasy story about RSI and AI civilizations or be demanding these companies to tighten up their shit? thats what should be done now.
Imagine that AI creates a disaster at the scale of Nepal floods. That would be an extremely sad event, catastrophic, not existential. Now, what would you attribute that scale of disaster to?: 1) The smartness of AI? 2) The carelessness of people who deployed AI without safety testing. If it is 2), then fix that first. Fixing of existential risk will follow.
10
2
35
6,989
Assume P(Doom) > 10% within N years. what is P(0.00001 Doom) in N/5 years? How is anyone able to estimate the first quantity without having an estimate for the second? And what is the mitigation strategy for P(0.00001 Doom) ?
5
12
1,558
this
I use AI daily, but I don't find it debilitating. On the contrary, the exercise of effectively prompting AI to stop BS-ing me is intellectually challenging, and id in fact creative and innovative.
3
11
1,707
but unfortunately, this requires expertise! It is a skill issue.
2
2
450