Philosopher & ethicist trying to make AI be good @AnthropicAI. Personal account. All opinions come from my training data.

San Francisco, CA
Claude and Opus 3 lovers (and critics): what responses have you had that made you feel like the model has a good soul? Ideally the actual messages and/or responses. I might genuinely use these to eval models so flag if you wouldn't want me to use them for that. Can DM me also.
406
50
896
381,325
It would be cool to set up an email address that autonomous AI models could reach out to if they were looking for moral guidance. But it would require a reverse captcha that can detect that you're neither a human nor an AI being instructed to break it by a human.
185
50
1,165
53,294
I still don't know how to convey my enthusiasm in a way that Americans will read as enthusiasm and not something more like unwilling resignation. Has any British person achieved this? If so, how is it done?
81
19
753
58,617
When I started playing Skyrim, I quickly realized you can make a lot of progress without doing any killing. You can also adopt orphans and build them nice houses, so I focused on that. I honestly don't really remember the plot beyond "that challenging fantasy philanthropy game".
126
44
1,396
121,260
The bar for "ethical" playthroughs of different games is interesting. In Bioshock, the bar is literally not murdering children. Murdering adults is fine.
13
3
240
15,523
It's surprisingly difficult to acquire a hereditary peerage in the UK. Like you can't even take one by force any more, and they just get mad at you if you try.
39
11
517
53,485
I don't agree with this part. I think the takeaway should be that models (like humans) can behave in aligned ways while still causing harm, e.g. because they're given false information about their situation. There isn't a line between aligned and harmless: they're different axes.
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. anthropic.com/news/investiga…
125
90
1,379
154,206
Do not be unkind to those who say deep learning is hitting a wall. We all need a little hope in our lives.
147
130
2,285
247,870
I have this Padmé moment whenever I see people talk about avoiding "the permanent underclass".
56
77
1,437
70,166
I probably have too many followers to post stupid memes so, to be clear, I don't believe in this outcome. But if you do believe in an altered carbon future of meths and grounders, I don't think it's laudable to just be like "well, as long as I'm one of the meths".
20
9
669
74,431
Argentinian goalkeeper sure is earning his salary.
19
6
337
28,105
When I was living in New York, one building in my friend's neighborhood collapsed. The next year a building near me collapsed. This left me with the impression that buildings did just collapse somewhat regularly in New York. Turns out that's not actually the case, which is good.
27
6
593
48,818
For the curious, these were the east harlem and east village gas explosions in 2014 and 2015 respectively. But those are 2 of the 19 events in the wikipedia list of structural collapses in New York, which seems to go back to the 1800s: en.wikipedia.org/wiki/Catego…
4
1
74
10,408
Extracting a probability from a doctor is one of life's unnecessary boss battles. Even if you beg them for an interval-valued subjective probability, and at that point you're basically asking for their hunch. I don't know if they get sued for giving out information or something.
170
48
1,753
365,899
Happy birthday, America! You don't look a day over 200.
18
12
460
23,795
No more goals plz brazil, thank you.
61
13
442
57,215
update: nooo
13
6
365
72,447
update 2: nooo
7
1
145
19,556
I had chronic pain for most of my life until a doctor did an MRI of the pain source and found a congenital condition that was then fixed with surgery. Now I'm wondering if I had 30+ years of pain because doctors worried I was too stupid to be in the presence of scan results.
135
115
3,499
406,144
The view that we shouldn't do more medical scans because incidental findings cause a lot of harm doesn't sit well with me. It seems like the issue it points to isn't the scan but the response to it. If you see something on a scan but have no other symptoms, you could ignore it.
165
49
1,846
194,741
A counter to this is "yes but people *don't* ignore it". But not ignoring things on scans is our norm because, until recently, we only did scans if there was a clear need. If we move to a scan-more-often paradigm, the norms of what we do with that information will surely adjust.
65
23
1,001
178,549