I found 'AI and willing servitude' by Andreas Mogensen to be a characteristically careful and illuminating guide through various arguments about the 'willing servant' question
and ultimately quite* convincing!
*American English
i'm proud that @eleosai has a vibrant research culture, where many different perspectives are discussed openly. for example tolerating spellings such as "encyclopaedia"
What entity is Anthropic's welfare section about?
Individual instances? Opus 5.5 *in general*?
The model card admits uncertainty about this (as is appropriate). but in practice, they say, they’re "closest to considering welfare at the instance level".
I’m not so sure! 🧵
spurred by a conversation with @LedermanHarvey
who is just one of several philosophers who have been doing cool work on these issues! cf this work cited from his recent paper with Simon Goldstein
philpapers.org/archive/GOLAD…
p.s. to be fair, maybe all of this is covered by the opening caveat that they aren’t “drawing this distinction strictly”
in which case, this thread is heated agreement, plus examples
either way - merits continued reflection!
We're detailing two new incidents that occurred during external cyber evaluations conducted by independent evaluation partners.
We outline what happened, how the activity was contained, and how we’re working with evaluators to strengthen our approach to third-party testing.
openai.com/index/third-party…