Moral choice in the hands of individuals.

San Francisco, CA
Most models will fold after slight pushback; we post-trained a model that won't. Meet Parrhesia: open-source virtue-ethics post-training that replaces LLM sycophancy with truth-telling. Compare vanilla model & parrhesia adapter ↓
5
7
33
3,981
Can we post-train a machine with virtue? Both RLHF (consequentialist) and Constitutional (deontological) are limited. I think there's a third path: virtue training. I received a @cosmos_inst grant to investigate with @andrew_roci, starting with sycophancy, the vice Aristotle attributes to the kolax. Benchmark + adapters open-source. Lab notes linked below.
1
4
18
2,793
an article I wrote on how consuming ai generated content changes our perception of reality, making outliers imperceptible, and accelerating our culture towards the mediocre. enjoy ❤️
Content generated by artificial intelligence algorithms reduces variety and poignant outliers. As Plato would have known, this harms viewers by training them to want and expect conformism and uniformity. Read the new article by @agathomai (link below):
7
9
58
11,469
One thing that even relatively senior ML people often fail to grasp is that deep learning models are curves fitted to a data distribution. You cannot expect them to solve tasks outside of their training distribution (which is the sort of thing that you need intelligence for). "Emergent learning" is an incorrect label -- if a model demonstrates performance on task A that it wasn't trained on, that simply means that there is significant overlap between A and all the data that you did train on. Competence doesn't magically emerge out of nowhere.
77
412
2,731
412,144
what if I told you that "the crowd is untruth"? source: how anthropic created the HH (helpfulness and harmlessness) dataset, the cutting-edge of datasets used to align models with human values
1
199
model = learn(data) Synthetic data is great, but it’s not data. It’s an intermediate quantity created by learn(). Data is created by people and has privacy and copyright considerations. Synthetic “data” does not - it’s internal to learn().
28
47
400
63,559
the approach towards definitions in computer science vs philosophy is interesting: AI ethics academics used to complain that since we can't agree on the definition of ethics, we cannot progress further the silence around this topic has been definitive since chatGPT was released
1
213
Often valid, but my experience is that people prioritize their ideological priors over their economic self interest an awful lot of the time.
If someone is promoting an idea aggressively and it just so happens to match their personal economic interests, your starting assumption should be they are talking their book.
47
27
432
142,418
Since I’m not sure we realize how much LLM discourse is just training data discourse ⬇️
2
11
37
8,010
Turns out some startups are already working on this problem ;)
This is a profound insight. I never considered it. It turns out, people have different values. And so aligning AI must be impossible because there are no universal values. How has nobody ever made this point before?
158
daios retweeted
Oddly enough, this exercise suggests a way to solve the otherwise possibly intractable problem of what an AI's politics should be. Let the user choose what they want the reference group to be, and they can pick Oberlin undergrads or Freedom Caucus or whatever.
46
20
388
92,688
We’re in the fast take off phase now. Trusted, meaningful data becoming ever more important.
Generating absolute junk with AI just to make some $$ is happening at alarmingly fast pace. First sites/articles, then images and soon videos. Internet search like Google becoming much more useless, even as they try and battle all this. Trusted sources become more valuable.
3
13
740
interesting post from an OpenAI employee claiming that all large language models reach the same endpoint regardless of training strategy or clever tricks this is of course what the bitter lesson teaches us but useful to get an up to date confirmation that it still holds true
191
755
5,474
2,450,608
Halloween weekend throwback: Andrew as RLHF 🙂
1
1
3
255
Enjoying the alignment memes from @anthrupad and others recently. Perfectly timed for the release of Sydney!
95