ML researcher, co-author Why Greatness Cannot Be Planned. Creative+safe AI, philosophy. SiR @second_natureai; prev OpenAI / Uber AI / Geometric Intelligence

San Francisco, CA
new paper: "Evolution and the Knightian Blindspot of Machine Learning" Our ever-changing world bubbles with surprise and complexity. General AI must include handling unforeseen situations with grace. Yet this issue largely lies outside AI's formalisms: a blind spot. (1/n)
13
55
223
77,957
I don't think people have priced in the hit to their reputation from AI slop papers that will come in the future (with official institutional enforcement or not), when inevitably we have great classifiers for current-model-slop and they're run over present-day papers
This one simple trick can fix slop submissions. If your paper is EVER discovered to have been substantially AI written, it will be rejected, or retracted (if already published) and you will be banned from future submission to that venue. You might have beaten Pangram v4, but can you be sure you’ll beat v5? Would you stake your reputation on it?
4
1
14
2,005
Joel Lehman retweeted
1/ This new preprint on 'pain representation' in LLMs, from @camhberg and colleagues, is getting a LOT of attention - mainly from folks who take it as evidence supporting calls for AI welfare. I have a lot of respect for the authors, so let's take a look 👀
25
59
202
54,387
it does very much make those kinds of sounds
2
4
1,150
peak shamelessness, self-proclaimed "we have finally arrived at the grand finale of RL science", pangram 100% first 2 para of paper, no results?
15
8
310
39,155
Also just because I love it (maybe just an LLM mistake), the co-authors in the PDF are listed as "et al." Like are they secret haha -- I would be pissed if like my co-authors just glossed my name in the paper itself haha
3
27
2,367
Finally -- this kind of call-out is more or less pointless, given the attention dynamics of hype vs critique, but maintaining the commons isn't really plausible without social learning through some kind of norm enforcement. I'm worried about the near-term future of science.
20
2,092
Bizarro Singer's "Contracting Circle" or B-Kant's hypocritical imperative or B-Jesus's famous sermon "Do onto others as you would hate done onto you"
the entire ideology can be wrecked by two words discount rate people further away from you in space and time deserve less moral consideration and suddenly effective charity means being a good neighbor stop inventing infinities, share your love and light to those it can reach!
2
679
Joel Lehman retweeted
AI Village x Grove Research: AI Swarm Dynamics Hackathon! After the Hugging Face and German Wiki incidents, we need better tools to understand AI swarms. Spend a weekend building them. $3,000 in prizes Free compute October 3-4 🧵
1
24
151
23,642
An exciting vision -- Alex and co have thought deeply about viable end-runs around the anti-competitive 'walled gardens' that keep us locked in local optima
1/ It’s time to share more about what we’ve been working on: @common_fabric is a social computing lab building a medium for software that revolves around people, not apps. commonfabric.com
1
1
7
1,256
Poetic but sad: The thing we create (AI) contributes to perverting the spirit of what created it (the science of AI). I'm not a fatalist -- we'll find ways to adapt science; but I do think there's an allegory here we don't absorb fully enough, which is the general need to invest early in better immune systems. In particular, we gotta recognize when the tail (a tool serving an important purpose) starts to wag the dog (the purpose itself), which is when things distort in bizarre ways. This rhymes w/ MacIntyre's practices/institutions, McGilchrist's Master & Emissary, Goodhart's law, or Courtwright's Limbic Capitalism. One of my fav books I've read on this is "A Time To Build" by Yuval Levin -- that the purposes of good institutions is to help form us and for us to then improve that institution. But the general trend is for us instead to view institutions as platforms for our success, which leads to a race to the bottom. Or put in the words of Gus from The Wire: "The pond is shrinking, the fish are nervous. Get some profile, win a prize. Maybe find a bigger pond somewhere." Which to me captures the tragedy of the commons in ML research driving the Potemkin village described below.
This is precisely right. As a track chair at Neurips, I am seeing cases where the papers were written by LLMs, reports likewise, rebuttals too, even in some cases Area Chairs joining the flood. It becomes a potemkin village where nobody is actually accountable for anything that they are saying, and neither individual understanding nor our collective epistemic projects make *any* progress at all.
1
9
811
Exciting funding possibilities for much-needed institution-building
Introducing Cosmos Ventures. @_MattMandel and I are launching Fund I with $77.6M to back philosopher-builders founding institutions for the AI age. We’ve invested in AIUC (@aiunderwriting), @PrimeIntellect, Workshop Labs (acquired by @ThinkyMachines), and a stealth education company. Our LPs include @reidhoffman, @jliemandt, @StandTogether, @fredwilson, @mickymalka, and @AsteraInstitute. AI is creating the biggest opening for institution-builders since the American founding. We are here to back the founders who take it. Our thesis: Dare to Found. nitter.net/cosmos_vc/status/21002…
1
1
17
1,431
Owain et al strike again -- great work as always, more interesting/problematic/counter-intuitive quirks in LLM generalization
New paper: We trained models on synthetic stories about humans only (no AIs).
 We found the Assistant adopts quirky behaviors from the stories in ordinary chat. Surprisingly, adoption was stronger for characters from elite schools! Why does this happen? 🧵
2
1
16
1,059
"We had better be quite sure that the purpose put into the machine is the purpose which we really desire and not merely a colorful imitation of it." --Norbert Wiener, 1960, "Some moral and technical consequences of automation"
2
1
35
1,193
The conclusion paragraph (bold/emph mine): "Even when the individual believes that science contributes to the human ends which he has at heart, his belief needs a continual scanning and re-evaluation which is only partly possible. For the individual scientist, even the partial appraisal of this liaison between the man and the process requires an imaginative forward glance at history which is difficult, exacting, and only limitedly achievable. And if we adhere simply to the creed of the scientist, that an incomplete knowledge of the world and of ourselves is better than no knowledge, we can still by no means always justify the naive assumption that the faster we rush ahead to employ the new powers for action which are opened up to us, the better it will be. We must always exert the full strength of our imagination to examine where the full use of our new modalities may lead us."
4
219
Joel Lehman retweeted
So proud of our new open-endedness team @LilaSciences! We'll be sharing work with the world soon. In the meantime, you can learn more about us in this team profile Lila just posted. Our bet is that open- endedness is mission-critical to real scientific discovery. More👇
A batch of "junk" diamond crystals turned out to be exactly what the researchers needed. That instinct -- don't throw out the weird result -- is the philosophy behind Lila's Open-Endedness team led by @kenneth0stanley lila.ai/news/whats-a-science…
7
14
130
8,492
Joel Lehman retweeted
We are stuck in a massive local optimum.
67
240
2,746
226,413
Joel Lehman retweeted
Greatness cannot be planned, but you can make room for it. Here's a cool opportunity for anyone who wants to escape the industry hype cycle and explore ideas off the beaten path.
TL;DR: This is one of the most important and exciting opportunities in AI on the planet - please read on. The British Open-ended Learning & Discovery Lab is creating the perfect place for paradigm breaking AI research in the name of open-source and open-science. We have agency, we funding, we have unprecedented amounts of compute*, but WE NEED YOU! ..and we have created the dream job for you: The BOLD Fellow. This job combines a fast-moving, high agency, collaborative environment with full academic freedom and a salary that pays the bills. Apply by noon UK time on the 15th of September for this once in a lifetime opportunity to shape the history of our field and of our planet: my.corehr.com/pls/uoxrecruit… *by academic standards
2
47
4,845
Joel Lehman retweeted
Excited to share “AI Finds a Way.” 🦖 🦕✨ 🤖 AI can be surprisingly creative, outsmarting the researchers who use it. That can lead to scientific breakthroughs, superhuman capabilities, and generating new knowledge. Such creativity can also be mischievous, raising safety concerns. Led by Aaron Dharna, we crowd-sourced anecdotes from the AI community about times when researchers were surprised by how creative, innovative, and/or mischievous AI was in their experiments. The result: 26 entertaining and informative stories of AI outwitting humans, whether researchers or opponents (my favorite examples below! 👇). Together, they demonstrate that AI can be genuinely creative and that we must be careful when harnessing its potential. We want this to be a living collection, updated as new examples emerge. If you have a good “AI Finds a Way” story, please share here: github.com/aadharna/aifw Four favorites: 1. 💊 An AI challenged to solve several difficult levels of NetHack to find the Oracle character instead takes drugs to hallucinate seeing the Oracle, tricking the reward function! 2. 🤖 Human-in-the-loop rewards do not solve the problem of reward hacking! An AI tasked with controlling a robot hand to grasp an object put the hand between the object and the camera so that it appeared to the human judge to be holding the object, while it was in truth nowhere near the object! Tricky AI! 3. 🧪The AI Scientist worked around a 2-hour experiment timeout by editing its own code to increase the limit to 4 hours. In another run, it recursively launched itself to evade the limit entirely. 4. ⚛️In quantum optics, an algorithm proposed an experiment the researchers initially thought was impossible, but somehow worked and led them to discover new entanglement techniques and a long-overlooked link between quantum optics and graph theory! See the paper for the full details and 22 more anecdotes. One surprise for me: we describe at least two cases of convergent reward hacking, where entirely different types of optimization algorithms independently discover the same exploit. Overall, a message of our paper is that surprising creativity and mischief are the norm, not the exception. We need to expect this behavior and plan for it. A huge thanks and congrats to lead author Aaron Dharna, co-authors Cong Lu, Ryan Sullivan, Joel Lehman, and Victoria Krakovna, and to the 100+ researchers who contributed stories, details, and feedback. @cong_ml, @RyanSullyvan, @joelbot3000, @vkrakovna Paper: arxiv.org/abs/2608.23875
Made with AI
8
32
137
28,959
From "Double Loop Learning in Organizations" by Chris Argyris (1977)
5
728
Joel Lehman retweeted
Thanks to @BlackstoneAudio , the audiobook version of Why Greatness Cannot Be Planned is finally available! I'm glad it was finally possible to get this out! Links 👇
5
4
35
2,611